Researchers at the Institute of Automation under the Chinese Academy of Sciences (CASIA) published PhiZero, a world model that reasons in a learned discrete ‘physical language’ before rendering video. The approach collapses the number of tokens needed to represent a four-second clip by a factor of about 175…
Source







and then