Li Fei-Fei Explains Atlas, World Labs’ New-View Prediction Model
Summary
In an a16z interview, Li Fei-Fei and World Labs co-founders explain Atlas, the company’s new multimodal world model. Its core operation is novel-view prediction: given one or more images, videos or scene descriptions with camera poses, Atlas predicts what a scene should look like from another position in space and time. The model combines generation and reconstruction in one architecture, processing text, images, video, depth and camera information while producing RGB frames or 3D outputs. The team says Atlas can reconstruct spaces from sparse captures, reducing inputs from hundreds or thousands of images to roughly three to 25 images in some demonstrations, and can create controlled bullet-time sequences from three iPhones. Compared with World Labs’ earlier Marble model, which primarily output Gaussian-splat 3D worlds and was limited in input context, Atlas uses novel-view prediction as its basic representation and can accept substantially more visual context. The founders describe this as a step toward spatial intelligence, where models generate, understand, simulate, edit and interact with persistent environments. They see applications in film, games, design, architecture and construction, as well as robot training through real-to-sim conversion and randomized simulation data. Atlas already includes limited dynamic content, but the released post-training emphasizes static scenes; stronger dynamics, editing and interaction remain future priorities. The founders also argue that generative novel-view prediction could be as fundamental to intelligence as next-token prediction, though that claim remains a hypothesis for future models to test.