World Labs, the spatial intelligence company founded by AI pioneer Fei-Fei Li, introduced Atlas on September 1, 2026, describing it as its next-generation "omni" world model. In a company blog post, the team said Atlas is pretrained from scratch to natively operate on text, images, video, and 3D — a single model built to generate, reconstruct, and simulate worlds rather than specialize in any one of those tasks.
The launch is a milestone for the world models thesis that Li has championed since founding the company: that the next frontier of AI lies not just in language, but in systems that understand how worlds appear, behave, and evolve. World Labs frames the technology as the foundation for rendering imagined worlds for creative users, simulating the real world in high fidelity, and helping robots plan actions. For more on this and other frontier research efforts, explore our AI industry coverage.
One Model, a Shared Spatial Context
Architecturally, Atlas is what World Labs calls a multimodal autoregressive diffusion transformer. Like a large language model, it first encodes its inputs into a context and then generates outputs conditioned on that context. The difference is spatial: each input image is grounded at a 3D position in space, forming what the company calls a spatial context, and Atlas generates what comes next while staying 3D-consistent with everything it has seen — imagining what lies beyond it.
That design enables capabilities that would be awkward to bolt onto conventional video generators. Two unrelated reference images can be placed into the context at different 3D positions, and Atlas will generate a world that smoothly interpolates between them, inventing doorways, hallways, and transitions that plausibly bridge the two scenes. The company also said the model is built to scale, with performance improving as training compute increases.
Pixel-Perfect Camera Control and Long-Form Video
Atlas's marquee generation feature is camera-controlled video. From one to six input images, the model can generate up to one minute of video at 1440p resolution, following precisely specified camera paths. World Labs emphasizes that camera geometry is a native input type — going beyond coarse text prompts — so creators can frame every shot and control every motion. In demos, the team hand-designed camera paths through a scene to produce a coherent one-minute clip, describing the experience as being "in the director's chair" rather than pulling the lever of a slot machine.
The model also generates images and 360-degree panoramas from text, with the ability to follow complex prompts, render text, and produce a variety of visual styles.
Reconstruction That Challenges Specialist Models
On the reconstruction side, World Labs makes a bold claim: Atlas reconstructs real-world scenes from as few as two or three ordinary images, "outperforming state-of-the-art results by models specially trained only for 3D reconstruction." The company calls novel view synthesis from sparse inputs a decades-old fundamental problem in 3D computer vision.
Atlas can also ingest more than a hundred input images for faithful recreation of real environments. In one demonstration, the team rebuilt Stanford's Main Quad piece by piece from two to twenty-five ground-level photos, then generated aerial paths flying high above the campus — filling in views no camera ever captured.
Critically for practical use, Atlas does not stop at 2D frames. It natively operates on image frames and 3D depth maps, outputting worlds as point clouds or 3D Gaussian splats that render on-device at high resolution and frame rates. That is the same representation used in Marble, World Labs' first product, so Atlas outputs plug directly into the company's existing pipeline.
From Bullet Time to Robot Training
Atlas's simulation capabilities may matter most for robotics. The company showed that footage from as few as three ordinary cell-phone cameras — mounted with tripods and clamps that fit in a backpack — is enough for Atlas to reconstruct a scene and enable "bullet time" reframing from impossible angles, no specialized capture studio required.
For Real-to-Sim workflows, Atlas reconstructs spaces from short phone videos and then generates the RGB and depth data a simulated robot's body-mounted cameras would observe along a path — with the world and the robot's view of it coming from the same model. The company demonstrated navigation across large scanned environments using 24 frames each, plus manipulation scenarios where simulated tasks can be varied by changing objects, positions, and scenes once a task is simulated.
Why It Matters
World models are increasingly viewed as a complement to large language models: LLMs reason about descriptions of the world, while world models aim to model the world itself — its geometry, physics, and dynamics over time. If Atlas's claims hold up under external scrutiny, applications span VFX, gaming, design, engineering, and robot training, where synthetic but physically plausible environments could dramatically cut the cost of teaching machines to act.
The company behind the model is notable in its own right. World Labs was founded by Fei-Fei Li — the Stanford professor whose creation of ImageNet helped ignite the deep learning era — alongside Justin Johnson, Ben Mildenhall, and Christoph Lassner, and its investor list includes a16z, Nvidia, AMD, Intel, Adobe, Salesforce, and Samsung, among others. Marble, the company's first product, lets users create persistent 3D worlds from images, video, text, and 3D layouts, and World Labs said Atlas will power future versions of it.
Atlas is not yet generally available: the company is taking requests for early access. Still, the release marks one of the most concrete steps yet toward AI systems that don't merely describe the physical world, but hold a consistent, manipulable model of it.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →