World Labs introduced Atlas on September 1, 2026, calling it a next-generation omni world model for spatial intelligence. The company said it pretrained Atlas from scratch to work natively on text, images, video, and 3D. The architecture is a multimodal autoregressive diffusion transformer that packs those inputs into a shared spatial context, then generates what comes next while staying consistent in 3D.

Atlas is meant to cover world generation, reconstruction, and simulation in one model. World Labs said it will power future versions of Marble, the company’s existing 3D world product, and opened a request list for early access.

Generate, reconstruct, then simulate

For camera-controlled generation, Atlas takes one or more reference images and produces new views at camera positions the user specifies. The company says it can emit up to one minute of video at 1440p. Camera geometry is a native input, not a text hint, so a path can be designed shot by shot. Atlas can also place unrelated reference images in 3D space and interpolate a world between them.

On reconstruction, Atlas builds a scene from one image or from dozens. World Labs says two or three views are often enough for a faithful result, and that the model can also consume more than a hundred images when the job is an exact location rather than an imagined fill. It outputs 2D frames and explicit 3D depth, which the company is pointing at robotics, games, design, and visual effects.

Space-time simulation lets Atlas reframe input video and support real-to-sim work for robots. It can also generate images and 360 panoramas from text. World Labs said performance has improved with training compute and that it expects that trend to continue.

Decoded Take

Atlas is World Labs trying to collapse three product categories (novel-view video, photogrammetry, and robot sim) into one model people request access to. That is a research flex and a distribution problem. Marble already exists. If Atlas only makes prettier Marble worlds, it is a version bump after a $1 billion raise. If sparse-view reconstruction and real-to-sim actually land in a robotics or VFX pipeline, the omni claim starts to matter. Watch whether early access names a camera-control API or just a waitlist, whether explicit 3D outputs are meshes customers can edit, and whether the one-minute 1440p demo survives outside the company’s own camera paths.