World Labs Atlas: What the Omni World Model Means for AI Agencies
What happened. On September 1, 2026, World Labs — co-founded by Stanford AI professor Fei-Fei Li — unveiled Atlas, its next-generation world model [1][2]. World Labs calls it an "omni model" that operates natively on text, images, video, and 3D: it generates up to one minute of 1440p video with pixel-perfect camera control from one or more input images, reconstructs real scenes as explicit 3D, and turns a few phone cameras into a "bullet time" capture studio [1]. For agencies selling video and 3D work, Atlas is a new tool category: a camera you direct, a scene you can rebuild, and a pipeline that connects both.
What Is the Atlas Omni World Model?
Atlas is a multimodal autoregressive diffusion transformer that combines text, images, video, and 3D into a shared spatial context, with each image grounded at a 3D position [1][2]. Instead of generating a flat 2D clip, it learns where things are in a scene, so it can move a camera through that scene, generate views no camera actually filmed, and output 3D structure another tool can use [1]. Atlas builds on Marble (public since November 2025) and the January 2026 World API; the company says Atlas will power future versions of Marble [2][1]. World Labs has raised $1.23 billion — $230 million at founding in 2024, then a $1 billion February 2026 round from Autodesk ($200M), AMD, NVIDIA, and Fidelity, at a Forbes-reported $5 billion valuation [2][8].
Why Agencies Should Care: Video and 3D Work
Most AI video tools are text-to-video: type a prompt, get frames. Atlas treats the camera as a native input. World Labs notes competing models "do not accept cameras as a native input format," so camera paths must be written into text prompts — and that Atlas "outperforms recent video models at camera-controlled generation," with the advantage growing as camera trajectories get more complex [1]. Client briefs are shot lists, not prompts: hero product moves, flythroughs, reframes, and virtual walkthroughs are all camera directions.
- Up to 1 minute of 1440p video from one to six input images, with hand-designed camera paths [1].
- Explicit 3D output — point clouds or 3D Gaussian splats, the same representation Marble uses — so a generated world can drop into downstream 3D tools [5].
- Bullet-time reframing from as few as three phone cameras, plus Real-to-Sim workflows for robotics clients [1].
This is the world model AI video category taking commercial form — we track the full field in our AI agent tools roundup.
Capabilities
- Camera-controlled generation. Images and videos from one or more images with pixel-perfect camera control; up to 1 minute at 1440p [1].
- Spatial reconstruction. Real scenes from one to dozens of input images, with novel-view frames and explicit 3D; faithful reconstructions from as few as two or three images, and spatial context that holds over a hundred images [1][5].
- Space-time simulation. Multi-phone footage into a "bullet time" studio, plus Real-to-Sim for robotics training [1].
- Image and 360-panorama generation from text or image prompts [1].
Pricing and Availability
Atlas is entering early access with select partners, with a request-access form on the launch post [1][5]. There is no public pricing as of September 1, 2026; the World Labs platform is login-gated and Marble's public page shows no pricing [10][9]. Treat Atlas as partner-access today, a priced product later — the same arc Marble followed.
Agency Use Cases for Atlas
- Virtual tours. The Stanford Main Quad demo generates aerial flythroughs from two to twenty-five ground-level photos — a template for real-estate, hospitality, and campus clients [1].
- Product and spatial visualization. Reconstruction from two or three images, scalable to environments with over a hundred [1].
- Camera-controlled 3D video. Hand-design a path for up to a minute of 1440p video; reframe footage from three to five phones into impossible angles [1][6].
- Real-to-Sim robotics. RGB and depth from simulated robot-mounted cameras for navigation and manipulation training [1].
Caveat: these are World Labs capability demos, not published customer case studies [5].
Atlas vs Sora vs Veo in 2026
No mainstream outlet has published a head-to-head as of September 1, 2026 [5], so the comparison rests on World Labs' own claims and on architecture — treat performance numbers as vendor-reported:
- Camera control. Atlas takes camera geometry as a native input; Sora- and Veo-class models need camera paths described in text [1].
- 3D output. Atlas reconstructs real scenes and outputs explicit 3D (point clouds, 3D Gaussian splats); text-to-video models output 2D frames only [5].
- Benchmarks. In World Labs' camera-follow test with third-party human raters, Atlas beat MiniMax H3 75% of the time and Seedance 2.5 94% of the time; every figure is World Labs' own measurement [5].
- Still use Sora and Veo for pure text-to-video at scale. For the per-model breakdown, see our AI video tools comparison.
Limitations to Know
- No paper, arXiv entry, model card, or code shipped with the launch [5].
- Videos are generated from one to six input images, constraining longer, multi-scene clips [1].
- Marble-era guidance (per Forbes): best on interiors and realistic 3D; illustration-style prompts and exteriors are less reliable [8].
Building an agency video or 3D offer? Map the world-model category before you quote the stack.
Open the AI agent tools roundup →FAQ
What can I make with Atlas world model?
Camera-controlled video up to 1 minute at 1440p from one to six input images; explicit 3D reconstruction (point clouds or 3D Gaussian splats) from as few as two or three images; virtual tours, 360 panoramas, and bullet-time reframes of phone-shot video.
What is World Labs Atlas?
Atlas is World Labs' next-generation world model, unveiled September 1, 2026: an omni model pretrained to operate natively on text, images, video, and 3D, with pixel-perfect camera control, up to one minute of 1440p video, and explicit 3D reconstruction.
Is Atlas available and how much does it cost?
Atlas is in early access with select partners via a request-access form. No public pricing exists as of September 1, 2026.
How does Atlas compare to Sora and Veo in 2026?
No published head-to-head exists as of September 1, 2026. World Labs says Atlas outperforms recent video models at camera-controlled generation and accepts camera geometry as a native input, while Sora- and Veo-class models need camera paths in text — and Atlas outputs explicit 3D, not just 2D frames.
What is a world model in AI?
An AI system that learns spatial and temporal understanding of scenes — where objects are and how they move — so it can generate novel views, simulate motion, and reconstruct 3D structure instead of producing flat 2D clips.
Sources
- World Labs: Atlas — A World Model for Spatial Intelligence (Sept 1, 2026)
- Cryptobriefing: World Labs unveils Atlas, an omni world model (Sept 1, 2026)
- Radiance Fields: World Labs Announces New World Model, Atlas (Sept 1, 2026)
- David Pantera (World Labs product): Atlas short film demo (Sept 1, 2026)
- Forbes: World Labs $1 Billion Bet To Advance Spatial Intelligence (Mar 2026)
- World Labs Platform
- Marble product page
Accuracy note: September 1 facts (omni architecture, 1440p up to one minute, pixel-perfect camera control, one-to-six input images, explicit 3D output, bullet-time, early access, no public pricing, $1.23B funding) come from World Labs' Atlas announcement, corroborated by Cryptobriefing and Radiance Fields; the $5B valuation and interiors-vs-exteriors limitation come from Forbes (March 2026). No mainstream Atlas-vs-Sora/Veo head-to-head existed at research time; the 75%/94% figures are Radiance Fields' reading of World Labs' chart. Refresh triggers: public pricing, general availability, third-party benchmarks, or a named customer case study.