
Type a sentence: a foggy pier at dawn, gulls over grey water, a lighthouse blinking in the distance. Wait a beat, and the screen stops being a screen. You are on the pier. You take a step and the boards run out ahead of you; you tilt your head and the lighthouse beam swings past; you turn to check the shoreline behind you. When you turn back, the lighthouse is exactly where you left it, still blinking, in the same spot.
Here is the strange part. That pier did not exist a minute ago, and it exists nowhere but on your screen. No artist modeled it, no film crew shot it, no game studio coded its waves or its gravity. An AI conjured the whole place the instant you asked, and it is now painting each new moment on the fly, reacting to where you walk and what you look at, the way a video game reacts to a controller. Except nobody wrote this game's rules.
That is the idea behind a "world model." To see what is new about it, hold it against the two moving pictures you already know. A film is fixed: you press play and watch it happen, the same way every time. A video game is a place you can actually move through, but people hand-built every wall, rule and law of physics inside it. A world model is the third thing: a place you can move through, conjured from a sentence, that nobody built by hand and that remembers itself as you explore.
Not a movie you watch. Not a game someone coded. A place an AI dreams up around you, and holds steady while you look around. In barely two years that trick has gone from a lab curiosity to one of the best funded bets in AI, because a machine that can imagine a consistent world can also let a robot or a self-driving car practise inside one. Here is what these systems can really do today, and where the illusion still cracks.
Here is what happened
Google fired the starting gun. In December 2024, DeepMind showed Genie 2, which turned a single photo into a playable 3D world lasting under a minute. On August 5, 2025 came Genie 3, a real-time, interactive world model that renders explorable scenes at 720p and 24 frames a second, holds them together for several minutes, and remembers changes for about a minute: paint a wall, wander off, come back, and the paint is still wet. By February 2026 Google turned it into a product, Project Genie, for its AI Ultra subscribers. DeepMind calls it a stepping stone toward artificial general intelligence.
Fei-Fei Li is chasing it from another angle. The Stanford scientist behind ImageNet, the dataset that helped launch the deep learning boom, founded World Labs in 2023. In November 2025 it shipped Marble: type a sentence or upload a photo, and it builds a persistent, downloadable 3D scene you can edit and view in a Quest or Vision Pro headset. Unlike Genie, which improvises as you move, Marble builds the whole place upfront, closer to a stage set than a live show. World Labs has raised more than 1.2 billion dollars.
Nvidia is aiming it at machines. It launched its Cosmos world models at CES in January 2025 for robots and self-driving cars. Israeli startup Decart had already shown, in October 2024, a world model running a real-time, playable version of Minecraft called Oasis, learned purely from watching gameplay footage.
The field is crowded and global. Runway, Wayve, Tencent and Alibaba have all shipped world models since. Tencent open-sourced a 3D world generator downloaded millions of times, and Alibaba has shown a rival system that reportedly runs for 24 hours straight on a single graphics card.
How it works
Predict the next moment. A world model is an AI that has watched enough video of real life, or of games, to learn how things tend to behave, then predicts what happens next. Feed it the current scene plus an action, you turn left, you push a box, and it predicts the next frame, then the next, thousands of times a second. It is never given the rules. It infers them from data, the way it infers that dropped objects fall, because it has seen that pattern, not because anyone coded gravity.
Not the same as an AI video clip. A video generator, like OpenAI's Sora, predicts a fixed sequence of frames and hands you a clip to watch, no different in spirit from a film. A world model predicts one step at a time in response to your input, and keeps a working memory of what it built, so a room stays a room after you turn your back. DeepMind says Genie 3 started keeping objects in place on its own, without being trained to. It fell out of learning to predict well.
A practice ground for robots. The base system is a "foundation model," one large model trained on a huge spread of video, then adapted for narrower jobs, the same approach behind ChatGPT. Once trained, it can act as an endless simulator: practice scenes for a robot or a self-driving system before either goes near a real sidewalk. Carrying a skill from that simulation into a physical machine, called sim-to-real, is still the hardest step in the whole pipeline.
Why it matters
A "ChatGPT moment for robotics." Nvidia CEO Jensen Huang argues that just as language models unlocked chatbots, world models unlock machines that can act in physical space. His company's Cosmos platform, trained on 20 million hours of video, already counts Uber, XPENG and robotics firms Skild AI and Neura Robotics as early users, generating synthetic training scenes instead of relying only on scarce real footage.
Already reshaping self-driving. Britain's Wayve has scaled its GAIA world models to a 15-billion-parameter GAIA-3, whose simulated tests now track real road results closely enough to cut failed synthetic tests roughly fivefold. Scrappier rivals such as comma.ai chase the same payoff with far less hardware: cameras instead of costly sensor stacks.
The field's biggest names are betting on it. Yann LeCun left his post as Meta's chief AI scientist in November 2025, after twelve years, to build his own world model lab, raising more than a billion dollars within months. He predicts a new architecture takes over from today's chatbots within three to five years. Fei-Fei Li argues intelligence was never really about language alone: she calls perceiving and moving through space "the fundamental native ability" that let humans and animals act long before anyone spoke a sentence.
The honest catch
Strip away the demos and this is real research, not a finished product, and the cracks show fast.
Short memory. Genie 3 holds a scene together for a few minutes, not hours, and details you are not looking at can drift or vanish over time. It also has no sound yet.
Visual glitches. Reviewers have caught these systems freezing objects that should move, rendering water and snow oddly, and even showing people walking backward.
Heavy compute, narrow access. The models need enterprise-grade graphics chips to run, sim-to-real transfer is unproven at scale, and the best versions sit behind research previews or paid subscriptions, not open to everyone.
EDITOR'S TAKE
World models are the clearest answer yet to a real gap in AI: today's chatbots have read most of the internet but never opened a door. The new part is not the graphics, it is the persistence, a system that remembers a room instead of just rendering one. The honest read is that this sits where large language models sat around 2019: impressive in a demo, not yet reliable enough to trust with anything that matters. Watch three things this year: whether Project Genie's access widens past paid subscribers, whether Wayve and Nvidia's simulated robots keep matching real-world results as they scale, and whether LeCun's new lab or a Chinese open-source rival ships something that actually unseats the leaders. The wow is real. So is the gap between a demo and a holodeck.
Quick questions
What exactly is a world model?
A world model is an AI trained to predict what a scene will look like next, given what it currently looks like and what action you take inside it. That is different from memorizing facts or text: it is learning the rough physics and logic of a place, like the fact that water flows downhill or a door swings on a hinge. Because it predicts one step at a time, you can steer it in real time, unlike a normal video clip. Google DeepMind's Genie, Fei-Fei Li's World Labs, Nvidia's Cosmos and several Chinese labs are all building versions of this idea, aimed at everything from games to robots.
How is this different from a tool like Sora that also makes AI video?
Sora and similar tools generate a fixed clip: you type a prompt, wait, and get a finished video you can only watch back, the same as any film. A world model like Genie 3 or Marble instead keeps generating in response to what you do, moment by moment, so you can turn around, backtrack or linger, and the scene keeps up with you. The two technologies share a lot of the same underlying research, and OpenAI itself has described Sora as a step toward "world simulators." The practical difference today is interactivity and memory: a world model tries to remember the room you just left, a video generator does not need to.
Can I actually try one of these myself?
Yes, though access is limited and mostly paid. Google's Project Genie is available to Google AI Ultra subscribers in the United States, a subscription that costs 250 dollars a month and includes other Google AI tools. Fei-Fei Li's World Labs offers Marble on a freemium basis, with a free tier plus paid plans for higher-resolution exports, and it works with Quest and Vision Pro headsets. Decart's Oasis and several open-source Chinese models, including Tencent's Hunyuan3D world model, can be tried for free or run on capable home hardware, though expect rough edges, since none of this is a finished consumer product yet.
Sources
Genie 3: A new frontier for world models: DeepMind's own announcement, with specs on resolution, memory and object permanence.
Genie 2: A large-scale foundation world model: DeepMind's earlier Genie 2 release, for comparison.
NVIDIA launches Cosmos world foundation model platform: Nvidia's own announcement and Jensen Huang quotes on physical AI.
Wayve launches GAIA-3: Wayve's release on its self-driving world model and simulated-versus-real test results.
Fei-Fei Li's World Labs speeds up the world model race with Marble: Coverage of World Labs' first commercial product launch.
Decart's AI simulates a real-time, playable version of Minecraft: Early proof that a world model could run a live, playable game.
Frontier Signal explains frontier technology in plain English. Company and agency figures should be independently verified. This is general information, not investment or professional advice.

