“World model” has become one of the most-used — and least-agreed-upon — terms in AI. It means one thing in computer vision, another in robotics, another again in reinforcement learning and generative AI. If the field is going to build toward spatial intelligence, it needs to be precise about what these systems actually do.

Rather than defining world models by how they are built, it is more useful to classify them by function. Three distinct roles emerge.

Renderers

Renderers produce visually plausible images and video optimized for human perception. They are commercially mature and can be strikingly convincing — but they are not physically accurate, which means they cannot be relied upon to guide engineering or robotics.

Planners

Planners generate action predictions for embodied agents. This capability is still nascent; most robotic demonstrations remain confined to heavily constrained laboratory setups, with limited validation in the messy conditions of the real world.

Simulators

Simulators are the bridge between the two. By capturing geometry, physics, and dynamics, they can both generate pixels and predict the outcomes of actions — something neither renderers nor planners can do alone.

Why simulation matters

Getting simulation right unlocks trillion-dollar opportunities across digital twins, robotics training, autonomous vehicles, and drug discovery. Real obstacles remain: 3D training data is scarce, the sim-to-real gap is stubborn, and AI-generated geometry still produces artifacts.

This is the problem World Labs is built to solve. Our first offering, Marble, takes multimodal prompts and produces explorable 3D environments — represented both as Gaussian splats for visual fidelity and as physics-compatible collision meshes for interaction.