Radical Blog

A Functional Taxonomy of World Models

By Dr. Fei-Fei Li, Co-Founder and CEO of World Labs

http://image%202

This week, we share an excerpt from an essay by Radical Scientific Partner Fei-Fei Li, Co-founder and CEO of Radical Ventures portfolio company World Labs. Fei-Fei proposes a functional taxonomy of world models: renderers output pixels for human eyes, optimizing for visual fidelity without explicit grasp of 3D structure; simulators output state, the physically faithful representation that humans and programs can compute on; and planners output actions, deciding what an agent should do given an observation and a goal. With that framing in hand, Fei-Fei makes the case for why simulation is the linchpin of the three. Read the full essay here

Computer vision, robotics, reinforcement learning, and generative AI each claim to be building world models, and each means something quite different. A video model that produces gorgeous but physically impossible flames, a language model improvising a playable game, and a physics engine that faithfully simulates combustion all go by the same name.

The ancient Greeks could never agree on what the world was made of, whether fire, water, or indivisible atoms, because “world” was never a single thing. It was always a stand-in for whatever totality a given thinker needed to reason about. AI has inherited the same problem, at exactly the moment when the field needs precision.

Three Functions of the World Model

The first kind of world model is a renderer. The renderer is by far the most commercially mature. A number of image or text-to-video products are expanding in the consumer or enterprise markets rapidly. Google’s Nano Banana model has put renderer-quality image generation in the hands of potentially hundreds of millions of users. The technology is real, and the markets are real. Yet renderers optimize for visual plausibility rather than physical accuracy, and that ceiling matters. Their outputs are beautiful, but they cannot be trusted to design a building or train a robot.

The planner is the most intriguing and the most nascent, closely connected to the rapidly evolving field of robotic learning. The field has produced robotic demos in the last two years that look impressive in videos, but candor is required about what those demos actually show. Almost all have been confined to heavily constrained laboratory setups, with narrow object sets and short task horizons. None have been validated at the complexity, variability, or duration that real-world deployment demands. The gap between a compelling demo reel and a robot that reliably works in a kitchen, a warehouse, or an operating room remains vast. The commercial bets are nonetheless substantial. A wave of well-funded entrants is racing to ship general-purpose planning systems, while the largest infrastructure players are positioning planning atop broader simulation stacks. A robot that can plan is a robot that can work, and the entire industry is racing to be the one that gets there first.

Simulation is the bridge between the two. If language is an abstraction of the world and pixels are a projection of it, then geometry, physics, and dynamics are the world itself. A simulator must work at that level: the structural backbone from which both visual appearance (for renderers) and action consequences (for planners) can be derived.

A model that masters simulation can project its understanding into pixels for human consumption, and into action predictions for embodied agents. A model that masters only rendering, or only planning, cannot do either. The commercial surface area is enormous. NVIDIA’s Omniverse alone targets what the company estimates as more than a trillion dollars of addressable market in factories, warehouses, supply chains, and digital twins. Robotics training, autonomous vehicle testing, architectural visualization, engineering, and drug discovery all depend on something simulation-shaped.

The hardest open problems in the field live there too. Three-dimensional data with explicit geometry, material properties, and physical annotations is orders of magnitude scarcer than the internet video that renderers train on. The sim-to-real gap, which is the difference between how things behave in simulation and how they behave in reality, persists. Generative simulators introduce a new risk on top of that: AI-generated geometry can look correct while containing self-intersections or wrong scale that produce nonsensical physics. Multi-physics simulation at scale, where rigid bodies, deformable objects, fluids, and cloth all interact, remains orders of magnitude more expensive than single-domain simulation.

At World Labs, Marble is our first move into this territory. It takes multimodal prompts (text, image, video, or spatial sketch) and generates explorable 3D environments, outputting Gaussian splats for visual exploration alongside collision meshes a physics engine can operate on. But Marble is only the first chapter of a much longer arc being written across the field as the lines between rendering, simulation, and planning begin to collapse.

For the full taxonomy and where world models go next, read Fei-Fei’s essay in full here.

AI News This Week

  • The Impending, Inescapable Deluge of AI  (New York Times)

    The global AI infrastructure boom is accelerating at a staggering pace as tech giants race to capitalize on the technology’s scaling laws. The sheer magnitude of this physical build-out, projected to top $1 trillion globally by 2029, is so massive that Rob Wachen, co-founder of Radical Ventures portfolio company Etched, calls it “the largest scale infrastructure build-out in the history of humanity.” This immense influx of computing power is essential to unlocking the next frontier of artificial intelligence, according to Google’s Chief Scientist Jeff Dean. Dean argues that unprecedented scale is required not only to deploy autonomous AI agents capable of solving complex multistep problems, but to eventually achieve recursive self-improvement, where AI models can “fully automate the loop.”

  • Could Agentic AI Bring American Manufacturing Back?  (Forbes)

    Radical Ventures portfolio company P-1 AI is positioning its agentic engineering tool as an answer to a shrinking pool of American engineers, a shortage that aerospace and industrial firms say is pushing work offshore and stretching delivery timelines. Co-founder and CEO Paul Eremenko frames the company’s agent, Archie, as a junior teammate rather than a replacement, with one Archie per five-to-ten-person engineering team handling mechanical, electrical, thermal, and systems design. P-1 board member and former GE CEO Jeff Immelt argues that tools like Archie neutralize the wage arbitrage behind decades of outsourcing, restoring the business case for domestic production. 

  • The Rise of Million-Dollar Companies With Just One Employee  (WSJ)

    AI tools are lowering the bar to starting a company so far that a growing class of founders launch and scale without hiring anyone. Stripe data shows thousands of solo operators on its platform now clear $1 million in revenue, with their ranks doubling between 2023 and 2025, while the number crossing $10 million nearly tripled. New business applications in the information sector rose almost 45% over the past year, even as that sector posted the sharpest drop in applicants planning to hire. 

  • Why ‘Workforce Orchestrator’ is the Next Hit Job  (Financial Times)

    A new job class is taking shape as companies fold AI agents into human teams. The “hybrid workforce orchestrator” designs and directs mixed teams of people, agents, and robots, deciding which tasks stay with humans, which move to agents, and which skills warrant investment. Behavioral scientist Gleb Tsipursky argues it functions as governance rather than workforce planning, since an orchestrator effectively decides whose work an organization values. Adjacent roles he anticipates include AI coaching, algorithmic accountability, and an AI ombudsperson for employees contesting opaque automated decisions.

  • Research: Is Progressive Disclosure All You Need for Long-Context Agents?  (UC Davis/Zhejiang University/University of Hong Kong)

    Researchers have conducted the first controlled study of “progressive disclosure,” an AI technique that provides agents with brief summaries of massive documents that they can selectively expand to answer queries. While evaluating this method against raw-document navigation and hybrid retrievers, researchers found it offers little benefit for single-book tasks if the AI agent is already highly capable. However, when scaling up to massive tasks across multiple books where standard navigation completely fails, progressive disclosure becomes a asset for maintaining accuracy. The study also highlights the importance of simplicity, noting that adding multiple layers of progressive disclosure often degrades rather than improves model performance. Ultimately, the research validates this technique as a powerful tool for scaling AI to handle vast datasets, provided it utilizes a straightforward, single-layer design.

Radical Reads is edited by Ebin Tomy (Analyst, Radical Ventures)