Generalist AI is the first company to demonstrate clear scaling laws in robotics, with its GEN-1.5 model continuing to show smooth performance improvements after eight months of pretraining.

Last week, Radical portfolio company Generalist AI released its latest robot foundation model, GEN-1.5. Researchers from OpenAI, Google DeepMind, Waymo, Nvidia, Prometheus and other top labs hailed GEN-1.5 as the “GPT-3 moment for robotics.”

What exactly does it mean for GEN-1.5 to have achieved a GPT-3 moment for robotics?

The 2020 OpenAI paper that introduced GPT-3 to the world was titled “Large Language Models Are Few-Shot Learners”. The significance of this title was that GPT-3, once initially pretrained, could learn to complete novel tasks competently based on just a few examples (via either in-context learning or efficient fine-tuning), rather than needing to be shown thousands or tens of thousands of examples. This made GPT-3 a “general intelligence” in a way that no language model that had come before it had been.

GEN-1.5 represents a parallel milestone for robotics. It stands as the first time that a robotic AI model has the ability to learn to complete novel tasks not contained within its training data based on only brief demonstrations: a few seconds of sensorimotor examples, or a single demonstration recorded in simulation (even though the model was not trained on any simulation data), or a person demonstrating a task live with their own hands in front of the robot’s cameras. Generalist coins a memorable term for this: physical prompting, analogous to prompting an LLM.

Generalist’s GEN-1.5 paper title is thus a fitting homage: “Embodied Foundation Models Are One-Shot Learners”.

GEN-1.5 (like GPT-3) possesses other startling and unexpected emergent capabilities, like the ability to use tools it has never seen before. One example: after being taught a clean-up task using a brush, the model is able to complete the same task using a banana and then a dust-pan, using these tools in different ways than it had been using the brush in order to achieve the same outcome, even though it had never seen these tools before. Such “improvisational intelligence” is possible only because the model has internalized and absorbed rich, generalized sensorimotor knowledge about the physical world.

Over the past half-decade, large language models have taken the world by storm as their ability to automate digital tasks has grown increasingly generalized and sophisticated. Robot foundation models are poised to follow a similar trajectory, but in this case for the physical world. Just like LLMs over the past few years, these models are beginning to improve at breathtaking speed. Buckle up!

Learn more by reading Generalist’s technical blog on GEN-1.5 on their website.

AI News This Week