Radical Blog

Teaching AI to Ask the Right Questions

By Aaron Rosenberg, Partner

In this week’s feature, Partner Aaron Rosenberg examines Radical portfolio company Inherent‘s new AI agent, Faraday, an AI scientist that outperforms far larger frontier agents at replicating research papers.

Radical Ventures portfolio company Inherent has introduced Faraday, a 27-billion-parameter AI agent that combines the capabilities of frontier coding agents with a layer of scientific judgment. Trained via long-horizon reinforcement learning, Faraday outperforms both Claude Opus 4.8 and GPT-5.5 at replicating AI research papers – spanning fundamental ML research to AI for biology, materials science, and weather forecasting – and represents a step towards AI Scientists capable of innovation.

Reproducibility underpins scientific progress. The replicability of experimentation ensures the reliability of existing results and provides a basis for further lines of inquiry. Replication also typically illuminates previously underspecified details and thus requires hypothesis-driven exploration similar to the kind of open-ended research that leads to the discovery of new knowledge. Human researchers first build their judgment and sense of “taste” that later drives original work through the practice of replication. Coding agents seem well suited to this task, especially when experiments can be run in silico (without physical assays). Yet paper replication requires inferring missing details and navigating the unknown: rather than optimizing for a fixed objective, scientists combine domain knowledge with an intuition about which questions to ask; scientific discovery is ultimately a creative act.

To train that intuition, Inherent built Replica, a scalable suite of RL tasks that require an agent to replicate a figure from a published paper under fixed time and compute constraints, without ever seeing the original plot. Since papers only report what worked, not the winding process that eventually produced such success, this form of faithful replication represents a difficult test and strong proxy for research judgment. When a full experiment cannot fit the budget, the agent has to design a faithful scaled-down version, a decision that itself demands research taste.

Rather than optimizing Faraday to write better code, Inherent trained Faraday to leverage coding agents as tools. As a 27-billion-parameter model supervising far larger ones, Faraday demonstrates the returns to training a compact layer of scientific intelligence, rather than scaling a single, monolithic model. What’s more, the skills Faraday learns compound as the coding agents themselves improve. As those agents grow more capable (and more expensive), knowing how to direct them efficiently only becomes more valuable.

This work, Inherent’s first publication, also demonstrates the company’s commitment to keeping humans firmly in the loop. With an eye to safety, the team is investigating how their methods might advance scalable oversight and mitigate risks associated with autonomous agents. If you are interested in learning more, you can access the paper here.

To learn more about how Inherent is betting on recursive self-improvement at the organizational level, listen to the latest episode of Radical Talks featuring Co-Founder and CEO Tantum Collins. Now on YouTube, Spotify, or Apple Podcasts.

AI News This Week

  • The Rise of the 1 AM Job Interview  (Wired)

    Radical Ventures portfolio company Ribbon, which builds voice-AI recruiting software, has surfaced a shift in how people job-hunt. According to company data, 24% of its AI interviews happen between 10 pm and 2 am, rising to 35% for its manufacturing clients. Co-founder Arsham Ghahramani frames this as a new job interview option for candidates who cannot free up time during the day, including parents, hourly workers, and people tied to their current shifts. 

  • What Are Companies Getting for All That A.I. Spending?  (NYT)

    A new field its practitioners call “tokenomics” has emerged to measure the return companies get on the tokens they buy to run AI. Token spend currently has little pricing transparency and no centralized exchange or standard for what a token should cost or accomplish. Firms are moving from encouraging maximal AI use to scrambling to rein in runaway bills, and tools are being built to map token spend to concrete outcomes like features shipped. Some companies are testing metrics like “bionic head count,” which converts AI spend into salary-equivalent units to weigh output against margin. 

  • AI agents are Checking the Scientific Literature — and Spotting Decades-old Errors  (Nature)

    Researchers are turning AI tools on the scientific record itself, using them to audit papers, databases, and reference books at a scale humans cannot match. A theoretical chemist found that an AI model flagged boiling-point values in a 75-year-old reference database as wrong, and checking the original literature confirmed the model rather than the long-accepted numbers. Other efforts are scanning conference papers, rerunning experiments, and comparing results against what authors reported.

  • AI for Science Needs Reasoning, Not Just Data  (MIT Technology Review)

    Eric Schmidt and Suhas Mahesh argue that the next phase of AI-accelerated science will come from AI agents that model the iterative reasoning of research itself. An agent pairs a reasoning engine with tools it can call, letting it draft hypotheses, critique them, and refine the strongest candidates the way a working scientist does. They see agents lowering the cost of experimentation and improving reproducibility, since every step is logged, along with a lab’s accumulated institutional memory. 

  • Research: Can AI agents conduct open-ended AI research?  (Princeton/UK AISI/et al.)

    Given six days and thousands of dollars in compute to take on the central question from an unpublished NeurIPS paper, frontier agents handled the full engineering pipeline unaided, running hundreds of experiments, managing GPUs, and compiling complete drafts, which the authors read as early evidence that agents can already do the engineering that underpins AI research. The open frontier is the judgment-heavy core, knowing when to abandon a weak approach, when a result clears the bar, and how to use the full compute and time available, all areas where the authors point to concrete paths for stronger scaffolds and models. That research agenda complements this week’s feature on Inherent’s Faraday, which trains scientific judgment for the adjacent task of replicating existing papers.

Radical Reads is edited by Ebin Tomy (Analyst, Radical Ventures)