Inherent, founded by DeepMind alumni, says its AI ‘teammate’ just outperformed Anthropic and OpenAI at replicating research

← Back to the feed

Inherent, founded by DeepMind alumni, says its AI ‘teammate’ just outperformed Anthropic and OpenAI at replicating research

TechCrunch · 3 hours ago

London AI lab Inherent, founded by former Google DeepMind staff, says its Faraday agent outperformed Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 at independently reproducing findings from published scientific papers. The company argues that the result matters because replication is a core scientific skill and a step towards its broader aim of building AI agents that can help make new discoveries, though the performance claim has not been independently verified in the article.

Faraday reportedly uses Qwen 3.6, a 27-billion-parameter model, rather than a much larger frontier model, and relies heavily on reinforcement learning to develop what Inherent calls “research taste”: choosing worthwhile experiments and designing them effectively. Inherent emerged from stealth only weeks earlier with a $50 million seed round, has roughly 12 employees working in King’s Cross, and says it uses OpenAI’s GPT-5.5 Codex for coding rather than building its own coding tool.

  • Inherent says Faraday beat larger rivals at replicating scientific research.
  • The agent uses a 27-billion-parameter Qwen model.
  • The company aims to build collaborative AI scientific agents.

New here? Start with this

AI systems are increasingly being tested on scientific work, rather than only on writing, coding or answering questions. One important task is replication: checking whether a published study’s findings can be reproduced by following its methods and running further experiments.

Inherent is a small London-based artificial intelligence company founded by former Google DeepMind employees. Its Faraday system is described as an AI “teammate” designed to plan and carry out research tasks, using an existing Chinese-made Qwen model alongside training methods intended to improve its choices about which experiments to run.

The company’s comparison involves systems from Anthropic and OpenAI, two leading developers of large AI models. If such tools can reliably support replication and eventually propose useful experiments, they could affect how researchers conduct scientific work; however, company-reported results need independent assessment to establish how broadly they apply.

Both sides, in good faith

The strongest fair case each way — we don't pick a winner.

The case for

Supporters argue that strong performance on independently reproducing published results would be a meaningful advance because replication requires reading literature, selecting methods, setting up sensible experiments and interpreting outcomes. They may also see the use of a relatively compact model and reinforcement learning as evidence that better agent training and experimental judgement, rather than sheer scale alone, can make AI more useful to working scientists and broaden access to capable research assistance.

The case against

Sceptics argue that a company’s own, unverified benchmark comparison is not yet reliable evidence that its system surpasses leading models, particularly when the task selection, prompts, tools, scoring and degree of human support may materially affect results. They may also contend that reproducing known findings is importantly different from generating robust new discoveries, and that claims about an AI research teammate should await independent evaluation across varied fields, including checks for reproducibility, safety and scientific validity.

AI Technology UK World

Read the full article at the source →