Wednesday, July 29, 2026

News

Google DeepMind Researcher: LLMs Can't Make Einstein's Creative Leap

ResearchPatryk Raba
Google DeepMind Researcher: LLMs Can't Make Einstein's Creative Leap
Fot. Underwood and Underwood, New York, Wikimedia Commons (Public domain)

Tom Zahavy of Google DeepMind has published a position paper presented at the ICML 2026 conference arguing that today's large language models cannot make the abductive creative leap that allowed Einstein to formulate general relativity.

Contents
  1. Three Types of Reasoning
  2. The Einstein Example
  3. A Voice From Inside the Lab
  4. What the Author Proposes
  5. What It Means for Practitioners

A researcher who works day to day on some of the most advanced AI systems has publicly declared that those same systems cannot do something Albert Einstein did more than a hundred years ago without a computer. Tom Zahavy of Google DeepMind, a co-author of the AlphaProof work, presented a position paper titled "LLMs can't jump" at the ICML 2026 conference, arguing that today's large language models are structurally incapable of creating new scientific theories from scratch.

Three Types of Reasoning

Zahavy builds his argument on a distinction between three types of reasoning: induction, deduction and abduction. Induction is statistical pattern recognition, something today's language models are already very good at. Deduction is formal proof from established premises, and models like the systems DeepMind builds for mathematics are rapidly getting better at it.

The third type is abduction: the ability to generate an entirely new explanatory hypothesis where no established premises previously existed. According to the author, this is precisely the mechanism that remains beyond the reach of current AI systems. A model can prove a theorem from given axioms, but it cannot invent those axioms from nothing on its own.

The Einstein Example

The paper's central case study is Einstein's path to general relativity. Zahavy describes scientific discovery as a cyclical process: an intuitive leap from sensory experience to new axioms, followed only afterward by logical deduction of conclusions from those axioms. In Einstein's case, the observational data available at the time he formulated the theory was scarce, yet he still managed to derive a model of gravity that was mathematically and physically coherent.

According to the author, this case undermines a claim popular in AI circles that creativity is essentially a form of data compression, meaning statistical generalization from large sets of observations. Since Einstein arrived at a groundbreaking theory with minimal data, data compression alone does not explain where his theory came from.

LLMs are structurally incapable of the abductive 'leap' needed to establish first premises - Tom Zahavy, Google DeepMind

A Voice From Inside the Lab

The paper draws attention mainly because its author is not an outside critic of the AI industry but a researcher at one of the labs investing most heavily in advancing the capabilities of language models. Zahavy co-created AlphaProof, DeepMind's system that in 2024 achieved a score equivalent to a silver medal at the International Mathematical Olympiad, so it is hard to accuse him of lacking insight into what today's models can actually do.

The publication arrived at a moment when AI labs, including DeepMind itself, are publicly declaring ambitions reaching toward artificial general intelligence and superintelligence within a few years. Zahavy's position calls into question one of the pillars of those announcements: the assumption that scaling current architectures will, on its own, lead to the ability to generate new scientific theories.

What the Author Proposes

Zahavy does not stop at criticism. As a way forward, he proposes building physically consistent, multimodal world models, meaning systems that would learn not just from text but from direct, multi-sensory contact with a physical reality governed by the laws of nature. In his view, it is this lack of sensory grounding, not the transformer architecture itself, that accounts for models' inability to make the abductive leap.

In other words, according to the author the problem will not disappear with the next, larger versions of today's language models trained mainly on text from the internet. What would be needed is a fundamentally different path of learning, closer to how humans build physical intuition through experience rather than by reading descriptions of that experience.

What It Means for Practitioners

For companies and research teams currently using language models to support scientific work, Zahavy's paper is a signal of caution against marketing promises of an "AI scientist" independently generating breakthrough theories. This does not undermine the usefulness of models in deduction, meaning proving theorems, verifying hypotheses or processing huge datasets, but it suggests that formulating entirely new theoretical frameworks will, for now, remain the domain of humans.

The paper was accepted as a poster at ICML 2026, one of the most important machine learning conferences, meaning it passed peer review from the scientific community rather than being merely the voice of a lone researcher published outside official channels. Discussion around the text, including on social media and researchers' blogs, suggests Zahavy's argument has touched on a significant dispute within the AI community over the limits of today's architectures.

Share: