# Zahavy's "LLMs Can't Jump" as Resource for Generating Philosophy
[[Tom Zahavy]]'s "LLMs Can't Jump" (2026, [[Google DeepMind]]) argues that LLMs are structurally incapable of the abductive "Jump" required for scientific invention. The paper is a resource for the [[Sessions/Generating Philosophy|Generating Philosophy]] project — not as an opponent, but as a contrastive foil whose diagnosis of LLM failure illuminates why philosophy is different.
## Zahavy's Framework
Zahavy adopts [[Charles Sanders Peirce|Peirce]]'s tripartite distinction:
- **Induction** (Case + Result → Rule): statistical pattern-matching. LLMs have mastered this.
- **Deduction** (Rule + Case → Result): formal derivation from axioms. AI is conquering this ([[AlphaProof]], etc.).
- **Abduction** (Rule + Result → Case): the creative leap to new explanatory hypotheses. LLMs cannot do this.
The bottleneck is [[Albert Einstein|Einstein]]'s E→A Jump: the translation from Sense Experience to a System of Axioms. Zahavy argues this requires **embodied simulation** — [[manipulative abduction]] in [[Lorenzo Magnani|Magnani]]'s sense — grounding abstract symbols in physical sensation. Einstein's "happiest thought" (the equivalence of gravity and acceleration) came from simulating the *feeling* of a falling observer, not from compressing data or deriving theorems.
LLMs are "high-dimensional [[Chinese Room|Chinese Rooms]]" that manipulate symbols without access to physical referents.
## Structural Complementarity with the Generating Philosophy Argument
Zahavy's failure diagnosis is domain-specific. The bottleneck — sensory→symbolic translation — applies where the object of study is external material reality. In philosophy, this bottleneck doesn't exist (or narrows dramatically), because philosophy's "sensory experience" is already symbolic: the [[Space of Reasons]], inferential relations, argumentative structures.
The [[Notes/Philosophy as self-grounding domain|self-grounding thesis]]: where Zahavy's physicist needs to feel what it's like to fall in order to generate the Equivalence Principle, the philosopher works with concepts, arguments, and their relations — material that IS the training data.
Zahavy's anti-compression argument also helps. He argues [[Jürgen Schmidhuber|Schmidhuber]]'s "creativity as compression" fails for physics because Einstein had no dataset to compress (the Newtonian loss function was near-zero). But in philosophy, the creative material IS richly represented in training data. The [[Notes/Dialectical saturation thesis|dialectical saturation thesis]] — that philosophical corpora are saturated with argumentative patterns, evaluative standards, and success/failure signals — is precisely the kind of claim that Zahavy's framework leaves room for.
## Planned Uses (Options A + B + C Combined)
**Section 0 (Introduction):** Pair Zahavy with [[Luciano Floridi|Floridi]] as the two strongest recent critiques from different angles — Floridi epistemological (stochastic core, abductive appearance), Zahavy cognitive-architectural (no embodied simulation for the Jump). Both are right about their domains. The paper asks whether their conclusions generalise to philosophy.
**Section 2 (Abduction and Philosophy):** After presenting [[Timothy Williamson|Williamson]] on abduction, use Zahavy to show what abduction looks like in physics (embodied, sensory, requiring the Jump) vs what it looks like in philosophy (comparative theory evaluation, textually checkable). The [[Chinese Room]] objection presupposes a gap between symbol and referent; in philosophy, that gap narrows to near-zero because philosophical referents — theoretical virtues, dialectical structures, inferential relations — are themselves constituted by relations between concepts as expressed in text.
**Section 3 (Learning the Game):** Use the anti-compression argument to motivate why philosophy's data-richness matters. Zahavy shows compression fails where data is scarce and the creative leap goes beyond available information. It does NOT fail where the creative material is redundantly encoded in the training corpus. Philosophy, unlike physics, also has **internal error signals** (logical inconsistency, dialectical inadequacy, failure of integration) — unlike Einstein's situation where the Newtonian loss function was near-zero.
## Zahavy's Own Domain-Specificity Admission
In the conclusion: "this proposal is specifically tailored to the physical sciences, where the object of study is external material reality. In abstract domains such as Mathematics or Computer Science, the Sense Experience (E) may be grounded in high-dimensional topology or have other goals such as generality or minimality."
He explicitly restricts his argument. He doesn't mention philosophy, but the implication is clear: his bottleneck doesn't straightforwardly apply to domains where the object of study isn't external material reality.
## Complications
1. **Status as source.** A [[Google DeepMind]] position paper, not peer-reviewed philosophy. Zahavy holds a BSc in physics and electrical engineering; his doctoral work is in AI/Deep RL. The philosophical terminology (especially around abduction) is less precise than Floridi's. Better as a contrastive resource than as a section-length interlocutor.
2. **Domain-specificity cuts both ways.** Using Zahavy risks highlighting where his critique DOES apply to empirically-engaged philosophy — philosophy of perception, philosophy of physics, perhaps ethics involving moral psychology. The stress test note's Position C ([[Notes/Stress Test - Floridi's Concession as Resource|domain-dependent position]]) already flags this tension.
3. **The Chinese Room point could boomerang.** If evaluative standards can be absorbed from philosophical corpora, someone could argue they can be absorbed from physics corpora too. The response: in physics, *applying* evaluative standards to the right candidates requires sensory grounding; in philosophy, it doesn't. But this IS the claim that needs defending.
4. **Over-committing to self-grounding.** Leaning on Zahavy risks pushing toward the strong self-grounding thesis ("philosophy needs no external grounding at all"), which the paper does NOT argue — see [[Stress Test - Philosophy as Self-Grounding Domain]], where Position B (moderate, bounded self-grounding) is identified as most defensible. The paper's actual claim is about philosophy's evaluative standards being substantially text-internal, not that philosophy has no external grounding.
---
Source: `Learning/LLMs Can't Jump by Zahavy 2026.pdf`