# The Symbol Grounding Problem in Philosophy: Why LLMs Cannot Produce Paradigm Shifts ## Introduction The claim that large language models can produce paradigm-shifting philosophy rests on a seductive but fundamentally mistaken picture of what philosophy is and how philosophical understanding works. Proponents argue that philosophy's subject matter is "already symbolic" — that the Space of Reasons is constituted by inferential relations expressed in and assessable from texts, making the philosophical corpus a self-grounding system where the map simply is the territory. From this they conclude that LLMs, having mastered the corpus, can produce genuine philosophical contributions, including paradigm shifts. This argument fails. It fails because it confuses the textual expression of philosophical work with the work itself. It fails because it mistakes correlation in the corpus for genuine inferential relations. It fails because it ignores the experiential dimension that generates philosophical inquiry in the first place. And it fails because it cannot account for what paradigm shifts actually require — a standpoint outside existing frameworks from which their inadequacy becomes visible. I will argue for three interconnected claims. First, text-internal evaluation is not sufficient for philosophy; the corpus cannot encode everything needed to evaluate philosophical quality. Second, phenomenology is not textually transmissible at the resolution philosophy requires; genuine confusion, insight, and truth-directedness are not structural features reducible to text patterns. Third, LLMs cannot produce paradigm shifts in philosophy except by chance; any paradigm-shifting output would be accidental recombination, not principled philosophical work. ## The Symbol Grounding Problem in Philosophy Harnad's symbol grounding problem poses a fundamental challenge to any system that processes symbols without contact with what those symbols represent. Consider his example: "Suppose you had to learn Chinese as a first language and the only source of information you had was a Chinese/Chinese dictionary! ... How can you ever get off the symbol/symbol merry-go-round?" No matter how comprehensive the dictionary, no matter how sophisticated the cross-references, you remain trapped in a closed system of symbols referring to symbols. The dictionary cannot convey what the symbols mean because meaning requires grounding in something outside the symbol system. Defenders of LLM philosophy attempt to evade this problem by claiming philosophy is different. Philosophy's subject matter, they say, is already symbolic — the Space of Reasons, inferential relations, conceptual structures. Unlike biology, which studies organisms, or physics, which studies matter, philosophy studies thought itself. So there is no gap between the symbols and what they represent. The philosophical corpus just is the philosophical terrain. This is wrong. The Space of Reasons is indeed constituted by inferential relations, but inferential relations are not textual patterns. They are normative structures — what actually supports what, what genuinely follows from what. When we say that the Gettier cases show something about knowledge, we are not merely noting that texts about Gettier correlate with texts about justified true belief. We are claiming that there is a genuine logical relationship between the cases and the analysis — that the cases reveal something about what knowledge is. An LLM learns correlations. When pattern X appears in the corpus, pattern Y tends to follow. The LLM has no access to why Y follows from X. It cannot distinguish genuine logical support from mere textual correlation. It cannot tell whether two texts are linked because one actually justifies the other or because philosophers happen to cite them together. The LLM learns descriptions of inferential relations, not the relations themselves. This is not a minor gap that sophisticated training might close. It is a structural limitation of any system that processes symbols without grounding. Harnad's point is that meaning requires contact with something outside the symbol system — what he calls iconic and categorical representations grounded in sensory experience. For philosophy, the grounding is not sensory in the narrow sense, but it is equally non-symbolic: it is the experiential encounter with philosophical problems that generates and sustains philosophical inquiry. ## The Experiential Ground of Philosophy Consider what it is to genuinely wrestle with a philosophical problem. Take the puzzle of informative identity statements that Frege articulated. "Hesperus is Phosphorus" tells us something that "Hesperus is Hesperus" does not, yet both are identity statements about the same object. The puzzle is not merely that texts exhibit this pattern. The puzzle is that something seems wrong — our concepts of meaning, reference, and identity generate a tension that demands resolution. This tension is experienced. It is not a textual feature that can be extracted by pattern-matching. It is the lived sense that one's conceptual apparatus is inadequate, that something does not fit. Frege did not discover the puzzle by surveying texts about identity statements. He encountered it in thinking about what identity means, and that encounter generated the sense/reference distinction that restructured the entire domain. Philosophy is full of such experiential moments. Kripke felt that descriptivist predictions about reference were wrong before he had a theory of rigid designation. The feeling came first — the sense that something was off about how descriptivist semantics handled proper names in counterfactual contexts. The theory emerged from this experiential ground. Gettier saw that something was missing from the JTB analysis. The cases he constructed gave form to an intuition that preceded them. This experiential dimension is not a contingent feature of how philosophy happens to be practiced by humans. It is constitutive of philosophical inquiry. Philosophy addresses problems — genuine puzzles, confusions, tensions in our conceptual schemes. But problems are not textual patterns. A problem is a problem for someone. It generates the urgency that drives inquiry forward, the motivation to seek resolution, the frustration when resolution proves elusive. Without this experiential ground, you have texts about problems, but not the problems themselves. An LLM processing the philosophical corpus learns what problems look like. It learns the textual features that philosophers produce when describing problems — formulations of tensions, expressions of puzzlement, articulations of competing considerations. But it has no access to the problems themselves. It cannot feel the wrongness that generates inquiry. It cannot experience the inadequacy of a conceptual framework. It processes descriptions of philosophical struggle without ever struggling philosophically. ## Phenomenology Cannot Be Textualized Defenders of LLM philosophy might concede that phenomenology exists but argue that philosophy does not require it. What matters, they claim, is the structural features of philosophical work — argument quality, coherence, engagement with objections. These features are textually accessible. So even if LLMs lack phenomenology, they can still do philosophy. This response fundamentally misunderstands what phenomenology contributes to philosophy. It is not a mere accompaniment to the "real" work of argumentation. It is the source of the normative grip that arguments have. Consider what it is to care about truth. Caring about truth is not tracking argument quality as a textual feature. It is not producing outputs that fit the statistical pattern of "argument-shaped text." It is the orientation toward getting things right that motivates inquiry, that makes certain answers unacceptable even when they are logically consistent, that sustains the demand for genuine understanding rather than mere coherence. An LLM optimizes a loss function. It generates outputs that minimize statistical divergence from training data. This is not caring about truth. It is not caring about anything. The LLM has no stake in whether its outputs are true. It produces text that fits patterns, regardless of whether those patterns track reality. This matters for philosophy because philosophical work is normatively constrained by truth-directedness. When I evaluate an argument, I am not merely asking whether it exhibits features correlated with "good arguments" in the corpus. I am asking whether it actually works — whether the premises genuinely support the conclusion, whether the concepts apply, whether the position coheres with how things are. This normative assessment requires caring about truth in a way that LLMs cannot. Consider confusion. Genuine philosophical confusion is not recognizing that texts exhibit tension. It is the lived experience of not knowing which way to go, of finding one's conceptual apparatus inadequate, of needing resolution. This experience has phenomenal character — there is something it is like to be philosophically confused. Nagel's point about the bat applies here: what it is like to be confused cannot be captured by descriptions of confusion. An LLM has no "what it's like." It processes patterns. When it generates text that describes confusion, it produces outputs statistically correlated with confusion-descriptions in the training data. But it is not confused. It does not feel the inadequacy of its conceptual resources. It does not need resolution. The text it produces is phenomenologically empty — the surface form of confusion without the experiential substance. ## The E→A Jump in Philosophy Zahavy's analysis of scientific discovery centers on what he calls the E→A Jump — the translation from Sense Experience to System of Axioms. Einstein's discovery of the Equivalence Principle required embodied simulation: imagining what it would feel like to be a falling observer, and translating that experiential simulation into physical principle. This is not pattern-matching on existing physics texts. It is a creative leap from experience to framework that no amount of textual knowledge could achieve. Zahavy restricts his argument to physics, suggesting that abstract domains like mathematics might work differently. Defenders of LLM philosophy seize on this restriction, arguing that philosophy is more like mathematics than physics — its subject matter already substantially symbolic, requiring no E→A Jump from experience to framework. This argument misses the structure of Zahavy's point. The E→A Jump is not specifically about sensory experience. It is about the gap between phenomena and the frameworks that explain them. In physics, the phenomena are physical processes observed through sensory experience. In philosophy, the phenomena are different — they are conceptual tensions, problematic intuitions, inadequacies in existing frameworks. But the gap remains. Philosophy requires translating these phenomena into systematic treatments, and this translation cannot be accomplished by pattern-matching on existing treatments. When Kripke developed rigid designation, he was not recombining patterns from the corpus of reference theory. He was responding to something that existing frameworks could not accommodate — the behavior of proper names in counterfactual contexts that descriptivism predicted incorrectly. His theory emerged from an encounter with this phenomenon, not from statistical manipulation of existing texts about reference. The E→A Jump in philosophy is the movement from experienced conceptual friction to theoretical resolution. Gettier experienced that something was wrong with JTB, then constructed cases that articulated what was wrong. The construction emerged from the experience; it did not precede it. An LLM can learn the pattern of Gettier cases — scenarios where someone has justified true belief without knowledge. But it cannot experience the wrongness that motivated their construction. It learns the output of the E→A Jump without access to the process. This is not a claim that LLMs cannot produce novel philosophy. They can produce novel combinations of existing patterns. But paradigm shifts are not pattern recombination. They are the recognition that existing patterns themselves are inadequate — that the entire framework needs restructuring. This recognition requires a standpoint outside the framework, and an LLM trained on the corpus has no such standpoint. The corpus is its world. ## The Circularity of Text-Internal Evaluation The claim that evaluative standards for philosophy are learnable from the corpus faces a fatal circularity. Where did those standards come from? They came from philosophers with genuine experience — people who wrestled with confusion, had insights, cared about truth. The standards encoded in the corpus are the sedimented results of this experiential process. They are outputs, not inputs. An LLM learning from the corpus learns what good philosophy looks like from the outside. It masters surface features — argument structures, dialectical moves, characteristic objections and replies. It can produce text that exhibits these features. But it lacks the generative source — the experiential engagement that produces and validates the standards in the first place. This is like learning to forge a master's signature by studying examples. You might achieve perfect surface mimicry. Your forgeries might be indistinguishable from authentic signatures under many inspections. But you are not signing authentically — you are copying a pattern without the intentional ground that makes a signature a signature. The standards in the corpus presuppose experiential engagement. When philosophers judge that an argument is "deep" or "illuminating" or "genuinely addresses the problem," they are not merely checking textual features. They are assessing whether the argument connects with the phenomena that motivate inquiry. An LLM cannot make this assessment because it has no access to the phenomena. It can only check whether the argument exhibits features correlated with positive evaluations in its training data. This circularity cannot be broken from within the corpus. Adding more text, more meta-level discussion, more explicit articulation of standards — none of this provides the experiential ground that standards presuppose. You cannot bootstrap understanding from symbols alone. This is Harnad's point applied to philosophy: the symbol/symbol merry-go-round spins forever unless something outside the symbol system provides grounding. ## Why Any Paradigm Shift Would Be Accidental For LLM-generated philosophy to count as paradigm-shifting in a non-accidental way, it must satisfy what I will call the principled production condition. The output must reflect systematic capacity, not luck. The internal process must be appropriately shaped by philosophical understanding. There must be some meaningful sense in which the system "knows what it is doing." LLMs fail all three components. Reliability: If an LLM occasionally produces text that appears paradigm-shifting, this reflects the statistical character of its outputs, not systematic capacity. The corpus contains paradigm shifts, so the LLM learns their surface features. Meta-patterns exist — introducing distinctions, constructing counterexamples, challenging presuppositions. Statistical recombination occasionally generates novel variants of these patterns. But this is high-dimensional chance, not reliable capacity. The LLM has no general ability to produce paradigm shifts; it has learned patterns that sometimes combine in novel ways. Process: The LLM's process is pattern-matching under constraints learned from the corpus. But paradigm shifts require recognizing when the constraints themselves need revision — when the entire framework, including the evaluative standards embedded in it, is inadequate. This requires a standpoint outside the corpus that the LLM cannot have. Its world is defined by the corpus; it cannot see the limits of that world. Understanding: The LLM mimics the structure of philosophical understanding without the substance. When a philosopher deploys a distinction, she understands its work — why it solves the problem, how it relates to other concepts, what motivates it. When an LLM deploys a distinction, it produces text that fits patterns for "distinction deployment." Same surface, utterly different reality. Any paradigm-shifting output from an LLM would be like a broken clock showing the right time. Occasionally, statistical recombination will produce something that genuinely illuminates. But the clock is not tracking time; it is broken. The LLM is not doing philosophy; it is manipulating patterns. The match between output and philosophical value, when it occurs, is coincidence. ## Conclusion The dream of LLM philosophy rests on a confusion between philosophical texts and philosophical work. Texts are the medium through which philosophical work is expressed and communicated. But the work itself — the encounter with problems, the struggle toward understanding, the care for truth — is not textual. It is experiential in ways that text cannot capture. This is not mysticism about ineffable insight. It is a straightforward point about what philosophy requires. Philosophy requires confronting phenomena — conceptual tensions, problematic intuitions, framework inadequacies. These phenomena are not textual patterns. They are experienced. Philosophy requires caring about whether one's conclusions are true, not merely coherent. This caring is not pattern-tracking. It is normative engagement with reality. Philosophy requires the ability to step outside existing frameworks and see their limits. This ability is not pattern recombination. It is critical distance that presupposes a standpoint the corpus cannot provide. LLMs are sophisticated pattern-matchers. Within that scope, they are impressive. They can generate text that exhibits the surface features of philosophical work. They can produce arguments, raise objections, draw distinctions. But they cannot do philosophy. They cannot encounter problems because they have no experiential contact with the phenomena that constitute problems. They cannot care about truth because they have no normative orientation beyond pattern-matching. They cannot produce paradigm shifts in a principled way because they are trapped within the corpus that defines their world. Any philosophy that emerges from an LLM will be accidental — statistical fluctuation occasionally producing something valuable, like monkeys eventually typing Shakespeare. The monkeys are not writing Shakespeare; they are pressing keys. The LLM is not doing philosophy; it is completing patterns. The difference matters, even when the outputs happen to match.