# Lipton Ch10: Prediction and Prejudice — Relevance to Generating Philosophy ## 1. Chapter Summary Chapter 10 of *Inference to the Best Explanation* takes up a puzzle that Lipton treats as both epistemological and methodological: does a theory deserve more credit when it successfully predicts a datum than when it was built to accommodate that same datum? The chapter opens by defining the two cases sharply. In accommodation, "the scientist constructs a theory to fit the available evidence." In prediction, "the theory is constructed and, with the help of auxiliaries, an observable claim is deduced but, unlike a case of accommodation, this takes place before there is any independent reason to believe the claim is true." The question is whether this temporal or epistemic difference matters to the degree of support the theory enjoys. Lipton distinguishes a weak advantage thesis from a strong one. The weak thesis holds that predicted data "tend to provide more support than accommodations, because either the theory or the data tend to be different in cases of prediction than they are in cases of accommodation." The strong thesis is more demanding: "a successful prediction tends to provide more reason to believe a theory than the *same* datum would have provided for the *same* theory, if that datum had been accommodated instead." Most of the chapter is devoted to defending the strong thesis, which is the philosophically interesting claim. The case against the strong thesis is laid out with care. Lipton notes that "the content of theory, auxiliary statements, background beliefs and evidence, and the logical and explanatory relations among them, are all unaffected by the question of whether the evidence was accommodated or predicted, and these seem to be the only factors that can affect the degree to which a theory is supported by evidence." He sharpens this with the twin scientists thought experiment: two scientists independently produce the same theory, one accommodating a datum the other predicts. If the strong thesis holds, the predictor should have more reason to believe the theory. But if the twins meet, "it seems clear that they should leave the meeting with a common degree of confidence in the theory they share" — and there is no principled way to say what that level should be. Three popular defences of prediction over accommodation are then rejected. The first, that accommodating theories are "ad hoc," is dismissed as either begging the question or merely naming the problem: "To assume that accommodating theories are ad hoc in the sense of poorly supported is to commit what might be called the '*post hoc ergo ad hoc*' fallacy." The second, that only predictions can "test" a theory, confuses the theory with the theorist: "a theory will not be refuted by evidence it accommodates, but that theory would have been refuted if the evidence had been different." The archer analogy — we want the bullseye drawn before the volley — is about evaluating the archer's skill, not the arrow. The third, that an overarching IBE favours prediction because accommodation pre-empts the truth explanation, fails because "to assume that accepting the accommodation explanation makes it less likely that the theory is true is once again to beg the question against accommodation." Lipton's own positive account is the "fudging explanation." The thought is that accommodation creates a distinctive epistemic liability: "When data need to be accommodated, there is a motive to force a theory and auxiliaries to make the accommodation. The scientist knows the answer she must get, and she does whatever it takes to get it." This may result in "an unnatural choice or modification of the theory and auxiliaries that results in a relatively poor explanation and so weak support." In prediction, "there is no motive for fudging, since the scientist does not know the right answer in advance." The crossword puzzle analogy clarifies the structure: generating an answer to a clue without looking at intersecting letters, and then finding the letters match, gives more reason to think the answer is correct than looking at the intersections first and building the answer around them. "It is only in the case of accommodation that the intersecting letters could possibly pull you away from the best answer to the clue, since it is only in this case that these letters are used in the process of generating the answer." Two forms of fudging are distinguished: theory fudging, where "the accommodated evidence is purchased at the cost of theoretical virtues," and auxiliary fudging, where the cost is "epistemic relevance." Both yield inferior explanations. Lipton insists that the fudging need not be conscious or deliberate, and that "a certain amount of fudging is not bad scientific practice." What matters is the structural motive. The chapter's most important philosophical move is the distinction between actual and assessed support. Lipton argues that "we need to distinguish between actual and assessed inductive support, between the extent to which the data actually render the theory probable and the scientist's judgment of this." The strong advantage thesis "does not claim that the actual support for the theory is different, only that our assessment of this support ought sometimes to be different." Support is "translucent, not transparent" — even the scientist herself cannot simply inspect whether her theoretical system has been fudged, because "fudging need not be a conscious process." The indirect evidence that accommodation provides about the likelihood of fudging therefore remains epistemically relevant to the scientist's own assessment. Lipton argues that the "tradition of stipulating an artificial division between the context of discovery and the context of justification" is what has obscured this: "what the fudging explanation shows is that this is relevant to the question of evaluation." The chapter closes by resolving the twin scientists puzzle. When the accommodator meets the predictor and discovers they produced the same theory, "this shows that she almost certainly did not fudge to make those accommodations." The predictor, ignorant of the data, "had no motive to fudge his theoretical system to get those results; consequently, the fact that he came up with just the same system provides strong independent evidence that the accommodator did not fudge either." Successful predictions also grant "retrospective epistemic stature" to earlier accommodations. ## 2. Connections to the Generating Philosophy Project The parallel between Lipton's prediction/accommodation framework and the question of whether LLMs can generate genuine philosophy is structurally precise. An LLM trained on the philosophical corpus has, in Lipton's terms, *accommodated* that corpus: the model was constructed to fit the available evidence (the texts it was trained on). The worry that LLMs merely reproduce training data maps directly onto the worry that accommodating theories are "ad hoc" — built to fit, and therefore not genuinely tested. And the question of whether LLMs can produce novel philosophical claims that turn out to be good — the Move 37 / tail novelty thread — is precisely the question of whether an accommodating system can generate genuine *predictions*. The mapping has several specific dimensions worth tracing. **Fudging and the "mere reproduction" charge.** The critics of LLM philosophy (Floridi, Zahavy) worry that LLMs cannot produce genuine abductive reasoning because they are merely recombining absorbed patterns. This echoes the intuitive case against accommodation: the fit between theory and data is unimpressive because the theory was built to fit. Lipton's analysis shows that this intuition, while containing a germ of truth, is not straightforward. The "post hoc ergo ad hoc" fallacy — assuming that because a theory was designed to accommodate data it is therefore poorly supported — is exactly the fallacy committed when one assumes that because an LLM was trained on philosophical texts its philosophical outputs are therefore epistemically worthless. The fact that the model was trained on the data does not, by itself, settle the question of whether its outputs are good philosophy. To assume otherwise is to name the problem rather than solve it. Lipton is explicit that the fudging concern is about a *structural motive*, not an inevitable outcome. Accommodation creates the *possibility* of fudging but does not guarantee it. Similarly, training on philosophical texts creates the *possibility* that an LLM merely reproduces patterns without tracking philosophical quality, but this does not follow necessarily. The project's existing argument — that philosophical norms are textually manifest and that the training corpus is filtered for quality (Section 3's "Learning the Game") — is structurally analogous to Lipton's point that "if the scientist is ever justifiably certain that she would have produced the same theoretical system even if she did not know about the evidence she accommodated, that evidence provides as much reason for belief as it would have, had it been predicted." If the philosophical training data encode not just content but evaluative structure (which is what the dialectical saturation thesis claims), the "accommodation" involved in training may not carry the fudging liability. **When accommodation is as good as prediction.** Lipton identifies several conditions under which the asymmetry between prediction and accommodation diminishes or disappears. These conditions map onto features of the LLM/philosophy case that the project has already been developing. First, Lipton notes that "we should expect the difference between prediction and accommodation to be greatest for complex and high level theories that require an extensive set of auxiliaries, and to decrease or disappear for simple empirical generalizations." This is because high-level theories offer more room for auxiliary fudging. Philosophical arguments, as Lipton's own framework suggests, sit in an interesting position here. The "auxiliaries" of a philosophical argument are its premises, definitions, and inferential steps — all of which are typically *explicit* in the text, not hidden. This reduces the scope for undetectable fudging. The project's emphasis on philosophy as textual all the way down (Section 2's argument) is thus directly relevant: the textual medium constrains the space of fudging in a way that empirical science's reliance on hidden auxiliary assumptions does not. Second, Lipton observes that "in a case where we convince ourselves that there is really only one possible explanation for the data that is, given our background beliefs, even remotely plausible... fudging is not an issue and accommodation is no disadvantage." The dialectical saturation thesis can be read as claiming something similar: if the argumentative patterns encoded in the corpus are the *only* adequate responses to the dialectical situations they address, then the LLM's having been trained on them is not a mark against the quality of its outputs. This is speculative, but the direction of argument is clear. If what the LLM has learned are genuine logical and argumentative constraints rather than arbitrary stylistic habits, the accommodation worry loses force. **The role of construction.** Lipton's analysis of how scientists "construct" theories to fit data has a direct analogue in how LLMs construct philosophical arguments. The fudging explanation holds that the risk lies specifically in the *process of generation*: "Only accommodated data can influence the process of generation, and this is the difference that the fudging explanation exploits." For an LLM, all of its training data influenced the process of generation (of the model's parameters). But there is a disanalogy worth noting. In Lipton's framework, the scientist knows the specific datum she needs to accommodate and adjusts her theory accordingly. An LLM does not "know" what specific argument it needs to produce in advance of producing it. It is trained on the whole corpus at once, not datum-by-datum with the goal of fitting each. This means that the specific fudging mechanism Lipton describes — where the scientist knows the right answer and works backwards — does not apply straightforwardly. The LLM's situation is arguably more like what Lipton calls "prediction" at the level of individual outputs: when prompted, the model does not know in advance what its output will be, and it generates philosophical claims that may or may not be good. Whether those claims are good is then subject to independent evaluation by the philosopher. This connects to the distinction between the model's *training* (which accommodated the corpus) and its *inference* (which produces novel outputs from novel prompts). Training is accommodation; inference is closer to prediction. The project could use this framing to argue that the philosophically relevant moment is not training but production: what matters is whether the output, evaluated on its own merits, constitutes good philosophy. The provenance question — which the project has already argued is "not the right kind of variable in philosophical evaluation" — receives additional support from Lipton's own framework, which insists that "the actual support the theory enjoys is precisely the same, whether the datum was accommodated or predicted." The difference is in our *assessment*, not in the objective quality. **Actual versus assessed support: translucency and philosophical evaluation.** The chapter's distinction between actual and assessed support may be the most directly useful element for the project. Lipton argues that support is "translucent, not transparent" — even experts cannot simply inspect whether a theoretical system has been fudged. Applied to LLM philosophy: even a competent philosopher reading an LLM-generated argument cannot simply *see* whether the argument is the result of genuine philosophical competence or a sophisticated pattern match. This is precisely the worry. But Lipton's own framework suggests that the response is not to throw up one's hands. Instead, one should attend to the structural features that make fudging more or less likely, and adjust one's assessment accordingly. The project's existing move — that philosophical evaluation is of the artefact, not the producer (Section 0), and that expert evaluation can detect the relevant quality markers (the appearance/reality collapse thesis) — can be strengthened by Lipton's framework. If the worry about LLM-generated philosophy is a worry about translucent support (we cannot directly observe whether the model has genuine philosophical competence), the remedy is the same as Lipton's remedy for scientific accommodation: look at the structural features of the output. Does it exhibit the marks of fudging — arbitrary conjunctions, unmotivated epicycles, poor fit with the philosophical background? Or does it exhibit the marks of quality — precision, explanatory power, responsiveness to objections? The evaluation is fallible, as Lipton insists all such evaluations are. But it is not arbitrary. **Retrospective epistemic stature and iterative LLM use.** Lipton's argument that successful predictions grant "retrospective epistemic stature" to earlier accommodations has a natural application to iterative philosopher-LLM collaboration. If an LLM, initially trained on the corpus (accommodation), goes on to produce a philosophical claim that the philosopher independently verifies as good — that is, a claim that withstands scrutiny, addresses genuine objections, and advances a live debate — this functions as a kind of prediction. And if such predictions succeed, they grant retrospective epistemic stature to the model's general philosophical competence, just as successful scientific predictions validate the theoretical system that also accommodated earlier data. The inner speech / LLM coupling thread, which treats the philosopher-LLM system as a collaborative extended cognitive system, could deploy this: the iterative process of generating, evaluating, revising, and re-generating is a process of converting accommodations into tested predictions. ## 3. Deployment Suggestions The prediction/accommodation distinction could be deployed in the paper in at least two ways, which need not be exclusive. First, it could appear in Section 1 alongside the existing discussion of Lipton's generation/selection distinction. Section 1 already uses Lipton's framework to introduce abduction. The prediction/accommodation analysis provides a further layer: the critics' worry is structurally a worry about accommodation (that LLMs merely fit the existing data), and Lipton's own analysis shows that this worry is more nuanced than it first appears. The "post hoc ergo ad hoc" fallacy label could be deployed directly against the assumption that training-on-text automatically vitiates philosophical quality. Second, it could appear in Section 3 ("Learning the Game") as part of the positive case. The dialectical saturation thesis claims that the philosophical corpus encodes evaluative structure, not just content. Lipton's analysis of when accommodation carries no fudging liability — when the theoretical system is tightly constrained, when the auxiliaries are explicit, when there is plausibly only one good explanation — describes exactly the conditions that the saturation thesis attributes to philosophy. This would allow Section 3 to argue not only that LLMs have learned the game (positive claim) but that the fact of their having been trained on the corpus does not constitute a mark against them (defensive claim). The actual/assessed support distinction could also be deployed in Section 2, where the paper argues that philosophical evaluation is of the artefact. Lipton's distinction offers philosophical vocabulary for the claim that provenance (whether the text was produced by a human or an LLM) is relevant to our *assessment* of quality but not to the *actual* quality. This clarifies the project's existing move without requiring it to take on the full epistemological apparatus of IBE. ## 4. Divergences and Limitations Three divergences between Lipton's framework and the LLM case deserve attention. First, Lipton's prediction/accommodation distinction operates within a framework where there is a fact of the matter about the world that the theory either gets right or wrong. Philosophy's relationship to external facts is different. The project has already argued (Section 2) that philosophy is "textual all the way down" and that evaluation is of the artefact rather than of correspondence to external phenomena. This means that the prediction/accommodation distinction does not map perfectly: there is no philosophical equivalent of Mendeleyev's predicted elements being independently detected. The closest analogue is a philosophical argument surviving sustained critical scrutiny — but this is a different kind of "independent verification" than empirical discovery. The project should acknowledge this disanalogy even while deploying the structural parallel. Second, Lipton's fudging explanation depends on the idea that the scientist *knows the right answer* when accommodating and therefore has a motive to work backwards. LLMs do not "know" answers in this sense. The model's training process is not goal-directed in the way a scientist's theory construction is. This means the specific mechanism of fudging — conscious or unconscious adjustment to get a known result — does not transfer directly. What transfers is the structural concern: that a system built to fit data may overfit to surface features rather than tracking the deeper structure. The project should be explicit that it is deploying the structural analogy rather than the specific mechanism. Third, Lipton's framework treats prediction and accommodation as properties of the relationship between a theory and individual data points. LLMs are trained on entire corpora, and the relevant "accommodations" are not datum-by-datum adjustments but statistical regularities across billions of tokens. The granularity is different. Whether this difference matters to the philosophical point is not obvious, but the project should not assume a frictionless mapping. ## 5. Passages to Flag Several passages are particularly valuable for direct engagement or quotation in the paper. On the "post hoc ergo ad hoc" fallacy: "To claim, however, that a theory that is ad hoc in this sense is therefore poorly supported begs the question... To assume that accommodating theories are ad hoc in the sense of poorly supported is to commit what might be called the '*post hoc ergo ad hoc*' fallacy." This is directly deployable against critics who assume that training on philosophical text automatically vitiates the outputs. On the theory/theorist confusion: "It is only in the case of prediction that the scientist runs the risk of looking foolish. This, however, is a commentary on the scientist, not on the theory." This supports the provenance-irrelevance argument already developed in the project. On knowledge as the operative distinction, not time: "what makes the difference between accommodation and prediction is not time, but knowledge. When the scientist doesn't know the right answer, she knows that she is not fudging her theoretical system to get it." This could be used to argue that LLMs, which do not "know" in advance what philosophical argument they will produce in response to a novel prompt, are in a structurally predictive position at inference time. On the conditions under which accommodation carries no disadvantage: "if the scientist is ever justifiably certain that she would have produced the same theoretical system even if she did not know about the evidence she accommodated, that evidence provides as much reason for belief as it would have, had it been predicted." This connects to the saturation thesis's claim that the argumentative patterns are constrained enough that the model would converge on them regardless. On the translucency of support: "we need to distinguish between actual and assessed inductive support, between the extent to which the data actually render the theory probable and the scientist's judgment of this." This provides philosophical vocabulary for the artefact-evaluation thesis. On the context of discovery bearing on justification: "what the fudging explanation shows is that this is relevant to the question of evaluation." Lipton's own argument against the strict discovery/justification distinction is useful for the project, which argues that the process of LLM generation is relevant to but does not determine the evaluation of outputs. On retrospective epistemic stature: "accommodations are most suspect when a theory has yet to make any tested predictions, but that the accommodations gain retrospective epistemic stature after predictive success." This supports the iterative collaboration model. On the limits of the fudging explanation with tightly constrained systems: the passage noting that the prediction/accommodation difference should "decrease or disappear for simple empirical generalizations" and for theories with "a tight and simple mathematical structure" suggests that in domains with strong internal constraints (as the saturation thesis claims philosophy is), accommodation is less epistemically suspect. --- *La distinzione tra previsione e adattamento rivela che il dubbio sull'originalit\u00e0 filosofica degli LLM ripete, in forma nuova, un pregiudizio epistemico che Lipton ha gi\u00e0 smontato.*