# Scratch Pad ## scraps from llms tools substack essay # Generating Philosophy - Section 3 Moves - Move 1: Williamson characterises philosophical abduction as ranking theories both by their fit with the evidence and by what he calls 'the intrinsic virtues of a good theory.' > "Apart from its relation to E, the more T has the intrinsic virtues of a good theory, the better (ceteris paribus). It should be elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated. It should be informative and general. In brief, it should combine simplicity with strength." (Williamson 2024, p. 354) - 1a: For Williamson, philosophical evidence need not be restricted to fresh empirical data, and philosophical explanation need not be causal. > "Nothing in this account requires the evidence propositions, the explananda, to be of some special kind. Any known truths will do." (Williamson 2024, p. 355) > "Nor does anything in the account require the explanations to be causal. They may be constitutive instead." (Williamson 2024, p. 355) - 1b: Williamson explicitly argues that an armchair methodology can remain abductive. > "What matters here is that mathematics is a precedent for a successful discipline with an 'armchair' methodology that still has a key role for abduction. Thus it would be myopic to assume that an abductive methodology for philosophy implies its assimilation to the experimental sciences." (Williamson 2024, p. 358) - Move 2: On this picture, the evidence base for philosophical abduction consists in arguments, counterexamples, thought experiments, distinctions, and the results of prior inquiry — and these enter the practice as written text. - 2a: Abduction ranks only those candidate explanations that have actually been articulated. > "Of course, we rank only those potential explanations that have been thought of." (Williamson 2024, p. 355) - 2b: The philosophical archive is not a random sample of text. It consists of texts that have been through peer review, that get cited and taught, and that have repaid sustained attention over time. This filtering process is noisy — bad philosophy gets published, mediocre work is sometimes rewarded — but noisy filtering still shapes the distribution in favour of arguments that are clear, non-ad-hoc, and explanatorily useful. - Move 3: An LLM trained on that archive learns a distribution that has already been shaped by the standards of the practice. Philosophical standards can be present in the model as latent regularities, without being represented as explicit rules. - 3a: Floridi is right that the immediate mechanism is probabilistic rather than deliberative. > "LLMs seem to perform a kind of zeroth-order abduction … given a prompt, they generate a plausible continuation … In reality, their operation is driven by maximising the probability of the sequence …" (Floridi et al. 2024, extracted lines 384–389) - 3b: What matters in philosophy is that the probability landscape is not neutral: the corpus on which the model is trained has already been shaped by the evaluative habits of the discipline. - 3c: The model can therefore carry philosophical standards as latent structure, in something like the way a language model carries grammar: not as an explicit rulebook, but as a learned regularity governing what well-formed continuation tends to look like. - Move 4: Lipton's distinction between likeliness and loveliness clarifies why statistical continuation and philosophical quality can converge without being the same thing. If the likeliest continuation just *were* the loveliest explanation, we would have a trivial claim — that the model is doing philosophy because the model is doing philosophy. Lipton's point is that likeliness and loveliness are distinct standards, but they can pick out the same explanation in a given case. > "Let us turn now to the second distinction. It is important to distinguish two senses in which something may be the best of competing potential explanations. We may characterize it as the explanation that is most warranted: the 'likeliest' or most probable explanation. On the other hand, we may characterize the best explanation as the one which would, if correct, be the most explanatory or provide the most understanding: the 'loveliest' explanation. The criteria of likeliness and loveliness may well pick out the same explanation in a particular competition, but they are clearly different sorts of standard. Likeliness speaks of truth; loveliness of potential understanding." (Lipton 2004, p. 59) - 4a: In a corpus filtered for illuminating, unified, and non-ad-hoc philosophical writing, the likeliest continuation can tend — imperfectly but non-accidentally — to track lovelier argumentative forms. - 4b: That does not mean every continuation will be good philosophy. It means that philosophical merit is not external to the distribution the model has learned. - Move 5: Zahavy argues that LLMs cannot make the 'Jump' from sensory experience to formal axioms that he takes to be required for scientific discovery. He frames this in terms of abduction, drawing on Einstein's letter to Maurice Solovine: discovery involves an intuitive leap from experience (E) to axioms (A), followed by logical deduction. > "How do we fundamentally discover new things? In a letter to Maurice Solovine, Albert Einstein conceptualized discovery as a cyclical process involving an intuitive 'jump' from sensory experience to axioms, followed by logical deduction. While Generative AI has mastered Induction (statistical pattern matching) and is rapidly conquering Deduction (formal proof), we argue it lacks the mechanism for Abduction — the generation of novel explanatory hypotheses." (Zahavy 2026, opening) - 5a: For Zahavy, the bottleneck is that LLMs lack the embodied simulation that would allow them to ground abstract symbols in physical sensation. Einstein, on his account, used thought experiments involving felt experience — the sensation of falling, the feel of acceleration — to formulate the equivalence principle. LLMs manipulate symbols without access to the physical referents that give those symbols meaning. Zahavy calls them 'high-dimensional Chinese Rooms'. - 5b: Zahavy explicitly restricts this argument to the physical sciences. > "Finally, we emphasize that this proposal is specifically tailored to the physical sciences, where the object of study is external material reality. In abstract domains such as Mathematics or Computer Science, the Sense Experience (E) may be grounded in high-dimensional topology or have other goals such as generality or minimality." (Zahavy 2026, p. 7) - 5c: One response for philosophy is to say that most analytic philosophy is closer to mathematics or computer science than to physics. Formal epistemology, philosophical logic, philosophy of language, metaphysics, ethics: these fields work through argument, counterexample, and conceptual analysis, not through phenomenological observation. The E→A jump that Zahavy worries about presupposes that the raw material is pre-linguistic sensory experience. Most of analytic philosophy does not have that structure. - 5d: Some philosophy does seem to rely on phenomenological contact. Philosophy of perception, for instance, often involves claims like 'colours appear spread across the surfaces of objects' or 'the visual field has a certain spatial structure'. And if one takes a Husserlian view, phenomenology is foundational to much larger swathes of the discipline. - 5e: The response here is that LLMs, while lacking first-person phenomenological experience, have access to vast amounts of text describing such experience. The claim that colours are spread across surfaces is implicit in enormous amounts of written material. An LLM trained on that corpus can work with articulated descriptions of phenomenological observations, even if it cannot have the experiences itself. - 5f: This does not show that LLMs *can* do philosophy. It shows that Zahavy's argument, which targets the physical sciences, does not straightforwardly apply to philosophy. - Move 6: Philosophical novelty often consists in introducing new distinctions and new ways of organising familiar material, not in leaping from sensation to axioms. > "However, enumerative induction is inadequate for systematic philosophical theorizing, which often requires introducing new distinctions at a more abstract level not given in the data." > "A glance at Dummett's own writings will show that he often introduces new meaning-theoretic distinctions, like that between 'assertoric content' and 'ingredient sense,' which cannot simply be read off the data …" (Williamson 2024, p. 353) - 6a: This kind of novelty is not simple induction, but neither is it Zahavy's Einstein case. It is novelty in conceptual organisation. - 6b: A model trained on philosophical texts can, in principle, learn patterns for drawing, testing, and revising higher-order distinctions, even if there are limits to how radical that novelty will be. - 6c: Novelty does not need a separate standard here, because a theory that merely restates what is already known fares badly on informativeness and generality — virtues Williamson already includes. - Move 7: The claim that LLM philosophy is 'just statistics' confuses mechanism with level of assessment. > "If these suggestions are along the right lines, then arguing that Inference to the Best Explanation is wrong because Bayesianism is right is like arguing that thinking about technique cannot help my squash game because the motion of the ball is governed by the laws of mechanics." (Lipton 2004, p. 108) - 7a: Token prediction describes how the output is generated. - 7b: Philosophical evaluation asks whether the output is unified, non-ad-hoc, dialectically responsive, and illuminating. - 7c: The stochastic description and the philosophical description can both be true at once. - Move 8: Once that much is in place, the remaining question is not whether philosophical standards can be present in the model at all, but under what conditions they are actually drawn out in output. - 8a: That is the point at which Section 3 should end and Section 4 should begin. --- # What's Happening *Active threads and today's activity — updated by /harvest* ## Active ## Sessions - 16:59 - "Revise Substack essay plan for LLMs piece" — Base directory for this skill: /Users/nickyoung/.claude/skills/contemplate # ... - 22:18 - "Review conversation JSON and Substack draft" — If you look in the downloads folder, you'll see that the most recent file in ... - 22:22 - "If you look in the downloads folder, you'll see..." — If you look in the downloads folder, you'll see that the most recent file in ... - 22:50 - "If you look in the downloads folder, you'll see..." — If you look in the downloads folder, you'll see that the most recent file in ... ## Actions ---