```dataview LIST WITHOUT ID file.link FROM -"windsurf" WHERE (file.cday = date(this.file.name) or file.mday = date(this.file.name)) AND !startswith(file.folder, "windsurf") SORT file.cday ASC ``` # [[Generating Philosophy with Artificial Intelligence]] 1 May 2025 #paper/generatingphilosophy #draft > “Forty-two," said [[Deep Thought]], with infinite majesty and calm. > It was a long time before anyone spoke. > Out of the corner of his eye Phouchg could see the sea of tense expectant faces down in the square > outside. > "We're going to get lynched aren't we?" he whispered. > "It was a tough assignment," said [[Deep Thought]] mildly. > "Forty-two\!" yelled Loonquawl. "Is that all you've got to show for seven and a half million years' work?" > "I checked it very thoroughly," said [[the computer]], "and that quite definitely is the answer. I think the > problem, to be quite honest with you, is that you've never actually known what [[the question]] is." > > \-- [[Douglas Adams]], [[The Hitchhiker]]'s Guide to the Galaxy In *[[The Hitchhiker]]'s Guide to the Galaxy*, humanity asks an AI to do some philosophy. An AI named [[Deep Thought]] is constructed and instructed to provide "The Answer to the [[Ultimate Question]] of Life, the Universe, and Everything" (REF). Humanity builds this computer only to receive the answer '42'—an answer which, while apparently correct (at least according to [[Deep Thought]]), means next to nothing at all due to humanity’s failure to formulate the [[Ultimate Question]]. In 2025 we are in a position to think about [[the relationship between]] philosophy and AI *for real*. Can LLMs enhance [[philosophical understanding]]? They are happy to dispense philosophical wisdom if we ask them to, but should we listen? I suspect many would doubt that they are, notoriously prone to 'hallucinations', and the banal, hyperbolic essays churned out by ChatGPT have become the scourge of undergraduate teaching. In this paper, I argue for cautious optimism: the creators of such systems sometimes describe them as *reasoning engines*, and I suggest this description is broadly accurate. Although LLMs themselves do not reason in the way that people do, they can simulate human reasoning, and through this simulation, they are capable of enhancing the [[philosophical understanding]] of their users. * In this paper.... ## 1 Understanding as Representing Dependence Networks To assess how LLMs might enhance philosophical understanding, we first require a definition of this understanding. In this paper, we adopt the framework proposed by Dellsén et al. (2024). Their approach, explaining understanding in terms of representing dependence networks, resonates with the broader philosophical aim, articulated by Sellars (1962, p. 1), to grasp 'how things in the broadest possible sense of the term hang together in the broadest possible sense of the term'. Sellars conceived philosophy as pursuing a synoptic vision, integrating diverse entities, concepts, events, norms, and theoretical posits into a coherent whole by mapping their causal, logical, conceptual, and explanatory interconnections. Dellsén et al. offer a specific way to construe this ‘hanging together’ for evaluating philosophical progress, focusing on the representation of dependence relations. Their account aims not simply to analyse the ordinary meaning of 'understanding', but rather "to explicate it in a way that makes it suitable for being deployed in an account of philosophical progress" (Dellsén et al. 2024, p. 674). The central claim holds that the degree to which a subject understands a phenomenon corresponds to the accuracy and comprehensiveness of their representation of the network of dependence relations relevant to that phenomenon. These dependence relations are conceived as objective, worldly relations often underpinning explanations. As Dellsén et al. (2024, p. 675) note, drawing on Kim (1994), dependence relations constitute the ontological correlates of explanation. While causation is paradigmatic in empirical science, other candidates include constitution, grounding, mereological dependence, truthmaking, conceptual containment, and supervenience; the framework remains agnostic about the precise list, assuming understanding involves representing whichever relations obtain in the domain. Furthermore, understanding a phenomenon X requires grasping its position within a wider network; it involves representing both how X depends on other phenomena and how further phenomena depend on X (Dellsén et al. 2024, p. 675). Such representation includes facts about existing dependencies – positive dependencies – and non-dependencies – negative dependencies, e.g., that X lacks dependence on Y. Understanding, on this view, admits of degrees relative to accuracy and comprehensiveness. Accuracy concerns how correctly the representation maps the dependencies that actually exist; comprehensiveness concerns the extent to which the representation includes relevant phenomena and their dependence relations (or lack thereof). These criteria can be in tension – sometimes motivating trade-offs like idealisation or abstraction (Dellsén et al. 2024, p. 675). This account is described as "robustly factive," requiring the representation to correspond to actual dependencies, yet also "epistemically undemanding" (Dellsén et al. 2024, pp. 674, 676). Understanding X, in this specific sense, does not require having justification for, or even belief in, the propositions represented. Dellsén et al. argue this feature is appropriate: "an explication of understanding that does not imply justification is arguably better suited for spelling out a plausible understanding-based account of philosophical progress" (2024, p. 676). Their motivation appears to be accommodating early-stage theorising or preliminary hypotheses where confidence might be low, yet the represented structure could still be largely accurate. This specific notion – understanding as accurate and comprehensive representation of dependence networks, without requiring justification – guides the subsequent analysis. While richer notions of understanding exist, this particular account provides a clear target for assessing the potential contribution of information-processing tools like LLMs. Enhancing philosophical understanding, therefore, involves refining this mental representation of the dependence network. Encountering arguments or information prompts evaluation: would incorporating this input yield a more accurate or comprehensive representation? For example, exploring theories of time might refine an initial map linking temporal passage simply to change. Distinguishing subjective passage from objective ordering enhances accuracy. Incorporating dependencies linking temporal theories to metaphysics (presentism, eternalism) or physics increases comprehensiveness. Mapping negative dependencies (e.g., B-series ordering not depending on subjective experience) or specifying dependence types (conceptual, metaphysical) improves clarity. The individual adjusts their model – adding, removing, or altering represented dependencies. The resulting revised representation constitutes enhanced understanding by virtue of improved accuracy or comprehensiveness, irrespective of justification for every element. Having defined this target concept, we turn now to how LLMs might assist humans in this process. ## 2. LLMs and Philosophical Understanding To begin with, someone who is optimistic about LLMs and philosophy can help themselves to some low hanging fruit. Dellsén et al.'s account allows for very trivial enhancements of a person's understanding. If you tell me what the English translation of *Cogito Ergo Sum* is, and what Descartes meant by that, you have, if I didn't know these things before, enhanced the comprehensiveness of my representation of dependency relations. LLMs are extremely good at answering these sorts of low level fact based questions. At this point, it might be objected that LLMs hallucinate and so should not be trusted for even these low-level fact-based questions. But this underestimates these systems. They excel at well-worn facts and frequently said bits of knowledge. Hallucinations more often emerge when they are asked a question which isn’t common knowledge (we will come back to this). A similar low-level capacity that LLMs excel at is *restructuring* text and thereby the ideas contained within it. Transformations include: simplification/sophistication, metaphors and analogies, or extrapolating specific examples from general categories. EXAMPLES? Again, we might think these are trivial and perhaps even temporary enhancements of philosophical understanding, but they are enhancements nonetheless. ## The Prospect of Significant Philosophical Generation A further question is whether LLMs can generate _novel_ outputs capable of significantly enhancing their users' philosophical understanding. By 'novel' and 'significant', I mean simply to mark out a range of philosophical contributions—from proclamations on the meaning of life, the universe, and everything, to thought-provoking article in an academic journal. Although I will leave these terms vague, most philosophers will have a practical sense of what counts as novel or significant. The issue is whether the outputs from these models can prompt genuine improvements or adjustments in a user's philosophical views The idea that their outputs can increase understanding is likely to face resistance. An initial objection one might have is that these systems do not themselves understand anything at all. They are stochastic parrots, incapable of doing anything other than arranging words in statistically plausible patterns. This may well me true: despite being based on neural networks the architecture of llms is a long way from that of a human mind\[^2\] \[^3\]; however, we should not assume that they need to reason in order to produce philosophically illuminating outputs. Butlin and Viebahn suggest that fine-tuned LLMs may produce outputs with a descriptive function, defined as "the function of conveying information to an observer or consumer system, so as to cause the consumer to behave as though some condition holds" (p. 4). While they argue this alone does not suffice for genuine assertion—since current LLMs cannot be sanctioned—they nonetheless highlight the significance of this capacity. Without fine-tuning, LLMs merely predict statistically likely sequences of words. Fine-tuning, however, involves additional training explicitly aimed at producing outputs that are accurate, relevant, and informative. By selectively rewarding outputs consistent with external sources or positively evaluated by humans, fine-tuning enhances the reliability of the information communicated. The significance of acquiring a descriptive function through fine-tuning lies in the generation of outputs reliable enough to aid users in refining their representation of dependence relations. In short, fine-tuning equips LLM outputs with a descriptive function, analogous to the way a thermometer is designed to indicate temperature. It might be then that fine-tuned LLMs should be given a hearing philosophically, indeed about anything else. Even if the system itself lacks understanding, and even it is not as consistently accurate as a thermometer, it has been made to provide correct answers. Correct answers are the sort of thing that would be useful for making one’s representation of dependency relations more accurate and more comprehensive. ## Generating Structured Output: Arguments and Chain-of-Thought Let us return to Deep Thought. What makes the answer ‘42’ so unsatisfying philosophically? Even if it were correct, it tells us nothing useful. Part of the problem, as the story suggests, is that the answer arrives without the question it answers. Without knowing the question, we cannot place the answer within any context or see how it relates to anything else. But even beyond that, the answer ‘42’ is inert because it stands alone, revealing none of the reasoning, calculations, or conceptual connections that might support it. It offers no insight into _why_ 42 might be the answer. Simple descriptions, like ‘the sky is blue’, face the same issue; on their own, they lack the structure needed to deepen understanding. To be philosophically useful, an output needs to show its working – the steps taken, the reasons considered. Had Deep Thought done this, its answer might have been informative. Reconsider the Deep Thought example: isolated pronouncements ('42', or perhaps 'Naive realism is false'), even if accurate, often lack utility for deepening one's grasp of dependencies. Indeed, such context-free outputs might fail to meet Butlin and Viebahn's condition for having a descriptive function in a robust sense; lacking structure or context, they may not reliably "cause the consumer to behave as though some condition holds" by guiding belief or action effectively. Philosophical understanding, particularly the mapping of dependence networks, is typically enhanced not by isolated statements but by _arguments_. Arguments function to reveal the purported logical, conceptual, or evidential links between claims, providing the structure needed to evaluate and refine one's representation network. This raises the question: can LLMs produce outputs structured like arguments, outputs which _would_ possess a descriptive function relevant to enhancing understanding? Historically, the ability of LLMs to provide reasoned arguments has been questioned. Earlier models often produced outputs where the reasoning was weak, inconsistent, or confabulated. [Optional: Add brief examples here]. However, the development of prompting techniques like Chain-of-Thought (CoT) marked a shift. CoT was designed specifically to elicit intermediate reasoning steps from LLMs, thereby improving performance on tasks requiring sequential reasoning. Common methods include few-shot prompting (providing examples of step-by-step reasoning) or zero-shot prompting (using triggers like "Let's think step by step"). It is also the case that many current advanced models (e.g., Gemini 2.5 Pro, ChatGPT o3) appear to incorporate CoT-like structured reasoning capabilities more implicitly, often generating step-by-step explanations or arguments without explicit CoT triggers, suggesting this approach is somewhat automated or integrated via system prompts or further fine-tuning. Technically, CoT leverages the underlying architecture of autoregressive Transformers. These models predict the next token based on the preceding sequence, using self-attention mechanisms to weigh the importance of different parts of the context. CoT prompting guides this process: instead of predicting the final answer directly, the model is steered towards generating tokens representing intermediate steps. Each generated step is appended to the context. The model's attention mechanism then considers this augmented context (including the steps generated so far) when predicting the subsequent token. This iterative process allows the model to condition later parts of its output on earlier parts, effectively using the generated text as an explicit "scratchpad" to manage the reasoning process. It is acknowledged that this process is a _simulation_ of reasoning, driven by the model learning statistical patterns from vast amounts of text data containing examples of arguments and explanations. LLMs do not manipulate internal symbolic logic representations or possess genuine comprehension. However, this simulation can be highly effective. LLMs utilising CoT achieve strong performance on various reasoning and problem-solving benchmarks (e.g., arithmetic reasoning, commonsense reasoning). Their simulated reasoning is often sufficient to solve complex tasks correctly. The key point for philosophical understanding is that the _outputs_ generated via CoT – the simulated arguments – are precisely the _type_ of structured text that facilitates the mapping of dependence relations. These generated arguments can be complex and detailed, providing substantive material for philosophical analysis. [Placeholder for examples of complex CoT philosophical arguments]. Furthermore, the utility of these outputs parallels engagement with human-generated philosophical texts. An argument does not need to be perfectly sound or entirely accepted by the reader to enhance their understanding. Evaluating an argument, even one containing flawed steps, prompts the reader to critically examine the purported dependencies and refine their own representation of the conceptual network. This process of refinement applies equally when the text being evaluated is a CoT output from an LLM. The effectiveness of this simulation is further supported by the observation that CoT capabilities emerge strongly with increased model scale (parameters in the hundreds of billions or trillions) and are enhanced by refinement techniques like Instruction Fine-Tuning (IFT) and Reinforcement Learning from Human Feedback (RLHF), which improve the coherence, reliability, and instruction-following of the generated reasoning. Therefore, CoT connects directly to philosophical methodology. It generates outputs structured like arguments, mirroring the sequential steps of premise-setting, inference, and conclusion. This structure, providing a potential pathway rather than just an endpoint, lends itself to analysis. It offers a valuable artefact for philosophical scrutiny due to its effectiveness as a simulation, even though it originates from a non-understanding source. **4. Enhancing Understanding via Chain-of-Thought Outputs** The structured outputs generated via CoT prompting can directly enhance philosophical understanding as defined by the Dellsén framework – the accurate and comprehensive representation of dependence networks. The value lies not in assuming the LLM 'understands' the dependencies it articulates, but in how the generated text can prompt refinements in the _user's_ representation. • **Improving Accuracy:** By presenting a step-by-step derivation, CoT outputs allow the user to scrutinise the purported dependence relations involved. If the LLM generates an argument linking concept A to concept C via an intermediate step B, the user can evaluate the claimed dependence of C on B, and B on A. Identifying a weak or fallacious link in the chain prompts the user to _correct_ their own mental model, potentially removing an inaccurately represented dependency. Conversely, a valid step might reinforce or confirm a dependency the user held tentatively. • **Increasing Comprehensiveness:** The intermediate steps generated by CoT can introduce new nodes or relations into the user's dependence network. The model might highlight a conceptual distinction, an implicit premise, or a relevant factor that the user had not previously considered, thereby expanding the scope of their representation. It might articulate dependencies between elements that the user had considered disparate. Furthermore, by tracing a line of reasoning, CoT can help delineate _negative_ dependencies – showing, for instance, why a certain conclusion does _not_ follow from a given premise under a specific interpretation, thus refining the boundaries of the represented network. • **Clarifying Dependence Types:** Philosophical understanding involves grasping _how_ things hang together. The step-by-step structure can help disambiguate the _nature_ of the proposed dependencies – are they logical, causal, conceptual, constitutive, or mereological? While the LLM might not explicitly label them, the context provided by the intermediate steps can allow the user to better infer the type of relationship being posited, leading to a more precise representation. Crucially, within the Dellsén framework, this enhancement does not require the user to have justification for, or even belief in, every proposition generated by the LLM. The framework is "epistemically undemanding" in this sense. The value derives from the potential of the generated structure to prompt a re-evaluation and refinement of the user's existing representation, leading to a potentially more accurate or comprehensive map of the relevant dependence network, even if parts of the generated chain are ultimately rejected upon reflection. ## Scrap?? - This distinction between the system's internal state and its output's potential function resonates with the analysis by Butlin and Viebahn (forthcoming) regarding AI assertion. # Prompt for analytic style ## Task Rewrite the TEXT below in the style of analytic philosophy while preserving 100% of its semantic content. ## Style Guidelines Transform the writing to be: - Precise and economical with language - Clear and straightforward in expression - Logically structured with minimal ornamentation - Free of unnecessarily complex vocabulary where simpler terms suffice - Rigorous in argumentation without pretentiousness Good analytic philosophy writing is dry, direct, and prioritizes clarity over stylistic flourish. Remove pompous or pretentious elements while maintaining all substantive points.