# [[Generating Philosophy with Artificial Intelligence]] (1000 word version) #paper/generatingphilosophy #draft >“Forty-two," said [[Deep Thought]], with infinite majesty and calm. It was a long time before anyone spoke. Out of the corner of his eye Phouchg could see the sea of tense expectant faces down in the square outside. "We're going to get lynched aren't we?" he whispered. "It was a tough assignment," said [[Deep Thought]] mildly. "Forty-two!" yelled Loonquawl. "Is that all you've got to show for seven and a half million years' work?" "I checked it very thoroughly," said [[the computer]], "and that quite definitely is the answer. I think the problem, to be quite honest with you, is that you've never actually known what [[the question]] is." In _The Hitchhiker's Guide to the Galaxy_, [[Douglas Adams]] considers whether AI can provide philosophical answers. Humanity constructs [[Deep Thought]], a computer tasked with determining life's ultimate meaning, only to receive the answer '42'—an answer that is (apparently) correct but, due to their failure to formulate a proper question, means next to nothing at all. In 2025 we are in a position to think about [[the relationship between]] philosophy and AI *for real*. Can LLMs enhance [[philosophical understanding]]? On one hand, they have been described as *reasoning engines*—precisely the sort of thing philosophers might want to seek out. On the other, they are notoriously prone to 'hallucinations', and the banally hyperbolic essays churned out by ChatGPT have become the scourge of undergraduate teaching. - In this paper.... ## 1 Understanding as [[Representing Dependence Networks]] To assess how LLMs might enhance [[philosophical understanding]], we first require a definition of this understanding. In this paper, we adopt the framework proposed by Dellsén et al. (2024). Their approach, explaining understanding in terms of [[representing dependence networks]], resonates with the broader philosophical aim, articulated by Sellars (1962, p. 1), to grasp 'how things in the broadest possible sense of the term hang together in the broadest possible sense of the term'. Sellars conceived philosophy as pursuing a synoptic vision, integrating diverse entities, concepts, events, norms, and theoretical posits into a coherent whole by mapping their causal, logical, conceptual, and explanatory interconnections. Dellsén et al. offer a specific way to construe this ‘hanging together’ for evaluating [[philosophical progress]], focusing on the representation of [[dependence relations]]. Their account aims not simply to analyse the ordinary meaning of 'understanding', but rather "to explicate it in a way that makes it suitable for being deployed in an account of [[philosophical progress]]" (Dellsén et al. 2024, p. 674). The central claim holds that the degree to which a subject understands a phenomenon corresponds to the accuracy and comprehensiveness of their representation of the network of dependence relations relevant to that phenomenon. These dependence relations are conceived as objective, worldly relations often underpinning explanations. As Dellsén et al. (2024, p. 675) note, drawing on Kim (1994), dependence relations constitute the ontological correlates of explanation. While causation is paradigmatic in empirical science, other candidates include constitution, grounding, mereological dependence, truthmaking, conceptual containment, and supervenience; the framework remains agnostic about the precise list, assuming understanding involves representing whichever relations obtain in the domain. Furthermore, understanding a phenomenon X requires grasping its position within a wider network; it involves representing both how X depends on other phenomena and how further phenomena depend on X (Dellsén et al. 2024, p. 675). Such representation includes facts about existing dependencies – positive dependencies – and non-dependencies – negative dependencies, e.g., that X lacks dependence on Y. Understanding, on this view, admits of degrees relative to accuracy and comprehensiveness. Accuracy concerns how correctly the representation maps the dependencies that actually exist; comprehensiveness concerns the extent to which the representation includes relevant phenomena and their dependence relations (or lack thereof). These criteria can be in tension – sometimes motivating trade-offs like idealisation or abstraction (Dellsén et al. 2024, p. 675). This account is described as "robustly factive," requiring the representation to correspond to actual dependencies, yet also "epistemically undemanding" (Dellsén et al. 2024, pp. 674, 676). Understanding X, in this specific sense, does not require having justification for, or even belief in, the propositions represented. Dellsén et al. argue this feature is appropriate: "an explication of understanding that does not imply justification is arguably better suited for spelling out a plausible understanding-based account of philosophical progress" (2024, p. 676). Their motivation appears to be accommodating early-stage theorising or preliminary hypotheses where confidence might be low, yet the represented structure could still be largely accurate. This specific notion – understanding as accurate and comprehensive representation of dependence networks, without requiring justification – guides the subsequent analysis. While richer notions of understanding exist, this particular account provides a clear target for assessing the potential contribution of information-processing tools like LLMs. Enhancing philosophical understanding, therefore, involves refining this mental representation of the dependence network. Encountering arguments or information prompts evaluation: would incorporating this input yield a more accurate or comprehensive representation? For example, exploring theories of time might refine an initial map linking temporal passage simply to change. Distinguishing subjective passage from objective ordering enhances accuracy. Incorporating dependencies linking temporal theories to metaphysics (presentism, eternalism) or physics increases comprehensiveness. Mapping negative dependencies (e.g., B-series ordering not depending on subjective experience) or specifying dependence types (conceptual, metaphysical) improves clarity. The individual adjusts their model – adding, removing, or altering represented dependencies. The resulting revised representation constitutes enhanced understanding by virtue of improved accuracy or comprehensiveness, irrespective of justification for every element. Having defined this target concept, we turn now to how LLMs might assist humans in this process. ## 2. LLMs and Philosophical Understanding To begin with, someone who is optimistic about LLMs and philosophy can help themselves to some low hanging fruit. Delsén et al.'s account allows for very trivial enhancements of a person's understanding. If you tell me what the English translation of *Cogito Ergo Sum* is, and what Descartes meant by that, you have, if I didn't know these things before, enhanced the comprehensiveness of my representation of dependency relations, and thereby my understanding. LLMs are extremely good at answering these sorts of low level fact based questions. At this point, it might be objected that LLMs hallucinate and so should not be trusted for even these low-level fact-based questions. But this underestimates these systems. They excel at well-worn facts and frequently said bits of knowledge. They struggle much more when they don't know the answer to a question. That is when they are most likely to start making things up. A similar low-level capacity that LLMs excel at is _restructuring_ text and thereby the ideas contained within it. Transformations include: simplification/sophistication, metaphors and analogies, or extrapolating specific examples from general categories. Again, we might think these are trivial and perhaps even temporary enhancements of philosophical understanding, but they are enhancements nonetheless. ## The Prospect of Significant Philosophical Generation A more interesting question, of course, is the extent to which LLMs can produce *novel* outputs which significantly enhance their users philosophical understanding, in the same way that reading an article in a top tier philosophy journal might.[^1] This idea, however, is likely to face resistance. An initial objection one might have is that these systems do not themselves understand anything at all. They are stochastic parrots, incapable of doing anything other than arranging words in statistically plausible patterns. This may well me true: despite being based on neural networks the architecture of llms is a long way from that of a human mind[^2] [^3], however, it is not obvious that that this them from producing text that has philosophical value – that is, output which enhances the philosophical understanding of those philosophers who read it and understand it. Butlin and Viebahn suggest that fine-tuned LLMs may acquire the capacity to produce outputs possessing a _descriptive function_: "the function of conveying information to an observer or consumer system, so as to cause the consumer to behave as though some condition holds" (p. 4). In simpler terms, a descriptive function refers to the model's ability to reliably communicate accurate, verifiable information that influences the user's understanding or actions. This capacity is developed during fine-tuning, an additional training stage where models are explicitly optimised to produce outputs valued for their accuracy, relevance, and informativeness. By selecting outputs based on criteria such as consistency with external sources or positive human evaluations, fine-tuning increases the likelihood that LLM-generated text will be correct and helpful. Put differently, the fine-tuning process instils this descriptive function in the LLM's outputs, much as a thermometer is designed with the function of indicating temperature. It might be then that fine-tuned LLMs should be given a hearing philosophically, indeed about anything else. Even if the system itself lacks understanding, and even it is not as consistently accurate as a thermometer, it has been made to provide correct answers. Correct answers are the sort of thing that would be useful for making one’s representation of dependency relations more accurate and more comprehensive. ## LLM Text, to be checked. While basic factual retrieval and text restructuring offer preliminary enhancements to understanding, the potential for LLMs to contribute significantly to philosophical inquiry hinges on their capacity for processes analogous to reasoning. One prominent capability in this regard is Chain-of-Thought (CoT) prompting. CoT involves instructing the LLM not simply to provide an answer, but to outline the intermediate steps or reasoning process that leads to that answer. Techniques such as adding "Let's think step by step" to a prompt encourage the model to externalise a sequence of inferences or considerations. From a mechanistic perspective, as outlined in the provided technical description, CoT does not fundamentally alter the LLM's next-token prediction objective. Instead, it guides the generation process to produce tokens that represent intermediate stages of computation or inference. When an LLM generates these "reasoning tokens," it modifies its own context. Subsequent tokens are then predicted based not only on the initial prompt but also on these explicitly generated intermediate steps. This allows the model to effectively condition its later outputs on its own earlier 'conclusions,' facilitating a more structured approach to problem-solving that mirrors aspects of human reasoning. By making intermediate computational states persistent and attendable within the sequence context, CoT effectively augments the model's ability to manage complexity and track inferential pathways. This capability is relevant to philosophical methodology. Philosophical arguments often involve sequential steps: establishing premises, drawing inferences, considering objections, and tracing implications. The structure generated by CoT outputs can mirror this process. While the LLM is still fundamentally engaged in sequence prediction based on patterns learned from vast datasets (which include examples of reasoning), the _output_ generated through CoT presents information in a format that lends itself to philosophical analysis. It provides not just a conclusion, but a potential pathway to that conclusion, which can be scrutinised for validity, coherence, and relevance. This externalisation of potential inferential steps represents a capacity to generate text _structured_ like reasoning, regardless of whether the system possesses genuine understanding or intentionality behind those steps. **4. Enhancing Understanding via Chain-of-Thought Outputs** The structured outputs generated via CoT prompting can directly enhance philosophical understanding as defined by the Dellsén framework – the accurate and comprehensive representation of dependence networks. The value lies not in assuming the LLM 'understands' the dependencies it articulates, but in how the generated text can prompt refinements in the _user's_ representation. • **Improving Accuracy:** By presenting a step-by-step derivation, CoT outputs allow the user to scrutinise the purported dependence relations involved. If the LLM generates an argument linking concept A to concept C via an intermediate step B, the user can evaluate the claimed dependence of C on B, and B on A. Identifying a weak or fallacious link in the chain prompts the user to _correct_ their own mental model, potentially removing an inaccurately represented dependency. Conversely, a valid step might reinforce or confirm a dependency the user held tentatively. • **Increasing Comprehensiveness:** The intermediate steps generated by CoT can introduce new nodes or relations into the user's dependence network. The model might highlight a conceptual distinction, an implicit premise, or a relevant factor that the user had not previously considered, thereby expanding the scope of their representation. It might articulate dependencies between elements that the user had considered disparate. Furthermore, by tracing a line of reasoning, CoT can help delineate _negative_ dependencies – showing, for instance, why a certain conclusion does _not_ follow from a given premise under a specific interpretation, thus refining the boundaries of the represented network. • **Clarifying Dependence Types:** Philosophical understanding involves grasping _how_ things hang together. The step-by-step structure can help disambiguate the _nature_ of the proposed dependencies – are they logical, causal, conceptual, constitutive, or mereological? While the LLM might not explicitly label them, the context provided by the intermediate steps can allow the user to better infer the type of relationship being posited, leading to a more precise representation. Crucially, within the Dellsén framework, this enhancement does not require the user to have justification for, or even belief in, every proposition generated by the LLM. The framework is "epistemically undemanding" in this sense. The value derives from the potential of the generated structure to prompt a re-evaluation and refinement of the user's existing representation, leading to a potentially more accurate or comprehensive map of the relevant dependence network, even if parts of the generated chain are ultimately rejected upon reflection. ## Scrap?? * This distinction between the system's internal state and its output's potential function resonates with the analysis by Butlin and Viebahn (forthcoming) regarding AI assertion. * They argue that for an output to be an assertion, it must possess a 'descriptive function' – the function of conveying information to a consumer. * Crucially, they contend that pre-trained LLMs, optimised solely to predict likely word sequences, produce outputs lacking this descriptive function. * Their outputs mimic assertion form, but their function is merely sequence completion, not information conveyance. * However, LLMs fine-tuned specifically for 'groundedness' or 'correctness', like LaMDA or Sparrow, undergo training that selects for accuracy and alignment with known sources. * This fine-tuning process, Butlin and Viebahn suggest, means the outputs of such systems may acquire a descriptive function, serving the purpose of providing trustworthy information, even if the systems themselves fail other conditions for genuine assertion, such as sanctionability. * This highlights that an output's functional properties, like potentially having a descriptive function, can be distinct from whether the generating system possesses understanding . One could argue that the origin of the text, whether from a thinking being or a non-understanding system, matters less than its content and its ability to make the human reader think philosophically. To take this argument further, one might use a concept like 'descriptive functions', possibly drawing on work by figures like Butler [Citation needed if specific work is intended]. The basic idea here is that even if an LLM output is not backed by the usual kind of authorial intention or understanding, the text itself can still function descriptively. It produces outputs designed to be useful, claiming to offer descriptions – perhaps unstated descriptions – of the world or parts of it. These outputs, because of their structure and content, could potentially contain or lead to philosophical insights, regardless of their source. This view, however, faces its own problems. A likely counter-argument might use the analogy of infinite monkeys eventually producing Shakespeare; if LLM outputs are just statistically generated sequences, why should their descriptive functions be taken seriously as potential sources of philosophical insight? This objection highlights the need for more study into why LLM-generated text, seen as having descriptive functions, might be considered important for knowledge or able to contribute meaningfully to philosophical discussion. Dealing with this requires a closer look at how LLMs work and the exact nature of these proposed descriptive functions. [^1]: There are of course degrees of significance, some philosophical works have a profound impact on the philosophical understanding of large swathes of people (Gettier, Kant etc.). [^2]: mention butlin, buckner and m etc. [^3]: Mention also the way they are trained being very different to how human's learn.