# here https://nickneek.github.io/presentations/ is a presentation that an LLM wrote for me based o...
## Retrieval Notes
- Session id: `993301df-5319-4009-9b77-ecdc013b32a1`
- Last activity: `2026-03-04T14:49:02.624Z`
## My Notes
<!-- Add your notes here. This section is preserved across syncs. -->
## Conversation
### User
here https://nickneek.github.io/presentations/ is a presentation that an LLM wrote for me based on this manuscript (below) I think that more slides coulkd be added, or extra information should be added to some slides, what do you think. Note, i am looking for stylisticallty appropriate augmentation. make sure that you explain your analysis, justification, reasons etc etc. for your answer BEFORE giving me your final answer. Please give me a long, deep, and detailed answer, with lots of options/ideas: as we are in the 'kicking ideas around' stage, I want as much feedback as you can give me so that i have a lot to consider.
don't make anychanges yet
BTW, if you cannot access the slides TELL ME don't dick about worryinng about it
Generating Philosophy Without Artificial Intelligence
Nick Young & Enrico Terrone
0. Introduction
“Forty-two,” said Deep Thought, with infinite majesty and calm. It was a long time
before anyone spoke. Out of the corner of his eye Phouchg could see the sea of tense
expectant faces down in the square outside. “We’re going to get lynched aren’t we?” he
whispered. “It was a tough assignment,” said Deep Thought mildly. “Forty-two!” yelled
Loonquawl. “Is that all you’ve got to show for seven and a half million years’ work?” “I
checked it very thoroughly,” said the computer, “and that quite definitely is the answer.
I think the problem, to be quite honest with you, is that you’ve never actually known
what the question is.”
— Douglas Adams, The Hitchhiker’s Guide to the Galaxy
In The Hitchhiker’s Guide to the Galaxy, humanity asks an AI to do some philosophy.
A computer named Deep Thought is constructed and instructed to provide “The
Answer to the Ultimate Question of Life, the Universe, and Everything.” Humanity
builds this computer only to receive the answer ‘42’—an answer which, while apparently
correct, means next to nothing at all due to humanity’s failure to know what
the Ultimate Question in fact is.
In 2026, humanity has reached a position in which it can actually ask machines
philosophical questions. One reason for optimism is that AI has had considerable
success in other domains. In February 2026, researchers working on gluon scattering
amplitudes gave GPT-5.2 worked examples for three, four, five, and six particles and
asked it to find the general formula. The model conjectured a formula, completed a
formal proof, and overturned a forty-year-old assumption (Guevara et al. 2026).1
The question is harder to answer than it might seem, because philosophy does not
have uncontroversial success conditions. What counts as a contribution depends on
1Other AI-assisted breakthroughs include protein structure prediction, which won the 2024 Nobel Prize
in Chemistry (Hassabis and Jumper, AlphaFold); solving a 30+ year challenge in quantum error correction
(Google Quantum AI, Willow chip); and discovering new symmetries in black hole event horizon equations
(Lupsasca with GPT-5).
what philosophy is, and conceptions of philosophy differ in ways that matter for the
question about AI.
Some conceptions locate philosophy in texts. Dellsén, Firing, Lawler, and Norton
(2024) argue that philosophical progress consists in putting people in a position to
increase their understanding—what they call the for-whom rather than by-whom
account—in which public utility of the work, not the internal states of whoever
produced it, determines whether progress has occurred. Bengson, Cuneo, and Shafer-
Landau (2022) characterise philosophical inquiry as theory construction evaluated
by criteria: accommodation of data, explanatory power, integration, theoretical
virtue, all of which are assessable by examining the theory itself. Williamson (2024)
defends an abductive methodology judging theories by their simplicity, elegance, and
explanatory power. On any of these accounts, a philosophical contribution is a text
exhibiting certain properties; the question who or what produced it does not enter
the evaluation.2
Other conceptions locate philosophy in the practitioner. For Hadot (1995), philosophy
is a practice of self-transformation; for the later Wittgenstein (1953), it is a form
of therapy; for Merleau-Ponty, it requires us to “slacken the intentional threads which
attach us to the world” (1945, p. xv) in order to examine them. What these views
share is a commitment: philosophy requires being a certain kind of subject, capable
of self-transformation, or therapy, or phenomenological observation—and machines
are not such subjects. On Nietzsche’s account (as Sorgner reads it), philosophers
are creators of values, expressing drives and psychophysiology that LLMs lack. To
a proponent of any such view, the question of whether LLMs can do philosophy is
closed before it opens.34
I do not take a position here on which conception of philosophy is correct. This paper
assumes the text-focused conception. If philosophical evaluation concerns proper2Pigliucci
offers a related formulation: philosophy “attempts to clarify things, or to analyze in order to bring
about understanding, not really to discover new facts, but rather to evoke rational conclusions arising from
certain ways of looking at a given problem or set of facts.” Whether such evocation requires a human evoker
is the question at issue.
3On transformative conceptions, what makes an activity philosophical is something that happens in the
practitioner rather than anything assessable in what she produces (Hadot 1995; cf. late Wittgenstein on
philosophy as therapy). Transcendental and phenomenological approaches presuppose having experience
(Kant 1781/1787; Merleau-Ponty 1945). World-view conceptions require the philosopher to live a human
life (Dilthey; see Overgaard, Gilbert & Burwood 2013: ch. 8). Jones (2006) holds that philosophy requires
entering an identity-conferring conversation within a community; Sorgner reads Nietzsche as requiring
biology and psychophysiology.
4The distinction between text-focused and practitioner-focused conceptions maps imperfectly but suggestively
onto the analytic/continental divide: analytic philosophy tends to emphasise texts and arguments
as the locus of evaluation, while continental traditions more often locate philosophical activity in lived
practice or self-transformation.
ties of texts—coherence, handling of objections, illumination of subject matter
—then whether LLMs can do philosophy is a question about the texts they produce.
Even within the text-focused framework, some argue that LLMs cannot produce
texts exhibiting the right properties. These objections do not concern what philosophy
is; they concern what philosophical reasoning requires. Floridi, Nobre, and
Taddeo (2024) argue that genuine abductive reasoning is beyond LLMs’ capacities;
Zahavy (2026) argues that they cannot make the creative leaps needed to propose
new theoretical frameworks. If either argument succeeds, LLMs cannot do philosophy
regardless of what we think philosophy is.
I argue that LLMs can produce philosophy exhibiting the relevant properties. Section
1 argues that philosophical evaluation concerns text-internal criteria. Section 2
presents objections from Floridi et al. and Zahavy. Section 3 responds. Section 4
considers what a demonstration would look like.
1. Philosophy in the Text
When Watson and Crick published their paper on the double helix in 1953, they
announced what they had found: a particular arrangement of nucleotides, with two
strands running in opposite directions, held together by hydrogen bonds between
complementary base pairs. The double helix existed before they described it; their
paper reported what was already there. Had Rosalind Franklin announced it first, the
finding would have been the same—the same arrangement, the same base pairings
—just differently attributed.
Quine’s “Two Dogmas of Empiricism,” published two years earlier, is not like this.
Quine made arguments: against the coherence of the analytic/synthetic distinction,
against reductionism about meaning. The arguments are the contribution. There is
no arrangement of facts the paper reports, no prior reality that someone else might
have found instead. A different philosopher reaching the same conclusions would
have had to make arguments. If the arguments differed, so would the contribution.5
What, then, are we evaluating when we evaluate a piece of philosophy? Not whether
a text accurately reports something prior to it, since there is nothing prior that it
reports. We evaluate the arguments themselves. But for what?
5Literature has a similar character: when we evaluate a novel, we assess prose, pacing, tension—features
internal to the text. Philosophy shares this constitutive character but differs in evaluative criteria: literature
is assessed aesthetically, philosophy for argumentative virtues.
Consider Lipton’s distinction between likeliness and loveliness. The likeliest explanation
is the most probable. The loveliest is the one that would provide the deepest
understanding if it were true. Lipton’s point is that we can assess loveliness independently
of likeliness.
Semmelweis investigated why women in the First Division of the Vienna maternity
hospital died at higher rates than those in the Second. He considered several potential
explanations: differences in birthing position, differences in the route taken
by the priest administering last rites, differences in exposure to cadaveric matter.
Each was a potential explanation—something that would explain the difference if
true. And Semmelweis could assess their loveliness without yet knowing which was
correct. The cadaveric explanation was lovelier because it unified the phenomenon
with known facts about infection. The priest explanation, even if true, would leave it
mysterious why the priest’s presence caused death.
Philosophical evaluation has this character. We assess arguments for properties that
can be judged from the arguments themselves, without first establishing that their
conclusions are correct. These properties include elegance and unity—a good theory
explains much with little, and its parts hang together rather than being a collection
of separate claims.
Williamson notes that these are the same theoretical virtues that guide theory choice
in science, but in philosophy they must be weighed without direct empirical test. He
compares the philosophical cycle of analysis, counterexample, and revised analysis
to overfitting in statistics. Each repair to accommodate a new counterexample risks
making the theory more ad hoc, more gerrymandered to the cases at hand.
Deep Blue plays good chess. Its moves respond effectively to threats, secure positional
advantages, and contribute to coherent strategic plans. But Deep Blue does
not play creatively—it searches exhaustively rather than intuiting the best move.
Gaut observes that this shows creativity and domain-specific excellence can come
apart. A move is good chess or it is not, regardless of whether it was found by creative
insight or brute computation. Whether a philosophical argument handles objections
well, draws distinctions at the right places, or illuminates its subject matter is
assessable in the same way. The question is what properties the argument has, not
how it came to have them.
Dellsén and colleagues argue that philosophical progress consists in putting people
in a position to increase their understanding. Suppose a scientist publishes an important
finding and then dies. Everyone who read the paper also dies, or forgets what
they read. Has the progress been lost? No. Progress occurred when the publication
made it possible for someone to understand, whether or not anyone actually did. The
materials remain publicly available; that is what matters.
The same applies to philosophy. A published argument constitutes progress if it
enables understanding, regardless of who or what produced it, and regardless of
whether anyone currently grasps it. Blind review operationalises this: referees assess
whether an argument handles objections and illuminates its subject matter without
knowing who wrote it. The practice treats authorship as irrelevant to evaluation.
If philosophical evaluation concerns properties of arguments—elegance, coherence,
illumination of subject matter—and these properties are assessable by reading the
arguments, then the production process is not evaluatively relevant. The question is
whether a text exhibits these properties, not what brought it into existence.6
2. LLMs and Abduction
I want to argue that the objections to LLM philosophy presented in the previous
section do not apply to the kind of philosophical work that Williamson describes.
Floridi et al. and Zahavy identify capacities that LLMs lack—hypothesis evaluation in
the one case, embodied simulation in the other—but these capacities are not required
for what Williamson calls philosophical abduction.
Consider first how LLMs produce their outputs. An LLM predicts the next token
in a sequence based on probability distributions learned from training data. When
prompted to explain why a car might not start on a cold morning, it generates text
that exhibits explanatory structure: it identifies a hypothesis (the battery), provides
a reason (cold weather reduces battery efficiency), and presents the explanation with
the connectives and qualifications that explanations typically have. But the LLM does
not select this explanation by comparing it with alternatives and judging it best. It
outputs the most probable continuation given its training. Floridi et al. put the point
this way:
Given a prompt, they generate a plausible continuation (a hypothesis or explanation)
based purely on learned associations. In reality, their operation is driven by maximising
6Different theorists articulate these criteria differently. Williamson emphasises elegance, unity, and non-
ad-hocness (2024, pp. 152–3). Bengson, Cuneo, and Shafer-Landau organise evaluative criteria into five
levels: accommodation, explanation, substantiation, integration, and virtue (2022). Dellsén and colleagues
cash out philosophical progress in terms of representing dependence relations accurately and comprehensively
(2024). The vocabularies differ, but all concern properties assessable from theories themselves.
the probability of the sequence… The model does not understand what an explanation
is, but it produces text that follows the typical phrasing and structure of explanations. It
does not reason about causes from scratch but outputs typical causes for typical effects
observed in the training data.
Floridi et al. call this zeroth-order abduction. The phrase marks an absence: what is
missing is the comparative evaluation that genuine abduction involves. In genuine
abduction—what Floridi et al. call strong abduction—one generates multiple hypotheses,
compares them, and selects the best. LLMs do not do this. They generate a
plausible continuation without evaluating whether that continuation is better than
alternatives they did not generate.
This matters for some questions. It matters, for instance, if we want to know whether
LLMs reason in the way humans reason. Floridi et al.‘s answer is that they do not: the
mechanism is stochastic, not inferential. But it matters less if the question is whether
LLMs can produce outputs that meet philosophical standards. For philosophical
evaluation concerns the output—whether the theory is elegant, unified, and handles
the evidence—not the process that generated it. Floridi et al. themselves note this:
“if an AI can generate the same explanatory hypothesis a human would, does it
matter that the process was different? From an epistemological standpoint, perhaps
yes—justification is significant—but regarding the content of the hypothesis and our
interpretation of it, maybe not.”
Zahavy’s objection cuts differently. His concern is not that LLMs fail to evaluate
hypotheses but that they cannot generate certain hypotheses at all. His paradigm
case is Einstein’s formulation of the equivalence principle. Einstein did not have
data sufficient to infer general relativity inductively; Newtonian mechanics faced no
empirical crisis, and the anomaly of Mercury’s perihelion was attributed to an undiscovered
planet rather than a flaw in Newton’s laws. Nor could Einstein deduce the
equivalence principle from prior axioms—it was itself a new axiom, something that
had to be formulated before deduction could begin.
How, then, did Einstein arrive at it? Zahavy’s answer is manipulative abduction: generating
hypotheses through embodied simulation rather than symbolic manipulation.
Einstein imagined himself inside a falling elevator. He simulated the sensations of
an observer in that scenario—objects released from the hand appearing to hover, the
floor rushing up to meet falling things—and abduced from that simulated experience
that gravity and acceleration must be the same phenomenon. The thought experiment
was not a logical exercise conducted in symbols but a sensory one conducted
in imagination.
LLMs, Zahavy argues, cannot do this. They operate entirely in the domain of symbols
—tokens, vectors, probability distributions—without access to the physical referents
those symbols represent. Zahavy quotes Harnad’s phrase: LLMs are “high-dimensional
‘Chinese Rooms’, manipulating the language of physics without access to the
physical referents that give that language meaning.” They can derive consequences
from axioms once those axioms are given in symbolic form, but they cannot make
the leap from sensory experience to new axioms that Zahavy takes to be constitutive
of scientific invention.
Zahavy limits his argument to physics: “this proposal is specifically tailored to the
physical sciences, where the object of study is external material reality.” But the
argument structure extends to phenomenological experience more broadly. If LLMs
lack subjective experience altogether, then philosophy that relies primarily on phenomenological
observation will be difficult for them. Similar considerations apply to
intuitions (Machery 2017) and to aesthetic experience. In the next section I argue that
these considerations are less damaging than they appear.
3. Thought Experiments and Armchair Abduction
Williamson characterises philosophical abduction in terms general enough to
encompass both science and mathematics. Theories are ranked as potential explanations
of a body of evidence, and the ranking depends on two things: how well the
theory fits the evidence, and how well it scores on what Williamson calls the intrinsic
virtues of a good theory:
Apart from its relation to E, the more T has the intrinsic virtues of a good theory, the
better (ceteris paribus). It should be elegant and unified, not arbitrary, gerrymandered,
ad hoc, or messily complicated. It should be informative and general. In brief, it should
combine simplicity with strength. (Williamson 2024, p. 354)
The evidence base for philosophical abduction is unrestricted. Williamson writes that
“nothing in this account requires the evidence propositions, the explananda, to be of
some special kind. Any known truths will do” (p. 355). In philosophy, this evidence
consists in arguments, counterexamples, thought experiments, and the distinctions
and results that prior inquiry has established—in short, the accumulated textual
record of the discipline. The explanations philosophy offers are typically constitutive
rather than causal: accounts of what something consists in, how concepts relate,
what follows from what.
Williamson draws an explicit analogy with mathematics: “mathematics is a precedent
for a successful discipline with an ‘armchair’ methodology that still has a key
role for abduction. Thus it would be myopic to assume that an abductive methodology
for philosophy implies its assimilation to the experimental sciences” (p. 358).
Philosophical abduction can be conducted from existing knowledge, without new
empirical observation, and evaluated by examining the theory itself rather than
comparing it to mind-independent facts.
This matters for the question of whether LLMs can produce good philosophy. The
philosophical corpus—the body of philosophy that has survived peer review, been
taught, been cited, and been anthologised—exhibits the intrinsic virtues Williamson
identifies. This is not accidental. Peer review is a filter: papers that fail to handle objections,
or draw arbitrary distinctions, or offer no illumination of the subject matter,
are rejected. What survives to be published and taught is a sample of what the discipline
judges good, where the standards of goodness track exactly the intrinsic virtues
Williamson describes.
An LLM trained on this corpus has learned the distribution. It has learned what
makes a philosophical explanation score well—not through explicit instruction, but
through exposure to a body of text that has been filtered by those standards over
centuries. When Floridi et al. describe LLMs as “engines of generative plausibility,”
they are describing systems that have absorbed, from the corpus, the evaluative
standards that philosophical abduction employs. Floridi et al.‘s diagnosis—that LLMs
produce plausible outputs without evaluating alternatives—is correct at the level of
mechanism. But what counts as “plausible” in philosophy is exactly what scores well
on Williamson’s virtues; and what scores well on those virtues is exactly what the
corpus encodes.
Turn now to Zahavy’s argument. Zahavy claims that scientific invention requires a
leap from sensory experience to formal axioms—what he calls the E→ A Jump—
and that this leap involves embodied simulation rather than symbolic manipulation.
Einstein imagined the sensations of an observer in a falling elevator, and from that
simulated experience abduced the equivalence of gravity and acceleration. LLMs,
lacking access to physical referents, cannot make this leap.
Does philosophy require something similar? Consider how philosophical thought experiments
actually work. Take Putnam’s Twin Earth case. Putnam asks us to imagine a
planet where the clear liquid in the lakes and rivers is not H2O but a different chemical
compound, XYZ, which is superficially indistinguishable from water. Oscar, on Earth,
and Twin Oscar, on Twin Earth, both use the word “water” to refer to the clear liquid in
their environments. They are molecule-for-molecule identical in their internal states,
yet—Putnam argues—they mean different things by “water.” Oscar means H2O; Twin
Oscar means XYZ. The conclusion: meaning is not determined by what is in the head.
Notice what this thought experiment does not require. It does not require anyone to
simulate the sensations of being on Twin Earth or drinking XYZ. The thought experiment
is articulated entirely in language, recorded in text, and does its intellectual
work at the level of concepts and propositions. Readers evaluate it by asking whether
the scenario is coherent, whether the conclusion follows, whether the argument illuminates
something about meaning—and all of these questions can be answered by
examining the text. The same is true of Jackson’s Mary, Searle’s Chinese Room, Parfit’s
teleporter, and every other philosophical thought experiment in the literature. They
are textual objects, and the work they do is textual work.
Zahavy’s model of creative invention—sensory experience, embodied simulation,
formal axioms—fits Einstein’s physics, where the object of study is external material
reality and the axioms must connect to that reality through grounded concepts. But
philosophical thought experiments do not make this demand. They are already articulated
in language; they enter the record as text; and their intellectual contribution
consists in the arguments they embody. Whatever private experiences philosophers
have in arriving at thought experiments, the thought experiments themselves—as
they enter the literature and do philosophical work—are linguistic objects.
But what of the broader objection—that LLMs lack phenomenological experience,
intuitions, and aesthetic response? We do not claim that LLMs have these capacities.
What they have is a training corpus containing extensive descriptions of human
experience. This is not first-hand access to experience but access to descriptions—and
descriptions of experience are what philosophical argument typically works with.
This bears on the question of novelty. Zahavy’s argument, if it worked, would suggest
that LLMs cannot produce genuinely new theories—that they are limited to recombining
existing materials. But what does philosophical novelty consist in? Williamson
notes that “enumerative induction is inadequate for systematic philosophical theorizing,
which often requires introducing new distinctions at a more abstract level not
given in the data” (p. 353). Dummett’s distinction between assertoric content and
ingredient sense, for instance, is not derived from data by induction; it is a conceptual
innovation, a new way of organising existing materials. Kripke’s framework of rigid
designators similarly reconfigured existing materials about reference, identity, and
modality. Lewis’s modal realism took the apparatus of possible-world semantics and
proposed treating it with full ontological seriousness.
These are innovations, but they are conceptual innovations—new ways of carving
up the problem-space, new distinctions at a higher level of abstraction. They do not
involve leaps from bodily sensation to formal axioms. An LLM trained on a corpus
that documents both the conceptual materials and the standards for combining
them has, in principle, the resources to produce similar reconfigurations. Whether it
will in fact do so is an empirical question. But the arguments that it cannot—Zahavy’s
arguments about embodied simulation and physical grounding—do not straightforwardly
apply when the materials and the standards are textual.
Lipton makes a related point about the relationship between different levels of
description:
If these suggestions are along the right lines, then arguing that Inference to the Best
Explanation is wrong because Bayesianism is right is like arguing that thinking about
technique cannot help my squash game because the motion of the ball is governed by
the laws of mechanics. (Lipton 2004, p. 108)
The ball’s motion is governed by mechanics, and an LLM’s outputs are generated
by stochastic processes. But the mechanical description of the ball does not make
thinking about technique pointless, and the stochastic description of the LLM does
not settle the question of whether its outputs meet philosophical standards. Floridi
et al. correctly identify the mechanics; philosophical evaluation concerns a different
level—the level at which we ask whether the arguments are clear, whether they
handle objections, whether they illuminate the subject matter. These are questions
about the output, not about the process that generated it, and they are the questions
that matter for determining whether what the LLM has produced is good philosophy.
References
Frankish, K. (2024). What are large language models doing? In A. Strasser (Ed.),
Anna’s AI Anthology: How to live with smart machines? (pp. 55–78). Xenomoi.
### Assistant
Let me first try to access the slides, read the current note, and check the required knowledge files.
### User
[Request interrupted by user]