# About
I'm a philosopher currently working on the philosophy of technology, AI, and aesthetics. In the past I have worked on perception and temporal experience. At the moment I'm a postdoctoral researcher on Enrico Terrone's ERC project, The Philosophy of Experiential Artifacts, at the University of Genoa.
My earlier research was on auditory perception. Against Hearing Sounds, argues that we don't hear sounds at all — we hear the material objects and events that produce sound waves, directly, without any sonic intermediary, making audition far more like vision than is usually supposed. Along the way I've defended related claims about what audition does acquaint us with: that to hear an event is to hear an object as extending through time, and that when sound waves reverberate we hear the empty space around us.
I've also written on temporal experience and agency. There I've argued that the common-sense belief that time passes is not grounded in perception, as is usually supposed, but in agentive experience: the feeling of being the source of one's own actions. That interest in agentive phenomenology carries over into my work on design, where I've used it to explain how an artifact can be beautiful in use — how the experience of acting with a thing, rather than just looking at it, can be aesthetically rewarding.
Most of my current work is on AI and Creativity, largely in collaboration with Enrico Terrone. We've argued that generative AI systems are best understood neither as agents nor as tools but as a new kind of artistic medium, one a user cultivates rather than controls — closer to a garden than a paintbrush. We've also developed an account of the aesthetics of engineering, on which the autonomous functioning of machines is a domain of aesthetic appreciation in its own right.
Two ongoing projects extend this work to large language models. The first argues that LLMs are capable of producing philosophy genuinely worth reading, and that the standard reasons for thinking otherwise — that there is no author behind the text, no inference, no consciousness — do not survive scrutiny. The second develops an aesthetics of LLMs themselves, on which models can be appreciated rather as we appreciate natural environments: by attending to the order in what they produce, guided by an understanding of the forces that shape it.
## United current section drafts — 2026-06-12
Source files: `section-1-draft.md`, `section-2-draft.md`, `section-3-draft.md`, `section-4-draft.md`. Old/superseded draft material below double separators was ignored.
# 1. The Challenge from Authorship — new iteration (v2)
In this section we address what we might call the *challenge from authorship*: the idea that philosophy is something that only persons, or at least minds, can produce. This view has not, to our knowledge, been explicitly defended in just this form, but it gives shape to an intuition that many philosophers may have: philosophy is a person-only domain. An imperfect comparison is with art. One might deny that an image generated by an AI system, at least in the familiar prompt-and-output cases, is an artwork, because no artist exercises the relevant kind of intentional control over its production.[^1] One might think, for similar reasons, that philosophy can only be done by people: no text produced by an LLM can be a work of philosophy, because no philosopher lies behind it.
The intuition draws support from how philosophy is taught and studied. Like the study of art, the study of philosophy is organised around individuals: undergraduates take courses on Kant's ethics or Lewis's metaphysics, and at more advanced levels there are specialists in, and conferences devoted to, the work of particular philosophers. The sciences are organised differently: a physics student is taught Newtonian mechanics from a current textbook, and the course loses nothing if Newton's own writing is never opened; a student of ethics is assigned the *Groundwork* itself, and a course that replaced it with a summary of its conclusions would be a worse course. In the sciences, that is, what a text contributes can be carried by other texts. In philosophy, the contribution and its original presentation are harder to prise apart, and this is some evidence that a philosophical work is bound to the activity of the particular person who produced it.
We will now try to make the intuition more precise, by considering how far Davies' *performance* theory of art transposes to philosophy. Davies writes:
> [T]he work – what the artist achieves – is the process eventuating in that product. Works themselves are neither structures nor objects simpliciter, nor are they contextualized structures or objects. They are, rather, intentionally guided generative performances that eventuate in contextualized structures or objects (or events, as we shall see) – performances completed by what I am terming a focus of appreciation. (2004, p. 97)
On Davies' view, when a painter paints a picture, the canvas is what we attend to, but it is not the work. The work is the artist's intentionally guided activity in producing the canvas; the canvas is the work's focus of appreciation, "the focus of our appreciative interest in the work" (2004, p. 150). Provenance, on this view, does more than supply context: facts about how the object came into being help determine what the work is and what is properly appreciated in it.
Davies supports the relocation of the work from the surface to the activity with cases in which two surfaces, or two texts, would look the same and yet differ as works, and the cases are of two kinds. In one kind there is no performance at all: an instance of the verbal structure of *Kubla Khan* might be generated by desert wind, or by the famous monkey at a typewriter. A theorist who identifies the poem with its verbal structure must then either count these as instances of Coleridge's work or explain why not (2004, p. 102). In the other kind there is a performance, but not the one the surface was taken to give access to: van Meegeren's *The Disciples at Emmaus* was presented as a newly discovered Vermeer, and what was appreciated under that description was not the achievement the canvas in fact issued from.[^2] Only the first kind bears on our question, since an LLM text would be a case of absent performance, not of misattributed performance. If Davies is right, the surface does not by itself settle the work.
Here is what the analogous proposal for philosophy would be. A philosophical text is not itself the philosophical work: the text is the product of a person's philosophising, and reading it is a way of engaging with that activity. The challenge this poses is constitutive: the activity is treated not as what causes a philosophical work to exist but as part of what the work is. If no one has philosophised, there is no work to which the text gives access, however the text reads — an LLM text would stand to philosophy as the wind-made *Kubla Khan* stands to poetry.
Should the transposition be accepted? We do not think it should. What entitles Davies to relocate the work into the performance is an evaluative fact about art: surfaces that look the same can differ in artistic value, as the forged Vermeer and a genuine one do, and provenance is part of what the difference consists in. The transposition therefore commits its defender to the corresponding claim about philosophy: that two texts containing the same argument could differ in philosophical merit. We can find no difference for the merit to consist in. If two texts contain the same argument, including the same inferential moves, the same considerations count for and against them: whether the argument is valid and whether the objections are answered are questions about the texts' contents, and two texts with the same contents receive the same answers. Their philosophical merit does not vary with the route by which the words came to be written.
The discipline's evaluative practice is built on the same denial. Journals strip author information from submissions before review because facts about authorship are treated as potential sources of distortion; if texts with the same contents could differ in merit, anonymising would discard evaluatively relevant information, and review would not be designed this way. The grounds for the judgement lie in the argument as presented, not in the history of its production. Nor is the author-centred teaching noted earlier in tension with this. That philosophy is taught through Kant rather than through summaries of Kant is a fact about where the discipline's contributions live, not about how they are evaluated: the *Groundwork* is assigned because reading it puts a student somewhere no summary has yet put one, which is a fact about the text.
Dellsén et al. (2024) hold that philosophical progress is "for-whom" rather than "by-whom": it consists in putting people in a position to increase their understanding, usually by making philosophical ideas publicly available (p. 679). On this account the discipline's success-conditions locate the contribution in the public text, and a view that locates the philosophy behind the text, in the process by which it came about, misplaces it. The public text is not a dispensable trace of philosophy; it is where the philosophical contribution becomes fully available.
A performance theorist can hold the line: no philosophising, no work, whatever the text contains. The position can be granted in full, because the thesis of this paper does not use the notion it restricts. Suppose the desert wind assembled not *Kubla Khan* but a sound argument against a familiar theory of perception. No one would deserve credit for it; there would be no achievement to admire, and no entry for anyone's bibliography. A reader who worked through it would nonetheless meet a thesis and the considerations advanced for it, and would be in a position to answer or extend it — the position a philosophical text puts its readers in when it is worth their time. Whether such a text is a *work* may then be reserved for texts with performances behind them; what cannot be reserved is the text's being worth reading, since everything that judgement answers to is on the page.
The challenge from authorship therefore fails, and the doubt it leaves standing is of a different kind. That a text came from an LLM cannot disqualify it; nothing so far shows that LLMs can produce such texts. If a parrot produced a sound sequence that was a sound argument, the argument would not be disqualified by its source — but parrots produce no arguments, because they lack the capacities arguing requires. The doubt about LLMs, in their current state, is of this kind: that they lack capacities that producing philosophy worth reading requires. Section 2 takes up the claim that they cannot perform the inference on which philosophical theorising runs; Section 3 the claims that they stand in no relation to the world and have no experience. A further challenge grants a worthwhile text and asks whose work it is, the model's or the prompting person's; it arises only if the capacity challenges fail, and we take it last.
[^1]: This is not to deny that systems of this kind can produce beautiful images; we return to image generation in Section 4.
[^2]: Han van Meegeren, the Dutch forger exposed in 1945; his *The Disciples at Emmaus* was authenticated as a Vermeer and celebrated before the forgery came to light.
---
# 2. The Challenge from Abduction — rewrite v2
Section 1 concluded that a text cannot be ruled out as philosophy worth reading simply because an LLM produced it. A doubt of a different kind remains: such systems, at least in their current form, may lack capacities that producing worthwhile philosophy requires. This section examines one such capacity challenge; the next examines two more.
Abduction, or inference to the best explanation, is reasoning from a body of evidence to the hypothesis that would best explain it. In a deductive argument the premises fix the conclusion: if all men are mortal and Socrates is a man, then Socrates is mortal, and there is no wriggle room. In abduction the evidence settles less. Now, imagine walking into your kitchen and finding the floor wet. What has happened? The wet floor does not determine the answer in the way the two premises determined Socrates' mortality: a burst pipe would have left the floor wet, and so would a spilled bucket. But, given that the window is open, the water is under the window, and it rained last night, rain coming through the window seems the most plausible answer.
To reason in this way, deciding what best explains a set of facts, is common in everyday life and in the sciences alike. A scientist chooses one theory over another when it explains the same data more simply: Copernicus's model of the solar system was preferred to Ptolemy's because it explained the observed planetary motions without the elaborate epicycles the older model required. Williamson argues that philosophy is continuous with the sciences, and that its theories should be chosen in the same way (2007; 2021, p. 351). In philosophy too there are data that a candidate theory must accommodate — intuitions about cases, and the phenomena of the domain itself — and there are rival theories that would each accommodate them at different costs. Choosing among the rivals is abductive: the theory to prefer is the one that would, if true, best explain the data, and the virtues by which the rivals are compared are explanatory virtues — a good philosophical theory is, in Williamson's words, "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated", and should "combine simplicity with strength" (2021, p. 354 %%check page%%). This conception of philosophy is widely held (Sider 2011; Paul 2012; Dellsén et al. 2024), though not universally (Bueno and Shalkowski 2020; Thomasson 2015), and we shall assume it in what follows: producing philosophy worth reading requires abduction.
If the capacity for abduction is what is required to produce worthwhile philosophy, we can ask whether LLMs possess it. Floridi et al. (2025) argue that they do not, describing what such models do instead as zeroth-order abduction:
> LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025, p. 9)
An LLM, on this account, has "a stochastic core and an abductive appearance" (2025, p. 3). The model is trained to predict which words are likely to follow which, and at each step it produces the continuation its training makes probable; it aims at the likely continuation, not at the truth. Floridi et al. explain the abductive appearance by what the training data have passed on: models, they write, have "absorbed patterns of human abductive reasoning as expressed in writing" (p. 10) — how explanations are typically phrased, which causes are typically offered for which effects. What is inherited, on their account, is the look of the reasoning, not the reasoning itself.
Floridi et al.'s position, then, is that nothing a model does is abduction of any kind. Explaining the wet kitchen floor involved two separable activities: you came up with candidate explanations — the burst pipe, the spilled bucket, the rain — and you settled which of them the open window and the position of the water favoured. Call the first *generating* and the second *weighing*. A model's text can exhibit both: presented with a scenario it will offer plausible candidates, and given candidates it will single one out, much as a human judge would. Floridi et al.'s claim is that the model performs neither activity.
Asked why a car might not start on a cold morning, a model replies that a weak battery is one possibility, since cold reduces a battery's efficiency; that thickened engine oil is another, since a cold engine is harder to turn over; and that, "[b]ased on your description, the battery is the most likely explanation" (2025, p. 11). Candidates are offered here, and one is singled out. On Floridi et al.'s reading, however, the offering is not generating: the model is not reasoning about causes from the user's case but reproducing the causes typically offered for this effect in the text it was trained on — weak batteries and thickened oil are what explanations of cold-morning car trouble standardly cite (p. 9). And the singling out is not weighing: the closing verdict is a conversational move the model has learned, reproduced because people end such explanations by naming the most likely cause%%this is weird and too specific%%, not a conclusion reached by setting one candidate against the other (p. 11); where an output marks a genuine point of difference between two hypotheses, that is something the model has seen stated, not something it has derived anew (p. 14). The look of each activity is on the page, and, on their account, neither activity occurred.
If producing philosophy worth reading depends on abduction, and LLMs do not perform abductive inference, then LLMs cannot produce philosophy worth reading.%%not so clear, depends on abduction is ambiguous which may be ok, but it ends sound like WE me and Enrico, are claiming that LLMs cannot produce philosophy worth reading because of these things.%% Floridi et al.'s support for the claim that models do not perform abductive inference is of two kinds.%%really unclear way of presenting this%% The first we have now seen: the model neither generates nor weighs, however much its text exhibits both. %%and you absolutely shouldn't be only now revealing that there is a second thing. yopu are still screqwing up the structure.%%The second concerns the model's relation to the world: its words are connected to no perception of anything, and a hypothesis, once produced, is never tested against the world (2025, pp. 7–9). What a system without experience of the world could know of it, and what philosophy it could ground, are questions we take up in the next section. Our concern here is the first kind of support.%%so weird to have this distinction made at this point in the section.%%
Salimi et al. (2026) survey empirical studies of abductive reasoning in language models, and their findings fit Floridi et al.'s picture. %%fit is very vague, and i think it has a knock on effect on the rest of the paragraph%% The studies test the two activities separately. To test weighing, a model is given a scenario together with several candidate explanations, and is scored on whether it picks the one human annotators judged best; to test generating, it must produce an explanation itself, with no candidates supplied. Models approach human accuracy at the first and do markedly worse at the second, and abduction overall is the form of reasoning at which they perform worst: median accuracy across the surveyed studies is roughly 43%, against 80% for deduction.[^1] That is the profile Floridi et al.'s account predicts: a model that has learned what explanations typically look like can recognise a plausible candidate when one is put in front of it, but it has nothing further to draw on when the explanation must be produced.
A system of this kind still has a use, and Floridi et al. say what it is: %%the first clause makes it sound like we believe this%%
> In essence, LLMs function like brainstorming assistants that toss out ideas without filtering for quality. After all, they work like statistical interfaces to an enormous amount of data accumulated for millennia by generations. A cautious human collaborator can sift through and assess them. (2025, p. 12)
The division of labour Floridi et al. propose leaves the weighing always with the person: the model supplies candidates, and assessing them remains the collaborator's work. If that is the most a model can do, nothing it produces would be philosophy worth reading in itself; it would be raw material for philosophy done by someone else. A list of unweighed candidates is no more worth reading than a bare pronouncement that direct realism is correct. %%compressed and so shit, i also think the idea could be conveyed better%%
The challenge needs a further step, however%%meta commentary as a crutch%%: from the fact that the model performs no inference to the conclusion that no good weighing can appear in anything it produces. %%this point could be made more clearly. see my publications%%Producing worthwhile philosophy depends on abduction, but the dependence runs through texts %%not clear because editorial vague language%%: theory choice in philosophy is conducted in the literature, in writing that sets a theory against its rivals and makes the case for preferring it, and a text is worth reading in virtue of what it makes available to its reader.%%example list as a crutch, put some CONTENT here instead%% Section 1, moreover, has already fixed where the grounds of such judgements lien%%not how i write%%— in the argument as presented, not in the history of its production. The absence of inference behind a text accordingly leaves open what the text contains.%%unclear%% Consider the car-battery reply again: it is not a set of ideas tossed out for a collaborator to sift, for it brings the cold morning to bear on each candidate, sets one against the other, and closes in favour of the battery. The sifting that the brainstorming picture reserves for the person is conducted on the page.n%%not how i write%%That nothing performed the comparison is conceded, and settles nothing about the reply. The complaint that the reply's comparison is hollow, its verdict unearned by the considerations cited, concerns a piece of text and is made good or refuted by reading it — and it will be true of some outputs and false of others, as it is of paragraphs written by people. What survives of the challenge is the claim that the weighing an LLM's text contains cannot be good. %%this paragraph is badly organised/structure, its length is a give away that it is a mess%%
Peter Lipton's account of inference to the best explanation says what goodness in such a weighing consists in. The best explanation, he argues, can be understood as the likeliest, the one most warranted by the total evidence, or as the loveliest, the one which, if correct, would provide the most understanding: "[l]ikeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). The two come apart: Newtonian mechanics is no longer the likeliest account of the observations it was built on, but it remains as lovely an explanation of them as it ever was (p. 60). Loveliness is the standard a philosophical text answers to. Whether the position a paper argues for is true may stay unsettled for a century%%surely it can never be settled, and also, why are you tyelling the reader this?%%, and a paper can repay attention though one rejects its conclusion%%not how i write%%; what its readers assess is whether the explanation it offers would, if true, give understanding, and give more of it than the rivals it considers. Because the assessment runs under "if correct", it does not wait on verification, and a reader can conduct it on the page. On Dellsén et al.'s account, philosophical progress consists in putting people in a position to increase their understanding (2024, p. 679); a lovely explanation puts its reader in exactly that position.
Loveliness is earned, on Lipton's analysis, in the comparison of rivals.%%I worry this is too blunt to be accurate or good.%% Explanation is contrastive: we explain why this rather than that, and to do so we must cite a difference between the two, something in the favoured case to which nothing in its rival corresponds (2004, ch. 3). In the kitchen, "rain rather than a burst pipe, because the window is open and the water lies under it" cites such a difference, since a burst pipe would have wet the floor by the pipe; "rain rather than a burst pipe, because the floor is very wet" has the same comparative shape and cites nothing that bears on the contrast, since the wet floor favours neither rival. Both sentences instantiate the form of a weighing, only the first contains one worth having, and no inspection of the form tells them apart. Distinguishing them requires understanding what is claimed and asking whether it decides between the candidates, which is what a reader of any philosophy paper does, and all that assessing a model's text asks of its reader. %%i don't know what is being attempted in this paragraph. %%
Nor is there a rule that would spare the reader this work. Lipton allows that our grasp of what makes one explanation lovelier than another is weak (p. 61), and that the standards are carried, in part, by past explanations that serve as exemplars and by prevailing styles of reasoning (p. 139). A system that had absorbed only the format of explanation could not on these terms be credited with good weighing — but possession of the format was never what earned a human philosopher's comparison its credit either, %%really deeply unclear%%and the bar that separates the two kitchen sentences separates human paragraphs and machine paragraphs alike. Where reference answers give out, the machine-learning literature itself assesses generated explanations in just this way, scoring them for consistency, parsimony and coherence as features of the output (Dalal et al. 2024; He et al. 2025), partial proxies for loveliness validated against human judgement rather than a procedure for it.
A model trained only to continue text respects constraints that were never stated for it, and Wolfram (2023) assembles the cases. A model trained on English respects English syntax, although no grammar was supplied to it: the syntax is carried by the writing itself, in which well-formed sentences predominate, and a system fitted to continue the writing comes to respect what the writing respects. Its sentences are, for the most part, meaningful rather than merely grammatical, although here there was no rule available even in principle, since no stated theory exists of what makes a sentence meaningful (p. 66). And a model can be expected to complete a syllogistic pattern correctly: not because anything was derived, but because the form pervades the writing, so that the continuation the corpus makes likely is the one the logic requires. In each case a structure is present in the output while the capacity that ordinarily produces that structure — knowing the grammar, grasping the meaning, performing the deduction — is nowhere in the system. The absence of a rule for loveliness, which left the reader's judgement informal, is therefore no obstacle on the production side. A system that wrote by applying stated rules would be halted exactly where no rule exists; but these systems were never given stated rules for anything, and what they acquire, they acquire from exemplars — which are, on Lipton's account, precisely where the standards of loveliness live. The precedent is narrower than the cases suggest, since a syllogism has a single correct completion and an abductive comparison does not: what carries over is not determinacy of output but the weaker point, and the only one needed, that a structure can be present in a text without the capacity that ordinarily produces it standing behind the text.
The corpus such models are trained on is general — most of it is not philosophy — but it contains the philosophical literature, and that literature supplies exemplars of the structure at issue. A philosophy paper is built as a displayed comparison: a position is stated, set against rivals, and defended through the objections that are taken to decide between them, so that few kinds of writing exhibit comparative structure so regularly. Wolfram's own cases stop at the sentence, and the extension past the sentence is ours, but it asks for nothing different in kind. His observations concern regularities in text rather than grammar in particular; the well-formed sentence is the standing case of local, word-by-word tendencies issuing reliably in a globally coherent nested structure; and an argument that states a candidate, sets out its rivals, and locates the difference that decides between them is a larger structure of the same nested kind — as much a recurring regularity of the writing as syntax is. That models produce coherent discourse across many pages, and not well-formed sentences only, indicates that what is absorbed does not stop at the sentence.
It may be said that this redescribes the statistics and nothing more: the model reproduces the regularities of its training text, and reproducing regularities is not weighing. Lipton met an objection of the same shape in another setting. Bayesian probability theory, it had been suggested, gives the mechanics of belief revision and so leaves explanatory considerations no work to do; he compares the suggestion to arguing that thinking about technique cannot improve one's squash game because the flight of the ball is governed by the laws of mechanics (2004, p. 108). A true description of the mechanism does not displace a true description of what is produced. And here the mechanism is the one Floridi et al. themselves describe: the patterns absorbed from writing, noted above, are patterns of reasoning as expressed in writing. The texts the model absorbed do not contain the phrasing of explanations detached from their organisation — which considerations are brought to bear on which rivals, and what is taken to decide between them, are in the writing — and a system that learns to continue the writing learns them with it. The look of the reasoning was never separable from the organisation that makes reasoning assessable on a page.
It may be objected that syntax is one thing and inference to the best explanation another: whatever structure next-word prediction carries, a system of this kind is too shallow for abduction, and the survey's figures read like confirmation. Wolfram's account, however, draws the division elsewhere: his toy network fails to balance long sequences of parentheses, a task that demands explicit counting with no shortcut, while managing whatever a person can judge at a glance, and tasks of the first kind it is, in his phrase, "too computationally shallow" to perform reliably (2023, p. 65). The line such systems fail at falls between exact procedure and holistic judgement, not between the simple and the sophisticated — and weighing, on Lipton's account, sits with judgement: no rule runs from the evidence to the loveliest explanation, and the assessment is informal, comparative, and a matter of degree. Read with that line in hand, the survey's record divides against the account it seemed to confirm. A model that can recognise explanations but has nothing to draw on in producing one should fail wherever production is demanded; instead the collapse is concentrated where abduction has been recast as the exact recovery of a single canonical missing premise under formal constraint, with the strongest model reaching 21.5% on the hardest such benchmark and most models scoring near zero, while on open-ended tasks, where the output is judged as an explanation, the strongest models' validity exceeds 90%.[^2] Failure tracks the demand for exact recovery — the parenthesis side of the line — and philosophical abduction does not live on that side.
The figures also record one way of using a model. Salimi et al. run every benchmark with "a single task-specific direct instruction template held fixed across models" (2026): one prompt, one pass, the answer scored. The same survey catalogues methods built to elicit better abduction — prompts that separate reading the observation, producing candidates, and comparing them; pipelines in which a generated explanation is criticised and revised over several passes — yet none of this reaches the headline numbers, which measure what a model does when handed a test and no method. A benchmark asks whether the model passes; it does not ask how the capacity under test is best drawn out, and the claim defended here concerns what such systems can produce, not their pass rate under one fixed condition. What the person eliciting good work from a model contributes, and whose philosophy the result then is, are the business of Section 4.
None of this returns to the model any capacity Floridi et al. deny it. The model infers nothing, weighs nothing, and tests nothing; what it produces is text, and the text can contain what its producer never did — a candidate stated, the live rivals organised, the difference that decides between them located. Whether a given text does this, and does it well, is settled by the reading any philosophy paper receives, under the same standard and no other. Floridi et al. come close to saying so themselves: asked whether anything turns on the process being different when the hypothesis produced is the same, they answer that for justification it perhaps does, "but regarding the content of the hypothesis and our interpretation of it, maybe not" (2025, p. 13). A good weighing of the positions a literature already contains is not yet a distinction that literature lacks, and whether a model can supply the second is among the questions Section 4 takes up. The model's relation to the world, meanwhile, remains as Floridi et al. describe it: a system that weighs nothing also perceives nothing, and what philosophy a system without experience of the world could ground is the subject of the next section.
[^1]: These benchmarks operationalise abduction in commonsense and formal domains with crowd-labelled or mechanically checkable answers, where philosophy has no answer key, and several score generated explanations against human-written references — a comparison nothing in this paper relies on. Performance also drops under small variations to a problem (Mirzadeh et al. 2025), and Floridi et al. note that beyond the training distribution "the facade can crack" (2025, p. 10); the paper's claim is a capacity claim — that such texts can be produced — and is untouched by variation in how reliably they are.
[^2]: Salimi et al.'s benchmark suite separates formally constrained missing-premise completion (ProofWriter, AbductionRules, NeuLR) from open-text explanation tasks (ART, e-CARE, UNcommonsense); the figures are from their Tables 3–6. They observe that exact-match scoring is demanding precisely where several different missing facts would each be plausible — the abductive situation — and that target structure and the size of the hypothesis space shape difficulty at least as much as subject matter does.
---
# 3. The Challenge of Connecting to the World and the Challenge from Experience — v4
In this section we address two challenges. Both are claims that LLMs, in their current state, lack something that producing worthwhile philosophy requires. The first is the challenge of connecting to the world: the model perceives nothing and tests nothing, and a system so placed can neither come by philosophy's starting points nor check its results. The second is the challenge from experience. Few would say that current LLMs are conscious, and we assume here that they are not; the challenge is that experience does work in philosophy — supplying the starting points of some philosophical reasoning, and constituting the subject matter of some philosophical inquiry — that a system without experience cannot do. We treat the two together because one account of how philosophy is related to the world answers both, and we take the challenge of connecting to the world first.
That challenge was deferred from the previous section. Floridi et al.'s second kind of support for the claim that models do not perform abductive inference was that the model stands in no epistemic relation to the world: its words are connected to no perception of anything, and a hypothesis, once produced, is never tested against the world (2025, pp. 7–9). These are two claims: the first concerns where a producer's materials come from — the model has nothing of its own to draw on in saying how things are — and the second concerns what happens to what it produces, since nothing is checked against how things are.
Zahavy (2026) gives the challenge from experience its fullest recent statement. Current models, he argues, are structurally incapable of the jump from sense experience to new axioms, and his case study is the thought experiment that gave Einstein the equivalence principle:
> Einstein's variation required inventing new axioms based on a physical intuition that did not yet exist in the mathematics. He envisioned a physicist inside an elevator being uniformly accelerated through deep space [...]. Inside this enclosure, the sensory experience reveals a specific pattern: when objects are released, the floor rushes up to meet them. To the physicist, the objects appear to fall with identical acceleration, regardless of composition. Thus, the simulation here was not a permutation of symbols, but a manipulation of perceptual experience. (Zahavy 2026, §5)
Einstein imagines a set of circumstances and attends to what would be experienced within them: everything released inside the elevator appears to fall with identical acceleration. On Zahavy's reconstruction the simulation supplies an observation, and an abductive step supplies the axiom — the simulated experience of acceleration was indistinguishable from the remembered experience of gravity, and Einstein concluded that the two are one phenomenon.[^2] Zahavy argues that LLMs have no access to the observation. A model can produce descriptions of elevators and of weightlessness, both of which are present in its corpus, but it has not undergone either, and a discovery whose premises are fixed by simulated experience is unavailable to a system that does not simulate experience.
We might think that philosophical thought experiments depend upon experience in the same way. Does Mary, released from her black-and-white room knowing every physical fact about colour vision, learn something when she first sees red (Jackson 1982)? Settling the question requires considering what the experience is like, and the argument proceeds from the verdict; a starting point of that kind is one a system that has never experienced anything appears unable to supply. The same absence appears to bar a model from philosophy whose subject matter is experience itself: work that interrogates what it is like to see red (Harman 1990), to feel anger (Goldie 2000), or to have an intuition take hold (Chudnoff 2011).[^3]
What either challenge shows depends on how philosophy is related to the world, and here the discipline differs from the sciences. Pigliucci (2017) offers an account on which philosophy is constrained by the world without investigating it in the way the natural sciences do:
> This means that the basic parameters that philosophers use as their inputs, the starting points of their philosophizing, their equivalent of axioms in mathematics and assumptions in logic (or rules in chess) are *empirical* data about the world. This data comes from both everyday experience [...] and of course increasingly from the world of science itself. Philosophy, I maintain, is in the business of exploring the sort of conceptually evoked spaces that Smolin is talking about, where the evocation is the result of whatever starting assumptions are made by individual philosophers working within a particular field and, crucially, of the constraints that are imposed by our best understanding of how the world actually is. (2017, pp. 79–80)
*Evocation* is a term Pigliucci takes from Smolin, for truths that are neither discovered, in the sense of corresponding to mind-independent states of affairs, nor invented, in the sense of being arbitrary constructs, and his example is chess: when the rules of a game are codified, a whole bundle of facts about it becomes demonstrable — objective facts, in that anyone who can demonstrate one demonstrates the same fact as anyone else — although chess did not exist before its rules were written down (Unger and Smolin 2015, p. 423; quoted at Pigliucci 2017, p. 78). Pigliucci's proposal is that philosophy ascertains evoked truths in this sense, with an addition that separates it from mathematics and chess alike: its starting points are constrained by how the world actually is. The addition is also what separates philosophy from fiction, on his account. A novelist's worlds are invented rather than evoked — nothing about them is rigid, since even the constraints the novelist adopts could have been otherwise — whereas philosophy "is in the business of doing empirically informed evoking, not inventing", so that its objects of study have rigid properties (2017, p. 80). Thought experiments fall under the same picture: when philosophers explore imagined scenarios, they do so "with an interest in figuring things out as far as this world is concerned" (2017, p. 80). A thought experiment, so understood, articulates a starting point, experiential or empirical, and the philosophical work proceeds within the structure the articulation evokes.
Philosophy's starting points are empirical, then, and a system that perceives nothing has no way of its own to them. But the data Pigliucci describes comes from everyday experience and, increasingly, from science, and the scientific kind reaches working philosophers in one form only: articulated. A philosopher of physics works from published results, not from laboratory apparatus, and nobody holds that one must have run the experiments oneself to philosophise about what they show. Much the same holds for the everyday kind: what the discipline retains of ordinary experience, it retains as the literature's accumulated descriptions of how things seem. For the worldly materials philosophy actually uses, written access is the profession's normal condition rather than a deficiency, and a corpus is such access. On this point the model stands where every philosopher already stands with respect to nearly all of the empirical data they use.
The other half of the connection challenge was that the model never tests what it produces against the world. Testing of that kind is not what philosophical theses await. The elevator and Mary differ in this respect. Einstein's evoked structure yielded a hypothesis whose status was then a matter for experiment — the equivalence principle was confirmed by measurement, and would have been discarded otherwise — though, as Zahavy notes, even its formulation was "independent of immediate external verification" (2026, §5). In Mary's case there is no measurement to await. The evoked structure is itself the object of inquiry: the question is what the landscape contains, and because what is evoked has rigid properties, there are facts of the matter about its contents, which anyone can demonstrate and no one can choose. Checking a philosophical thesis is demonstration of this kind — establishing what does and does not hold within the structure the starting point evokes — and it is conducted in the literature, in the assessment the previous section described. A producer that cannot check a hypothesis against the world accordingly lacks nothing that a philosophical text requires of its producer; the checking philosophy uses is exploration of the evoked landscape, and exploration of that kind is done in writing.
The challenge from experience concerns philosophy that starts from, or is about, what something is like, where access through text can seem not to be access of the relevant kind at all. Yet no competent discussant of the knowledge argument has undergone Mary's transition. What every discussant works on is the case as articulated: a description, in public language, of an experiential situation and of what that situation is taken to establish. Responses to Jackson address this articulated structure — Lewis's reply modifies what is taken to follow from Mary's situation, not what her situation is like from the inside. Phenomenology enters philosophical argument in articulated form generally: the experience a description draws on need not be undergone by the people working on the description, which is how philosophers come to work on the experiences of the blind and of non-human animals without first-hand access to either. LLMs have no raw phenomenology of their own, but no text corpus contains raw phenomenology either. What a corpus contains is articulated phenomenology, and the fitting to a literature that put displayed comparison into a model's text in the previous section puts the literature's articulations of experience there on the same terms.
Merleau-Ponty's discussion of self-touch turns on phenomenological *discovery*. When one fingertip touches another, one finger plays toucher and the other touched; the roles can reverse, but not simultaneously, so that at any instant the body is split between touching and touched. Suppose the asymmetry was first identified by Merleau-Ponty himself, by sustained attention to his own embodied experience. The asymmetry would then be a starting point out of reach of any model not trained on Merleau-Ponty or his interlocutors — one the model could not have produced for itself, because producing it required a body and attention to what having one is like. Nothing in this prevents a model from working philosophically on the description once articulated, but it concedes a route: some articulations are originated through first-person attention, and a model has no experience to attend to.
The concession concerns how articulations come into existence, not what they are good for once they exist. First-person attention is one way an articulation comes into being; it is not what gives an articulation its philosophical use, which on Pigliucci's account lies in what it evokes. An articulation can also be arrived at by working from the articulations a corpus already contains, generating new ones by extension and recombination. Whether a candidate articulation succeeds — whether what it evokes sustains philosophical work — is assessed as the previous section's texts were assessed: publicly, on the page, under the standard any philosophical text faces, whatever its origin. Whether current models can in fact deliver by this second route — an articulation the literature lacks, a toucher-and-touched arrived at without a body — is the question of novelty, and Section 4 takes it up.
Both challenges move from an absence in the producer to a verdict on the product, and both fail at that step. The model perceives nothing and undergoes nothing, and one route to new articulations is closed to it accordingly. But philosophy's empirical starting points reach its practitioners in articulated form; its theses are checked by demonstration within the structures they evoke rather than by inspection of the world; and its phenomenology is used in argument as articulated content, on which a system trained on the literature can work as the literature's human discussants do. Abduction entered philosophy, in the previous section, as a weighing displayed in text; the world and experience enter it as content articulated in text. What a model produces in this territory faces the reading any philosophy paper receives, and nothing about its production settles in advance how that reading must go.
[^2]: Zahavy, following Magnani, calls the process *manipulative abduction*: hypothesis generation through the manipulation of a model — here a simulated experience — rather than of symbols (Magnani et al. 2009; Zahavy 2026, §5). It is abduction in the previous section's sense: the equivalence principle is inferred as the best explanation of the simulated observation, the simulation supplying an explanandum that no search over existing text would have produced. What experience contributes, on this picture, is not the inference but its starting point.
[^3]: The materials need not be sensory: the feeling of understanding something is sometimes used to motivate the claim that thought itself has a phenomenology (Pitt 2004).
---
# 4. The Challenge from Observation — v2
Sections 2 and 3 argued that a model which performs no abductive inference, connects to nothing beyond text, and has no experience can nonetheless produce text bearing the properties those capacities ordinarily produce. Anyone can put the conclusion to an immediate test. Ask a current model a philosophical question — what is the correct theory of consciousness? is there a meaning to life? — and what comes back is not philosophy worth reading. At best it is a bland and hedging summary of the positions in the field; ask for the meaning of life and there is a fair chance of being told that it is forty-two.[^1]
The observation supports a final challenge, and it is a challenge to this paper as a whole. Three sections have argued that nothing about a model's provenance or its missing capacities rules its text out from producing philosophy worth reading. Yet the models are here, the questions have been put to them, and the philosophy has not appeared. If the preceding arguments were sound, one would expect at least some of what comes back to be worth a philosopher's attention, and almost none of it is. Something, the challenge concludes, must be wrong: whatever the arguments said, the observed record is the record of systems that cannot produce worthwhile philosophy.
The observation is accurate, and it reports less than it seems to: it reports what models produce under one use — a bare question, put once, answered in one pass. How these systems are built explains why that use yields what it does. A model is first fitted to a vast general corpus and trained to continue text, so its response to a bare philosophical question is the likely continuation of such a question in writing at large, and the likely continuation of "what is the meaning of life?" in a general corpus is not an analytic tract. It is the sort of text that follows the question at large: a survey of views, a consoling generality, a joke. The model is then further shaped to converse as a helpful assistant, and the shaping presses the same way, since a person employed to be helpful to all comers would not answer the question with a tract either. The survey is not a ceiling the systems have hit; it is the likely continuation of exactly what was given them.
The use that generates the observation treats the model as an oracle: a system whose answers are its measure, so that asking is all the eliciting there is.[^2] The empirical record tells against the assumption. The survey of abductive benchmarks discussed in Section 2 runs every test with a single fixed instruction and scores the answer, while cataloguing, in the same pages, methods that alter what models produce — prompts that separate the stages of a task, pipelines in which an answer is criticised and revised over several passes (Salimi et al. 2026). What a model returns depends on what it is given, and the observation samples one point in that space, the bare question. It therefore cannot discriminate between the two hypotheses at issue — that the capacity defended in the preceding sections is absent, and that it has not been elicited. Both predict the observed record, and an argument against this paper needs the first; the observation supports it no better than the second.
The challenge has a natural escalation. If philosophy worth reading comes out of these systems only when a philosopher directs the process — supplies the framing, sets the constraints, presses for development — then the philosophy, it will be said, is the philosopher's. The model is an instrument in the production, as a typewriter is, and crediting it with the result is crediting the dummy with the ventriloquism. Section 1's challenge held that a model's text is not philosophy tout court; what stands here is narrower, that the philosophy in such a text is not the model's.
A prompt articulates a starting point, as a thought experiment does. A prompt that sets out a position and the rivals it must beat stands to the model as Jackson's two paragraphs stand to the profession: a starting point handed over for development. What an articulated starting point does, on the account already in place, is evoke a structure with rigid properties — there are facts about what holds within it, demonstrable by anyone and chosen by no one, and they outrun whatever has been stated, just as the facts about chess outran the rules the moment the rules were written down. Most of what a starting point evokes, no one has ever said.
What the model contributes is the development, and the mechanics are the ones Section 2 drew from Wolfram: a model produces a reasonable continuation of the text it has been given, where what counts as reasonable is relative to the corpus it was fitted to (2023). A prompt is part of the text the model has been given. An articulated starting point therefore changes what there is to continue — the reasonable continuation of a stated position under stated constraints is not the reasonable continuation of a bare question — and the model makes use of what the prompt states in everything that follows: tell one of these systems something once, Wolfram observes, and it is used thereafter (2023). The continuation that results states consequences of the starting point that the starting point does not state. Section 2 said what it is for such a text to go well — the comparison it displays cites differences that bear, and would, if correct, give understanding — and whether a given continuation goes well is read off the continuation.
Nothing in this makes the development a transcription. An evoked structure contains more than any text states: the rules of chess settle every fact about chess, and do not settle which theorems get written down, in what order, or to what depth, so that two writers working from the same rules produce different books, both correct, neither dictated by the rules. The mechanics mirror the structure, since the same prompt, run twice, yields different continuations (Wolfram 2023). The starting point underdetermines the development, and the gap between them is where the model's contribution lies: were there one text the prompt fixed, the output would transcribe what the person had already settled, and the instrument description would be true. The gap also leaves room for error. A development can state what does not hold in the evoked structure — a chess writer can publish a false theorem, a philosopher can misdraw the consequences of their own thought experiment, and a model can do both, along with its characteristic failure of stating fluently what nothing supports. The errors are found on the page. And an error is attributable only to a developer: no one blames the rules of chess for a false theorem, and no one's typewriter has ever made a mistake of content. Three contributions, then, and three owners: the articulated starting point is the person's; the structure it evokes, and the facts that hold there, are no one's; the text that develops them is the model's.
Much in the instrument picture is true. The person writes the prompt and the prompt is authored; the person chooses which continuations to pursue and when to stop; without the person, there is the survey. What the picture adds to these truths is a description of the model — a device, like the typewriter, that fixes only what its user has already settled — and the description is what the account above denies. Every word of the novel was the author's before the typewriter touched it; the consequences a model's text states were nobody's before the text stated them. We have pressed the same point against the same picture elsewhere: an image generator returns pictures that no one put into it, and a tool of which that is true is not a tool like a brush (Young and Terrone 2025). What the user of a typewriter settles is the text; what the writer of a prompt settles is a starting point.
The account invites an obvious enrichment of the prompt. State the position, name the rivals, list the objections and the lines along which they are to be met, and at some point, it will be said, the prompt contains the philosophy and the model is expanding what the person wrote — so that where a model's output is good, one should suspect a prompt rich enough to have done the work. But enriching a prompt enlarges the starting point without converting it into the development. A game with more rules is a bigger game, not a book of its theorems, and however much the prompt states, the consequences the output draws were not among the statements. There is a genuine limiting case — a prompt that states the comparison and the verdict, so that the continuation only rephrases — and it is identified the way everything in this paper is identified: set the output against the prompt and ask what the text states that the prompt did not. A text that states nothing beyond its prompt is a paraphrase, and owed to the person; a text that states what the prompt left unstated is a development, and the unstated part is not the person's. Which of the two a given output is, is settled by reading them together.
[^1]: While preparing this paper we asked GPT-5.5 for a detailed overview of the positions an analytic philosopher might take on the meaning of life. What came back was a competent, hedged survey of the field; what did not come back was an argument for any position in it. %%add date of test%%
[^2]: That these systems are mischaracterised as oracles — with the corollary that no benchmark of single-pass answers should be expected to probe the upper limits of what they can produce — has been argued from inside the practitioner literature (Janus 2022).
*Draft note: the section currently ends at the rich-prompt reply; the close and the novelty question (Section 3's hand-off) are deliberately unwritten pending design. Citation flags: Janus 2022 is a pseudonymous LessWrong post — confirm citation practice; Williamson absent from this section by design.*