## 1. What Exactly Has Floridi Conceded?
The first task is to parse Floridi's concessions with care, because the counterargument's weight depends entirely on their scope. The paper contains several formulations that differ in what they attribute to LLMs, and these differences matter.
The narrowest concession is purely syntactic. Floridi notes that LLMs learn connectives and typical phrasings:
> "It also learns common patterns of explanation and argument, such as how 'because' often introduces an explanation, and that scientific questions are answered with specific explanatory forms." (Floridi et al., p. 8)
If this were all Floridi conceded -- that LLMs have learned that "because" precedes explanations and "therefore" precedes conclusions -- the counterargument would have very little to work with. Syntactic connectives are plainly decorative in the sense that matters: knowing where to place "therefore" tells you nothing about whether the inference it flags is any good.
But Floridi concedes substantially more than this. Consider the passage about what training data encodes:
> "When their output exhibits an apparent abductive quality -- often reinforced by interface design -- this effect is due to the model's training on human-generated texts that encode reasoning structures." (Floridi et al., Abstract)
"Reasoning structures" is doing considerable work here. Floridi does not say "reasoning vocabulary" or "reasoning-associated syntax." He says *structures* -- and the plural suggests he means something beyond the placement of connectives. The text elaborates:
> "LLMs have effectively absorbed patterns of human abductive reasoning as expressed in writing." (Floridi et al., p. 8)
"Patterns of human abductive reasoning" is a phrase that sits between two possible readings. On a deflationary reading, it means something like: the typical order in which humans present their reasoning in text -- hypothesis first, then evidence, then evaluation. On a more generous reading, it means something closer to: recurring inferential structures, including the relationship between evidence and hypothesis, the way competing explanations are weighed, the criteria by which one explanation is preferred over another.
Floridi's own examples suggest the generous reading is closer to his intent. When he discusses the car-not-starting example, he notes that the LLM "mimics that of a mechanic or knowledgeable friend performing IBE: listing hypotheses and selecting one as the most plausible" (Floridi et al., p. 10-11). The model has not merely learned to write "because" and "therefore" -- it has learned the *pattern* of listing competing hypotheses, weighing them, and selecting one. That is a dialectical template, not a syntactic regularity.
The broadest concession appears in the discussion of commonsense causal knowledge:
> "Essentially, the LLM's extensive training on language has endowed it with a vast repository of commonsense causal knowledge, although not explicitly structured. It 'knows' that slippery floors cause falls, that not eating causes hunger, that polls predict elections, and so on -- because it has processed countless expressions of these relations." (Floridi et al., p. 17)
And when he discusses abductive benchmarks:
> "LLMs perform remarkably well, often at a near-human level [on the Abductive Natural Language Inference challenge]. Such findings already suggest that LLMs, despite lacking explicit reasoning, recognise patterns that align with human explanatory preferences." (Floridi et al., p. 4-5)
"Patterns that align with human explanatory preferences" is a significant phrase. It implies that the model has absorbed not just structures but *evaluative dispositions* -- preferences about which explanations are better. This goes beyond syntax, beyond dialectical templates, and into something approaching inference-pattern recognition with built-in quality signals.
So what has Floridi conceded? I interpret the text as conceding at least the following:
1. **Syntactic patterns**: the placement and use of reasoning-associated vocabulary.
2. **Dialectical templates**: the structure of hypothesis-generation, evidence-weighing, and explanation-selection as they appear in text.
3. **Inference patterns**: recurring abductive structures, including the relationship between evidence and hypothesis.
4. **Evaluative alignment**: patterns that track human explanatory preferences -- that is, not just *how* explanations are structured, but something about *which ones succeed*.
What Floridi denies is that any of this constitutes genuine reasoning. The source text says explicitly:
> "We are not claiming that LLMs hold literal beliefs or follow Peirce's method of hypothesis internally. Instead, we argue that the output structure of LLMs often resembles that of an abductive reasoning process... and this resemblance is not random but systematic, resulting from training on human explanations." (Floridi et al., p. 15)
The resemblance is "not random but systematic." The question is whether "systematic resemblance to abductive reasoning, learned from training on human explanations" can do more philosophical work than Floridi thinks.
## 2. The Decorative vs. Constitutive Distinction
The counterargument kernel turns on whether the absorbed structures are *decorative* (surface presentation that could clothe any content, good or bad) or *constitutive* (the very thing that makes philosophical practice what it is). I want to develop three positions here honestly. %%this sentence is not the correct register for an analytic philosophy paper%%
**Position A: The structures are constitutive.**
The claim: in philosophy, the argumentative structures just *are* the method. There is no deeper layer of "substance" beneath them. When a philosopher makes a distinction to resolve an apparent tension, raises an objection, offers a repair, integrates a claim with background commitments -- she is *doing philosophy*. The doing just is the enactment of those structures under appropriate constraints. If an LLM produces text that enacts the same structures under the same constraints, and the result satisfies the discipline's publicly checkable standards, then the text is philosophy.
Bengson, Cuneo, and Shafer-Landau provide independent support for this reading. Their account of method is explicit that criteria serve a "dual role: they provide instructions for the construction of a theory, given the data, while also serving as standards by reference to which the merits of theories are evaluated" (Bengson et al., p. 77). The same norms govern building and assessing. And they insist that these criteria are not esoteric:
> "We endorse the method not because it makes a philosopher's job easy; indeed, it is quite demanding. Nor are we drawn to its constituent criteria because they revolutionize philosophical thinking; on the contrary, all of them are familiar from the way many philosophers go about their business." (Bengson et al., p. 107-108)
The criteria are familiar *from practice*. They are extracted from what philosophers actually do. And the "friendliness datum" -- the observation that implementing philosophical method involves advancing arguments, raising objections, offering replies, providing clarification, developing explanations, and displaying sensitivity to logic, science, and common sense (Bengson et al., p. 80) -- is telling. Bengson et al. write:
> "Whatever philosophical method is, it is something that is friendly to these activities. By this we mean that, in the paradigm case, implementing philosophical method involves engaging in such activities." (Bengson et al., p. 80)
The activities *are* the paradigm case of implementing the method. If a system produces text that enacts these activities -- that advances arguments, raises objections, offers replies, develops explanations -- and does so in a way that satisfies the criteria (accommodation, explanation, substantiation, integration), then on Bengson et al.'s own account, the text implements the method. Whether the system "understands" what it is doing is a question about the system's internal states, not about the method's being enacted.
This is the strongest form of the constitutive reading: method is enacted in the activities, the activities are textually manifest, and satisfaction of the criteria is publicly checkable.
**Position B: The structures are necessary but not sufficient (decorative+).**
Floridi would likely resist Position A by arguing that the structures are necessary but not sufficient. You need structures *plus* truth-directedness *plus* verification capacity *plus* grounding. On this view, the structures are not purely decorative -- they matter -- but they are like the formal properties of a legal argument: necessary for the argument to *count as* legal reasoning, but insufficient without genuine engagement with the law's substance.
This position has real force. Floridi's point about hallucination is precisely that the structures can be satisfied while the content is fabricated:
> "The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations." (Floridi et al., p. 8)
> "An explanation can be coherent and convincing (even optimal by IBE criteria) and yet still false. LLMs lack an epistemic compass to navigate that distinction." (Floridi et al., p. 20)
The decorative+ position says: structures get you the *form* of philosophy, but philosophy also requires that the form be *truth-directed* -- that the accommodation be of genuine data, that the explanations track real explanatory relations, that the substantiation involve actual epistemic support. An LLM might produce text that *looks like* it satisfies these criteria without actually satisfying them, because it has no way to distinguish genuine satisfaction from its appearance.
**Position C: Domain-dependent.**
I speculate that the most nuanced position is domain-dependent. The structures might be constitutive for some philosophy but not all. In formal philosophy -- philosophical logic, decision theory, the evaluation of argument forms -- the structures may be genuinely constitutive, because the domain's subject matter just is inferential and logical relations. When the model produces a valid counterexample to a schema, the counterexample either works or it does not, and this is checkable from the text. But in philosophy of perception, say, where claims need to be grounded in empirical findings about perceptual processing, the structures alone are insufficient -- you also need the structures to be hooked up to the right empirical content, and the model's ability to do this depends on the quality and currency of its training data in ways that go beyond structural competence.
This position has the advantage of precision but the disadvantage of giving up the generality the constitutive reading promises.
## 3. Floridi's Natural Counter-Move
Floridi's most natural response to the constitutive reading would be something like: "I conceded form. You are now claiming form is substance. But form without epistemic substance is *exactly* what I mean by 'abductive appearance.' You have not refuted my position -- you have restated it while changing the evaluative sign."
This is a powerful counter. Let me develop it in its strongest version.
Floridi could point to hallucination as the decisive test case. When an LLM produces a philosophically structured argument that relies on a fabricated source, the structures are all present -- hypothesis, evidence, weighing, conclusion -- but the output is worthless. The structures were satisfied *decoratively*. If structures can be satisfied decoratively in this case, what principled reason is there to think they are satisfied constitutively in the cases where the output happens to be good? The difference, Floridi would say, is not in the structures but in whether the content is true, and truth-tracking is exactly what the model lacks.
He could also invoke the "zeroth-order abduction" concept:
> "LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations." (Floridi et al., p. 8)
Zeroth-order abduction is abduction *without commitment* -- generating hypotheses without the epistemic backing that would make them genuine inferences. The constitutive reading, Floridi might say, confuses zeroth-order abduction (hypothesis generation) with first-order abduction (hypothesis generation plus evaluation plus commitment). Philosophy requires the latter.
Can this counter be resisted? I think it can, but it requires a specific move. The resistance depends on distinguishing two things:
(a) Whether the *model* has epistemic substance (truth-tracking, verification, commitment).
(b) Whether the *text* satisfies publicly checkable philosophical norms.
Floridi's counter works if the locus of evaluation is the model. But if the locus of evaluation is the text -- the artefact -- then the question becomes: does this text accommodate the data, explain the phenomena, substantiate its claims, integrate with background commitments? These are features of the text that competent readers can check. The hallucination case fails this test *precisely because* competent readers can identify the fabricated source. The structures are not satisfied decoratively in that case -- they are not satisfied at all, because the Substantiation Criterion (to use Bengson et al.'s term) fails when the supporting evidence is invented.
This is the move that makes the constitutive reading work: relocating evaluation from the producer to the artefact. But Floridi could respond that this just pushes the problem back -- who is doing the checking? If it is humans, then the model is not doing philosophy independently; it is producing drafts for human evaluation. If that is all the constitutive reading establishes, it is less interesting than it initially appears.
## 4. Developing What Floridi Left Undeveloped
The most striking passage in Floridi's paper is where he raises but does not develop the question of whether process matters:
> "In particular, if an AI can generate the same explanatory hypothesis a human would, does it matter that the process was different? From an epistemological standpoint, perhaps yes -- justification is significant -- but regarding the content of the hypothesis and our interpretation of it, maybe not." (Floridi et al., p. 12-13)
Floridi splits the question into justification (where process matters) and content (where it may not). But he develops neither side. Let me try.
**The best version of "maybe not" (content-focused):**
Philosophy, unlike empirical science, evaluates contributions primarily through text-internal features. When a referee assesses a philosophy paper, she asks: Is the argument valid? Are the premises defensible? Does the theory accommodate the data? Is the explanation illuminating? Are objections handled? These are all features of the text. The referee does not ask: "What was the author's subjective experience while writing?" or "Did the author have genuine understanding of the concepts?" She asks whether the *paper* meets the standards.
Williamson's account of philosophical methodology supports this. He argues that philosophy should use "a broadly abductive methodology" (Williamson, "Widening the Picture," p. 356), and that the criteria for evaluating philosophical theories include simplicity, fit with evidence, explanatory power, and elegance. These are all properties of theories, not of theorists. Williamson explicitly notes that the abductive method "can be applied in philosophy, whether or not it should be" (p. 356), and that abductive criteria work by ranking theories, not by inspecting the cognitive processes of those who propose them.
If philosophical evaluation is fundamentally about whether the artefact satisfies public criteria -- and if both Bengson et al. and Williamson provide independent reasons for thinking it is -- then the process by which the artefact was produced is irrelevant to its philosophical quality. This is the "maybe not" position fully developed: for the content of the hypothesis and our evaluation of it, provenance does not matter.
**The best version of "perhaps yes" (justification-focused):**
The justification-focused response says: even if the text looks good, the lack of genuine epistemic backing makes the output epistemically unsafe. A human philosopher who produces the same text has justified beliefs backing her claims; the model does not. This matters because justification is not just a property of the text -- it is a property of the epistemic agent's relationship to the text. A lucky guess that happens to be correct is not knowledge, and a text that happens to satisfy philosophical criteria but was produced without understanding is not genuine philosophy.
This position draws strength from the analogy with testimony. If someone tells you something true but has no justification for believing it, you do not gain knowledge from their testimony (on standard accounts). Similarly, if a model produces a philosophically adequate text but has no epistemic standing with respect to its claims, the text might not transmit philosophical understanding to readers in the way genuine philosophy does.
But I want to note a weakness in this position: it is not clear that philosophy *works* through testimony in this way. When I read a philosophy paper, I do not gain knowledge by trusting the author's authority. I gain understanding by following the argument. If the argument is valid, the premises defensible, the theory well-integrated -- I have the philosophical goods regardless of the author's epistemic states. This is what distinguishes philosophy from, say, news reporting, where the reporter's epistemic position matters because I cannot independently verify the claims.
## 5. Independent Motivation for the Constitutive Reading
Can the constitutive reading be motivated independently of the LLM debate? I think so, drawing on Bengson et al. and Williamson.
Bengson et al.'s account of method makes the constitutive reading natural. Their "friendliness datum" says that implementing philosophical method *just is* engaging in certain activities -- advancing arguments, raising objections, offering replies, developing explanations. If method is constituted by these activities, and these activities are textually manifest patterns, then the patterns are constitutive of the method. This follows from Bengson et al.'s own framework without any reference to LLMs.
Williamson's argument for abductive methodology in philosophy provides further support. He argues that philosophy should proceed by ranking theories according to abductive criteria -- simplicity, fit with evidence, explanatory power, non-ad-hocness. These criteria are *intrinsic to theories*, not to the cognitive processes of theorists. When Williamson describes the deductivist methodology as leading to "stalemate" because opponents simply reject premises as "question-begging" (Williamson, p. 364), his proposed remedy is explicitly abductive: rank theories by their theoretical virtues, rather than trying to deduce conclusions from self-evident premises. The virtues are properties of the theories themselves.
The independent argument, then, is: if philosophical method consists in the enactment of certain publicly checkable activities (Bengson et al.) evaluated by criteria intrinsic to theories rather than to theorists (Williamson), then the method is constituted by structures that are in principle learnable from texts and assessable in texts. The constitutive reading is not an ad hoc move in the LLM debate -- it follows from mainstream philosophy of philosophy.
## 6. The Role of Evaluative Data
A final consideration strengthens the constitutive reading. Floridi treats the training data as containing reasoning structures, but he underestimates what "structures" includes. Philosophical corpora do not just contain arguments; they contain *evaluations of arguments*. They contain referee reports, response papers, objection-and-reply sequences, editorial decisions, citation patterns. A model trained on this material has learned not just how arguments are structured, but which arguments the discipline treats as successful and which it treats as failures. It has absorbed what Bengson et al. call the criteria -- accommodation, explanation, substantiation, integration -- not as explicit rules, but as distributional regularities in how the philosophical community responds to philosophical claims.
Floridi's phrase "reasoning structures" underestimates the richness of what has been absorbed. The model has not just learned forms; it has learned which forms succeed and which fail, which moves are treated as adequate replies and which are dismissed, which explanations are accepted and which are challenged. This evaluative layer is part of the training data, and it makes "reasoning structures" considerably more substantial than Floridi's framing suggests.
Whether this is enough to close the gap between "abductive appearance" and genuine philosophical competence remains, I think, genuinely open. The strongest version of the constitutive reading says: if the model has absorbed the structures *and* the evaluative standards *and* can produce text that satisfies those standards as judged by competent readers, then saying it "merely appears" to do philosophy is a verbal move, not a philosophical objection. The strongest version of Floridi's position says: evaluative alignment learned from text is still just correlation, and correlation without epistemic substance is exactly the illusion he warned about.
I am genuinely uncertain which position is correct. The constitutive reading is more plausible than I initially expected, and it has independent support from Bengson et al. and Williamson. But Floridi's counter -- that form without truth-directedness is exactly what he means by appearance -- has real bite. The resolution may depend on whether philosophy's truth-directedness is itself something that can be partially captured by structural and evaluative patterns in text, or whether it requires something that no amount of textual pattern-learning can provide.
---
*Resta aperta la domanda se il confine tra apparenza e realta, nel caso specifico della filosofia, sia poi cosi netto come Floridi presuppone.*