# please find the most recently added pdf to my learning folder
## Retrieval Notes
- Session id: `3c47d30f-b4ed-490e-ae4b-d88f8feb3a61`
- Last activity: `2026-01-26T09:37:26.169Z`
## My Notes
<!-- Add your notes here. This section is preserved across syncs. -->
## Conversation
### User
please find the most recently added pdf to my learning folder
### Assistant
The most recently added PDF to your Learning folder is:
**What Kind of Reasoning (if any) is an LLM actually doing? On the Stochastic Nature and Abductive Appearance of Large Language Models** by Luciano Floridi, Jessica Morley, Claudio Novelli, and David Watson
Added today (Jan 26) at 10:21 — 346KB, so a journal article length.
Floridi is always worth reading on AI philosophy. Want me to extract and summarize it, or create a reading note?
*Sembra un articolo pertinente al tuo lavoro sull'estetica generativa dell'IA.*
### User
I believe there is some information in your config files as to best practicies when talking to me about academic texts. can you find it for me and repeat verbatim
### Assistant
Yes — from `~/.claude/CLAUDE.md` under **Behavioral Requirements**, here are the two relevant entries verbatim:
> **Source quotation**: When discussing PDFs, books, or academic sources, quote extensively and directly from the text. Use block quotes (>) for passages of 2+ sentences. Include page numbers or line numbers when available. The user needs to see the author's own words to evaluate interpretations. Don't paraphrase when you can quote. Err on the side of quoting too much rather than too little.
> **Epistemic discipline**: Always distinguish what a source says from your interpretation from your speculation. Use phrases like "the text says," "I interpret this as," "I'm speculating that." If you don't have a quotation to back a claim, either find one or flag the claim as your own thinking. Never present interpretation as textual fact. This applies equally to conversation and to notes I help write or draft.
So: lots of direct quotes, block-quoted when substantial, with page numbers when possible — and always flagging what's text vs. interpretation vs. speculation.
*Queste regole servono a mantenermi onesto quando parlo dei testi che leggi.*
### User
ok great. with this in mind, please give me a thorough, section by section summary of the paper.
### Assistant
Here's a thorough section-by-section summary with extensive quotations.
---
## Paper Overview
**Authors:** Luciano Floridi, Jessica Morley, Claudio Novelli (Yale/Bologna), David Watson (King's College London)
**Central Thesis:** LLMs have a "stochastic core" but produce outputs with an "abductive appearance." The text argues:
> Such LLMs generate text based on learned associations rather than performing abductive inferences. When their output exhibits an apparent abductive quality – often reinforced by interface design – this effect is due to the model's training on human-generated texts that encode reasoning structures. (Abstract)
---
## Section 1: Introduction
The paper opens by noting that LLMs "produce fluent language and often seem to reason, explain, and converse in a human-like manner" (p.1). This prompts the guiding question:
> What kind of reasoning (if any) is an LLM truly undertaking? Is it genuinely following logical rules or scientific inference methods, or is it doing something fundamentally different that merely appears to be (human) reasoning? (p.2)
The authors' main argument:
> LLMs occupy a conceptual space "between" traditional stochastic processes and human-like abductive reasoning. On the one hand, their internal processes are entirely stochastic: during training, they gather statistical correlations from text, and during generation, they produce words based on learned probability distributions. They lack explicit representations of meaning, everyday relevance, truth values, or causality as a reasoning agent would. (p.2)
Yet:
> Their outputs appear to share a phenomenological similarity to human reasoning. This effect is deliberately achieved through interface design, which encourages users to interpret outputs as explanations, commonsense reasoning, or analogies, but it also relates to the abductive patterns present in the data used to train the models. The result is a compelling illusion of genuine and structured inferential reasoning. (pp.2-3)
**Scope note:** The paper addresses "current, mainstream Large Language Models (LLMs) based on the token-completion paradigm" (p.1), acknowledging alternatives like Byte-Level Models and neurosymbolic systems exist.
---
## Section 2: Abduction and Inference to the Best Explanation
This section establishes the philosophical terminology. On Peirce's *abduction*:
> Peirce coined the term "abduction" to describe inference from effect to hypothesised cause. In a classic example, coming home to find the lawn wet, you might abduce that it rained earlier. This is not certain (someone might have run a sprinkler), but it provides a plausible explanation. Abduction thus contrasts with *deduction* (which reasons forward, in this case from cause to effect with certainty) and with *induction* (which generalises from many wet-lawn observations to a potentially probabilistic rule). (p.3)
On Harman's *Inference to the Best Explanation* (IBE):
> IBE can be understood as a form of abduction that adds a comparative evaluation step: multiple candidates are generated, then weighed by criteria such as simplicity, coherence with background knowledge, scope of explanation, and so on. The "best" explanation is then inferred as the most likely to be true. (p.3)
The authors distinguish *weak* and *strong* abduction (following Calzavarini & Cevolani 2022):
> Weak abduction—hypothesis generation without strong commitment—and strong abduction—inferring the most probable or best hypothesis. Weak abduction involves constructing a plausible story from the facts. Strong abduction entails choosing the best explanation among alternatives, which aligns more closely with IBE proper and may require comparative judgment or additional evidence. (p.4)
They note that LLMs "seem to perform at least weak abduction" and "can even seem to carry out a form of strong abduction when all candidate hypotheses are explicitly provided" (p.4), performing well on benchmarks like the Abductive Natural Language Inference challenge.
---
## Section 3: Probability, Statistics, and Stochasticity
This section clarifies technical terms. On stochastic processes:
> A data-generating process that includes random variables and/or probabilistic transition rules is said to be "stochastic". A stochastic process, such as a coin toss, is inherently random—though not necessarily in an unconstrained way. (p.6)
The authors note that philosophers debate whether IBE reduces to Bayesian reasoning (Lipton, Poston, Dellsen) or is distinct from it (Douven). They remain agnostic but observe:
> Probabilistic inference often aligns with abductive reasoning in scientific discovery and everyday thinking. Reichenbach (1938) and subsequent philosophers of science described inference as comprising two parts: the *context of discovery*, where abduction or IBE generates hypotheses; and the *context of justification*, where we test those hypotheses, almost always via statistical inference. (p.6)
Key claim about LLMs:
> Interestingly, LLMs seem to perform only the first part. They generate candidates (explanations, answers) but do not genuinely validate them against reality (unless they are specifically augmented by other systems, which only proves the point). They aim to model the conditional distribution of tokens in text, not to evaluate truth. (pp.6-7)
---
## Section 4: LLMs as Stochastic Engines of Text
This is the paper's core analytical section. On how LLMs work:
> During training, an LLM processes enormous amounts of text and optimises a model (usually a neural network transformer) to predict the next token (word or sub-word) based on the preceding context. The result is essentially a complex probability distribution: for any particular sequence of tokens/words, the model can assign likelihoods to potential continuations. (p.7)
The famous Shanahan illustration:
> When we prompt an LLM with a question like "Who was the first person to walk on the Moon?", we are not directly accessing a knowledge base or reasoning about the Moon landing. In reality, we are asking: given the statistical distribution of words in its training data, what is the most likely continuation of the prompt "The first person to walk on the Moon was..."? The model outputs "Neil Armstrong" because that is the most statistically common completion in its training data for that sentence prefix. (p.7)
Why do outputs *look* like reasoning? Two factors:
> (1) *latent knowledge* and (2) *emergent pattern completion*.
On latent knowledge:
> Through exposure to billions of words, an LLM acquires a broad range of information about the world. It "knows", in a statistical sense, many facts, relationships, and even commonsense truths, simply because these are reflected in language use. It also learns common patterns of explanation and argument, such as how "because" often introduces an explanation, and that scientific questions are answered with specific explanatory forms. (pp.7-8)
On pattern completion:
> If a chain of reasoning often solves a problem in a text, the LLM may generate such a sequence. A notable improvement is that prompting LLMs with "let's think step by step" often leads them to produce a logical chain of thought, which enhances accuracy on multi-step problems. The model is not suddenly performing real deduction; instead, the prompt triggers an output mode that mimics how humans outline reasoning steps, which strongly correlates with correct solutions in the training data. (p.8)
The paper engages the debate about whether LLMs have "implicit world models":
> Some researchers argue that LLMs develop an implicit world model and can perform limited reasoning within it, thus exhibiting emergent reasoning as scale increases. Others maintain that any reasoning success is simply a superficial pattern-matching trick and would fail with slight variations in problems, hence calling successes "luck" or artefacts. (p.8)
On hallucinations:
> An illustrative example is the phenomenon of AI "hallucinations," in which an LLM invents a non-existent source or confidently offers a fabricated statement or explanation. [...] This tendency shows that the abductive style of LLM outputs is a double-edged sword: the model proposes an explanation or answer because that is what fluent, human-like responders do, and because models are trained to be 'helpful', projecting certainty so as not to undermine their perceived credibility. (p.9)
Key limitation:
> LLMs, unless enhanced with tools or human oversight, currently lack these safeguards by default. [...] In reality, their operation is driven by maximising the probability of the sequence [...]. The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. (pp.9-10)
---
## Section 5: The Phenomenology of Plausibility
This section examines *why* LLM outputs feel like reasoning to users:
> When users interact with an LLM-based AI, such as a chatbot or assistant, they often perceive the AI's responses as if they were created by an intelligent mind reasoning through the question. [...] What underpins this phenomenology? In large part, it is because the LLM's training on human language enables it to mimic how humans communicate explanations and reasons. Human-written text in its training data often results from IBE. (p.10)
The car example illustrates the mechanism:
> "Why might my car not start on a cold morning?" The LLM might respond: "It could be due to a weak battery, as cold weather reduces battery efficiency, making it harder to deliver the necessary current..." [...] Yet, the LLM lacks actual understanding or mental grasp of cars; it strings together probable sentences about car troubles. (p.11)
On LLM limitations:
> In less common situations, LLMs can falter or produce a confident-sounding explanation that is subtly incorrect. [...] Unlike a human doctor, who carefully weighs evidence (or at least can and should), the LLM "does not know what it does not know"—it has no awareness of its own ignorance—nor does it necessarily detect subtle inconsistencies. (p.11)
The authors introduce the term *over-abduction*:
> In terms of IBE, it is as if the model always chooses an explanation, even when none is justified—it cannot "resist" explaining because generating a plausible and preferable continuation is its task. This could be termed over-abduction: a human reasoner might say, "I'm not sure; more information is needed", while the LLM often makes a guess regardless. (p.12)
On sycophancy:
> This is the tendency of LLMs to generate outputs that prioritise alignment with user beliefs or preferences over factual accuracy. Because users frequently prefer convincingly-written sycophantic responses to correct responses, they are not minded to 'fact-check' if the output supports their explanation. (p.12)
But there's a positive side—LLMs as hypothesis generators:
> LLMs can generate hypotheses that a person might not have considered, effectively broadening the scope of abductive search. [...] In this way, the LLM functions as an abduction generator, supporting the human reasoner during the discovery phase. It then becomes the human's task to carry out the justification phase, evaluating those hypotheses and developing the correct reasoning. (p.12)
Evocative metaphor:
> To adapt a common metaphor, they are the new box that can help think outside the old box. (p.12)
On the question of whether process matters if outputs are identical:
> If an AI can generate the same explanatory hypothesis a human would, does it matter that the process was different? From an epistemological standpoint, perhaps yes—justification is significant—but regarding the content of the hypothesis and our interpretation of it, maybe not. (pp.12-13)
---
## Section 6: Objections
The authors consider five objections:
### Objection 1: LLMs don't reason at all, so comparing them to abduction is misguided
Reply:
> We are not claiming that LLMs hold literal beliefs or follow Peirce's method of hypothesis internally. Instead, we argue that the output structure of LLMs often resembles that of an abductive reasoning process, which is frequently reinforced by interface design, and this resemblance is not random but systematic, resulting from training on human explanations. (p.15)
### Objection 2: If LLMs are just stochastic parrots, why do they sometimes outperform humans?
Reply:
> When an LLM surpasses humans on a task, it could be because it has encountered many examples during training and has effectively learned patterns that humans might find unintuitive. [...] Nonetheless, this does not amount to understanding; it is more like a student who has seen many example solutions and can pattern-match to solve a new problem in the same format. But if the format is slightly altered [...] humans can adapt while the LLM may fail. (pp.15-16)
### Objection 3: Abduction requires common sense and causality, which LLMs lack
Reply:
> The LLM's extensive training on language has endowed it with a vast repository of commonsense causal knowledge, although not explicitly structured. It "knows" that slippery floors cause falls, that not eating causes hunger, that polls predict elections, and so on—because it has processed countless expressions of these relations. [...] What it lacks, however, is an experiential or embodied grounding of that knowledge. (p.16)
### Objection 4: LLM outputs are often incoherent, not like good explanations
Reply:
> Our analysis has focused on situations where LLMs succeed in delivering explanation-like answers. However, it is essential to remember that this requires a sufficiently capable model and often requires careful prompting. [...] With ample data and parameters, the model captures more of the coherence present in human discourse. (p.17)
### Objection 5: This analysis only applies to token-completion models
Reply:
> At the time of writing, the most successful model in the GPT series, GPT 5.1, remains a token completion model. [...] Importantly, its "reasoning" capability does not fundamentally distinguish it from a token completion model; rather, it is an advanced feature implemented using the token completion mechanism itself. The model generates internal, hidden tokens that function as a scratchpad before producing the final user-facing output. (p.18)
---
## Section 7: Limitations
The authors acknowledge:
> We have treated "LLMs" somewhat generally, focusing mainly on the latest large models as of 2025, with the GPT series as a reference point. [...] We remain agnostic about future possible systems. (pp.18-19)
They also flag ethical implications they don't fully explore:
> If a model provides a convincing but incorrect explanation, users may trust it unnecessarily, thereby spreading misinformation. That is arguably an epistemic harm of the abductive illusion: it can mislead us into treating informed guesses as if they were knowledge. (p.19)
Practical upshot:
> One should regard its output more as the opinion of an anonymous forum poster—possibly correct, possibly incorrect—rather than an expert. (p.19)
---
## Section 8: Conclusion
The paper's core formulation:
> We can briefly describe LLMs as fundamentally stochastic, with surface-level abductive appearances. Recognising this duality helps clarify some debates: we can agree with sceptics that no human-like understanding occurs internally, while also explaining why these models are so successful and attractive: they leverage the informational richness of human language and thus effectively stand on the shoulders of our collective knowledge and reasoning. (p.20)
On promise:
> They represent the novelty of "engines of generative plausibility": never before have we had systems capable of producing human-like, plausible text at scale. [...] They are the best interfaces we have today for accessing, querying, and managing the immense accumulation of human content. (p.20)
On danger:
> If one conflates the surface with the core—if one assumes the LLMs genuinely "know what they are talking about"—one can be misled. [...] An explanation can be coherent and convincing (even optimal by IBE criteria) and yet still false. LLMs lack an epistemic compass to navigate that distinction. As users or deployers of LLMs, we must provide that compass externally. (p.20)
Final sentence:
> These powerful tools have increased our epistemic responsibilities. (p.21)
---
## My Assessment
This is a clear, philosophically careful paper that stakes out a moderate position between "LLMs are stochastic parrots" and "LLMs genuinely reason." The central framing—stochastic core, abductive surface—is useful.
I'd flag a few things worth noting:
- The paper is explicitly dated to "as of 2025" and acknowledges GPT-5.1, so it's engaging with very recent models
- The Floridi et al. team is sympathetic to LLMs as *tools* (hypothesis generators, brainstorming partners) while being sceptical about their outputs as *knowledge*
- The "over-abduction" concept (LLMs can't resist explaining) is helpful for understanding hallucination
The paper doesn't break new empirical ground but synthesizes the philosophical situation clearly. It would pair well with the more technical benchmark papers it cites (Webb et al. on analogical reasoning, Balepur et al. on abductive NLI).
*Floridi resta uno dei pochi filosofi che riesce a parlare sia ai tecnici sia agli umanisti senza perdere rigore.*