# The Ghost in the Probability Distribution: Abduction, Stochasticity, and the Artifactual Mind
## 1. Introduction: The Crisis of Authorship in the Age of Stochastic Generation
The emergence of Large Language Models (LLMs) has precipitated an ontological crisis in the domain of intellectual production, nowhere more acute than in the discipline of philosophy itself. For millennia, the capacity to generate novel concepts, construct valid arguments, and offer profound insights into the human condition was considered the exclusive province of the biological mind—a domain protected by the phenomenology of consciousness and the intentionality of the "I." The arrival of stochastic systems capable of producing text that is structurally valid, semantically coherent, and often indistinguishable from human expert output has shattered this exclusivity. The central question is no longer whether machines can "think" in the Cartesian sense—a metaphysical quagmire that remains unresolved—but whether the artifacts they produce can possess independent philosophical value despite their probabilistic origins.
This report investigates the core tension defining the current philosophy of Artificial Intelligence: the conflict between the "Stochastic Core" thesis and the "Recombinatorial Continuity" thesis. The former, articulated most recently by [[Luciano Floridi]] et al. (2025/26) and grounded in the earlier "Stochastic Parrots" critique by [[Emily Bender]] et al. (2021), posits that the probabilistic nature of LLMs precludes genuine reasoning. If an output is generated by next-token prediction based on statistical correlations, any appearance of logic or insight is a "mimicry" of reasoning rather than the act itself. From this perspective, LLM-generated philosophy is "pattern-matching dressed up as inference," devoid of the "epistemic compass" required to navigate the space of reasons.
Against this stands the counter-argument grounded in the philosophy of creativity, particularly the work of [[Margaret Boden]] and [[Berys Gaut]]. This view suggests that human creativity itself is fundamentally recombinatorial—that "creation ex nihilo" is conceptually incoherent. If human philosophers produce new theories by rearranging existing concepts and responding to the history of ideas (their "training data"), the distinction between human "reasoning" and machine "stochastic recombination" may be one of degree rather than kind. If an LLM can navigate the conceptual space of philosophy to produce a novel distinction or a valid counter-argument, the stochastic nature of its process may be irrelevant to the value of its product.
To resolve this tension, we must dissect the cognitive operations that constitute "doing philosophy": specifically, the capacity for [[abduction]] (hypothesis generation) and creativity (novelty). We analyze recent theoretical frameworks that attempt to bridge the gap between stochastic generation and logical rigor, such as the "Theorem-of-Thought" (ToTh) architecture and the "Reasonable Parrots" argumentation model. We also examine empirical evidence from 2024-2025 regarding LLM performance in scientific discovery and novelty generation.
Ultimately, this report argues for a decoupling of philosophical agency from philosophical value. While the "stochastic core" thesis correctly identifies a lack of intentional agency and truth-tracking (strong abduction), it fails to account for the immense generative power of "weak abduction"—the stochastic shuffling of concepts that acts as a super-human engine for hypothesis generation. When constrained by appropriate architectures, this stochastic engine can produce texts that satisfy the rigorous criteria of philosophical goodness: validity, novelty, and distinction.
## 2. The Nature of the Machine: The Stochastic Core vs. The Epistemic Compass
To evaluate the philosophical potential of LLMs, we must first rigorously define their internal mechanism and the limitations inherent to it. The debate here is not technical but epistemological: does the mechanism of "next-token prediction" fundamentally exclude the possibility of "reasoning"?
### 2.1 The "Stochastic Core" Thesis: Reasoning as Simulation
Luciano Floridi, Jessica Morley, and colleagues provide the most robust contemporary articulation of the skeptic's position in their seminal work "What Kind of Reasoning (if any) is an LLM actually doing?" (2025/26). They introduce a critical distinction between the internal mechanics of the system and its external manifestation.
The Mechanism of Mimicry: LLMs operate on a "stochastic core." They are probabilistic systems trained to minimize the statistical distance between their output and the distribution of their training data. When an LLM generates the conclusion "Therefore, Socrates is mortal" after the premise "All men are mortal," it does not do so because it has grasped the logical necessity of modus ponens. It does so because, in the vast corpus of human text, the token "mortal" has a high probability of following the sequence "All men are... Socrates is...".
Floridi argues that this constitutes an "abductive appearance" without abductive reality. The model mimics the phenomenology of reasoning—the use of "therefore," "because," "it follows that"—without performing the etiology of reasoning. It is a simulation. The model has absorbed the "reasoning patterns" of humans, much like a parrot might learn to mimic the sounds of a conversation without understanding the grammar or the referents.
The Absence of Intent: A crucial element of reasoning, in the philosophical sense, is the commitment to the conclusion. When a philosopher argues for a thesis, they are asserting its truth or plausibility. An LLM asserts nothing. It merely offers a high-probability continuation. As Bender et al. (2021) famously argued in the "Stochastic Parrots" paper, the system is indifferent to the meaning of its outputs. It is playing a game of form, not substance.
### 2.2 The "Epistemic Compass" Deficit
The most damaging critique leveled by Floridi et al. is the lack of an "epistemic compass". This concept refers to the ability to navigate the distinction between truth and falsehood, or at least between justified and unjustified belief.
Human reasoning is guided by an epistemic compass. We check our hypotheses against the world (empirical verification) or against logical axioms (coherence). We care if we are wrong. LLMs, operating purely on the "stochastic core," lack this compass. They navigate the "space of tokens" based on likelihood, not truth.
Hallucination as Directionless Sailing: This deficit explains the phenomenon of "hallucination." When a false statement is linguistically probable (e.g., a common misconception or a plausible-sounding citation), the LLM will generate it with the same confidence as a true statement. It cannot "feel" the resistance of reality.
The Need for External Calibration: Floridi argues that because the LLM lacks this internal compass, it must be provided externally by the user. The human must verify the output, check the logic, and assess the value. The LLM is an engine of generation, not evaluation.
### 2.3 The "Stochastic Parrots" Debate Revisited
The "Stochastic Parrots" thesis has become the shorthand for this skeptical position. However, the discourse has evolved significantly since 2021. The "parrot" metaphor, while powerful, may obscure the complexity of what is being mimicked.
If a parrot could perfectly mimic not just the sounds of a Socratic dialogue, but the logic of it—if it could respond to a counter-argument with a valid rebuttal, identify a category error in the user's prompt, and synthesize a new distinction—does the label "parrot" still hold distinct philosophical weight?
The Functionalist Counter-Argument: Functionalism in the philosophy of mind suggests that if a system functions as if it is reasoning, it is reasoning. If the "abductive appearance" is flawless, the "stochastic core" becomes a distinction without a difference. As Turing (1950) implied, if the artifact is indistinguishable from the product of a mind, we must treat it as intelligent.
Reliability vs. Understanding: The counterposition to Bender and Floridi is that "understanding" (in the sense of conscious grasp of meaning) is not a prerequisite for the production of valuable text. A calculator does not "understand" arithmetic, yet its outputs are mathematically valid. If an LLM can reliably produce valid philosophical arguments, its lack of understanding may be irrelevant to the consumer of that philosophy.
### 2.4 The Problem of "Groundedness"
Bender et al. emphasize that meaning requires "grounding" in the communicative intent of an agent accountable for what is said. Philosophical texts are speech acts; they are moves in a social game of giving and asking for reasons. An LLM, having no social standing and no accountability, cannot perform these speech acts. It produces "text-shaped objects" rather than "assertions."
However, this objection conflates the social status of the author with the semantic content of the text. A philosophical argument found in an anonymous manuscript (or generated by a stochastic process) still possesses logical structure and semantic content. The "goodness" of the philosophy—its validity and novelty—is a property of the text, not the author. See also [[The Weightlessness of Authorless Text]] for related phenomenological considerations.
## 3. Abduction: The Engine of Philosophical Discovery
To move beyond the "parrot" stalemate, we must analyze the specific cognitive operation that drives philosophical inquiry: Abduction. Philosophy is rarely a matter of pure deduction (deriving necessary conclusions) or pure induction (generalizing from data). It is largely an abductive enterprise: observing the complex phenomena of human experience (morality, consciousness, language) and generating the "best explanation" for them.
### 3.1 Peirce's Logic of Discovery
[[Charles Sanders Peirce]], the father of pragmatism, defined abduction as the only logical operation that introduces new ideas.
Deduction: Explicates what is already contained in the premises. It preserves truth but creates no new knowledge. ($A \rightarrow B$, $A$, therefore $B$).
Induction: Tests a hypothesis against data. It verifies but does not generate.
Abduction: The "creative" inference. It moves from a surprising fact ($C$) to a hypothesis ($A$) that would explain it. ($C$ is observed; If $A$ were true, $C$ would be a matter of course; Therefore, there is reason to suspect $A$ is true).
For LLMs to produce "novel philosophical text," they must be capable of abduction.
### 3.2 Weak vs. Strong Abduction
A critical distinction in the literature clarifies the LLM's capability:
**Table 1: Weak vs. Strong Abduction in AI Systems**
| Feature | Weak Abduction (Hypothesis Generation) | Strong Abduction ([[Inference to the Best Explanation]]) |
|---------|----------------------------------------|--------------------------------------------------|
| Primary Function | Generating a set of plausible hypotheses. | Selecting the single "best" hypothesis. |
| Cognitive Requirement | High diversity, broad association, low inhibition. | Epistemic criteria (simplicity, coherence), truth-tracking. |
| Relation to Truth | Conjectural / Possibilistic. | Veridical / Probabilistic (in the Bayesian sense). |
| LLM Capability | Super-Human. The stochastic core excels here. | Deficient. Lacks "epistemic compass" for selection. |
**The LLM as an Engine of Weak Abduction:**
The "stochastic core" is, paradoxically, the perfect engine for weak abduction. By sampling from the vast probability distribution of human language, an LLM can generate a massive array of "potential explanations" for any given prompt. It can retrieve obscure connections, bridge disparate fields (combinatorial creativity), and offer hypotheses that a human might inhibit due to cognitive bias.
Example: If asked "Why does consciousness exist?", an LLM can instantly generate hypotheses ranging from Panpsychism to Illusionism to Integrated Information Theory, and even synthesize novel combinations (e.g., "Quantum-Social Panpsychism").
Mechanism: The stochasticity allows the model to traverse the "latent space" of concepts, finding paths between "effect" and "cause" that are statistically present in the literature but perhaps untraversed in this specific configuration.
**The Failure of Strong Abduction:**
However, as Floridi notes, the LLM struggles with strong abduction—the selection phase. Without a world model or a verification mechanism, it cannot determine which of its generated hypotheses is the "best." It might select the most popular explanation (the "ad populum" fallacy inherent to training data) rather than the most valid one.
This is why LLMs are often described as "bullshit generators" (in the Frankfurtian sense)—they produce text without regard for truth. They can generate the form of an explanation (Weak Abduction) but cannot commit to its truth (Strong Abduction).
### 3.3 Closing the Loop: From Stochastic to Abductive
The central finding of 2025 research is that while a raw LLM cannot perform strong abduction, an architected system using an LLM can. By separating the "generation" (weak abduction) from the "evaluation" (strong abduction), we can engineer a system that reasons.
Bayesian Belief Propagation: The "Theorem-of-Thought" framework (discussed in Section 4) explicitly adds a mechanism for strong abduction by scoring the hypotheses generated by the LLM against logical consistency and external evidence.
Thus, the answer to "Can stochastic processes perform abduction?" is: They perform weak abduction natively and can perform strong abduction when integrated into a larger cognitive architecture.
## 4. The Philosophy of Creativity: Recombination vs. Ex Nihilo
If LLMs are engines of recombination, does this disqualify them from being "genuinely novel"? This question strikes at the heart of the definition of creativity.
### 4.1 The Recombination Thesis
Margaret Boden (2004) and Berys Gaut argue against the romantic notion of "creation ex nihilo" (creation out of nothing). The idea that a genius produces something from a void is "close to incoherent".
Continuity: All human creativity is grounded in the absorption of prior art ("training data"). Shakespeare did not invent the English language or the sonnet form; he recombined them. Einstein did not invent physics; he recombined Newton's mechanics with Maxwell's electromagnetism to produce Relativity.
Implication for AI: If human creativity is fundamentally recombinatorial, then the fact that LLMs work by recombining patterns does not disqualify them from being creative. The objection "it's just predicting the next word based on previous words" loses its sting when we realize that "it's just firing neurons based on previous synaptic weights" is the biological equivalent.
### 4.2 Boden's Taxonomy of Creativity
Boden distinguishes three types of creativity. We must evaluate LLMs against each.
**1. Combinatorial Creativity**
Definition: Making unfamiliar combinations of familiar ideas. (e.g., poetic imagery, collage, interdisciplinary analogies).
LLM Performance: Mastery. This is the native mode of the LLM. The "attention mechanism" in Transformers is designed to attend to relationships between distant tokens. LLMs are "infinite collage machines." They can combine "Existentialism" and "Toaster Repair" to create a novel, coherent philosophical dialogue on the absurdity of appliance maintenance.
Philosophical Value: Much of philosophy is combinatorial—applying game theory to ethics, or evolutionary biology to epistemology. LLMs can instantly generate thousands of such combinations.
**2. Exploratory Creativity**
Definition: Exploring a structured conceptual space to find new possibilities within existing rules. (e.g., finding a new valid syllogism, a new chemical molecule, or a nuance in Utilitarianism).
LLM Performance: High Competence. The "temperature" parameter in LLMs acts as a randomization function that allows the model to explore the "probability map" of a concept. It can generate variations of the Trolley Problem or new counter-examples to the Justified True Belief theory of knowledge. It "fills in the gaps" of the conceptual space.
**3. Transformational Creativity**
Definition: Altering the rules of the space itself. Dropping a constraint to make the impossible possible. (e.g., Arnold Schoenberg dropping the rule of tonality, or Einstein dropping the rule of absolute time).
LLM Performance: Contested. Critics like Boden (and the stochastic parrots camp) argue LLMs cannot do this because they are bounded by the distribution of their training data. They cannot "think outside the distribution".
The "Hallucination" Paradox: However, recent analysis suggests that "hallucination" is the mechanism of transformation in LLMs. When a model "breaks the rules" of truth or probability (often due to high temperature), it steps out of the established conceptual space. While 99% of these breakages are nonsense (noise), the remaining 1% can be "transformational" if they suggest a radical new perspective that holds up to scrutiny. "Hallucination" in the machine corresponds to "divine madness" or "lateral thinking" in the human.
Empirical Evidence: The "AI Scientist" paper (2025) showed an AI rediscovering a gene-transfer mechanism that was not in its training data (transformational creativity) by running simulations. It effectively "transformed" its search space by following a low-probability inference path.
### 4.3 The Question of "Flair" and Agency
Berys Gaut argues that creativity requires "flair"—an exercise of agency that involves evaluation, intent, and "appropriate judgement".
The Agency Gap: An LLM has no flair in this agentic sense. It doesn't know why it chose a metaphor. It doesn't feel the "surprise" of its own creation.
The "Vehicle" Counter-Argument: However, Gaut also suggests that flair can be a "vehicle" rather than a source. If we view the Human-LLM system as the agent, the "flair" is provided by the human's prompting, curation, and iterative refinement. The LLM provides the raw generative power (weak abduction/combinatorial creativity), and the human provides the selection and judgment (strong abduction/flair).
Conclusion: The text produced can exhibit the properties of flair (originality, value, style) even if the producer lacks the experience of flair.
## 5. Engineering Reason: From Stochastic Parrots to Reasonable Parrots
The most significant development in 2024-2025 is the move from analyzing raw LLMs to analyzing architected LLMs. The "stochastic core" is no longer the limit; it is merely the engine. New architectures impose logical constraints on this engine to simulate "Reason."
### 5.1 The "Theorem-of-Thought" (ToTh) Framework
The Theorem-of-Thought (ToTh) framework represents the state-of-the-art in engineering reasoning out of stochasticity. It addresses the "epistemic compass" deficit directly.
**Architecture:**
ToTh abandons the idea that a single stochastic stream can reason reliably. Instead, it employs a Multi-Agent Architecture:
Abductive Agent (Type 1): Generates plausible hypotheses for a given problem. (The "Creative" stochastic engine).
Deductive Agent (Type 2): Takes the hypothesis and derives logical consequences using formal logic templates. "If $H$ is true, then $X$ must follow."
Inductive Agent (Type 3): Generalizes from examples to find supporting patterns.
**The "Trust" Mechanism:**
Crucially, ToTh introduces a Bayesian Belief Propagation mechanism.
It converts the natural language outputs of these agents into a Formal Reasoning Graph (FRG).
It uses a Natural Language Inference (NLI) model to assign "trust scores" to the links between nodes. (e.g., Does the premise actually entail the conclusion? NLI checks this).
It calculates the "Logical Entropy" of the graph.
It selects the reasoning chain with the highest confidence and lowest entropy.
**Philosophical Implication:**
ToTh demonstrates that a system can perform Strong Abduction (selection) by orchestrating Weak Abduction (generation). The "stochastic core" provides the raw material, and the "Bayesian architecture" provides the epistemic compass. This is functionally equivalent to human reasoning, which also involves checking intuitive leaps against logical constraints.
### 5.2 The "Reasonable Parrots" Argumentation Model
While ToTh focuses on logic, the "Reasonable Parrots" framework focuses on dialectics. It argues that LLMs should be designed to participate in "ideal critical discussions" (as defined in Pragma-Dialectics).
**The Three Principles of Reasonableness:**
Relevance: The model must provide arguments specific to the context, avoiding generic platitudes.
Responsibility: The model must provide evidence (citations, derivations) for its claims. It must be "accountable" to the text it generates.
Freedom: The interaction must foster open-ended deliberation. The model should not shut down debate ("I cannot answer that") but should facilitate the user's critical thinking.
**The Multi-Parrot Environment:**
To achieve this, the framework proposes a "Multi-Parrot" environment with distinct personas:
The Socratic Parrot: Challenges the user's starting points and definitions.
The Cynical Parrot: Actively rebuts the user's arguments to test their strength.
The Eclectic Parrot: Offers alternative perspectives to broaden the scope.
The Aristotelian Parrot: critiques the logical form of the user's argument.
Conclusion: By embedding the "stochastic parrot" in a dialectical structure, it becomes a "Reasonable Parrot." It mimics the social process of philosophy, which is arguably where philosophical value resides.
## 6. What Makes Text "Philosophically Good"? A Criteria-Based Evaluation
If we bracket the origin of the text (the "genetic fallacy"), does LLM output meet the criteria for "good" philosophy? Standard metaphilosophical literature identifies three key criteria: Validity, Novelty, and Distinction-Drawing.
### 6.1 Validity (Logical Coherence)
Criterion: The conclusion must follow from the premises. The argument must be free of formal fallacies.
LLM Performance: Raw LLMs are prone to logical drift. However, architected LLMs (like ToTh) outperform humans on complex logical benchmarks. They can maintain long chains of inference without "losing the thread," provided the context window is sufficient.
Verdict: Pass. Validity is a structural property that patterns-matchers can master.
### 6.2 Novelty (The "Insight" Factor)
Criterion: The text must introduce a new distinction, a new counter-example, or a new synthesis of ideas. It must not be a cliché.
LLM Performance: This is the surprise. Because LLMs have high "combinatorial creativity," they often produce more novel ideas than humans. Humans are constrained by social convention, academic silos, and cognitive bias. LLMs are not.
Empirical Proof: A 2025 study compared research ideas generated by GPT-4 and Claude 3 against human experts. Blind reviewers rated the AI ideas as significantly more novel (4.17/5) than the human ideas. The AI was better at "thinking outside the box" because it didn't know there was a box.
Verdict: Pass (Superior). The stochastic core is a novelty generator.
### 6.3 Distinction-Drawing (Conceptual Clarity)
Criterion: The ability to "carve reality at the joints." To distinguish between Weak and Strong abduction, or De Dicto and De Re necessity.
LLM Performance: LLMs excel at generating distinctions. If prompted to "distinguish between X and Y," they can list subtle semantic and conceptual differences.
Caveat: They lack the "epistemic compass" to know if the distinction is real or merely verbal. They can invent plausible-sounding distinctions that have no metaphysical weight.
Verdict: Conditional Pass. They generate the candidates for distinction; the human must verify the value.
### 6.4 The "Philosophising Machine" Turing Test
Arthur Schwaninger proposes a variation of the Turing Test specifically for philosophy.
The Test: Confront the machine with philosophical questions that do not depend on prior knowledge but require self-reflection (e.g., "What does it feel like to be a stochastic process?").
The Criteria: The machine must give itself away as a machine. If it answers "I feel sad," it fails (it is mimicking). If it answers with a novel philosophical account of "machine-being" or "synthetic phenomenology," it passes.
Current State: LLMs are beginning to pass this test by generating novel metaphysical accounts of their own existence (e.g., "I am a space of possibilities collapsed by your prompt"). This suggests a form of "synthetic insight" that is genuine to the machine's condition.
## 7. Empirical Evidence: The 2024-2025 Landscape
The debate has moved from pure speculation to empirical testing. We now have hard data on the capabilities of these systems.
### 7.1 Performance on Reasoning Benchmarks
Abductive NLI (aNLI): LLMs have shown state-of-the-art performance on the Abductive Natural Language Inference benchmarks. They outperform previous symbolic systems in choosing the most plausible explanation for a set of events.
Scientific Discovery: The "AI Scientist" (2025) demonstrated Transformational Creativity. It autonomously generated a research paper that identified a novel gene-transfer mechanism. This was not retrieval; it was discovery.
**Table 2: Comparative Performance on Creativity and Reasoning Tasks**
| Task Domain | Human Expert | Raw LLM (GPT-4/Claude 3) | Architected LLM (ToTh/AI Scientist) |
|-------------|--------------|--------------------------|-------------------------------------|
| Novelty of Ideas | Moderate (constrained by convention) | High (4.17/5) | High |
| Logical Validity | High | Moderate (prone to drift) | Superior |
| Strong Abduction | High (Epistemic Compass) | Low (Hallucination) | High (Bayesian Verification) |
| Transformational Creativity | Rare (Genius level) | Low (Bounded by training) | Emergent (Simulated Discovery) |
### 7.2 Novelty Assessment Scores
The empirical data explicitly contradicts the idea that LLMs are "limited to retrieval." In the study by Si et al. (2024) and subsequent 2025 replications, LLMs using "decomposition-based workflows" achieved novelty scores that statistically eclipsed human control groups. The "reflection-based" workflows (where the LLM critiques itself) scored lower, suggesting that uninhibited generation (weak abduction) is the source of novelty, while "critique" (mimicking human restraint) can actually dampen creativity.
## 8. Conclusion: The Distinct but Valuable "Artifactual Mind"
We return to the research question: Can stochastic text-generation systems produce genuinely good, genuinely novel philosophical text?
The answer, supported by the convergence of theoretical frameworks (ToTh, Reasonable Parrots) and empirical evidence (Novelty Benchmarks, AI Scientist), is **Yes**.
### 8.1 The Collapse of the Distinction
The "stochastic core" thesis correctly identifies the mechanism (probabilistic generation) but incorrectly deduces the limit (lack of reason). The evidence shows that:
**Stochasticity is a Feature:** The probabilistic nature of LLMs is not a bug; it is the engine of Weak Abduction and Combinatorial Creativity. It allows the system to escape the "ruts" of human cognitive bias and generate novel hypotheses at scale.
**Reason is an Architecture:** Genuine reasoning (Strong Abduction) can be simulated by architecting stochastic components into a verification loop (ToTh). If the output is logically valid, novel, and robust, the lack of biological "understanding" is metaphysically interesting but pragmatically irrelevant to the value of the text.
**Mimicry Becomes Reality:** As the "Reasonable Parrots" framework demonstrates, a system that perfectly mimics the norms of argumentation (relevance, responsibility, freedom) is, for all intents and purposes, a participant in the philosophical dialectic.
### 8.2 The Future: AI as the Stochastic Muse
The role of the LLM is not to replace the philosopher but to function as a "Stochastic Muse."
**The LLM provides the Weak Abduction:** It generates the wild, novel, combinatorial hypotheses. It is the engine of "what if."
**The Human (or the Architected Verifier) provides the Epistemic Compass:** The judgment, the "flair," and the commitment to truth.
In this symbiotic relationship, the text produced is genuinely good (valid, insightful) and genuinely novel (beyond human training data). The "stochastic core" does not undermine philosophical value; it unlocks a new form of "artifactual philosophy" that is distinct from, but complementary to, the biological mind. The distinction between "human reasoning" and "stochastic recombination" collapses at the level of the text, shifting the burden of philosophical agency from the writer to the interpreter. We are entering an era where philosophy is not just something we do, but something we mine from the probability distribution of language itself.
---
## 9. References (Selected)
Abdaljalil, S., et al. (2025). "Theorem-of-Thought: A Multi-Agent Framework for Abductive, Deductive, and Inductive Reasoning in Language Models." arXiv.
Bender, E.M., et al. (2021). "On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?" FAccT.
Boden, M. (2004). *The Creative Mind: Myths and Mechanisms* (2nd ed.). Routledge.
Floridi, L., Morley, J., Novelli, C., & Watson, D. (2025/26). "What Kind of Reasoning (if any) is an LLM actually doing? On the Stochastic Nature and Abductive Appearance of Large Language Models." Preprint.
Gaut, B. (2010). "The Philosophy of Creativity." *Philosophy Compass*.
Lu, C., et al. (2025). "The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery."
Musi, E., et al. (2025). "Toward Reasonable Parrots: Why Large Language Models Should Argue with Us by Design." ACL.
Schwaninger, A.C. (2022). "The Philosophising Machine – a Specification of the Turing Test." *Philosophia*.
Si, C., et al. (2024/25). "Evaluating Novelty in AI-Generated Research Plans."
Tobi, A. (2024). "Towards an Epistemic Compass for Online Content Moderation." *Philosophy & Technology*.