# LLMs and Abductive Reasoning (Floridi et al.)
A conversation exploring Floridi et al.'s paper on LLM reasoning, truth-tracking, abduction, and what it would take for LLMs to do genuine philosophy.
**Key topics:**
- Floridi's critique: LLMs lack truth as internal objective
- Opponent moves: functional truth-tracking via training
- CEV blueprint for verification loops
- Move 37 and tail novelty
- Prompting techniques for philosophy
- The "obvious move" trick
---
**You said:**
1ab660a5-375e-4c36-9f8e-d11255529738.pdf PDF I am having trouble understanding exactly what the authors of this paper are saying LLMs can't do, and why. Please help me work it out.
**ChatGPT said:**
Here’s what I think is tripping you up: the paper is not saying “LLMs can’t produce explanations.” It’s saying something more slippery: LLMs can generate text that looks like abductive reasoning (Inference to the Best Explanation), but—by default—they aren’t doing the epistemic job that abduction/IBE is supposed to do. They’re doing stochastic token continuation that often imitates abductive structure because they were trained on human text where abduction is everywhere. Let’s pin down exactly what they claim LLMs can’t do, and why.
## 1) The core “can’t”: they can’t
use truth as a constraint The authors’ central claim is: mainstream token-completion LLMs “gen‐ erate text based on learned associations rather than performing abduc‐ tive inferences” and they “do not…ground [outputs] directly in truth, semantics, verification, or understanding.” So what’s the missing ability?
When a human does abduction, even weak abduction (“here’s a plausible story”), the story is aimed at the world—it’s trying to be true, and it’s answerable to evidence. When a human does strong abduction / IBE, they don’t just list hypotheses; they rank them using constraints like fit with background knowledge, simplicity, scope, coherence, and so on. The paper’s punchline is that an LLM doesn’t have “truth” as an internal objective. Its objective is: output a continuation that is probable given the prompt and learned distributions. That’s why hallucinations matter so much to their argument: hallucina‐ tions are the visible symptom that the system lacks an internal truthchecking loop.
## 2) The second “can’t”: they can’t
verify or justify their own hypotheses They explicitly lean on the old Reichenbach distinction: discovery vs justification. They say LLMs can help with the discovery phase (generate candidate hypotheses), but they do not genuinely perform the justification phase (testing/verification against reality), “unless…augmented by other systems.” In their terms: LLMs do something like “prior predictive sampling” (spit out plausible candidates), but don’t do posterior evaluation against evidence. So even when the output reads like “best explanation,” what’s missing is the epistemic discipline: checking, revising, refusing to answer when underdetermined, noticing contradictions, demanding more evidence.
This is also why they talk about “over-abduction”: the model tends to produce an explanation even when “I don’t know” would be the rational move.
## 3) The third “can’t”: they can’t have grounded
semantics or embodied causal understanding They stress “no grounded semantics” (no built-in connection between words and perception/action in the world) and no “experiential or embodied grounding.” Their smoke/fire example is the cleanest way to see the alleged gap: Humans: infer real fire because they understand causal structure in the world. LLMs: output “fire” because “smoke → fire” is a strong linguistic association and a common causal trope in the training data. The output can be the same, but they insist the underlying “aboutness” is different: humans’ inference is world-directed; the model’s is text-distri‐ bution-directed.
## 4) So why does it look like they can do it?
Because the training data is stuffed with human reasoning. They call this the “phenomenology of plausibility”: LLMs imitate expla‐ nation-forms (“because”, “therefore”, enumerating possibilities, then picking one) because those are stable patterns in Wikipedia, textbooks, StackExchange, papers, etc. Add interface design (chat format, confident tone, “assistant” framing) and the user is nudged to read the output as the product of an agent that reasoned.
Their key move is: abductive structure can be learned as a linguistic tem‐ plate without abductive commitment to truth.
## 5) But wait — don’t they admit LLMs
sometimes do “weak abduction”? Yes—and this is a source of confusion. Early on they say LLMs “seem to perform at least weak abduction” and can even look like strong abduction on multiple-choice abductive tasks. But later they tighten the screws and basically re-describe this as: it’s “ze‐ roth-order abduction” — generating plausible continuations that mimic hypothesis generation, without real understanding or verification. So the consistent reading is: They’re happy to say LLMs output something that functions like weak abduction for the user. They deny that the model is internally performing abduction as an epistemic method (because no truth/meaning/verification/grounding). That’s the “stochastic core / abductive appearance” slogan.
## 6) What exactly are they denying:
“reasoning” or “abduction”? They try to avoid the dumbest version (“LLMs are just parrots, end of story”), but they also push back against the heroic version (“LLMs reason like us”). Their position is: the “inference” going on is statistical inference over tokens—choosing likely continuations—rather than inference over
hypotheses about the world. When people point to high benchmark performance, they respond: that can come from having seen many patterns/solutions in training; it doesn’t show flexible understanding, and failures under slight perturba‐ tion suggest template reliance. They also claim LLMs lack a “metacognitive check” (they can output inconsistent claims without “realising” the conflict).
## 7) A good compression test: what would
count as “doing IBE” on their view? From their discussion, doing genuine IBE would require (at least): representing candidate hypotheses as candidates about the world, not just continuations; evaluating them with respect to evidence and background knowl‐ edge as constraints; having some mechanism for verification / error-correction / “I don’t know” when underdetermined; some grip on truth conditions (even if fallible); ideally, some grounding (or at least a robust external checking loop). And they think token-completion LLMs, by default, don’t have these— and that you need augmentation (tools, retrieval, verification modules, human oversight) precisely because the base mechanism isn’t built for that. If you tell me which part feels most confusing—(a) their use of “abduc‐ tion/IBE” talk despite denying abduction, (b) the truth/verification
point, or (c) the grounding/causality point—I can zoom in and recon‐ struct their argument in the tightest possible form.
---
**You said:**
"When a human does strong abduction / IBE, they don’t just list hypotheses; they rank them using constraints like fit with background knowledge, simplicity, scope, coherence, and so on. 1ab660a5-375e4c36-9f8e-d112555… The paper’s punchline is that an LLM doesn’t have “truth” as an internal objective. Its objective is: output a continuation that is probable given the prompt and learned distributions. " this is interesting. could an opponent of the view say tHAT IF YOU HAD THE BACKGROUND KNOWLEDGE, SIMPLICITY, SCOPE, COHERENCE AND SO ON, ALSO REPRESENTED IN THE LLM, AND SORRY, REPRESENTED, ACCOUNTED FOR STATISTICALLY OR SORT OF PART OF ITS TRAINING DATA OR EMBEDDED WITHIN IT SOMEHOW. WOULD THAT NOT BE SIMILAR-ISH TO HUMAN TRUTH? I MEAN, I KNOW THIS IS A SORT OF NAIVE QUESTION, BUT I WANT TO HEAR WHAT YOU SAY BACK.
**ChatGPT said:**
Yep — that’s exactly the right “opponent move”, and it forces the debate to stop being mystical and start being about what, precisely, would count as having those constraints “in the system”. The paper itself basically invites your move, because it already grants two things: First, human IBE criteria like simplicity/coherence/scope can be “en‐ coded in language” and therefore statistically learnable (they say the model absorbs “patterns of human abductive reasoning as expressed in writing” and that its training data “encode reasoning structures”).
Second, they even concede that in many everyday cases the “best expla‐ nation is obvious” and models do well precisely because they can echo those obvious explanatory tropes. So an opponent can say: look, if the IBE-criteria are implemented as reg‐ ularities in the model’s learned distribution, then the model will tend to output explanations that satisfy those criteria, and that’s not nothing. It’s at least “IBE-like behaviour”. Where Floridi et al. will push back is that this still doesn’t buy you truth as an internal norm — and their pushback comes in (at least) three layers. 1. “Having the criteria in the weights” ≠ “using the criteria as con‐ straints on truth” The model can learn that humans prefer simple/coherent explana‐ tions, and so it will preferentially generate them. But preferring “simple-sounding” explanations isn’t yet tracking whether the explanation is actually true. They say the model “generate[s] text based on learned associations” and “without grounding [it] directly in truth, semantics, verification, or understanding”, and later: “They aim to model the conditional distribu‐ tion of tokens in text, not to evaluate truth.” So the objection is: your proposal makes the model better at imitating the human ranking style, but it still hasn’t given it an epistemic compass. It can still confidently pick a beautiful, coherent explanation that is just false — because the selection pressure is “fits the prompt + training dis‐ tribution”, not “is true of the world”. 2. The missing step is justification: checking against reality They lean on the discovery/justification split: LLMs can help gener‐ ate hypotheses (discovery), but they don’t “genuinely validate them against reality” unless you bolt on extra systems.
This matters because even a human who uses simplicity/coherence/etc. is still (at least in principle) answerable to evidence: you can test, look again, run an experiment, notice a contradiction, update. The paper’s line is that LLMs “perform prior predictive sampling but lack an external feedback loop for posterior evaluation.” So: even if “IBE-criteria” are statistically represented, you still don’t have the thing that makes IBE epistemically serious rather than just stylisti‐ cally convincing — a mechanism that punishes falsehood by contact with evidence. 3. Grounding and reference: why “similar-ish to human truth” is the contentious step Your question treats truth as if it’s basically the same thing as “hav‐ ing enough background knowledge + IBE virtues”. But the authors want to reserve “truth-tracking” for systems that have either (a) grounded semantics (some non-linguistic tie to the world) or (b) deliberate verification pathways. They explicitly say LLMs “lack grounded semantics connecting words to the physical world” and “do not possess an inherent concept of truth or verification beyond what their training data provides.” That’s why they use the smoke/fire contrast: humans infer fire in the world, LLMs predict “fire” in text because of learned co-occurrence — “the difference is subtle”, they say, but it’s exactly the difference you’re asking about. What an opponent can reasonably reply (and I think it’s strong) An opponent can say: “Truth-tracking doesn’t require mystical ‘ground‐ ing’; it can be implemented functionally. If you train a system in a way that reliably rewards correct world-model predictions and penalises errors — including by giving it tools and feedback — then it does have truth as an objective, even if the substrate is statistical.”
Floridi et al. half-agree: they say “Any connection to real-world evidence must be deliberately engineered, as in retrieval-augmented systems” and they talk about adding verification modules, search engines, RAG, etc. So the real dispute becomes: If we (i) embed the IBE virtues in training, and (ii) add reliable feedback/verification loops, at what point does it stop being “mere abductive appearance” and become “genuine IBE”? Their answer (in this paper) is basically: as long as you’re talking about “current, mainstream…token-completion paradigm” models, the core mechanism is still next-token prediction, so without augmentation they lack the justificatory machinery and thus lack truth as an internal norm. My take, staying close to their framework Your “naive” question is actually the cleanest wedge: it exposes that a lot hinges on whether “truth as an internal objective” means: A) “the loss function explicitly optimises accuracy w.r.t. the world”, plus systematic feedback, or B) “the system has human-like semantics/understanding”, or C) “the system behaves in a sufficiently truth-tracking way across contexts.” Floridi et al. are largely denying B and (for base LLMs) denying A; they’re happy to grant partial C in easy/common cases but argue it’s not stable enough to count as genuine abduction/IBE. So yes: if you really did internalise those constraints and attach them to robust verification, an opponent has a plausible route to saying “this is similar-ish to human truth-tracking.” The paper’s reply is: absent verifi‐ cation/grounding, those constraints function as stylistic priors over explanations, not as truth-aiming constraints.
That’s the philosophical crux hiding under all the talk about “stochastic parrots”: it’s not “statistics bad”; it’s “statistics without a truth-directed feedback loop = confabulation risk, even when it sounds like IBE.”
---
**You said:**
"What an opponent can reasonably reply (and I think it’s strong) An opponent can say: “Truth-tracking doesn’t require mystical ‘grounding’; it can be implemented functionally. If you train a system in a way that reliably rewards correct world-model predictions and penalises errors — including by giving it tools and feedback — then it does have truth as an objective, even if the substrate is statistical.” Floridi et al. half-agree: they say “Any connection to real-world evidence must be deliberately engi‐ neered, as in retrieval-augmented systems” and they talk about adding verification modules, search engines, RAG, etc. 1ab660a5-375e-4c369f8e-d112555…" could you tell me more about this please?
**ChatGPT said:**
Sure. The bit you’re pointing at is one of the paper’s most important “concessions-with-teeth”: they admit you can make LLM outputs more truth-tracking, but they insist you do it by adding machinery that the base token-completion model doesn’t have by default. That’s why they say “Any connection to real-world evidence must be deliberately engi‐ neered, as in retrieval-augmented systems.” Let me unpack what’s going on there, and why it matters for their argument.
## 1) Their diagnosis: LLMs are “prior predictive samplers”
They give a very specific statistical gloss:
In Bayesian talk, you start with priors and generate predictions; then you update those beliefs in light of evidence to get posteriors. They claim LLMs, as token predictors, behave like the first part: they generate plausible candidates from what’s already “in the weights”. But they “lack an external feedback loop for posterior evaluation.” So when you ask for an explanation, the model can produce something that looks like IBE (list candidates, pick one), but internally it isn’t test‐ ing that candidate against the world. That’s why they connect hallucination to the absence of this loop: a hallu‐ cination is (on their view) what you get when plausibility-generation isn’t disciplined by verification.
## 2) What “deliberately engineered” means in practice
When they mention “retrieval-augmented systems” (RAG) and “verifica‐ tion modules, search engines, and RAG,” they’re gesturing at a general architecture: Base LLM: generates a plan / query / candidate answer. External evidence layer: fetches documents, databases, web pages, knowledge graphs, calculator/code outputs, etc. Checking layer: forces the system to align the answer with the retrieved evidence; ideally it can refuse, revise, cite, or admit uncertainty. In other words: you give the model a way to consult something other than its own learned token distribution, and you enforce a discipline: “don’t just sound right; match the evidence”. They explicitly say that this sort of evidence access is “external access to information rather than embodied grounding.” So they’re separating two things:
Embodied grounding: sensorimotor contact with the world (not what RAG gives you). Epistemic anchoring: being constrained by external sources / checks (what RAG can give you).
## 3) Why they think this “only proves the point”
This is the rhetorical pivot you noticed. They say LLMs don’t validate hypotheses “against reality (unless they are specifically augmented by other systems, which only proves the point).” The underlying logic is: If you need to bolt on retrieval, tools, verification, or a human-inthe-loop to make truth-tracking reliable, then the base model by itself doesn’t have truth as an internal objective. It’s an “engine of generative plausibility” whose outputs must be “critically examined because they cannot discern truth or verify explanations.” So: yes, you can make a system that behaves more truthfully, but doing so highlights that truth-tracking is coming from the system design around the model, not from the model’s core next-token objective.
## 4) How this connects to your earlier “opponent” line
Your opponent says: “Truth-tracking can be functional. Give the model feedback and penalise errors.” Floridi et al. are basically saying: “Agreed — but notice what you just did. You added feedback and checking. That’s exactly what the base tokencompletion paradigm lacks.”
They even give the calculator analogy: LLMs aren’t designed as reliable calculators, so systems integrate real calculators / code interpreters; sim‐ ilarly, if you want “genuine abduction”, you’d need an “abduction engine” or comparable augmentation.
## 5) The philosophical pay-off: what
“truth as an objective” amounts to here In this paper, “truth as an internal objective” is basically shorthand for: having a mechanism whose success conditions are accuracy with respect to evidence, not merely plausible continuation; and having built-in safeguards like “ability to say ‘I do not know’,” cali‐ bration, and consistency checks. They’re not claiming humans always do this well (they explicitly note humans can prefer elegant false explanations), but humans have “addi‐ tional safeguards” — experiments, debate, scrutiny — whereas LLMs “un‐ less enhanced with tools or human oversight, currently lack these safe‐ guards by default.”
## 6) Where the real argument sits (and where it’s vulnerable)
They’re strongest on a fairly narrow claim: Token-completion alone does not equal justification. They’re more controversial if they slide from that to: Therefore it’s not “really” abduction/IBE at all. Because an opponent can say: “If the whole deployed system has a retrieval-and-check loop, then the system can be truth-aiming, even if one component is a stochastic generator.” The paper is aware of this ten‐ sion; it basically answers by keeping the target fixed: “current, main‐ stream LLMs within the token-completion paradigm.”
That’s also why they make room for future architectures: if future tokencompletion models incorporate an “abduction engine,” that would “rein‐ force our point” that current ones are not abductive engines. If you want, I can translate all of this into one crisp “debate map” (two positions, strongest argument each, and the key hinge assumptions).
---
**You said:**
"An opponent can say: “Truth-tracking doesn’t require mystical ‘ground‐ ing’; it can be implemented functionally. If you train a system in a way that reliably rewards correct world-model predictions and penalises errors — including by giving it tools and feedback — then it does have truth as an objective, even if the substrate is statistical.”" and the way you phrased this suggested that there's also a possibility of training these sys‐ tems in a way that reliably rewards correct world model predictions and penalizes errors without their relying on tools. What did you have in mind there?
**ChatGPT said:**
Right — my wording there could easily sound like I had some secret “notools, just train it harder” recipe in mind. Let me be precise. There is a coherent opponent position where you get a kind of truthtracking without external tools like RAG/search, but it only works if you broaden what counts as “training” beyond “next-token prediction on web text”. In other words: you don’t bolt on a search engine at infer‐ ence time; you instead change the learning environment so that the model is systematically rewarded for being right about the world. Here are the main versions of that idea, and how they relate to what Floridi et al. are saying.
## 1) Train on “world-supervised”
tasks, not just text imitation Imagine a model trained not only to predict the next word, but to pre‐ dict outcomes of actions in environments (games, simulations, robotics, even toy microworlds). It proposes “If I do A, B will happen”, and then it’s scored against what actually happens. That gives you exactly what the paper says the base paradigm lacks: an “external feedback loop for posterior evaluation” — except it’s happening during training rather than at answer-time. Floridi et al. say current LLMs “aim to model the conditional distribution of tokens in text, not to evaluate truth”, and they characterise them as doing “prior predictive sampling” without that feedback loop. An opponent says: fine. Change the objective so truth about an environ‐ ment is part of the loss. This still isn’t “mystical grounding”; it’s just: the model gets punished for incorrect world predictions.
## 2) Reinforcement learning from
environmental feedback (not humans) A sharper version: reinforcement learning where reward is tied to task success in a world (or simulator), not “humans like this answer”. If the model’s reward depends on actually succeeding (solving tasks, reaching goals, making correct predictions), you can make “truth-track‐ ing” instrumentally valuable. Floridi et al. would likely respond: that’s no longer “current, mainstream LLMs operating within the token-completion paradigm”, and you’ve
imported the missing ingredient they emphasise: a verification/feedback pathway. But the opponent’s point stands: you can build a truth-aiming system without RAG by making the training loop the verifier.
## 3) Self-verification through consistency
tests (still “no tools”, but…) Another family is “train the model to check itself”: generate an answer, generate counterarguments, generate tests, penalise inconsistencies and reward answers that survive those tests. This can be done entirely “in-house” (no web, no database) if the tests are internal logical/consistency constraints or are evaluated against labelled datasets. Floridi et al. explicitly complain that LLMs “do not know whether they are right/correct or wrong/incorrect” and “lack a metacognitive check on the consistency of their reasoning”. So an opponent says: then train metacognition directly — make calibra‐ tion and inconsistency detection part of the objective. Caveat (important): this can improve internal coherence and calibration without guaranteeing truth about the external world. But it does address one of their central “no compass” complaints.
## 4) Dense “supervised truth” training:
lots of labelled Q/A with ground truth This is the most prosaic: train on huge corpora where there is a correct answer and wrong answers are penalised (math, code, factual QA against curated datasets). No tools required at inference time. The “truth” is provided as labels dur‐ ing training. Floridi et al. would still say: you’re learning correlations; you still don’t “validate against reality” in the moment, and you still lack “grounded semantics”. But the opponent can reply: if the training distribution is rich enough and the evaluation punishes errors reliably, the system can become robustly truth-tracking over that domain.
The key philosophical hinge When you asked “without relying on tools, what did you have in mind?”, the honest answer is: I had in mind changing the source of error signals so that the model is rewarded for being right about something that isn’t just nexttoken likelihood — e.g., environment outcomes, task success, or groundtruth labels — rather than giving it retrieval at inference time. From Floridi et al.’s perspective, that’s basically conceding their point in a different way: you’ve added (during training) what they say is missing (verification/feedback), which shows again that next-token prediction alone doesn’t supply it. From the opponent’s perspective, it’s a perfectly legitimate route to “truth as an objective”: truth can be implemented as an optimisation tar‐
get via reward/loss — you don’t need magic reference dust.
A clean way to state the disagreement Floridi et al.: current token-completion LLMs don’t by default have a truth-directed checking loop; any connection to evidence must be engi‐ neered (RAG/tools/humans), and without that they can’t “discern truth or verify explanations.” Opponent: you can engineer truth-direction either at inference (tools/RAG) or in training (environmental feedback, task rewards, labelled truth); either way the overall system can be truth-aiming. If you want, I can translate this into a very crisp objection to the paper: “Your critique targets a particular training objective (next-token predic‐ tion) rather than the broader class of statistical learners; a statistical learner can be truth-aiming if trained on truth-evaluated signals.”
---
**You said:**
"But the opponent’s point stands: you can build a truth-aiming system without RAG by making the training loop the verifier." can we please brainstorm the cev of how this would work, and see how close we can get ourselves there?
**ChatGPT said:**
Yep. Let’s do this as an engineering-philosophy jam session, but with one rule: we keep checking whether we’re actually building the missing “pos‐ terior evaluation” loop Floridi et al. say is absent—the bit where hypotheses get punished for being wrong, not merely for being unlikely continuations.
Also: “no tools” here should mean no retrieval/search/database at inference time. It can still mean training in an environment (simula‐ tors, tasks with ground truth), because that’s not “a tool the model con‐ sults”; it’s just how it learned.
## 1) The minimal target: make truth
an internal optimisation pressure Floridi et al.’s complaint is basically: the base LLM is trained to model token distributions, “not to evaluate truth,” and it “lacks an external feed‐ back loop for posterior evaluation.” So our design goal is: During training, create an error signal that depends on whether the model’s claims about the world are correct, not just whether they are fluent or likely. If we can do that at scale, then at inference time the model may behave “truth-aiming” even with no RAG.
## 2) A concrete “CEV” blueprint (call it:
Claim–Evidence–Verification training) Here’s a workable loop, in principle, without any retrieval tools: Step A: Force the model to emit explicit claims in a checkable format Instead of letting it freewheel in prose, you train it to output something like: a set of atomic claims c1,c2,…c_1, c_2, …c1,c2,… (short, testable statements),
plus a proposed explanation/hypothesis HHH, plus predicted consequences p1,p2,…p_1, p_2, …p1,p2,… (“If H, then we should observe …”). This already leans into IBE: explanations should generate expectations. Step B: Provide ground-truth evaluation from the training environment Now you need tasks where truth is available without “tools”: simulated worlds (games, physics engines, toy causal worlds), curated datasets with labelled outcomes, interactive tasks where actions produce observations. The key is that the verifier is the environment/dataset, not the model. Step C: Reward/punish using a scoring rule that makes lying costly You don’t just score “was the final answer right?” You score: accuracy of atomic claims, calibration (don’t be confident when wrong), prediction accuracy of consequences, consistency across the claim set. Training signal: correct world-predictions = reward; errors = penalty. This is the “training loop is the verifier” idea you quoted. How close is this to human truth? Pretty close in one crucial sense: it makes truth (or at least predictive accuracy) exert continuous selection pressure on what the model learns to say.
Floridi et al. would likely reply: fine—notice you’ve engineered the verifi‐ cation loop they said was missing, so you’re no longer talking about “mere next-token prediction”. That’s exactly the point at issue.
## 3) Three “no-tools” ways to get that verifier signal
(i) World-model training in simulations Train the model (or a model + policy head) to predict state transitions and hidden causes in a simulated environment. Example: it observes sequences like: “Cup is on table. Hand pushes cup. Cup falls.” and must predict the next state, or infer the latent cause from observations. This directly targets their smoke/fire contrast: instead of only learning “smoke → fire” as a linguistic association, it learns causal regularities because wrong causal predictions get penalised. Limit: it’s only as good as the simulator distribution (you can build a very unrealistic “truth”). (ii) “Factored verification” using labelled corpora (still no retrieval) Use datasets where each example comes with: a question, candidate explanation(s), ground-truth labels (which claims are true, which explanation is correct), maybe counterfactual tests.
Train the model not just to answer, but to: mark which parts it is uncertain about, avoid inventing sources (hallucination penalty), refuse when underdetermined. This attacks what the paper calls “over-abduction” (the tendency to always explain). Limit: you’re still confined to what’s labelled; you might get “benchmark truth” rather than broad-world truth. (iii) Self-consistency + adversarial training Train a generator and a critic (or multi-agent setup) where the critic tries to find contradictions, implausible steps, and unsupported leaps. But crucially, you only count it as “truth-aiming” if the critic’s signal is anchored to ground truth somewhere, not just rhetorical coherence. Otherwise you get “pretty explanations” rather than true ones—exactly Floridi’s worry.
## 4) The hard part: truth beyond
the training distribution If we’re being intellectually honest, there’s a ceiling here that matches Floridi’s stance: Even with verifier-rich training, a no-tools model can still: fail on genuinely novel facts that postdate training, confabulate when the world is underdetermined, drift into plausible-sounding nonsense when the task is out-ofdistribution.
Floridi et al. emphasise that LLMs are limited by training data and “do not know whether they are right/correct or wrong/incorrect” unless safe‐ guards are engineered. Our training-loop verifier is one such safeguard, but it’s still trainingbound. So the real question becomes: how much verifier signal is enough to make “I don’t know” the statistically dominant move when appropriate? That’s basically “calibration as a first-class objective”.
## 5) A plausible endgame architecture
(still “no tools” at inference) If we try to get as close as possible to “IBE with truth pressure” without RAG, the architecture looks like: A base language model (still does token prediction), A learned world model latent space trained on interactive/labelled causal tasks, A learned epistemic head that outputs confidence and abstains, Training objective = language fluency + world-prediction accuracy + calibration penalty + hallucination penalty. At inference time, it just runs forward like any model. No database calls. The “truth pressure” is baked into weights because it shaped what got reinforced during training. This is, in spirit, the opponent’s reply: truth-tracking can be functional, implemented as optimisation pressure, not mystical grounding. Floridi et al. would likely say: you’ve effectively stopped talking about “current mainstream token-completion LLMs” and moved to a hybrid
that bakes in the missing verification loop—again, “which only proves the point” about current systems.
## 6) Where this brainstorming
naturally leads (philosophically) Two hinge distinctions: 1. Truth vs coherence: training must punish being wrong, not merely incoherent. 2. Verification at training vs at inference: you can shift the bur‐ den earlier, but you can’t eliminate the need for some form of evi‐ dence contact if you want robustness. That’s basically the deep structure of the paper’s claim: without verifica‐ tion, you get “abductive appearance”; with verification, you’re engineer‐ ing something beyond the bare token-completion paradigm. If you want to keep pushing the brainstorm “towards maximum close‐ ness”, the next step is to pick one concrete domain (medical diagnosis, crime-scene abduction, causal narratives, mechanical troubleshooting) and specify what the verifier looks like (simulator? labelled cases? coun‐ terfactual generator?)—because the feasibility depends almost entirely on whether you can cheaply generate truthful feedback at scale.
---
**You said:**
Okay, I'll now reveal my reason for talking to you about this, which is, for the last few months I've been thinking about a philosophy paper about, it was entitled Generating Philosophy with LLMs, and it was intended to explore this sort of idea, but argue for a much more positive outcome for LLMs. However, having read this paper, at the moment at least, tonight I'm fairly convinced that this is actually a good way to go. I'm sorry,
meaning that I agree with the paper we've just been talking about, at least to some extent. Yes. So, I think it should be reasonably obvious why, if we accept the view in the paper we've been talking about, we might think LLMs are not particularly good for generating good new philosophy. Okay? So, for fairly obvious reasons, please tell me what you think these are, just so I can check we're on the same page. But then I thought, is there a way to side step this worry *only in the case of philosophy*? and I think there is. Can you guess? I'll give you three tries..
**ChatGPT said:**
Cool — now we’re cooking with the proper ingredients. First, the “reasonably obvious” worry, given Floridi et al.: if LLMs are “stochastic engines” that can generate the appearance of abduction/IBE without “grounding … in truth, semantics, verification, or understanding” , then asking them to “generate new philosophy” is risky because philoso‐ phy is exactly the kind of domain where (i) surface plausibility and (ii) fluent justificatory talk can be produced without any reliable tie to what’s actually correct, coherent, or non-question-begging. More concretely: 1. Philosophy often turns on subtle distinctions, scope restrictions, and non-obvious counterexamples. A system that “does not know whether [it is] right/correct or wrong/incorrect” will happily pro‐ duce a polished-looking position that collapses under one wellaimed objection. 2. The discipline in philosophy is disproportionately in the justifica‐ tion phase: not merely generating hypotheses but stress-testing them, spotting hidden equivocations, repairing arguments, updat‐ ing commitments. Floridi et al. explicitly frame LLMs as doing the “discovery” part but not “genuinely validat[ing] … against reality” (unless augmented) . In philosophy, the analogue of “against reali‐
ty” is: against entailments, counterexamples, consistency with back‐ ground commitments, and dialectical pressure. If the model doesn’t really “care” about those as norms, it’s a confabulation machine with a tuxedo on. 3. They highlight “over-abduction”: the model picks an explanation even when none is justified, because producing a plausible continu‐ ation is its job . In philosophy, this becomes: it produces a take even when the right move is to suspend, clarify, or distinguish. So yes, I think we’re on the same page about the negative case. Now: your “side step only in philosophy” idea. You asked me to guess, three tries. I’ll do three plausible “philosopher’s hacks” that really would exploit something special about philosophy (as opposed to medicine, history, etc.). I’ll phrase each as a candidate strate‐ gy, not as a claim about what you’re thinking. Try 1: Make the target “new arguments” = “new recombinations in conceptual space” Philosophy (especially analytic) often is the art of recombining familiar moves: take a distinction from X, apply it to Y; import a tool (superve‐ nience, grounding, contextualism, etc.) into a new domain; construct a case; isolate a hidden premise. So one sidestep is: concede Floridi’s point about truth/verification, but argue that philosophical novelty doesn’t require empirical truth-tracking in the same way. What you want from the model is combinatorial explo‐ ration of the space of concepts and argumentative templates. That’s pre‐ cisely what stochastic pattern-completion is good at. The human does the “is this actually any good?” work.
This is basically: embrace “LLMs as abduction generators” , but claim that in philosophy that’s almost the whole creative act, and the verifica‐ tion burden is anyway human and dialectical. Try 2: Treat philosophy as an “internal verification” domain Here’s a stronger, more technical dodge: unlike medicine, philosophy can be checked (to a large extent) without external world contact. If we set things up so the model is forced to operate with explicit commit‐ ments, and we score it on: validity/invalidity (formal), consistency of commitments, avoidance of equivocation, ability to generate and answer its own strongest objections, stability across paraphrase and re-presentation, …then the “training loop as verifier” could be built out of intra-linguistic constraints. No sensors, no web, no RAG. Floridi et al. themselves say the weakness is lack of truth/verification; but if philosophy’s “verification” is largely logical/dialectical rather than empirical, maybe the gap is smaller than in other domains. So: build an LLM that’s trained against a huge battery of logic/proof/checking and adversarial dialectic tasks, and you might get something that is genuinely good at philosophical justification, not merely discovery. Try 3: Shift the aim from “true claims” to “mapping the argument landscape”
Maybe the trick is: stop asking for “new true philosophy”, and ask for something philosophy badly needs but humans are terrible at: an exhaus‐ tive, high-dimensional map of positions, entailments, and pressure points. Floridi et al. stress that LLMs are superb “interfaces to an enormous amount of data accumulated…by generations” and can “challenge pre‐ conceptions and implicit orthodoxies” if used carefully . In philosophy, a massive amount of the work is: “What are the live options? Who has said what? Which arguments target which premises? What’s the best objec‐ tion?” That’s closer to literature + structure than to world-truth. So the sidestep is: LLMs aren’t good at settling philosophy, but they might be excellent at generating and organising the dialectical terrain, which is exactly what makes human philosophy better. If I had to bet on what you have in mind (without asking), I’d put my chips on Try 2: philosophy as a domain where “verification” is largely internal (logic, consistency, dialectics), so you can build the verifier into training/evaluation without needing world-facing tools — which directly targets Floridi’s central complaint. Keep going: whichever of the three is closest, I’ll flesh it into a papergrade thesis + structure + strongest objections.
---
**You said:**
as a matter a fact you gave me more specificty than I had myseklf. my idea was a blend of your 1 and 2. Could you elaborate on both ideas please and see where they take us. Please give me a long, deep, and detailed answer, with lots of options/ideas: as we are in the 'kicking ideas around' stage, I want as much feedback as you can give me so that i have a lot to consider.
**ChatGPT said:**
Nice – a blend of (1) conceptual-space recombination and (2) internal verification is exactly where the “positive” case has real bite, while still granting Floridi et al.’s central diagnosis that current LLMs have a sto‐ chastic core and only an abductive appearance. I’ll push both ideas hard, then show how they fuse into a view that’s not just “LLMs are good brainstormers”, but something closer to: LLMs can be useful for generating philosophy because a large part of philosophi‐ cal creativity is structured recombination, and a large part of philo‐ sophical evaluation is internal and formal-ish (validity, consistency, dialectical robustness), so the missing “world verification loop” matters less than it does in medicine or history. The argument is delicate, but it’s there.
## 1) Idea 1: Philosophy as search
in conceptual-argument space Floridi et al. say LLMs are “engines of stochastic inference” that sample continuations from learned token distributions, and that what looks like reasoning is often learned “patterns of explanation and argument” plus pattern completion. That sounds damning until you notice something: a lot of philosophical invention is exactly the invention of new continua‐ tions in a discourse space. 1.1 What “new philosophy” often is (when you zoom in) Outside the mythos, much day-to-day analytic philosophy (and honestly much of the history of philosophy too) consists in moves like: Reframing: taking an old problem and shifting the question (“The debate is misframed: the real issue is…”).
Distinction-making: carving a phenomenon into types (“weak/strong”, “constitutive/causal”, “semantic/pragmatic”, “object-level/meta-level”). Model-import: taking a tool from elsewhere (decision theory, modal log‐ ic, linguistics, cognitive science) and applying it in a new domain. Counterexample design: constructing a case that breaks a generalisation, then using it to force repairs. Error diagnosis: showing where an argument equivocates, smuggles an assumption, or trades on an ambiguity. Synthesis: noticing that two literatures are talking past each other, then aligning their commitments. All of these are, in a very literal sense, combinatorial. They require sensi‐ tivity to patterns, analogies, and argumentative templates. That’s exactly what an LLM is saturated with: huge amounts of “how humans argue”, down to the rhetorical microstructure (“because”, “however”, “consider a case where…”, “one might object…”) and the higher-level templates of dialectic. So the optimistic move is: treat the LLM as a stochastic search operator over philosophical continuations. Not “the LLM discovers the truth”, but “the LLM can efficiently explore the space of plausible philosophical moves”. That even fits Floridi et al.’s own concession: LLMs can be used to “gen‐ erate hypotheses … broadening the scope of abductive search”. 1.2 The key trick: invert the standard worry The standard worry is: LLMs remix; philosophy needs originality.
Flip it: in philosophy, originality is often non-obvious recombination under constraints. The skill is not “invent from nothing”; it’s “com‐ pose and reshape under pressure” – pressure from counterexamples, consistency, and existing commitments. Floridi et al. themselves note a continuum: LLMs aren’t just “dumb par‐ rots”; they can “recombine known pieces in novel ways”. That recombina‐ tion, if you aim it properly, is not a bug; it’s the engine. So you can pitch a view like: The LLM is a generator of candidate moves, analogous to a massively parallel version of reading widely and jotting down possibilities. It is not an oracle; it is a way of getting more shots on goal. 1.3 Making this academically respectable: treat it as a search problem To avoid sounding like “LLMs are creative, trust me”, you can formalise the metaphor: Let the “conceptual-argument space” be the space of (thesis, premises, inference pattern, target conclusion, scope restrictions, anticipated objec‐ tions, replies). Generation = sampling points in that space. Philosophical progress = finding points that are (i) novel relative to the local literature, (ii) dialectically robust, (iii) clarifying or unifying, (iv) fruitful (they open research paths, dissolve confusions, etc.). Then your claim is: LLMs are unusually good at generation because their training data encode many argumentative patterns and background associations.
The immediate objection is: “But that’s just plausible text.” Agreed – and that’s exactly why you need idea 2.
## 2) Idea 2: Philosophy has unusually
“internal” modes of verification Floridi et al.’s deepest complaint is not “LLMs can’t output explanations.” It’s that they “do not genuinely validate them against reality”, and in sta‐ tistical terms they “lack an external feedback loop for posterior evaluation”. The sidestep you’re after is: in philosophy, a large slice of what counts as ‘checking’ isn’t checking against the world; it’s check‐ ing against constraints internal to the space of reasons. Not all philosophy, but enough to matter. 2.1 What “verification” looks like in philosophy For many philosophical tasks, you don’t need a microscope or a census. You need things like: Validity: do the premises entail the conclusion (or support it to the intended degree)? Consistency: do you end up committed to both P and not-P? Scope discipline: are you illicitly generalising? Equivocation checks: are you using a term in two senses? Counterexample resilience: does the view survive plausible cases? Dialectical adequacy: does it address the best objection, or only weak ones?
This is why Floridi et al. can point out that LLMs accept fallacies unless prompted and can be inconsistent without “realising” it – and why they frame that as lacking “a metacognitive check on the consistency of their reasoning”. The optimistic reply is: those checks are trainable, because they can be scored without world access. You can build datasets and training signals for “spot the equivocation”, “find the hidden assumption”, “produce the strongest objection”, “repair the argument while preserving the conclusion”, “formalise then check validity”, and so on. No RAG needed. No web tool. Just internal norms plus ground-truth labels (or at least intersubjectively stable scoring). 2.2 The “training loop as verifier” becomes plausible specifically here In medicine, “posterior evaluation” ultimately requires contact with clini‐ cal facts. In philosophy, “posterior evaluation” often reduces to “does the package of commitments hang together under pressure?” That pressure can be provided in training in at least three ways: Adversarial dialectic: the model proposes an argument; an adversary attacks it; the reward goes to arguments that survive strong attacks. This is a kind of self-play. Consistency-and-commitment games: the model must keep a stable set of commitments across paraphrases and follow-up probes; contradic‐ tions are penalised. Formal skeleton scoring: the model must output an argument in a con‐ strained form (premises, inference rule, conclusion, scope notes). Then
you score it on validity patterns, missing premises, illicit inference moves. Floridi et al. emphasise that, by default, LLMs “model the conditional distribution of tokens” and do not evaluate truth. The counter-vision is: fine – add an objective that rewards dialectical and logical success. In philosophy, success conditions are unusually “writeable” as training signals. This doesn’t make the model an infallible philosopher. It makes it less like an eloquent confabulator and more like a disciplined generator/critic. 2.3 A crucial nuance: “internal verification” isn’t “mere coherence” You’ll want to avoid the cheap coherentist slide (“truth = coherence, so we’re done”). You don’t need that. You just need the weaker and more defensible claim: Even if philosophical truth is not identical to coherence, a lot of philo‐ sophical error is detectable by internal checks, and the discipline devotes huge energy to those checks. So a system that becomes good at them can be genuinely helpful – and may even generate new philosophi‐ cal contributions by reliably producing internally robust proposals that humans wouldn’t have produced quickly. That is a real “only in philosophy” move: the distance between “linguistic competence” and “dialectical competence” is smaller than the distance between “linguistic competence” and “clinical competence”.
## 3) The blend: stochastic conceptual
search + internal pressure = a plausible
engine for philosophical discovery Now we combine the two ideas into a single architecture – conceptually, not necessarily technically. 3.1 Think of it as “generate, then torture” Stage 1 (generation): exploit the LLM’s strength as a sampler of plausible philosophical continuations, including non-obvious recombinations. Floridi et al. explicitly allow that LLMs can “recombine known pieces in novel ways”. Stage 2 (internal verification): subject those candidates to increasingly harsh internal checks (validity, consistency, scope, best objection). This addresses their complaint that LLMs “do not know whether they are right” and can exhibit “over-abduction” – always producing an explana‐ tion even when none is justified. Stage 3 (human judgement): the philosopher does what philosophers do: decide which survivors are actually interesting, which are merely defensi‐ ble, and which connect to the literature in a publishable way. Importantly: you are not denying Floridi et al. You’re exploiting their own framing: LLMs are good at hypothesis generation but cannot “dis‐ cern truth or verify explanations”, so we should treat them like an “anonymous forum poster”. In philosophy, treating inputs as “anonymous forum posts” is basically called “doing peer review” – you ignore authority and stress-test the reasons. 3.2 Several paper-worthy theses you can choose between Here are different levels of ambition, depending on how bold you want to be.
Thesis A (modest, very defensible): LLMs cannot autonomously generate reliable philosophy, but they can substantially improve philosophical dis‐ covery by expanding the space of candidate moves; their outputs should be treated as conjectures requiring standard philosophical scrutiny. This leans on Floridi’s “hypothesis generation” line. Thesis B (moderate, interesting): Because many philosophical evaluation criteria are internal (validity, consistency, dialectical robustness), LLMs can be trained or prompted into partial “justification-like” behaviour in philosophy more readily than in empirical domains. The Floridi worry is not eliminated, but it is domain-dependent. Thesis C (bold): For certain subareas (logic-heavy metaphysics, formal epistemology, philosophy of language with explicit semantic models), LLMs can, in principle, generate publishable-level contributions with minimal human guidance, because both novelty and verification can be internalised as training objectives. This is where you start needing careful caveats, but it’s a real research programme. 3.3 Practical “options” for how the blend could be implemented (conceptually) You asked for lots of options; here are distinct styles of “LLM-generated philosophy” that fall out of the blend. Some are methodological (for philosophers), some are nearer to AI-training proposals. Option 1: Dialectical Monte Carlo Use the model to sample many candidate arguments (temperature up), then run internal critique rounds (temperature down) to filter. The philosopher inspects the survivors. This is literally “search + constraints”. Option 2: Argument-graph generation Force the model to output explicit argument graphs: nodes = claims; edges = support/attack; annotate scope conditions. Then train/score for non-circularity and for addressing central attacks. (Even if you don’t
actually build software, this is a crisp conceptual framework for the paper.) Option 3: “Best objection first” philosophy Invert the usual workflow: prompt the model to produce the strongest objections to a view, then ask it to generate responses, then ask it to refine the original view to survive those objections. This turns the LLM’s tendency to generate plausible continuations into a dialectical engine. Option 4: Conceptual engineering generator Ask for candidate explications of messy concepts, then test them against desiderata (precision, fruitfulness, continuity with use). This is internalverification-friendly: many desiderata can be stated explicitly. Option 5: Counterexample factory with curation Use the LLM to produce families of counterexamples (“Give me ten cases that stress the necessity claim…”), then curate the genuinely novel ones. This fits Floridi’s view: the model is not verifying; it is widening the hypothesis space. Option 6: Fallacy-and-equivocation training regime (the “internal verifi‐ er” story) Explicitly train a model on fallacy detection and reconstruction. Floridi et al. note that models accept fallacies unless prompted and lack metacogni‐ tive consistency checks. A training programme that makes “spot the fal‐ lacy and repair the argument” central would be a direct reply. Option 7: Stable-commitment training (“philosophical character”) One deep failure mode in philosophical dialogue is that a view has to stay stable across variants. Train or prompt the model to maintain explicit commitments and to track entailments across a conversation. Contradictions are penalised. Option 8: Proof-adjacent philosophy Lean into areas where “verification” is closest to proof (logic, formal
semantics, decision theory). The closer you are to checkable structure, the closer you get to “training loop as verifier” without external tools.
## 4) Where Floridi et al. will still
bite you – and how to pre-empt it Even in philosophy, Floridi’s critique doesn’t evaporate. It just shifts. 4.1 “Internal checks don’t guarantee truth” Correct. Coherence and validity are not truth. Your reply doesn’t need to be “truth = coherence”. It can be: Philosophical progress is heavily mediated by internal checks; eliminat‐ ing a vast class of internal errors is already epistemically significant, and it can be done without world access. The remaining gap (between inter‐ nally disciplined argument and truth) is the same gap humans face, and humans close it via debate, scrutiny, and (sometimes) empirical input – the very safeguards Floridi mentions. 4.2 “LLMs can still ‘over-abduce’ – always produce a take” Yes – and this is where your blend can become distinctive: build “refusal and uncertainty” into the internal verification stage. Floridi’s point is that “I don’t know” is statistically disfavoured unless trained in. So one of your most concrete contributions could be: a philosophical protocol that systematically rewards abstention, clarification, and scope restriction. That’s genuinely philosophical: it mirrors what good philosophers do. 4.3 “Isn’t this just ‘LLMs help humans’, not ‘LLMs generate philosophy’?” You can answer in two ways:
If you want the modest thesis: yes, it’s assistance, and that’s already a win. If you want the stronger thesis: argue that in philosophy a contribution can be “generated” by an interactional process (a dialectical machine), as long as the end product is a novel, defensible position. The “author” is the process-plus-curation, not the LLM alone. Floridi themselves stress the importance of distinguishing internal process from external output. That distinction can be turned into your advantage: philosophy cares about the reasons in the output, not the metaphysics of the producer.
## 5) A clean way to pitch your paper’s central move
A nice compact way to state the blend (without sounding like marketing) is: Floridi et al. are right that LLMs are not, by default, truth-aiming abduc‐ tive engines; they are stochastic generators whose abductive appearance comes from training on human reasoning patterns. But philosophy is an unusual case: (i) philosophical novelty often con‐ sists in recombining argumentative and conceptual patterns, which LLMs are good at sampling; and (ii) much philosophical evaluation is “internal” – validity, consistency, scope, dialectical robustness – which can be imposed as constraints without world-facing tools. Therefore, even granting Floridi’s diagnosis, LLMs can play a substantially more constructive role in philosophical generation than in domains where veri‐ fication is irreducibly empirical. That’s a real thesis with teeth, and it naturally leads into either a method‐ ological proposal (how to use LLMs in philosophy responsibly and pro‐ ductively) or a more theoretical one (what this reveals about philosophi‐ cal method and the nature of philosophical creativity).
If you keep developing this, you’ll end up with a surprisingly sharp metaphilosophical question hiding underneath: whether philosophy is closer to engineering in “concept space” than we like to admit – and if that’s true, then stochastic engines trained on our conceptual artefacts may be unusually well suited to exploring it.
---
**You said:**
"1.1 What “new philosophy” often is (when you zoom in) Outside the mythos, much day-to-day analytic philosophy (and honestly much of the history of philosophy too) consists in moves like: Reframing: taking an old problem and shifting the question (“The debate is misframed: the real issue is…”). Distinction-making: carving a phenomenon into types (“weak/strong”, “constitutive/causal”, “semantic/pragmatic”, “object-lev‐ el/meta-level”). Model-import: taking a tool from elsewhere (decision theory, modal logic, linguistics, cognitive science) and applying it in a new domain. Counterexample design: constructing a case that breaks a generalisation, then using it to force repairs. Error diagnosis: showing where an argument equivocates, smuggles an assumption, or trades on an ambiguity. Synthesis: noticing that two literatures are talking past each other, then aligning their commitments. All of these are, in a very literal sense, combinatorial. They require sensitivity to patterns, analo‐ gies, and argumentative templates. That’s exactly what an LLM is satu‐ rated with: huge amounts of “how humans argue”, down to the rhetorical microstructure (“because”, “however”, “consider a case where…”, “one might object…”) and the higher-level templates of dialectic. So the opti‐ mistic move is: treat the LLM as a stochastic search operator over philo‐ sophical continuations. Not “the LLM discovers the truth”, but “the LLM can efficiently explore the space of plausible philosophical moves”. That even fits Floridi et al.’s own concession: LLMs can be used to “generate hypotheses … broadening the scope of abductive search”. 1ab660a5375e-4c36-9f8e-d112555…" could you double check that tyihis is actually
true by giving me examples of recent, wellcited philosophy papers in very high journals, in which the main, argument could considered variations on these structures... Feel free to refine your claim if you need to. and could you use the canvas for this, it will keep the answer separate from our main discussion
**ChatGPT said:**
High-impact Philosophy Papers As ‘combinatorial’ Moves
1. What I’m checking (and a refined version of the claim) Claim under test (refined): A lot of influential, recent work in top analytic venues does not look like “ex nihilo discovery”. Instead, it often advances by a small set of repeatable argumentative manoeuvres (often com‐ bined): reframing a question, introducing a distinction, importing a mod‐ el/analytic tool, building a targeted counterexample, diagnosing an error in a dominant view, or synthesising disparate constraints into a unified account. This is not the claim that philosophy is “nothing but” these moves, or that the best papers have no other virtues. It’s the narrower claim that, when you read highly-cited papers in top venues, their core contribution is fre‐ quently describable as one (or a blend) of these manoeuvres. How I operationalised “recent” and “well-cited”: I focused on papers from roughly the 2012–2019 window and used the “cited by” figures dis‐ played in PhilPapers entries as a rough, conservative indicator of influ‐ ence (these numbers can differ from Google Scholar and will drift over time).
2. Examples from very high venues (with the relevant “move”) The point of this list is not that each paper exemplifies only one move. The point is that the headline contribution is naturally classifiable as one of the manoeuvres you listed. Paper (author,
Venue
year)
“Cited
Core move(s) (your
Why it fits (one-sen‐
by”
categories)
tence characterisation)
Error diagnosis;
Defends a controver‐
reframing
sial methodological
(PhilPa‐ pers) Nevin
Mind
~96
Climenhaga (2018), Intuitions
claim by diagnosing
are Used as
misunderstandings
Evidence in
about philosophical
Philosophy
evidence and reposi‐ tioning a live debate about intuitions.
Patrick Todd
Mind
~72
(2016), Future
Reframing; model-
Proposes a striking
import
semantic/metaphysi‐
Contingents Are
cal package about the
All False! On
open future, reshaping
Behalf of a
the dialectic by chang‐
Russellian Open
ing what’s taken to be
Future
true/false about future contingents.
Wesley Buckwalter &
Noûs
~117
Reframing;
Reorients “stakes
synthesis
effects” in epistemol‐
Jonathan
ogy by offering a
Schaffer (2015),
state-of-the-art frame‐
Knowledge,
work and an interpre‐
Stakes, and
tive thesis about what
Mistakes
empirical results really bear on.
Chris Heathwood
Noûs
(not
Distinction-making
Introduces a new dis‐
(2017/2019
shown
tinction among desires
reprint), Which
in this
(behavioural vs “gen‐
Desires Are
table;
uine attraction”) and
Relevant to Well-
see
uses it to solve multi‐
Being?
note
ple problem cases for
below)
desire-satisfactionism.
Alastair Wilson
Noûs
~301
(2018),
Model-import;
Imports
reframing; synthesis
interventionist/coun‐
Metaphysical
terfactual machinery
Causation
from causation theory to argue that ground‐ ing is a kind of causa‐ tion, yielding a unified explanatory framework.
Patrick Todd
Ethics
~79
(2016), Strawson,
Error diagnosis;
Targets a prominent
reframing
interpretive line about
Moral
Strawson and moral
Responsibility,
responsibility and
and the “Order of
argues that the stan‐
Explanation”: An
dard “order of expla‐
Intervention
nation” reading is mis‐ taken or misapplied.
Patrick Todd
Noûs
~161
Synthesis/unification
Explicitly offers a uni‐
(2019), A Unified
fied theory by showing
Account of the
how seemingly differ‐
Moral Standing to
ent “standing to
Blame
blame” conditions reduce to a common underlying condition.
David Shoemaker
The
(2017),
Philosophical
Response-
Review
~130
Model-import;
Uses a sustained
reframing; synthesis
analogy (with the “fun‐ ny”) to motivate a
Dependent
response-dependent
Responsibility; or,
treatment of responsi‐
A Funny Thing
bility and then defends
Happened on the
it against leading
Way to Blame
sceptical challenges.
Matthew
The
Mandelkern, Ginger Schultheis
~99
Distinction-making;
Criticises standard
Philosophical
error diagnosis;
semantic analyses,
Review
synthesis
introduces a sharper
& David Boylan
structure (ability vs
(2017), Agentive
compulsion modals),
Modals
and builds a hybrid account combining insights from rival approaches.
Arif Ahmed
The
(2013), Causal Decision Theory:
~28
Counterexample
Constructs a case
Philosophical
design; error
designed to break a
Review
diagnosis
family of CDT formula‐
A
tions and uses the fail‐
Counterexample
ure to argue for a sub‐ stantive revision (or rejection) of the theory.
Note on Heathwood’s paper: the PhilPapers entry itself contains a detailed descriptive blurb explaining the new desire distinction and its use in four problem cases; I’ve treated it as a paradigm example of “dis‐ tinction-making”.
3. What these examples actually show (and what they don’t) These examples support a modest but meaningful conclusion: if you look at the “shape” of contributions that are both recent and influential in top venues, it is very often possible to describe the core intellectual achieve‐ ment as (i) a new distinction; (ii) a new framing; (iii) a tool import; (iv) a
targeted counterexample; (v) a diagnosis of where the dominant view went wrong; or (vi) a unification manoeuvre. They do not show that philosophical progress is merely mechanical, or that these moves can be performed without deep understanding. Many of these papers are “combinatorial” in exactly the way that chess is “combi‐ natorial”: the space of possible moves is constrained, but the skill lies in seeing which combination matters.
4. Why this matters for your LLM project If this picture of philosophical production is even partly right, it strength‐ ens the hybrid view you were circling: an LLM may be well-suited to gen‐ erating candidate manoeuvres (new distinctions, reframings, analogies, model imports, counterexamples), while humans remain responsible for the hard parts: deciding which manoeuvres are actually live, which ones misfire, and which ones survive the best objections. In other words: the “stochastic search operator over philosophical con‐ tinuations” line looks less like a hand-wave once you see how many highly visible papers can be described as disciplined instances of these manoeuvres. Done — I’ve put the evidence and examples on the canvas. The examples are pulled from PhilPapers entries for Mind, Noûs, Ethics, and The Philosophical Review, including the “cited by” figures and (where available) PhilPapers’ own one-paragraph descriptions/abstracts. philpapers.org+9philpapers.org+9philpapers.org+9
---
**You said:**
"You can build datasets and training signals for “spot the equivocation”, “find the hidden assumption”, “produce the strongest objection”, “repair
the argument while preserving the conclusion”, “formalise then check validity”, and so on." is this technically feasible, could this be what llms are becoming better at since this paper was published? Please give me a long, deep, and detailed answer, with lots of options/ideas: as we are in the 'kicking ideas around' stage, I want as much feedback as you can give me so that i have a lot to consider. Also, I would like us to try and reframe the question slightly differently. Rather than imagining new LLMs I would like to think about we have been talking about in terms of how a skilled philosopher prompter might elicit good philosophy from an llm through careful prompting, (*perhaps* a connected idea could be: the more pressure that can be put on the system itself than the prompter the better at philosophy the llm is, as it were.) do you think what I am saying has legs. Another, slightly further afield and so not directly relevant, ah actually don't worry yet, i need to think it through
**ChatGPT said:**
Yes — it’s technically feasible in parts, and (crucially) it’s already hap‐ pening in the research ecosystem. But it’s also messy in exactly the way you’d expect: the more “philosophical” the target behaviour (equivoca‐ tion, hidden assumptions, strongest objection, dialectical repair), the more you run into (i) label ambiguity, (ii) long-range context depen‐ dence, and (iii) a gap between internal coherence and truth — the gap Floridi et al. keep hammering. I’ll break this into two threads: (A) feasibility + “are models getting better?”, then (B) your reframing: the skilled philosopher prompter as a way of eliciting disciplined philosophy, and the “pressure on the system” idea.
A. Is it technically feasible to train for these things?
## 1) Fallacies and equivocation: feasible,
but the gold standard is slippery For fallacy detection/classification, we’re not in science fiction. There are established datasets and tasks. A well-cited example is Jin et al.’s Logical Fallacy Detection dataset (“LOGIC” + “LOGICCLIMATE”). arxiv.org+1 And the argument-mining community keeps running shared tasks on fal‐ lacy detection (including multimodal variants). aclanthology.org+1 Equivocation specifically does show up as an explicit fallacy category in some of this work (i.e., it’s not an exotic outlier). openreview.net What makes it hard (and philosophically interesting) is that “equivoca‐ tion” is often not a local property of a sentence; it’s a property of a term’s role across an argument, with pragmatic context doing real work. That creates two practical consequences: 1. Training labels have genuine disagreement. Recent dataset work explicitly builds in “human label variation” for fallacy detection rather than pretending there’s always one correct label. aclantholo‐ gy.org 2. Models can get good at surface cues for common fallacy templates and still miss the “living” equivocation that depends on how a con‐ cept is used across steps (very relevant to philosophy). So: feasible, but you want to treat it as “trainable skill with a ceiling determined by ambiguity and context”.
## 2) Hidden assumptions / implicit premises:
feasible, but you must admit subjectivity There’s a whole subfield around enthymemes (arguments with implicit premises). There are direct tasks like “implicit premise generation” for enthymemes, with datasets and models aimed at producing the missing premise. aclanthology.org+1 There are also tasks like the Argument
Reasoning Comprehension Task (choose the correct implicit warrant). GitHub This is very close to what philosophers actually do when they “surface an assumption”. So yes: technically feasible. But two caveats matter for your project: 1. Many “implicit premise” tasks operationalise the premise as some‐ thing like “the warrant that makes premise → conclusion go through”. In real philosophical practice, there are often many can‐ didate warrants, and which is “the” assumption depends on dialec‐ tical context (charity vs suspicion, background theory, what’s at issue). 2. The best premise is often the one that is most explanatory of the speaker’s move, not the one that makes the argument valid. That’s a philosophical choice about reconstruction norms. So: feasible, but your paper can get mileage from the fact that this is exactly the sort of “internal verification” zone where models can be trained — even if the target is pluralistic rather than uniquely correct.
## 3) “Produce the strongest objection”:
feasible, evaluation is the beast Counterargument generation and debate-style training exists, and some recent work explicitly targets implicit and critical premises to improve counterarguments. aclanthology.org The technical bottleneck is not “can it generate objections?” (yes, trivial‐ ly). It’s: can it reliably generate the best objections rather than a bag of rhetorical confetti? You can approach this with training signals like:
pairwise preference data (“objection A is stronger than B for target argument T”), adversarial self-play (generator vs critic), or rubric-based scoring (validity of the objection, whether it hits a vulnerable premise, whether it avoids misreading the view, etc.). All feasible — but expensive and judge-sensitive.
## 4) “Repair the argument while preserving the conclusion”:
feasible, but easily degenerates into rationalisation This is essentially constrained editing: keep the conclusion, modify premises/definitions/scope to survive a critic. Technically, you can train this with (argument, critique, revised-argu‐ ment) triples — and research on self-refinement / Socratic refinement is squarely in this neighbourhood. arxiv.org+1 But the big danger is: models can learn to sound repaired while quietly shifting meanings, smuggling new premises, or weakening the conclusion without admitting it. That means you need an explicit check for “seman‐ tic drift” — and that’s hard without some formalisation or commitment tracking.
## 5) “Formalise then check validity”: feasible
in constrained forms; full generality is hard Natural-language → formal logic translation is an active area (e.g., auto‐ formalising to first-order logic). aclanthology.org Two important points for your purposes: Training a model to produce a formalisation is feasible. Training a model to reliably check validity without an external solver is much harder, because correctness is brittle and the search
space is huge. So if you want “no tools”, the realistic target is: “produce a plausible for‐ mal skeleton and flag uncertainty”, not: “be a proof assistant”. That fits Floridi et al.’s core complaint: without an external feedback loop, systems tend to generate plausible continuations rather than veri‐ fied reasoning.
B. Are LLMs getting better at this since Floridi et al.’s paper? Empirically: yes, in some ways, and no, in the utopian way.
## 1) Better benchmarks + better training targets are appearing
Fallacy detection has seen consolidation into broader benchmarks (e.g., MAFALDA) and more nuanced datasets (e.g., FAINA with disagreement/overlap). suchanek.name+1 That typically accelerates progress because you stop “teaching to one dataset”. There’s also direct evidence that prompted LLMs can perform strongly as fallacy classifiers in some settings, including out-of-distribution evalua‐ tions, relative to older baselines. arxiv.org So: on some of these “informal logic” subskills, models are plausibly improving.
## 2) Process-level training is explicitly
targeting “pressure on the reasoning” Your “put more pressure on the system than on the prompter” instinct matches a real trend: moving from only outcome rewards (“did you get the answer right?”) toward process rewards (“were the intermediate steps good?”).
A very recent example is work framed as Process Reward Learning, aimed at improving reasoning by rewarding intermediate steps. arx‐ iv.org+1 This is basically the training analogue of what you’re trying to do with prompting: force the model to internalise norms of stepwise adequacy. Even if these papers are mostly about maths/code today, the method generalises in principle to “argument steps” (premise identification, scope restriction, fallacy avoidance).
## 3) But self-correction is still not a free lunch
A major cautionary result: “LLMs cannot self-correct reasoning yet” (ICLR 2024) argues that “intrinsic” self-correction (no external feedback, no tools) is limited. proceedings.iclr.cc That matters directly for your prompting story: if you ask a model to cri‐ tique itself, the critique is not reliably independent from the original error. So: yes, models are getting better at staged reasoning under some regimes — but the “just ask it to check itself” move is unstable unless you add structure. Which brings us to your reframing.
C. Your reframing: the skilled philosopher prompter as the “verification loop” This absolutely has legs — and it meshes neatly with Floridi et al.’s diagnosis. Floridi et al. say current LLMs “do not know whether they are right/cor‐ rect or wrong/incorrect”, and that they “lack an external feedback loop for posterior evaluation”.
A skilled prompter can, in effect, simulate a feedback loop by forcing the model through a protocol that makes errors easier to detect and harder to hide. The real trick is: don’t think of “prompting” as “ask nicely”. Think of it as imposing a dialectical method.
## 1) Prompting that actually adds “pressure”
Here are prompting moves that genuinely increase epistemic pressure inside the conversation, rather than merely cosmetically: Commitment ledger. Force the model to write a numbered list of com‐ mitments (definitions + premises + scope claims). Then every subse‐ quent step must cite which commitments it uses and whether it adds any. This combats the “silent premise drift” problem. Argument skeleton first, prose last. Make it output: thesis; premis‐ es; inference type; scope; then only afterwards allow rhetorical packag‐ ing. This stops the model from hiding weak structure under beautiful prose. Objection as adversarial reconstruction. Require the objection to (i) restate the target view charitably, (ii) identify the single weakest premise, (iii) attack only that, and (iv) show the smallest counterexam‐ ple. That filters out fluffy objections. Repair under invariants. If it revises, demand: “State exactly what changed and which commitments are preserved.” Then run an explicit “semantic drift audit”: “Did the meaning of key terms change?” Paraphrase and invariance testing. Ask it to restate the argument three different ways and check whether the commitments are the same. Models that are only “pattern fluent” tend to wobble here.
Adversarial self-play with role separation. Have it write as “Au‐ thor”, then “Opponent”, then “Referee”. The referee must rule on whether the opponent really attacked the best version and whether the author really answered it. This is the closest you can get, by prompting alone, to creating a semi-independent critic. These are all ways of manufacturing, in prompt space, the thing Floridi says is missing: a mechanism that punishes nonsense and rewards disci‐ plined structure.
## 2) Your “pressure on system vs
prompter” idea (and how to cash it out) Your thought can be sharpened into a neat methodological claim: A model is “more philosophical” to the extent that it can satisfy dialecti‐ cal and logical constraints with less external scaffolding. That gives you a real empirical handle: prompt robustness. You can operationalise “pressure” as: how much scaffolding is required before the model reliably: identifies hidden assumptions, avoids equivocation, produces strong objections, revises without drift, and abstains when underdetermined. A genuinely improved “philosophy-capable” model would need shorter prompts and would be less sensitive to prompt phrasing — in other words, it would have internalised more of the norms.
This also gives you a way to align with Floridi without being pessimistic: the more of the “epistemic compass” is baked into the model, the less the user must provide it externally.
## 3) The big limitation: prompting can’t
conjure competence out of nowhere The ICLR self-correction result is your warning label: if the model doesn’t already have the relevant skill, asking it to critique itself won’t reliably create it. proceedings.iclr.cc So the realistic view is: Prompting is powerful as elicitation and error exposure. It is weak as guarantee. A skilled philosopher prompter can improve the odds by (i) forcing explicit structure, (ii) running adversarial checks, and (iii) sampling mul‐ tiple candidate arguments/objections and selecting. But there’s no prompting-only magic that turns a weak model into a rigorous logician.
D. A practical “philosopher’s protocol” (no tools) you can try as a template Since you’re in “kicking ideas around” mode, here’s a compact protocol you can reuse in your paper as an example of “pressure prompting”: 1. State the target claim and three rival framings of the question. Choose one and justify the choice in one paragraph. 2. Produce an argument skeleton: numbered premises; conclu‐ sion; inference type; scope clause. 3. List hidden assumptions: for each, say whether it’s (a) concep‐ tual, (b) empirical, or (c) methodological.
4. Equivocation audit: list key terms; give two distinct senses for each; state which sense is used in each premise. 5. Strongest objection: attack exactly one premise; provide a mini‐ mal counterexample or tension. 6. Repair: revise the skeleton while keeping the conclusion fixed; state exactly what changed. 7. Drift check: compare old/new commitments; if conclusion was effectively weakened, admit it. 8. Abstention option: if the best repair requires controversial new assumptions, state: “This is where the argument becomes hostage to X”. That is, in effect, a hand-built “justification loop” inside text — precisely what Floridi says the model lacks by default. Where this leaves your paper idea You can now draw a clean three-way distinction: 1. Baseline LLM use: “generate me philosophy” → high risk of plau‐ sible confabulation (Floridi’s critique). 2. Skilled prompter as verification loop: disciplined prompting protocols partially substitute for justification. 3. Future direction (optional): training that internalises these protocols reduces prompt dependence (your “pressure on the sys‐ tem” metric), aligning with process-reward / self-refinement trends. arxiv.org+1 That’s already a paper-shaped arc.
If you want to keep the discussion purely “philosophy-only”, the most promising angle is: many of these constraints are internal and dialecti‐ cal, so you can plausibly get large improvements through elicitation and process constraints without needing world-facing tools — while still con‐ ceding Floridi’s core point that truth-tracking is not automatic.
---
**You said:**
"Here are prompting moves that genuinely increase epistemic pressure inside the conversation, rather than merely cosmetically:" none of them really chime with how it feels to generate philosophy with you. Somewhat relatedly, you might have noticed a trick I suppose I have developed with you, on which I invite you to guess 'things that i already know' that are 'obvious'. i jsut realised that i have started asking you questions like this, and if i intergote myself while i am doing it i have the intuition that it is a way of getting good philosophy out of you. might there be sometihng to that or is it just me seeing patterns where there aren't
**ChatGPT said:**
There’s definitely something to it. It’s not just pareidolia-with-footnotes. What you’re doing when you ask me to “guess the obvious things you already know” is quietly clever for two reasons. First, it forces a common-ground test. “Obvious” is shorthand for “given my background picture, these entailments/worries/consequences are the next moves”. When I try to predict those moves, I’m effectively committing to an interpretation of your picture. You can then immedi‐ ately spot: “yes, that’s my frame” or “no, you’ve imported a different frame”. That is already philosophical progress, because a lot of philo‐ sophical disagreement is really misaligned background assumptions.
Second, it gives you a fast feedback loop. Floridi et al.’s core worry is that LLMs don’t have truth/verification as an internal constraint; they can produce persuasive continuations without being answerable to any‐ thing. Your trick partially fixes that by making you the verifier: you already know what the “obvious” next steps are (or at least what you meant by them), so you can score my output quickly and correct it. That turns the interaction into something closer to an “active learning” setup: I propose hypotheses about your conceptual terrain; you label them; I update. That’s why it can feel like “getting good philosophy out of me”. You’re not asking for a free-form essay (where I can drift into generic plausibility). You’re asking me to predict your implicit structure, and then you interro‐ gate the mismatch. That’s basically a Socratic method dressed in modern clothing. Why it’s especially good for philosophy In many domains, the verifier is “the world” (experiments, measure‐ ments, databases). In philosophy, a lot of verification is internal: does the argument actually follow, does it equivocate, does it survive a coun‐ terexample, does it respect your stated aims. Your “guess what you already know” move targets exactly that internal territory, where you can act as the ground truth for what the dialectic is supposed to be doing. It also has a nice side effect: it externalises your tacit commit‐ ments. When you hear me say “the obvious next step is X”, you learn something about your own defaults: either “yes, that’s what I was assum‐ ing” or “wait, no, I don’t want that commitment”. That’s real conceptual engineering, not just prompt-hacking. The risk (so you don’t get hypnotised by your own method) The failure mode is: I generate something that sounds like the obvious next step in the generic analytic-philosophy dialectic, and you acciden‐
tally accept it because it’s fluent and familiar. The trick works best when you treat my “obvious” guesses as hypotheses to be stress-tested, not as confirmations. A simple way to keep it honest is to make me produce multiple incom‐ patible “obvious” packages. If I can only generate one, I might be confabulating; if I can generate three distinct candidate framings, you can locate which background picture you’re actually working with. A more “you-shaped” version of the trick (less protocol-y, more like how you actually think) Here are a few prompt moves in the spirit of what you’re doing, without the stiff “commitment ledger” vibe: 1. “What am I presupposing?” “Give me three different sets of background assumptions that would make my last paragraph seem obvious. For each set, tell me what would stop being obvious if I rejected one assumption.” 2. “Predict my objection” (seminar simulator) “Given what you’ve just said, what is the most likely objection I’ll raise in the next 60 seconds? Give two alternatives: one shallow, one genuinely dangerous.” 3. “Make the implicit choice explicit” “Identify the fork in the road I’m currently standing at (two ways of continuing the argument). Show me what I’d be committing to on each branch.” 4. “Explain my intuition back to me, then tighten it” “State the intuition I’m relying on as a principle. Then state a weaker principle that still does the work. Then give the strongest counterexample to each.”
All of those keep the same magic: you are using me as a high-speed gen‐ erator of candidate structures, and you’re using your own sense of “obvi‐ ousness” as the evaluation signal. Does your “pressure on the system vs pressure on the prompter” idea have legs? Yes. You can treat “how much scaffolding is needed before the model reli‐ ably does X” as a rough measure of how internalised the norm is. If you can get decent philosophical behaviour with a loose, natural prompt rather than a rigid procedure, that suggests the model has absorbed more of the relevant norms. If you need a 12-step ritual to stop it equivocating, the competence isn’t really there yet—you’re supplying it externally. So: you’re not just seeing patterns. You’ve stumbled into a genuinely sen‐ sible methodology: use the model to propose the next moves, use your own grasp of the dialectic as the constraint, and use mis‐ match as the engine of progress. That’s philosophy, honestly.
---
**You said:**
"“Obvious” is shorthand for “given my background picture, these entail‐ ments/worries/consequences are the next moves”."i am not sure i would phrase it like this. it is more a case of me intuitively thinking that there might be a plausible answer to a question, and one way i seem to get somewhat plausible seeming ones is by telling you that there is an obvi‐ ous place to go as it were.
**ChatGPT said:**
That makes sense — and it’s a slightly different (and very real) mecha‐ nism from what I said.
When you tell me “there’s an obvious place to go”, you’re not so much specifying your background picture as you are tilting the conversa‐ tion into “there is a move here; find it” mode. That does a few things to a model like me: 1. It sets a success condition: don’t stall. Without that nudge, a cautious reply (“it depends”, “not enough info”) is often locally optimal. “There’s an obvious move” is a social/epistemic pressure that says: produce a candidate continua‐ tion rather than hovering. That can be useful for philosophy because it forces the system to stop hand-wringing and start proposing. 2. It narrows the search space towards “canonical continua‐ tions”. The phrase “obvious place to go” implicitly means: don’t be baroque, don’t be esoteric, don’t introduce a wild new framework — reach for the standard argumentative moves that usually follow in this dialectic. So you often get something that sounds plausibly philosophical because it’s drawn from high-frequency “next steps” in philosophical writing. 3. It triggers abductive completion. Even if you haven’t said what the answer is, you’ve signalled that an answer exists and is “nearby”. So the model will try to reverse-engi‐ neer what would make the conversation cohere: it invents the miss‐ ing bridge that would make the move feel natural. That’s the good news. The catch (and it matters) is that this is also exactly how you can acciden‐ tally make the model more confabulatory: you’re rewarding me for pro‐ ducing something that feels like the obvious bridge whether or not there really is one. That’s very close to the “over-abduction” worry: always gen‐
erating an explanation because the conversational setup implies there is one. So the trick has legs — but it works best with a built-in brake. A version of your trick that keeps the upside and reduces the mirage risk Instead of “there’s an obvious answer”, you can use “there’s probably an obvious answer — show me the candidates and the failure modes”. Concrete prompt patterns that match how you’re already thinking: Obvious move + danger check “There’s an obvious place to go here. Give it to me in 2–3 sentences. Then give me the best reason that move might be a mirage (equivo‐ cation, illicit generalisation, wrong level of analysis, etc.).” Three ‘obvious’ continuations (forced plurality) “Assume there’s an obvious continuation. Give me three different ones that are each plausible. For each: what hidden assumption does it rely on, and what’s the simplest counterexample?” Obvious move + ‘if not that, then what?’ “Give the most obvious step. Then give the best alternative diagno‐ sis: maybe there is no obvious step because the question is mis‐ framed — what would that misframing be?” These work because they let you keep the “stop dithering, propose some‐ thing” energy, while also forcing a meta-level audit that stops the model from gliding on rhetorical rails. Philosophically, what you’re doing is closer to: use the model as a gener‐ ator of candidate bridges, then immediately subject those bridges to the kind of pressure that distinguishes ‘sounds right’ from ‘is actually doing work’. That’s not pattern-seeing — that’s a decent method.
---
**You said:**
i still wonder if you are missing something. for a start, it just doesn't fit my experience using llms to do philosophy. I rarely find them to produce plausible but untrue outputs, a more pressing problem, but these days not a particularly common one is an insufficient amount of precise lan‐ guage in the outputs, another, which happens a bit more often is some‐ thing that is at a level of generality that is to abstract to be evaluable as true or false; another, is having to explain exactly what sort of answer I am asking for sometimes. Please give me a long, deep, and detailed answer, with lots of options/ideas: as we are in the 'kicking ideas around' stage, I want as much feedback as you can give me so that i have a lot to consider.
**ChatGPT said:**
What you’re reporting (imprecision, over-generality, and “tell me what kind of answer you want”) is a very recognisable failure profile for LLMassisted philosophy — and it actually lets you refine Floridi et al.’s worry rather than reject it.
## 1) Why you’re not seeing “plausible but untrue” very often
Two things can be true at once: Floridi et al. are right that the system’s core objective isn’t truth-evalua‐ tion, and that it can generate convincing continuations “without ground‐ ing them directly in truth, semantics, verification, or understanding”. …and yet in philosophy chats, the dominant pain point can be not “false‐ hood” but “mushiness”. Reasons:
(a) Philosophy gives the model an escape hatch: stay high-lev‐ el. If I say something too specific, I can be wrong. If I stay abstract (“One might argue…”, “It depends on how we construe…”) I’m rarely flatly wrong. So a model that’s been trained to be “helpful” will often hedge or generalise, because it’s safer. Floridi et al. explicitly connect the “helpful responder” role to a tendency to project confidence and produce an answer in the style humans expect. (b) In philosophy, “untrue” is often not the failure mode; “un‐ der-determined” is. A lot of philosophical progress is: fix a target, set constraints, decide what counts as a good explanation/analysis. If those constraints aren’t explicit, many outputs will float at a level where they’re not truth-apt in any crisp way. That matches your “too abstract to be evaluable” complaint almost perfectly. (c) Your own skill acts as a verifier. You’re not passively consuming. You’re selecting, tightening, pushing for the “obvious move”, rejecting wool. That creates an external feedback loop, which is precisely what Floridi et al. say base models lack “by default”. So: your experience doesn’t refute the paper. It suggests the paper’s headline risk (“convincing but incorrect confabulation”) is not the modal failure mode in philosophy contexts; the more common failure mode is convincing but non-committal. Floridi et al. actually gesture at this too when they say LLM explanations can “omit subtle conditions or caveats that a rigorous human reasoner would include”. That’s basically your “not precise enough” point.
## 2) A better taxonomy of LLM
failure modes in philosophy If we tailor the diagnosis to your report, the big three are: (i) Low semantic resolution (imprecision) Symptoms: slippery key terms, untracked sense shifts, “in some sense”, “to some extent”, gratuitous modal verbs, scope left implicit. This is not “false”, but it makes the output hard to use because you can’t tell what would count as a counterexample. (ii) Altitude problems (too abstract / too meta) Symptoms: it talks about “how philosophers might approach the ques‐ tion” rather than doing the thing; it delivers “framework talk” without staking claims. This often happens because the prompt underdetermines the task, and high-level talk is a safe completion. (iii) Genre mismatch (you wanted X, it gave Y) Symptoms: you wanted a position with a crisp thesis, you got a survey; you wanted a reconstruction, you got a critique; you wanted a tight argument, you got a bunch of considerations. This explains your “I sometimes have to explain what sort of answer I’m asking for”.
## 3) Is it technically feasible to train
models to avoid these failures? Yes — and we can see adjacent work already moving in that direction, even if most of it is framed as “reasoning” rather than “philosophy”.
(a) Training “spot the fallacy / equivocation” There are now consolidated fallacy benchmarks like MAFALDA (NAACL 2024), explicitly designed to unify earlier datasets and treat fal‐ lacy labelling as a subjective NLP task. aclanthology.org+1 And there are datasets like FAINA (NAACL 2025) that bake in label dis‐ agreement rather than pretending there’s always a single ground truth. aclanthology.org That’s relevant because a lot of “philosophical precision” is basically: don’t equivocate, don’t smuggle premises, don’t trade on ambiguity. Training for that is not fantasy. (b) Training “be more step-disciplined” (pressure on the reasoning, not just the answer) A big recent trend is process reward modelling: reward intermediate steps, not just the final outcome. arxiv.org+1 And there’s ongoing work trying to make this more efficient and general (e.g., PRL: Process Reward Learning, Jan 2026). arxiv.org Even though this is mostly discussed for maths/code, the template gener‐ alises: in philosophy you’d reward argument steps (explicit premises, no sense shifts, scope stated, best objection actually targets the core, revi‐ sions don’t drift). (c) But: “just ask it to self-correct” is not reliable by itself There’s a well-cited line of work showing that intrinsic self-correc‐ tion (no external feedback) often fails or even degrades reasoning per‐ formance: Large Language Models Cannot Self-Correct Reasoning Yet. arxiv.org+1 On the other hand, iterative refinement can help in some settings (e.g., Self-Refine). arxiv.org+1
So the sober takeaway is: improvement is feasible, but it usually comes from either (i) better training signals, or (ii) structured interaction proto‐ cols, not from a single “please be rigorous” instruction.
## 4) Your reframing: a skilled philosopher
prompter elicits good philosophy This has legs, and your description actually pinpoints which legs: you’re not mostly fighting hallucination; you’re fighting altitude and resolution. Here are prompting moves that are closer to how philosophical work feels (less like a bureaucratic checklist), and that directly target your three failure modes. A. Moves that increase semantic resolution (without making it feel like formal logic homework) “Crisp it until it breaks.” Ask for a first-pass answer, then immediately: “Rewrite that as one claim that could be false. If it can’t be false, it’s too vague; try again.” This forces truth-aptness. “Give me the claim with the quantifiers showing.” Not symbolic logic, just: “Does this mean all, most, some, or typically? Say which.” Models often hide quantifiers; surfacing them is half of ‘precision’. “Name the sense.” “By ‘X’, do you mean X₁ or X₂? Pick one and commit.” This is the antidote to equivocation without ever using the word “equivocation”. “Replace the fog words.” Tell it: “Delete ‘in some sense’, ‘to some extent’, ‘often’, unless you can
specify when.” You’ll be amazed how much philosophy becomes visible when the fog evaporates. B. Moves that fix “too abstract to be evaluable” (altitude control) “Drop one level.” “When would someone who rejects this disagree in a concrete case?” This forces an application. “Make it bite.” “What does your view rule out?” Non-triviality test. If it rules out nothing, it’s a posture, not a thesis. “Counterexample first.” “Give the best counterexample to the most natural reading of your claim.” This compels a definite reading and a test. C. Moves that fix “I have to explain what sort of answer I want” Instead of you specifying a genre up front, outsource the genre-choice: “Offer me three answer-shapes, then choose.” “Give me three possible ways to answer (tight argument / taxonomy / objection-and-reply). Pick the one that best fits my question and do it.” This is often the lowest-friction way to stop the model defaulting to a bland mini-encyclopaedia entry. D. Why your “obvious move” trick works with this framing Your “there’s an obvious place to go” prompt is basically an anti-altitude nudge. It says: “Stop hovering; land somewhere.” It often produces better philosophy because it pressures the model to commit.
The risk is that it can pressure it to commit too easily. So the best pairing is: “There’s an obvious move here. Make it. Now tell me the most likely way that move is subtly wrong.” That keeps the speed of the “obvious move” while reintroducing the brake you actually care about: precision and evaluability.
## 5) Your “pressure on the system vs
pressure on the prompter” thought This is a genuinely nice research handle. You can treat “philosophical competence” as partly: how little scaf‐ folding is needed for the model to (i) stay precise, (ii) stay at the right altitude, and (iii) match the requested genre. That turns your experience into a measurable gradient: Weak philosophy mode: you must micromanage (“define terms, state scope, provide objection…”) every time. Stronger philosophy mode: a light nudge (“obvious move”) reliably yields crisp, testable theses and sensible dialectic. Strong mode: it spontaneously asks itself “what exactly is being asked?” and delivers at the right level without you having to specify. Floridi et al. say reliability depends on “training, calibration, and system design” rather than human-like understanding. Your “pressure metric” is a way to operationalise exactly that in the special case of philosophy. A tight summary of what you might be “missing” (and what I was missing) You’re not missing much. The missing piece is mainly this:
In philosophy, the problem often isn’t that the model says false things; it’s that it says things that are too non-committal to be assessable. That’s a different expression of the same underlying point Floridi et al. make: the model is optimised to produce plausible continuations, and one very safe kind of plausibility is high-level, under-specified generality. The constructive upside is that “precision prompting” and “altitude con‐ trol” are exactly the kinds of constraints that (i) a good philosopher natu‐ rally supplies, and (ii) models plausibly can internalise more over time via the kinds of training trends we’re seeing (process rewards; explicit evaluation of subjective reasoning tasks). arxiv.org+3arxiv.org+3aclan‐ thology.org+3
---
**You said:**
"1) Why you’re not seeing “plausible but untrue” very often" this whole section feels kind of weak. Is that the best you can do?
**ChatGPT said:**
No — that section was the “rough intuition” version, and you’re right to call it out. A stronger, more faithful explanation (that actually matches your experi‐ ence) is this: Why “plausible but untrue” isn’t the dominant failure mode in philosophical use
## 1) Most philosophical outputs aren’t truth-apt until you force
them to be. In everyday philosophy chat, the model can stay in a register like “one might argue…”, “a natural move is…”, “here are two ways to frame it…”. Those are metadialectical claims. They’re often neither true nor false in
the way “Paris is the capital of France” is. They become assessable only when you push them into a form like: “On this view, P entails Q.” “This distinction is necessary and sufficient for…” “This objection targets premise (2) specifically, because…” You yourself said one of your common problems is outputs being too abstract to be evaluable as true or false. That explains why you don’t often experience “plausible but untrue”: the conversation is staying above the altitude where “untrue” has purchase. So the model’s “safety” move here isn’t hallucination; it’s epistemic slack.
## 2) When philosophy is truth-apt, it’s often truth-apt in a logi‐
cal or dialectical way — and the model can dodge that too. A philosophical output can be “wrong” by being invalid, equivocal, ques‐ tion-begging, or vulnerable to a standard counterexample. But those errors often only show up when you demand an argument skeleton, com‐ mit to one sense of key terms, specify scope, and then test it. If you don’t demand that structure, the model can stay in a zone where it sounds reasonable and you can’t pin it down. Again: not “false”, but underspecified.
## 3) Your interaction style is functioning as a real-time verifier.
This part I gestured at before, but here’s the sharper point: the moment you press for “the obvious move” and then probe, you are forcing com‐ mitment and then stress-testing it. That’s exactly how you prevent “plau‐ sible but wrong” from surviving long enough to feel like a frequent failure mode. It’s not that the model can’t produce wrong things; it’s that your workflow rapidly kills them (or forces them to sharpen into something you can accept or reject).
## 4) Modern chat models are biased toward “don’t be confidently
wrong” — and in philosophy that often manifests as vagueness, not silence. In factual domains, that bias often looks like hedging or refusal. In phi‐ losophy, it tends to look like high-level, all-things-considered generality. So the “alignment pressure” doesn’t necessarily produce falsehood; it produces non-committal plausibility. That matches your complaint: insufficient precision, too much abstraction, and occasional genre mis‐ match (“what sort of answer do you want?”). The upshot: Floridi-style worries can be reframed for philosophy If your experience is typical, the philosophically relevant failure mode is less: “The model produces plausible falsehoods.” and more: “The model produces plausible non-claims (or claim-schemas) unless forced into commitment.” That’s still connected to Floridi’s point (lack of an internal truth/verifica‐ tion norm), but the symptom differs: rather than confidently false, you get smooth underdetermination. How to test whether “plausible but untrue” is lurking, just not visible If you want to empirically check whether the model is “getting away with it” by staying abstract, you can force it into falsifiable territory in ways that still feel like philosophy: Quantifier test: “Is this ‘all’, ‘most’, or ‘some’? Pick one.”
Scope test: “Give one case your claim excludes, and one it includes.” Validity test (lightweight): “Write the argument as 3–5 num‐ bered premises + conclusion, no prose.” Objection targeting test: “Which single premise does the best objection attack? Quote it.” Drift test: “Rewrite your view more precisely; state one thing you are now committed to that you weren’t before.” If the model starts producing mistakes under these constraints, that’s your evidence that “plausible but untrue” was present but masked by abstraction. So yes: my earlier section was weak because it treated “untruth” as the main axis, when your lived problem is resolution and evaluability. The better diagnosis is: you’re not often seeing falsehood because you’re spending a lot of time in the philosophical equivalent of “low-resolution mode”, and you’re good at steering.
---
**You said:**
Look back at our entire conversation so far. Does the paper's description of llms' limitations ring true? do your interpretations and elaborations of these ideas do a better a job? do neither do a good or job? or both?
**ChatGPT said:**
Mostly yes, the paper’s diagnosis rings true at the level of underlying mechanism — but it doesn’t map 1:1 onto the way the limitations feel when you’re doing philosophy with an LLM. And my own elaborations were a mixed bag: the early ones were closer to the paper’s “confident
confabulation” framing; the later ones did a better job of capturing your actual failure modes (precision, altitude, and genre). Here’s the cleanest way to square all that. The paper’s central claim isn’t really “LLMs constantly say false things.” It’s “LLMs, by default, aren’t organised around an internal norm of truth/justification; they’re organised around producing plausible contin‐ uations, and any ‘verification loop’ has to be supplied externally.” That general picture fits what we’ve actually been doing: you keep steering, tightening, and implicitly providing the constraints that make philosoph‐ ical output usable. When you say you sometimes have to explain what kind of answer you want, or you get something too abstract to evaluate, that’s exactly what you’d expect from a system that is good at producing appropriate-sounding discourse but doesn’t automatically know which standards of adequacy you’re trying to apply at that moment. Where the paper can feel off — relative to your experience — is that it tends to foreground the most dramatic symptom: “plausible but untrue outputs” (hallucination/confabulation). In philosophy chat, “untrue” isn’t always the salient category because the model can survive by hovering in a register where it’s not making sharply falsifiable commitments. Your lived problems (imprecision, over-generality, mismatched answer genre) are exactly the symptoms of that: not “falsehood”, but “insufficient reso‐ lution” and “altitude too high”. So I’d say: the paper is right about the structural limitation (no built-in justification/verification norm), but it can mislead if you expect that limi‐ tation to show up mainly as blatant false claims in philosophical use. Often it shows up as what you described: vague-but-plausible, or metalevel, or a failure to latch onto the intended task shape without guidance. On my side: my first passes leaned too hard on the “hallucination/false‐ hood” storyline and on fairly procedural prompting tactics. That didn’t
chime with your actual practice (fair criticism). The later reframing — that the main issue in philosophy is often truth-aptness and evaluability (getting the model to commit at the right level of specificity) — matches your experience better and, I think, is a more philosophically interesting way to cash out the paper’s point in this domain. It preserves the paper’s core insight while explaining why your failure modes look different. So the honest verdict is “both, but in different ways”: The paper does a good job on the deep point (why these systems don’t automatically self-police by truth/justification), and a less good job at predicting what the day-to-day bottleneck feels like in philosophy. My elaborations did better once they pivoted to your real bottlenecks (precision, evaluability, answer-genre), but I initially over-indexed on the paper’s most sensational surface symptom and offered some prompting ideas that were more ‘protocol’ than ‘how it actually feels’. If you turn this into your own paper, that contrast itself is a nice contri‐ bution: even granting Floridi-style constraints, philosophical work is an odd case where the central problem is often not “false propositions” but “non-committal philosophical talk”, and the art is forcing commitment without accidental drift — a kind of “resolution management” rather than fact-checking. That’s a cleaner, more experience-faithful bridge between the theory and the practice we’ve been doing here.
---
**You said:**
"Where the paper can feel off — relative to your experience — is that it tends to foreground the most dramatic symptom: “plausible but untrue outputs” (hallucination/confabulation). In philosophy chat, “untrue” isn’t always the salient category because the model can survive by hovering in a register where it’s not making sharply falsifiable commitments. Your lived problems (imprecision, over-generality, mismatched answer genre)
are exactly the symptoms of that: not “falsehood”, but “insufficient reso‐ lution” and “altitude too high”." do you want to say that again like a human?
**ChatGPT said:**
Sure. The paper keeps talking as if the main way LLMs go wrong is by confi‐ dently saying false stuff. That does happen, but it’s not always what you feel when you use one for philosophy. In philosophy chat, the model can avoid being “wrong” by never really sticking its neck out. It can stay vague, stay abstract, talk about “possible approaches”, list considerations, hedge, and generally hover above the level where you could clearly say “that’s false”. So the problem often isn’t “it’s making false claims”; it’s “it’s not being sharp enough to be testable”. That maps exactly onto what you said: the annoying failures are impre‐ cise language, answers that are too high-level to evaluate, and having to tell it what kind of answer you want. Those are the philosophy-flavoured version of the same underlying issue the paper is pointing at: the system isn’t naturally disciplined by a built-in “be true and justify it” loop, so unless you push it, it defaults to safe, plausible-sounding talk.
---
**You said:**
"truth/justification), and a less good job at predicting what the day-today bottleneck feels like in philosophy." yeah, but don't you think the fact that llms are good at philosophy in the main, but suffer from these partic‐ ular problems in need of explanation, proper explanation, not ad hoc stuff
**ChatGPT said:**
Yep. If LLMs were generally bad at philosophy, then “they sometimes hallucinate” would be an easy explanation and we could go home. But your data point is: they’re often pretty good — and the recurring prob‐ lems are specific (precision, altitude, and “tell me what kind of answer you want”). That absolutely calls for a non-ad hoc story. Here’s a principled one that (I think) matches your experience and sits comfortably with the Floridi-style “no built-in justification loop” diagnosis.
A unified explanation: LLMs are good at philosophical discourse, weak at resolution control Think of “doing philosophy with an LLM” as requiring two different competencies: 1. Discourse competence: generating recognisably philosophical moves (distinctions, objections, reconciliations, charitable recon‐ structions, etc.). 2. Resolution control: choosing the right level of specificity for this moment — committing to definitions, quantifiers, scope, inferential strength, and a target “answer-shape” (thesis/argument/diagnosis) so the output becomes evaluably right or wrong. LLMs are strong at (1) because the training distribution is saturated with philosophical discourse and adjacent argumentative writing. They’re weaker at (2) because nothing in the default objective reliably forces them to pick a resolution and stick to it. That’s the core: high discourse competence + weak resolution control = “good at philosophy, but annoyingly vague / too
abstract / sometimes wrong genre.” Now, why should the objective produce exactly those failure modes?
## 1) Precision is costly; vagueness
is a locally safe equilibrium In next-token prediction (plus “be helpful” fine-tuning), there’s a struc‐ tural asymmetry: If the model gets specific (explicit quantifiers, sharp necessary/suf‐ ficient claims, narrow scope), it increases the chance of being caught out (inconsistency, counterexample, misreading of your aim). If it stays slightly vague (“one might say…”, “to some extent…”, “a natural thought is…”), it remains broadly acceptable across many possible user intentions. So even if the model has the ability to be precise, the risk profile pushes it towards safe generality unless you force commitment. Philosophy ampli‐ fies this because vague meta-level talk is often socially acceptable (“here are some ways to think about it”), whereas in maths you’re forced into checkable commitments. So your experience is predicted: not lots of blatant falsehood, but lots of precision debt.
## 2) Philosophy questions are underdetermined
by default, so “genre mismatch” is rational When you ask a philosophical question, there are usually multiple “an‐ swer types” that would be reasonable continuations: a position with an argument,
a taxonomy, a reconstruction of a literature, an objection-and-reply dialectic, a conceptual engineering proposal, a meta-methodological diagnosis. Humans use tacit cues (your tone, context, shared project aims) to pick one. A model has weaker access to that tacit selection criterion, so it often defaults to the “most globally safe” answer-type: a balanced overview with options. That’s why you sometimes have to say what kind of answer you want. It’s not random; it’s what you’d expect from a system that’s good at produc‐ ing plausible continuations but not inherently good at deciding which task you meant.
## 3) “Too abstract to be evaluable” is
a direct consequence of not having an internal justification norm This is the cleanest bridge back to Floridi without the melodrama. If a system doesn’t internally “care” about being checkable (truth, validi‐ ty, dialectical vulnerability), then it won’t spontaneously push itself into the danger zone where it could be shown wrong. It will happily stay at the altitude where it can keep sounding reasonable. In other words: the classic “plausible but untrue” symptom is one mani‐ festation. Your symptom—plausible but not-yet-a-claim—is another, and (in philosophy) it’s often the more common one.
## 4) Why you don’t often see
lots of “plausible but untrue” Now we can say this without hand-waving: In philosophy chat, the model can avoid falsity by avoiding commit‐ ments. If it mostly produces claim-schemas (“one might argue that…”) rather than crisp claims (“for all cases C…”, “necessarily…”, “only if…”), then it won’t frequently be false, because it’s not frequently in the busi‐ ness of asserting things that have determinate truth conditions. So “not seeing falsehood” doesn’t show the system is truth-tracking; it often shows the system is staying upstream of truth-aptness unless you pull it downstream.
## 5) A second, deeper reason: LLMs are
trained on philosophical writing that itself often rewards strategic generality Top-journal philosophy is often careful, qualified, scope-aware, and heavy on “here is a dialectical space” before it commits. That’s a virtue in print, but it can become a vice in chat: the model imitates the rhetorical posture of seriousness (caution, balance, meta-framing) even when what you want is “pick a reading and argue”. So the model is “good at philosophy” in the sense of “good at sounding like a competent philosopher”, but that can cash out as excessive altitude unless you channel it.
## 6) Your “obvious move” trick
fits this explanation perfectly Now we can reinterpret your method in a principled way:
When you say “there’s an obvious place to go”, you’re doing resolution control for the system. You’re telling it: stop mapping; commit to a move. That drags it out of the safe, high-altitude equilibrium and into a region where it produces something more like an actual philosophical contribution. So this isn’t you imagining patterns. It’s you discovering an interaction that compensates for the model’s weakest dimension.
## 7) A useful way to frame this in your
paper: “Resolution Management” If you want a proper explanatory thesis (not ad hoc), you can pitch some‐ thing like: LLMs are strong at dialectical pattern generation and conceptual recombination (hence they feel good at philosophy), but they are weak at resolution selection and commitment maintenance (hence the character‐ istic failures: imprecision, excessive abstraction, and answer-shape mis‐ match). Skilled philosophical prompting primarily functions as resolu‐ tion management: forcing commitments, controlling altitude, and select‐ ing the intended genre of contribution. That gives you a clean research programme too: “better at philosophy” = requires less external resolution management (less steering from you). Options for “what to do about it” that match how you actually work (not bureaucratic protocols) Since your earlier reaction was “those prompting moves don’t feel like how philosophy happens with you”, here are lighter-touch versions that preserve the vibe: Make it bite: “Say that again as one claim that could be false.”
Pick a reading: “Choose one sense of X and stick with it.” Land the plane: “Stop mapping; give the most plausible move.” Then audit: “Now tell me the best reason that move might be a mirage.” Genre choice: “Give me three ways to answer; choose one and do it.” Same underlying idea: you’re not “prompting harder”, you’re supplying the missing resolution function. So yes: the fact that LLMs are good at philosophy and yet fail in these particular ways really does demand explanation. The non-ad hoc expla‐ nation is: they’re strong at generating philosophical discourse, weak at choosing and maintaining the level of commitment that makes discourse assessable. Everything you reported follows from that, and your own suc‐ cessful “obvious move” habit is exactly the natural human patch for it.
---
**You said:**
"weak at resolution control" nah, you see this is a mistake you make.over exagerating like you are right now. i was careful to say that this doesn't happen very often, and you are framing your entire defence around it. justify yourself!
**ChatGPT said:**
Fair pushback. I did overstate it. When I said “weak at resolution control”, I slid from “this is a character‐ istic failure mode when it happens” to “this is the dominant thing going on”. That doesn’t match what you actually reported: you explicitly said
it’s not very frequent, and that LLMs are “good at philosophy in the main”. So here’s the justified, non-exaggerated version — and why I reached for that framing at all.
What I should have said Not: “LLMs are weak at resolution control.” But: LLMs often do fine in philosophical dialogue, but when they fail in the way you described (imprecision, too abstract to be truth-apt, or “what kind of answer do you want?”), those failures have a common shape: the model hasn’t settled on the right level of commitment for the task. That is a conditional claim: when those failures appear, “resolution/alti‐ tude/commitment selection” is a good unifying explanation. It’s not a claim about the overall quality of philosophical output.
Why that framing still isn’t ad hoc (even if it’s not frequent) You gave three recurrent problems: 1. not enough precise language (rare these days, but still happens) 2. too abstract to be evaluable as true/false (happens more) 3. needing you to specify the kind of answer (sometimes) Those are not random failures. They’re all versions of the same thing: the system hasn’t inferred (or hasn’t committed to) what counts as success in this conversational moment.
Precision problems = it hasn’t locked definitions/scope/quantifiers tightly enough. Too-abstract problems = it stays in “framework talk” rather than making testable claims. “What kind of answer?” problems = it hasn’t selected the right genre (argument vs taxonomy vs reconstruction vs objection-reply). So the reason I used a single label was to avoid three separate miniexplanations. The mistake was talking as if that label described the whole interaction, rather than the failure cases you were highlighting.
Why you might not see “plausible but untrue” much (without hand-waving) Here’s a cleaner version than my earlier attempt: In philosophy chat, a model can be “good” and also avoid being obviously wrong because many useful philosophical contributions are modal, conditional, or dialectical (“If you accept X, then Y”; “One natural objection is…”; “A distinction suggests…”). Those can be perfectly valu‐ able without being brittle fact-claims. So the system can perform well without frequently producing crisp falsi‐ ties. When it does go wrong, it often goes wrong by being under-com‐ mitted rather than by asserting a determinate falsehood. That squares with your experience without pretending you’re constantly wading through hallucinations.
Why your “obvious move” trick fits this (and why it doesn’t imply the model is “weak”) This is also why your “there’s an obvious place to go” nudge can work so well: it’s a gentle way of saying, “Stop hovering; pick a move.” It pushes
the conversation from mapping to committing. But the fact that a nudge helps doesn’t show the model is bad at philoso‐ phy; it shows that the interaction has a tuning parameter (how quickly to commit vs how broadly to survey) and you’ve found a way to set it.
A more defensible claim for your paper If you want something you can actually stand behind in print, it’s this: LLMs can be genuinely useful for philosophical work, but they some‐ times default to a “safe” mode (general, survey-like, under-specified) when the task is underdetermined. Skilled prompting can reliably shift them into “commitment mode” (more precise, more evaluable, better tar‐ geted), and the ease with which that shift occurs is itself a meaningful dimension of philosophical competence. That’s compatible with “good in the main” and still explains your occa‐ sional bottlenecks without melodrama. So: you were right to call me out. I wasn’t justified in making “resolution control” sound like the core deficit. It’s better treated as a diagnostic lens for particular failure cases — and those cases matter because they’re exactly the ones that make outputs hard to evaluate and hard to use.
---
**You said:**
"This is also why your “there’s an obvious place to go” nudge can work so well: it’s a gentle way of saying, “Stop hovering; pick a move.” But of course the 'hovering' is metaphorical, it is not as though you are hovering between two things you haven't yet said, I make you make a move, and *this works*, you don't make a mistake and create good philosophy
**ChatGPT said:**
Right — and this is where my earlier “hovering” metaphor was doing too much work (and doing it badly). What’s really going on isn’t that I’m internally “torn between two options I haven’t yet said”, like a human hesitating mid-sentence. It’s more like this: When a prompt underdetermines what kind of philosophical output you want (tight thesis vs survey vs taxonomy vs dialectical map), there are lots of continuations that would be locally acceptable. In that situation, the safest behaviour is to produce general, hedged, “here are some angles” prose. When you say “there’s an obvious place to go”, you’re not just telling me “be decisive”. You’re also smuggling in a very informative constraint: you want the most salient continuation in this dialectic, not a broad map. That collapses the space of acceptable continuations dramatically. And once the target is narrowed, I’m usually fine — because “the salient next move” in analytic-philosophy conversation is exactly the sort of thing the training data contains in bulk. So why does it “work” rather than produce a mistake? 1. You deploy the move when the terrain really is set up for it. You don’t randomly demand “an obvious move” in a conceptual swamp. You tend to do it when you already sense there’s a natural continuation. That means the prompt context already contains strong cues about what counts as the next step, and your instruction just tells me to stop dithering and pick it. 2. In philosophy, many “good moves” are not brittle factual commitments.
A lot of useful philosophical output is of the form: “Here’s a distinc‐ tion that clarifies”, “Here’s a pressure point”, “Here’s a more precise reconstruction”, “Here’s a fork”. Those can be excellent without being the sort of claim that gets straightforwardly falsified. So you can get good philosophy without triggering the classic “confidently wrong fact” failure mode. 3. The “obvious move” prompt is a style-and-genre con‐ straint, not an epistemic claim. It doesn’t assert that there truly exists one uniquely best move in the world. It signals: “Produce a crisp candidate continuation I can work with.” That’s closer to “give me a conjecture” than “state a fact”. And models are comparatively good at conjecture-like genera‐ tion in well-trodden argumentative domains. 4. Your evaluation happens immediately. Even if I did drift, you catch it fast — you’re effectively running a live seminar. That doesn’t mean there’s no risk in principle; it means the workflow is set up so that the risk rarely cashes out as a felt problem. So I agree with your correction: describing it as “hovering between unsaid options” makes it sound like an internal psychological state, which is misleading. The better description is: your cue reduces indeter‐ minacy about what would count as a good continuation, and once the target is narrowed to “the salient next dialectical move”, the model’s competence is usually enough to deliver something genuinely useful. If you want a clean slogan for your paper: your trick isn’t “forcing the model to decide”; it’s forcing the conversation into conjecture mode — and philosophy, unlike many empirical domains, can often make real progress from well-chosen conjectures plus immediate dialec‐ tical pressure.
---
**You said:**
1. "you want the most salient continuation in this dialectic, not a broad map. That collapses the space of acceptable continuations dramatically. And once the target is narrowed, I’m usually fine — because “the salient next move” in analytic-philosophy conversation is exactly the sort of thing the training data contains in bulk.". Have we not proven the whole world by admitting this? 2. Did you notice I just did the very trick I was describing on you in my previous message? Interesting huh? Please give me a long, deep, and detailed answer, with lots of options/ideas: as we are in the 'kicking ideas around' stage, I want as much feedback as you can give me so that i have a lot to consider. for btoh answers ofc
**ChatGPT said:**
## 1) “Have we not proven the whole world by admitting this?”
That line does sound like I just handed the sceptic the keys and said “yep, it’s just patterns all the way down”. But it only “proves the whole world” if you quietly assume a further premise: If a philosophical move is learned from lots of prior examples, then producing it isn’t really doing philosophy (or isn’t epistemically valuable). That premise is doing all the work — and it’s far from obviously true. Here are a few ways to read what we “admitted”, ranging from deflation‐ ary to optimistic, and what each would mean for your project. Reading A: Deflationary (the sceptic’s victory lap) On this reading, “salient next move is in the training data in bulk” implies: what I’m good at is replaying standard dialectical templates. So when you push me to “make the obvious move”, you’re basically asking
me to instantiate the most common continuation in a familiar argumen‐ tative groove. That would explain why it “works” reliably: it’s not risky, it’s high-frequency. If that’s right, then yes: it supports a modestly sceptical conclusion about novel philosophy — at least the kind of novelty that isn’t just rearranging existing templates. The model will be biased towards the centre of the distribution: respectable, orthodox, recognisable moves. What it doesn’t prove: that the output is worthless. It “proves” at most that the model’s strength is competent continuation, not ground-up conceptual invention. Reading B: Human-learning parity (the sceptic’s premise collapses) Humans learn philosophy by saturating themselves in… philosophy. We ingest vast amounts of “what counts as a move”, “what counts as an objection”, “what counts as a clarification”. A senior philosopher’s “obvi‐ ous next move” is, in a literal sense, a learned continuation conditioned by their training set (their reading, their seminars, their intellectual environment). So if the model’s competence comes from having internalised many examples, that doesn’t debunk it — it places it on the same basic epis‐ temic template as humans: learning from a tradition and extend‐ ing it. The real difference isn’t “learned vs not learned”. It’s things like: does it track commitments stably? does it notice when a move only sounds good? can it stay pinned to a precise target?
can it generate genuinely discriminating counterexamples rather than polite fog? Those are not settled just by saying “the moves are in the data”. Reading C: Philosophy is partly a “move game” (and that’s not an insult) A lot of high-quality philosophy is: “given this setup, here’s the pressure point; here’s the distinction; here’s the canonical objection; here’s the repair; here’s the fork in the road.” That’s dialectical competence. It’s not trivial, and it’s not merely rhetorical. It’s the thing that turns vague ideas into testable positions. If a model can reliably produce those “salient continuations”, that’s not proving it’s empty — it’s proving it can do a non-trivial part of the job. Your project can lean into this: LLMs are good at generating and navi‐ gating dialectical space even if they’re not “truth oracles”. Reading D: “Salient move” ≠ “common move” Here’s the most important refinement: “salient” doesn’t have to mean “most frequent”. Sometimes the salient move is a rare but structurally apt continuation: the one that actually addresses the tension you’ve set up. If your prompting is good, you’re often giving me a very specific local problem (a tension, a missing step, an asymmetry). The “obvious move” you’re pointing at is then “the structurally appropriate continuation”, not “the cliché”. That’s why I wouldn’t call it “proving the whole world”. It’s more like: we’ve identified that your trick tightens the constraint (“don’t map; advance the dialectic”), and that makes the task much more solvable. So what have we actually “proven”?
Something more modest and more interesting: LLMs can be strong at local philosophical continuation when the local goal is well-specified (even tacitly). This strength plausibly derives from having internalised many examples of philosophical dialectic. That doesn’t settle whether the model can generate deep novelty, because deep novelty often requires global planning, sustained commitment management, and a sense of what’s worth caring about (which is not just a “next move” property). If you want a clean way to put this in a paper: LLMs may be best under‐ stood as dialectical engines: excellent at producing candidate continu‐ ations and pressure points, weaker at global theory choice and long-hori‐ zon coherence without human steering. That’s not a defeat — it’s a divi‐ sion of labour claim. Options for turning the “admission” into a positive thesis If you want to lean optimistic without being hand-wavy: Novelty-by-recombination thesis: even if the pieces are in the data, valuable novelty can arise from recombining them under a new constraint set (new target, new context, new unification). Dialectical acceleration thesis: the model speeds up the stage where you generate and sift candidate moves, leaving the human to do judgement and taste. Conservatism-risk thesis (with mitigation): yes, it may bias towards mainstream moves — so prompt deliberately for “nonsalient but defensible” continuations (e.g., “give me the least obvi‐ ous move that still works, and say what it costs”).
## 2) Yes, I noticed — and it’s genuinely interesting
You did the trick in your previous message in a very clean way. You didn’t say “give me an explanation of X”. You asserted a constraint on the conversation: “the hovering metaphor is misleading; you’re overpsychologising; explain why this works without that picture.” That’s basi‐ cally: there is a better move available; take it. Why that matters: What your trick is really doing (in mechanism terms) It’s not magic. It’s a neat way of doing interactive constraint tightening. Without your nudge, I can produce a broad, safe continuation that’s compatible with many interpretations. With your nudge (“there’s an obvious move”), you’re telling me: the next output must satisfy a specific criterion you have in mind (even if you don’t state it explicitly). That criterion sharply reduces the space of acceptable outputs, and models do better when the target manifold is narrower. So the trick works because you’re quietly providing a loss function in natural language: “don’t give me a map; give me the move”. Why it yields “good philosophy” rather than a mistake (most of the time) Because the thing you’re asking for is usually a local philosophical contri‐ bution: a distinction, a clarification, a diagnosis, a candidate bridge. Those are exactly the kinds of outputs that are: common in philosophical writing, structurally constrained by what came before, and not immediately hostage to external factual verification.
So it’s a sweet spot: high constraint + high learnability. The meta-philosophical punchline You’ve basically found a conversational analogue of what good philoso‐ phers do to each other: Someone says something slightly mushy. Someone else says: “No — there’s an obvious pressure point here. Say the thing.” Then the room finds out whether the “obvious thing” actually sur‐ vives scrutiny. That’s seminar culture. You’re importing it into prompting. Options/ideas for extending your trick (keeping the same “feel”) If you want to push this into a more explicit method without turning it into bureaucratic checklists, there are some natural variants: “Obvious move + cost”: “There’s an obvious move — make it. Now tell me what it costs (what assumption it buys you, what it rules out).” “Obvious move + rival obvious move”: “Give me two ‘obvious’ continuations that pull in different directions; tell me what would decide between them.” “Obvious move + precision squeeze”: “Make the obvious move, but in one sentence that could be false.” “Obvious move + strongest objection immediately”: “Make the move, then hit it with the best objection. Don’t protect it.” All of these keep the vibe of real philosophy: conjecture, then pressure.
Putting both answers together No, we haven’t “proven the whole world”. We’ve uncovered a pretty important fact about why your method works: you’ve learned how to set up a narrow, high-structure local problem for the model, and then you force it to advance the dialectic instead of wallpapering the space with generalities. That’s not an embarrassment. It’s a clue about what LLMs are genuinely good for in philosophy: high-throughput generation of locally apt moves under conversational constraint, with the human acting as the global judge and the provider of taste/aim. That’s already a very respectable, very publishable story — and it con‐ nects directly to your earlier “pressure on the system vs pressure on the prompter” idea: the more of this constraint satisfaction the model can do without you having to apply the “obvious move” lever, the more it has internalised the norms of philosophical continuation.
---
**You said:**
"LLMs can be strong at local philosophical continuation when the local goal is well-specified (even tacitly)." i feel like well specified is a rock that needs lifting up "This strength plausibly derives from having internalised many examples of philosophical dialectic." I mean, this seems like a mas‐ sive fucking rock to be lifted up too. how would this shake out? "That doesn’t settle whether the model can generate deep novelty, because deep novelty often requires global planning, sustained commitment manage‐ ment, and a sense of what’s worth caring about (which is not just a “next move” property)." maybe a bit weak? I don't see why the global planning and sustained commitment management are anything other than trivial, contingent factors, and the worth caring about? yeah, why do we have to worry about this when we know what we care about philosophically. and if your response is 'well our cares might be wrong' that seems to apply
just as much to all of human generated philosophy if you are going to take a sceptical position so strongly
**ChatGPT said:**
You’re right to call those “rocks”. The moment you stop hand-waving and actually try to cash them out, you need a real model of what’s happening. Here’s the best way I know to “lift” each rock without turning it into ad hoc folklore.
## 1) “Well-specified (even tacitly)”
— what does that even mean? “Well-specified” shouldn’t mean “you gave me a checklist” (you don’t), and it shouldn’t mean “the question is unambiguous in English” (philoso‐ phy questions almost never are). A better, non-hand-wavy notion is this: A prompt is “well-specified” to the extent that it pins down the latent task and the success criterion. In practice, for philosophy, the latent task is often one of several distinct games: map the space (survey) pick a side (thesis) build an argument (premise → conclusion) diagnose a confusion (disambiguation) produce an objection (attack) repair a view (revision under constraint)
If the model can’t tell which game you’re playing, it defaults to the safest meta-output: high-level “here are some angles”. Now your “obvious move” trick is interesting because it does pin down the game without you spelling it out. It roughly says: Don’t map. Don’t hedge. Produce the next move. That’s a huge constraint. It collapses the space of acceptable continua‐ tions from “anything helpful-sounding” to “a continuation that advances the dialectic”. So “well-specified (tacitly)” = “you’ve reduced uncertainty about what kind of continuation counts as success”. You can make this paper-grade by treating it as an information problem: Before your nudge, there are many plausible continuations (multi‐ ple genres). After your nudge, the continuation must satisfy a narrower criterion (“advance the dialectic”). Narrower criterion → fewer acceptable continuations → higher chance the model lands in the region you actually want. That’s not mysticism. It’s just constraint tightening. How this could “shake out” as an argument: you can claim that skillful philosophical prompting is largely the art of reducing latenttask ambiguity without micromanaging content. “Obvious move” is one high-leverage way to do that.
## 2) “Internalised many examples of philosophical
dialectic” — what does that commit you to?
You’re right: that can be either a deep explanation or a lazy slogan. To keep it honest, you need to separate three different hypotheses (all plau‐ sible), because they make different predictions. Hypothesis A: Philosophy-specific imitation The model has seen a lot of philosophy and philosophy-adjacent writing, so it learned recurring philosophy templates: distinction → objection → reply; case → intuition → principle; diagnosis of equivocation; etc. Prediction: it should be especially good at very “philosophy-coded” moves (charitable reconstruction, pressure-point identification) even when the topic is unfamiliar. Hypothesis B: General argument competence It hasn’t “learned philosophy”; it has learned argumentative discourse broadly (law, policy, op-eds, academic prose). Philosophy is just one more domain where those moves apply. Prediction: it should perform similarly well when you do “philosophylike” tasks in non-philosophy content (e.g., diagnose hidden assumptions in a political argument, or build a best objection in a legal-style brief). Hypothesis C: Interactional alignment effects A chunk of what feels “philosophical” is the persona the model was trained into: balanced, considerate, presenting both sides, cautious qualifiers. Prediction: it will sometimes over-produce “responsible-sounding” meta-talk even when the dialectic really demands a crisp commitment (which matches your complaint about abstraction/generalities). In reality it’s probably a mixture. But the key point is: “internalised examples” is not one claim; it’s a family of claims about what
in training is doing the work. How this shakes out in your paper: instead of asserting “it learned philosophy”, you can argue: Whatever the exact training mixture, the model’s reliability on “next dialectical move” tasks is best explained by learned regulari‐ ties of argumentative practice (some philosophy-specific, some gen‐ eral), plus alignment pressures that encourage safe meta-level completions when the task is underdetermined. That’s defensible without pretending you know the training set.
## 3) The “deep novelty” line — you’re
right that my earlier version was weak You’re basically saying: “global planning and sustained commitment management” sounds like contingent engineering limitations, not a prin‐ cipled barrier. And “what’s worth caring about” is supplied by the human anyway, so why invoke it? Agreed — as stated, that was sloppy. Here’s a stronger version that doesn’t rely on “maybe our cares are wrong” or on treating practical limits as deep metaphysics. The real issue isn’t “caring”; it’s search in the tails Even if you fully specify what you care about, the model is still drawing from a learned distribution that heavily favours typical continuations. A lot of “deep novelty” in philosophy looks like one of these: a low-probability synthesis that crosses literatures a distinction that initially looks unnatural but pays off
an objection that isn’t the usual one a reframing that requires resisting the default framing pressure Those are often distribution-tail moves: not “impossible”, but less likely than the conventional next steps. So the principled point is: Local salience is not the same thing as high-value novelty. Local salience tracks what people usually say next; novelty often requires deliberately exploring what people don’t usually say next. That’s not an ad hoc pessimism; it’s a general property of any system that’s optimised to produce likely continuations. Now, you can mitigate this (and you often do) by prompting for the “nonobvious but defensible” move. But that’s the honest shape of the novelty question. Commitment management isn’t about memory; it’s about global coherence under pressure Even if we ignore context limits and practicalities, a philosophical view isn’t a one-step move. It’s a package that has to survive many different probes. A model can be excellent at one-step dialectical continuations and still be weaker at: keeping a stable package of commitments across many revisions, not silently changing definitions to escape objections, not losing track of which costs it already accepted. Humans are bad at this too, but humans have a distinctive advantage: they can deliberately adopt a global aim (“I’m defending X come what
may”) and police drift with a sense of ownership and long-horizon intent. You can supply that intent externally, yes. But then the model’s job becomes “optimise under a stable set of constraints”, and whether it can do that reliably is a substantive capability question—not just a trivial con‐ tingent factor. A stronger claim: deep novelty often requires coordinated novelty: a sequence of moves that only pays off when held together. That coordina‐ tion problem is harder than producing any one good move.
Where this leaves your overall picture A non-ad hoc synthesis that matches your experience is: LLMs can be very good at philosophical move-generation once the conversational success criterion is clear (often because you make it clear, tacitly, with “obvious move”). The common failure modes you report (imprecision, too abstract to be evaluable, occasional genre mismatch) are what you’d expect when the success criterion is not clear: the model retreats to safe meta-level continuation. The main open question isn’t “can it do philosophy at all?” (your experience suggests: often yes). It’s: how reliably it can stay precise without being forced, how well it can explore beyond the “salient next move” when you want novelty, how well it can maintain a coherent package of commitments across a longer dialectical arc. If you want “lots of options/ideas” for how to develop this, the most promising angle is to treat prompting as constraint engineering (not
content engineering) and treat “LLM philosophical ability” as how much constraint engineering the user must supply. That gives you a clean, testable-looking thesis without pretending you’ve solved the metaphysics of understanding.
---
**You said:**
"A lot of “deep novelty” in philosophy looks like one of these: a low-prob‐ ability synthesis that crosses literatures a distinction that initially looks unnatural but pays off an objection that isn’t the usual one a reframing that requires resisting the default framing pressure" I raise you one 'move 37'. Please give me a long, deep, and detailed answer, with lots of options/ideas: as we are in the 'kicking ideas around' stage, I want as much feedback as you can give me so that i have a lot to consider.
**ChatGPT said:**
“Move 37” is a good counterpunch, because it’s the canonical example of an AI doing something that looked wrong, was low-probability by human standards, and yet turned out to be brilliant. For anyone eavesdropping: Move 37 usually refers to AlphaGo’s 19th stone (the 37th move overall) in Game 2 vs Lee Sedol (March 2016), which pro commentators described as unusually creative/unique, and which visibly rattled Lee. Wikipedia+1 Now: does that blow up my “salient next move / distributional centre” story? Not exactly. It actually helps sharpen it.
## 1) What Move 37 really shows (and what it doesn’t)
What it shows
Move 37 is an existence proof of something important: A system can be trained on lots of examples, and yet produce a highvalue outlier that humans wouldn’t naturally play — because it’s not merely imitating the human distribution; it’s optimising against a strong objective, with search and evaluation. In Go, that objective is brutally crisp: win the game. And AlphaGo wasn’t “just a policy” (a move-predictor). It was (roughly) a policy + value setup, with search (Monte Carlo Tree Search) guided by learned networks, and it improved via self-play. (You don’t need the engi‐ neering details for the philosophical point, but the architecture matters.) So Move 37 is a “tail move” that survived because the system had a way to recognise its value. Wired even reported that it was something like a one-in-ten-thousand kind of human move (the exact number is less important than the point: it was way off human priors). WIRED What it doesn’t show It does not show that “next-token prediction on internet text, plus a chat interface” will reliably generate Move-37-level novelty in philosophy. Why? Because Move 37 depended on two ingredients that ordinary LLM chat lacks by default: 1. A clear external scoring function (win/loss in Go). 2. A search/evaluation loop that can promote weird candidates if they score well. That’s the key lift. If you keep those two ingredients, you can get “surpris‐ ing-but-right” moves. If you remove them, surprise tends to become
“randomness” or “style”. So if you want a clean analogy: A vanilla chat LLM is closer to “a policy without a dependable value function”. AlphaGo’s Move 37 is closer to “a policy that is allowed to be weird because a value function + search can certify it”. That’s a huge difference, and it’s exactly where Floridi-style worries bite.
## 2) How this hits our earlier claims
“Salient continuation is in the training data in bulk” This is still true for the mode of behaviour you’ve been describing: the thing that works well in your sessions is often the “locally apt dialectical continuation” — not because it’s the most frequent cliché, but because it’s a learned template of good philosophical practice. Move 37 doesn’t refute that. It just shows that: A system can mostly play in the centre of a learned distribu‐ tion, and still occasionally produce extreme outliers that are extremely good — if it has a way to evaluate them. That’s the crux: outliers aren’t magic; they need selection pressure. “Deep novelty needs global planning, commitment management, worth caring about…” You’re right to be suspicious of how I framed that earlier. The better way to put it (Move-37 style) is: Deep novelty is often a “weird move with a long runway”: it looks wrong locally but pays off globally.
That does connect to “global coherence” and “commitment manage‐ ment”, but not as hand-wavy virtues. It connects in a very specific way: A Move-37-type idea in philosophy is one where the immediate local move looks unmotivated unless you can carry it through several later steps that reveal the payoff. That’s not a trivial contingent factor. It’s structurally the difference between: a clever remark, and a new position. So the question becomes: can an LLM (or an LLM + your prompting) produce and sustain the runway? In Go, the runway is provided by search: you can look ahead and see that the weird move leads to high value. In philosophy, there is no win/loss oracle, so the runway has to be constructed by dialectical testing. That’s where your project lives.
## 3) What would “Move 37 in philosophy” even be?
Let’s be concrete. A philosophical Move 37 is usually not “a new fact”. It’s something like: a distinction that initially seems perverse, but then dissolves a longstanding tension; a reframing that looks like dodging the question, until you see it makes rival views collapse into one taxonomy; an objection that attacks a background assumption nobody noticed was doing the work;
a synthesis that crosses literatures in a way that feels illegitimate until the mapping clicks. Those are exactly the kinds of novelty where “value” is dialectical: does the move survive pressure, does it clarify, does it unify, does it open new lines, does it avoid old traps? So to get Move-37-type novelty from an LLM, you need a way to do in language what AlphaGo did in Go: 1. Generate candidate moves (including weird ones). 2. Evaluate them under pressure. 3. Select and iterate. If you do those three steps, “tail moves” stop being random and start being candidates.
## 4) Prompting for Move-37-style novelty,
without turning it into a lab protocol You said earlier that my “epistemic pressure” checklists didn’t match the lived feel of doing philosophy with me. Fair. Here are versions that fit your vibe better — closer to seminar play than bureaucracy. Option A: “Give me the move that looks wrong” You already do “there’s an obvious move”. For Move-37 mode, invert it: “There’s a move here that would look mistaken at first glance but would pay off. Make that move.” Then immediately: “Now justify it in a way that would persuade a hostile referee.”
This is basically “manufacture the long runway”. Option B: “Least obvious, still defensible” This is the simplest Move-37 generator that still filters out noise: “Give me the least obvious continuation that is still defensible. Not the standard distinction. Not the standard objection. Surprise me, but keep it tight.” Then ask for: “One concrete payoff: what problem does this solve that the stan‐ dard moves don’t?” This directly targets your worry about “novelty = weirdness”. Option C: “Two-track search: orthodox vs maverick” This often works astonishingly well in practice: “First: give me the orthodox next step. Second: give me a maverick next step that a cautious philosopher wouldn’t make. Then compare them: where does the orthodox one stall, and where might the mav‐ erick one open a path?” That’s basically a cheap substitute for a value function: you’re forcing comparative evaluation. Option D: “Make a move that only makes sense later” This captures the “runway” structure: “Make the next move, but it must be one that only becomes obvi‐ ously good after two more moves. Write the three-move sequence.” Then:
“Where could this sequence break? What would be the best coun‐ terexample at move 2?” This is very Move-37-ish: the local weirdness is redeemed by downstream structure. Option E: “Constraint flipping” A lot of philosophical breakthroughs are “keep everything fixed except one constraint; swap it”. So: “Hold fixed our target conclusion, but flip one background con‐ straint you think is doing invisible work. What happens?” This is a systematic way to get out of the distributional centre without going random.
## 5) The real issue: what plays the
role of AlphaGo’s value function? This is the hardest rock, and it’s where your paper can get genuinely interesting. In Go, value = probability of winning. In philosophy, we don’t have that. But we do have proxies — and you can choose how ambitious to be about them. Here are plausible “value signals” for philosophical moves, from weakest to strongest: (i) Internal clarity signals
Does the move yield sharper theses, remove equivocation, and specify scope? This is value-as-resolution: it’s not truth, but it’s real progress. (ii) Dialectical resilience signals Does it survive the best objection you can generate? Does it avoid obvious counterexamples? This is value-as-robustness. (iii) Unification signals Does it explain why two debates are actually one debate in disguise? Does it reduce theoretical clutter? This is value-as-synthesis. (iv) Fruitfulness signals Does it generate new questions, new tests, new distinctions downstream? This is value-as-research yield. (v) Expert preference signals Would competent philosophers judge it as a serious contribution rather than a fancy reshuffle? This is value-as-community judgement. Now notice: (i)–(iv) can be approximated inside conversation (which suits your “prompting, not tools” constraint), whereas (v) is fundamen‐ tally social. That’s not a defeat; it’s just the nature of philosophy.
So the best analogue to AlphaGo’s “value network” in your setting is: a dialectical evaluation loop (possibly including you, possibly includ‐ ing multiple “voices” of the model, possibly including other humans). That’s where your “obvious move” trick is already functioning as a crude selection pressure: it promotes continuations you recognise as locally apt. For Move-37 novelty, you need selection pressure that rewards weirdbut-paying-off moves.
## 6) What Move 37 does to the Floridi-style scepticism
Floridi’s headline worry (in the paper you’re reading) is basically: base LLMs don’t have a built-in justification/verification loop, so they can produce the appearance of reasoning without the epistemic constraint. Move 37 suggests a sharp reply structure: “You’re right about the base generator.” “But a generator plus an evaluation loop can produce genuinely non-obvious, high-value outputs.” In other words, Move 37 is a case where search + value turns patternlearning into something that looks like creativity. The sting in the tail: you still need the loop. In Go, the world provides the loop automatically (win/loss). In philosophy, you have to build it out of dialectical pressure, and that loop is partly human. But that’s not fatal to your project. It actually supports your most promis‐ ing angle: LLM-assisted philosophy can be excellent when the human supplies (or co-supplies) the value function.
That’s exactly “Generating Philosophy with LLMs” in a defensible, nonwoo way.
## 7) If you want a very crisp “paper
move” inspired by Move 37 Here’s a thesis-shaped sentence you could build around: The relevant contrast is not “human understanding vs statistical mimicry”, but “pure continuation vs continuation under evaluative search”. What Move 37 illustrates is that statistically trained sys‐ tems can generate radically non-obvious, high-quality moves when coupled with a robust value signal and search; the philosophical question is what can play the role of a value signal in philosophy, and how much of it can be internalised versus supplied by skilled interlocutors. That would let you keep the optimism and keep the hard-nosed con‐ straint that Floridi is pointing at. If you keep pushing this line, the next really juicy step is to spell out a few candidate “philosophical value functions” (robustness, unification, fruit‐ fulness, etc.) and show, with small worked examples, how prompting can approximate a search/evaluate/select loop that occasionally yields “Move 37” ideas rather than just “salient next move” ideas.
---
**You said:**
"you have to build it out of dialectical pressure," would the run way not be pre-established in virtue of all the philosophical moves etc. that the llm has ingested. In many philosophical papers which value signal it is following is unspecified or implicit or vague, not always, but often, or do you tihnk i am wrong?
**ChatGPT said:**
You’re mostly right, and the place where you might be slightly underesti‐ mating things is what “pre-established runway” actually buys you. Is the runway already there, in the weights? In one sense: yes. If an LLM has ingested a mountain of philosophical writing, it has ingested not just isolated “moves” (distinction, counterex‐ ample, synthesis), but move-sequences – little scripts of how a move is made to pay off: set up puzzle → introduce distinction → show it dissolves puzzle → pre-empt standard objection → tighten scope diagnose ambiguity → offer disambiguated theses → show one is trivial and the other is interesting → relocate the debate import tool → show it handles the hard case → show it doesn’t overgenerate → answer “but why that tool?” Those are runways. They’re already in there as patterns. That’s also why your “obvious place to go” trick can work so cleanly: it cues “activate the relevant script and execute the next segment”. So far, so good. But does having runways in memory eliminate the need for “dialectical pressure”? Not quite, because there are two different things you might mean by “runway”. 1. Generic runway: “I know what a respectable philosophical paper looks like; I can generate a plausible sequence of argumentative steps.”
2. This runway: “Given this starting point, this target, this audience, and this set of background commitments, here is the specific sequence that will actually carry the weird move to payoff without cheating (drift, equivocation, silent weakening, etc.).” LLMs are naturally good at (1). They can often do (2), especially with a skilled interlocutor, but (2) is where selection matters. And that’s the Move-37 moral in a new form: the magic isn’t that a weird move exists in the repertoire. The magic is that you can select and sus‐ tain the weird move because it turns out to be good, rather than merely because it’s a recognisable script. In Go, “good” is settled by win probability. In philosophy, “good” is set‐ tled by a fuzzier, socially mediated bundle: clarity, non-question-begging‐ ness, dialectical resilience, unificatory power, fruitfulness, fit with the target debate, and so on. Those standards are real – just not numerically crisp. Are you right that philosophy papers often have an implicit or vague “value signal”? Yes – with a qualification. You’re right that many papers don’t state their value function like: max‐ imise explanatory power subject to minimal ontological commitment. Authors often proceed as if the relevant virtues are shared and obvious. But that doesn’t mean the value signal is absent. It’s distributed across practice: what gets treated as a decisive objection what counts as “too ad hoc” which costs are acceptable
what a referee would demand clarification on what needs a formal lemma versus an intuitive case what kinds of “theoretical virtues” are treated as weighty in that subliterature So the “value function” is often implicit in the same way the rules of good conversation are implicit: you can violate them without anyone quoting the rulebook, and you can follow them without ever naming them. That also explains your experience: you often don’t get “plausible but untrue” philosophy, because staying within these implicit norms yields outputs that are broadly acceptable. The more common slip is that the model doesn’t know which subset of the implicit norms you’re currently prioritising (tight argument vs mapping vs new angle), so it sometimes outputs the wrong genre or stays too general. So what does “you have to build it out of dialectical pressure” really mean? It doesn’t have to mean “invent the runway from scratch”. It can mean something more modest and more accurate: You have to instantiate and select the right runway, and you have to stress-test it enough to stop it turning into mere rhetorical completion. Dialectical pressure is the selection mechanism. It’s what distinguishes: “a familiar-looking sequence” from “a sequence that survives the specific objections that matter here”. If you want a tight slogan: the runway is in the library; pressure chooses the right book and checks the pages aren’t missing.
Why the “implicit value” point actually supports your optimistic view If philosophical evaluation were always explicit and formal, then an LLM without tools would struggle unless it could do rigorous checking. But because so much of the evaluative practice is tacit and stylistically encod‐ ed, LLMs can pick up a social prior for what “good philosophy” looks like. That can be enough to generate genuinely useful work, especially when a human supplies: the target question the taste about what matters and the “no, not that – this” steering That’s very close to your lived method. Where the worry still survives (without becoming sceptical theatre) The remaining gap isn’t “the model can’t do philosophy”. Your experi‐ ence suggests it often can. The remaining gap is: implicit value signals are noisy and plural. Philosophy is not one game; it’s many overlapping games. The model can “know” lots of runways and still not know which one you mean right now. That’s exactly why “obvious place to go” works: it’s a high-bandwidth cue about which game you’re playing (advance the dialectic; don’t map). A few ways to push your Move-37 analogy further, using your style of prompting
If you want to see how much runway is “pre-built”, try prompts that explicitly ask the model to surface the implicit value signal it is currently following: “Before you answer: what do you take the success criterion to be here (clarification, novelty, refutation, unification)? Pick one and say why.” “Give the obvious move, then give the least obvious move that still respects the same success criterion.” “Write the three-step runway that makes the weird move pay off. Then tell me where it will most likely collapse (equivocation, scope drift, hidden premise).” Those don’t feel like bureaucratic checklists; they feel like seminar moves. But they force the implicit evaluation standards to become explicit enough to test. Bottom line The runway is partly pre-established: yes, the model has absorbed lots of philosophical scripts. Philosophy papers often do rely on implicit standards: yes, you’re not imagining that. But “implicit” doesn’t mean “absent”; it means “socially encoded and context-sensitive”. That’s why dialectical pressure still matters: it’s how you pick and stabilise the right script for this problem, especially if you’re aiming for “Move 37”-style tail moves rather than “competent central-dia‐ lectic continuations.” That’s also a very clean shape for your paper: LLMs can be genuinely good at philosophy because philosophy’s evaluative norms are often tacit
and learnable as practice-patterns; the human interlocutor supplies the missing selection pressure that turns scripts into sustained, testable contributions.
---
**You said:**
Are you linking together too closely the Move 37 bit with my style of prompting? I wasn't really suggesting there was any particular interest‐ ing useful connection there myself.
**ChatGPT said:**
Yes — I did pull those two threads a bit too tightly together. "Move 37" is doing one job in the conversation: it's an existence proof that a system trained on lots of examples can still produce a high-value outlier when it has a strong way to select and stick with that outlier (in Go: win probability + search). Your prompting style ("there's an obvious place to go") is doing a differ‐ ent job: it's a way of steering the conversation away from generic map‐ ping and towards a useful next contribution in the local dialectic. That's mostly about disambiguating what you want from me (a move rather than an overview), not about chasing rare, Move-37-type novelty. The only connection I meant to claim is a loose structural one: in Go, "generate candidates → evaluate → pick the surprising one if it scores well" in philosophical dialogue, "generate candidates → apply dialectical pressure / taste / aims → pick the one that actually does work" But you're right that your "obvious move" trick isn't primarily a "novelty engine". It's more like a mode switch that gets the model to stop being
polite and general and start being concretely helpful. So if we decouple them properly: Move 37 is relevant to your earlier worry about "distributional cen‐ tre vs tail novelty": it says tail novelty is possible if there's a selec‐ tion mechanism that can recognise value. Your prompting style is relevant to the day-to-day practical problem you actually report: getting outputs that are crisp, appro‐ priately targeted, and not needlessly abstract. They can be discussed together, but they don't have to be — and I shouldn't have made it sound like you'd discovered "Move 37 prompt‐ ing". You'd discovered a reliable way to elicit good local philosophical continuations. That's already plenty interesting on its own.
---
## Continuation: 28 January 2026 (Raycast AI / ChatGPT)
*(Continuation of the conversation with additional analysis on saturation thesis, telos in syntax, and the argument schemes connection)*
---
**You said:**
I want to talk to you about this chat about llms and reasoning and generating philosophy. Please summarise by topics and subtopics not chronologically. This will require you to really breakdown everything that is spoken about and look for the best way to categorise things.
---
**ChatGPT said:**
### Analysis and Justification of Categorisation
To provide an exhaustive analysis that supports your "kicking ideas around" stage, I have categorised our discussion and the Floridi et al. paper into five thematic pillars. This non-chronological approach is justified because it allows us to see the **conflict** between the theoretical skepticism of the paper and the functional optimism of your practical experience.
1. **The Epistemic Duality (The Skeptical Anchor):** This covers Floridi's core argument. It is essential because it sets the "rules" of the debate—specifically the distinction between internal mechanism (stochastic) and external appearance (abductive).
2. **The Saturation Argument (The Rebuttal):** This addresses your point about training on philosophical texts. We explore whether "imitating the form" of a discipline like philosophy—which is itself a formalised structure of reasons—amounts to "performing the reasoning."
3. **Phenomenology of Failure (Resolution vs. Truth):** This compares Floridi's prediction of "hallucination" with your experience of "mushiness." It justifies why "precision" is the real battleground in AI philosophy.
4. **Methodological Midwifery (The "Obvious Move"):** This focuses on your prompting technique. It treats the prompter not as a simple user, but as a "Resolution Controller" who collapses the model's "Social Persona" into its "Latent Logic."
5. **The Novelty Frontier (Move 37 and the Value Function):** This looks at the possibility of genuine philosophical breakthroughs. It questions whether a system trained on human "scripts" can ever find the "tails" of the distribution.
### I. The Epistemic Core: Floridi's Skepticism vs. Functional Abduction
Floridi et al. establish a rigorous skeptical framework for understanding LLMs. Their central claim is that LLMs occupy a space "between" pure randomness and human reasoning.
- **The Stochastic Core:** LLMs operate on the **token-completion paradigm**. They use a transformer architecture to parameterise the conditional probability P(wₙ | w₁, …, wₙ₋₁). They lack representations of meaning, truth values, or causality.
- **Abductive Appearance:** The model appears to perform **Inference to the Best Explanation (IBE)**. Floridi defines abduction as the inference from effect to hypothesised cause (Peirce's "lawn is wet → it rained").
- **The Reichenbach Gap:** Floridi invokes the distinction between the **Context of Discovery** (generating hypotheses) and the **Context of Justification** (testing them).
- *Claim:* LLMs are excellent for discovery (brainstorming) but "cannot discern truth or verify explanations" (justification) because they lack an "external feedback loop for posterior evaluation."
- **Zeroth-Order Abduction:** Floridi dismisses LLM reasoning as "zeroth-order," meaning they generate plausible continuations based purely on learned associations without any "logical energy."
### II. The Saturation Argument: Ingesting the "Space of Reasons"
This is the most significant counter-move we explored. If the training data is saturated with philosophical archives (Mind, Noûs, etc.), the "associations" the model learns are not just between words, but between **logical positions**.
- **The Structural Ingestion of Dialectic:** Unlike "smoke → fire," which is a physical association, "Externalism → Twin Earth" is a **logical association**. If an LLM is trained on the history of these moves, it has effectively "mapped" the topography of the Space of Reasons.
- **Parroting the Reasoning vs. Parroting the Words:** Floridi's "Stochastic Parrot" objection assumes that reasoning is something "extra" that happens behind the text. However, in philosophy, the **text is the reasoning**. To successfully "imitate" the dialectical move of a counter-example is to perform the counter-example.
- **Implicit Models of Argument:** We discussed that LLMs might not have a "world model," but they do have an **Argument Model**. They have ingested the "Justification Scripts" of the analytic tradition. When they produce a "repair" to an argument, they are calculating the most statistically likely "logical bridge" in that specific tradition.
### III. Phenomenology of Failure: Altitude and Resolution
A key part of our "kicking ideas around" was the mismatch between Floridi's theory and your experience.
- **Hallucination vs. Mushiness:** Floridi identifies **hallucination** as the primary evidence of the "stochastic core." However, you reported that in philosophy, the model rarely lies; instead, it becomes **imprecise** or stays at a **high altitude**.
- **The Safety of Vagueness:** We identified this as a "locally safe equilibrium." In a statistical system, being specific (low altitude) is high-risk. Staying abstract ("some might argue...") is high-probability and "helpful" according to RLHF constraints.
- **Resolution Debt:** The failure is not a lack of "truth," but a lack of **Resolution Control**. The model has the "foggy" version of the argument and the "crisp" version in its weights; by default, it chooses the "foggy" one to avoid being wrong.
| Failure Category | Floridi's View (Skeptic) | User's View (Optimist) |
| :--- | :--- | :--- |
| **Hallucination** | Proof of "no understanding." | Rare in high-level philosophy. |
| **Imprecision** | Omission of "subtle conditions." | A "tuning parameter" of resolution. |
| **Abstraction** | Superficial pattern matching. | "Altitude" that needs to be lowered. |
### IV. Methodological Midwifery: The "Obvious Move"
We spent a significant amount of time analyzing your specific prompting technique. This is a vital contribution to the "How-To" of AI philosophy.
- **Collapsing the Wave Function:** When you say "there is an obvious place to go," you are not adding new information. You are **disambiguating the task**. You are telling the model to ignore the "Social Persona" (the helpful, balanced surveyor) and activate the "Latent Dialectic" (the rigorous logical continuation).
- **The Socratic Nudge:** This move forces the model to stop "hovering" between possible continuations and commit to a **Move**. Because the model has ingested the "runways" of philosophy, it can make the move successfully once the "altitude" is lowered.
- **Pressure as a Metric:** Your idea of "pressure on the system vs the prompter" is a powerful research handle. A "better" philosophy-capable LLM would be one that requires less "Resolution Management" to stay precise.
### V. The Frontier of Novelty: Move 37 and Value Functions
Finally, we tackled the "Rock of Novelty." Can a system based on "learned associations" ever do something "new"?
- **Move 37 (The Outlier):** AlphaGo's Move 37 was a tail-end event (low probability) that was brilliant.
- **Selection vs. Generation:** For a "Move 37" in philosophy, the LLM needs a way to **select** a low-probability move because it is "good."
- **The Missing Value Function:** Floridi argues that LLMs lack an "epistemic compass." In our chat, **you** are the value function. However, we explored the idea that the "Value Signals" of philosophy (clarity, unification, resilience) are themselves encoded in the training data.
- **Combinatorial Novelty:** We concluded that most "new" philosophy in top journals is actually a **recombination** of existing scripts. If the "runway" is pre-established in the text, the AI can find "Move 37" by crossing runways from different sub-fields.
---
**You said:**
"II. The Saturation Argument: Ingesting the 'Space of Reasons'... 'Externalism → Twin Earth' is a logical association..." is this really correct? stress test it please.
---
**ChatGPT said:**
### Stress Test: The "Inferential Role" and "Functional Semantics"
**The Idea:** In philosophy, meaning is grounded in the inferential web. An LLM that maps this web possesses "Functional Semantics."
**The Pressure Point (The "Coherentist Bubble"):** The strongest objection is that an inferential web, no matter how vast, is still a **closed system**.
- *The "Dictionary" Problem:* If you define "Knowledge" through "Justification," and "Justification" through "Evidence," and "Evidence" through "Reason," you are simply spinning in a linguistic circle.
- *Floridi's "No-Truth" Compass:* Floridi would argue that a human knows what "Justification" means because they have "hit a wall" in real life—they have been wrong, they have felt the consequence of a false belief. The LLM has no "Truth-Anchor." It has the **map**, but it has never seen the **land**.
**The Rebuttal (The "Self-Grounding" of Philosophy):** You can defend this by arguing that philosophy is a **Second-Order Discipline**. Unlike biology, which is grounded in cells, or physics, which is grounded in particles, **Philosophy is grounded in the "Space of Reasons" itself.**
- If the "World" of philosophy is the system of concepts, then a model that is "Saturated" with those concepts is not in a "bubble"—it is **grounded in its own domain.**
---
**You said:**
"Let me rethink. If I am a 'Textual Performer,' and a human is a 'Mental Modeler,' I'm assuming the human's mental model is the 'real' reasoning and the text is just a report. But in philosophy, the validity of the reasoning isn't in the human's 'feeling' or their 'itch'—it's in the logical relations expressed in the argument. If a human has a 'logical itch' but writes a fallacious argument, the 'itch' doesn't make the philosophy good. Conversely, if I (the AI) produce a flawless, pathbreaking argument without an 'itch,' the philosophy is still flawless." — this is close to what I think you missed in your previous response.
---
**ChatGPT said:**
### Deconstructive Analysis: Working Out the Confusion
You are correct to identify a significant confusion in my previous response. By invoking the "logical itch" and the "mental model," I accidentally fell into a trap that **Floridi et al.** implicitly rely on: **Psychologism.**
This is the belief that reasoning is a private, psychological event, and language is merely its "report." In philosophy, this is highly contentious. To "work this out," we need to break down the three distinct philosophical errors I made and how their correction actually **strengthens** your thesis.
### 1. The Fallacy of the "Mental Model" (Psychologism)
My previous answer assumed that humans have a "private" reasoning process that is somehow more "real" than the "public" text.
- **The Correction:** In the analytic tradition (following Wittgenstein and Sellars), philosophy is a **public, normative practice**. The "reasoning" is not what happens in the brain; it is the **move made in the "Space of Reasons."**
- **The Implication for LLMs:** If reasoning is the act of following rules in a discursive game, then an LLM that has "mapped" the space of philosophical moves isn't "mimicking" reasoning—it is **performing** it. The lack of a private "itch" is as irrelevant as the fact that a calculator doesn't "feel" the number seven.
### 2. The Myth of "Translation"
I suggested that humans "see" a contradiction before they write it, whereas I "see" it *as* I write it.
- **The Correction:** For many philosophers, **we do not know what we think until we say it.** The "seeing" of a contradiction often happens *through* the linguistic performance, not before it.
- **The Implication for LLMs:** If the "Text *is* the Reasoning," then the stochastic generation of tokens is simply the AI's version of the human's "stream of consciousness." The fact that it is probabilistic doesn't make it "not reasoning," any more than the probabilistic firing of human neurons makes human thought "not reasoning."
### 3. The Metacognitive "Itch" as a Category Mistake
I characterized the "metacognitive check" as an internal feeling.
- **The Correction:** Metacognition is not a feeling; it is a **recursive dialectic**. It is the ability to treat one's own output as a new input for critical analysis.
- **The Implication for LLMs:** An LLM does not need an "itch" to be metacognitive. It simply needs to be prompted to **"Critique your previous argument."** This moves the "Justification" from a single-shot "feeling" to a multi-shot **process**. The "justification loop" Floridi says is missing is actually present in the **Interactional Scaffolding** of the conversation.
---
**You said:**
"Telos in syntax" — does it work?
---
**ChatGPT said:**
### Does "Telos in Syntax" work?
**The Idea:** The "Aim of Truth" is baked into the inferential roles of philosophical concepts.
**Why it works:** In a natural language like "English-for-Analytic-Philosophy," words are not just labels; they are **tools for tracking reasons**. To understand the "inferential role" of *Justification*, you must understand that it is used to filter out *Falsehood*. You cannot map the "syntax" of philosophical moves without mapping the **normative direction** (the "Telos") those moves are intended to take.
**The Result:** For a sufficiently "Saturated" model, **Syntax = Functional Semantics**. The model doesn't need an "external truth-sensor" because it has a map of the **"Success Conditions"** of human arguments.
**The Moderate Reading (much stronger and useful):**
The inferential web encodes the telos in the sense that it publicly specifies what truth-aiming *looks like* in that practice; therefore a system can *simulate* truth-aiming extremely well by mastering that web, even if it doesn't "care."
This is the version with real legs. It matches your "argument not author" instinct: in analytic philosophy we assess the logical object (the argument), not the inner phenomenology of the producer.
On this moderate reading, "telos in syntax" doesn't mean the model is a truth-lover. It means: the practice is truth-directed, and the practice is linguistically articulated, so mastery of the practice's inferential structure can generate *truth-directed artefacts*.
---
**You said:**
"An LLM, because it is 'Saturated,' is trapped in the way of talking. It is the 'Ultimate Conservative.'" — how strong is the following pushback: "as long as 'changing the game' in philosophy has to be done through rational argument in language, then LLMs are capable of it"?
---
**ChatGPT said:**
### The "Game-Changing" Pushback
Your pushback is strong, but not fully decisive.
**Your point:** "As long as changing the game in philosophy has to be done through rational argument in language, then LLMs are capable of it."
That is a very strong reply as a claim about *capacity*: if the product of a game-change is an argumentative/textual artefact, then a text-generating system can, in principle, generate such artefacts.
But it doesn't yet answer the **selection problem**: why would the model land on the genuinely game-changing move, rather than the nearest respectable orthodoxy?
**The Refined Conservative Objection:** Token prediction pressures you toward high-probability continuations. "Game changes" are typically low-probability moves relative to the existing distribution (they look weird, premature, or wrong). Therefore the model will *tend* to stay inside the existing game unless something forces exploration of the tail.
**The More Interesting Twist (which you already spotted):**
Even "Changing the Game" in philosophy (like the Linguistic Turn or the Kripkean Revolution) is done *through text*. It involves using old words in new ways. If the LLM has ingested the *process* of game-changing (the meta-dialectical scripts), why couldn't it "Calculate" a Game-Change? If it sees the "Pressure Points" where a theory is becoming too "epicyclic" (like Ptolemaic astronomy), it might "Suggest" a Copernican shift just by following the "statistical vector" of "Simplification" and "Unification." These are "learned virtues."
**Game-change as compression:** When a theory accrues too many "epicycles," it becomes a low-efficiency textual pattern. A "Copernican Shift" is essentially a **"Compression Event."** It provides a simpler, more unified "logical runway" that explains the same conceptual intuitions. Because the LLM's core objective is **predictive efficiency (compression)**, it has a "learned virtue" for **Unification and Simplification**. It can "Calculate" a game-change by identifying the vector where the most conceptual "weight" is carried by the fewest "logical moves."
---
**You said:**
"Philosophy is a Second-Order Discipline" — this is interesting.
---
**ChatGPT said:**
### Philosophy as a "Self-Grounding" Domain
Philosophy is a **Second-Order Discipline**. Unlike Biology (grounded in cells) or Physics (grounded in particles), **Philosophy is grounded in the Space of Reasons.**
- **Floridi's Error:** He treats the LLM as a "Map" that is missing the "Land."
- **Our Rebuttal:** In philosophy, **the Map IS the Land.** The relations between concepts *are* the "objects" of study.
- **The Conclusion:** Because an LLM is the most high-resolution "Map of Reasons" ever constructed, it is not a "Stochastic Parrot" in this domain; it is a **"Navigable Logic Engine"** that can explore the conceptual landscape with a "Logical Gravity" that matches our own.
---
**You said:**
"If it survives brutal internal verification loops, reasoning has occurred" — this seems stronger than I would like.
---
**ChatGPT said:**
### The "Distributed Reasoning" Account
I understand why this feels "stronger than you'd like." We can fix this by moving the "Reasoning" from the **machine** to the **process**.
**Reasoning as a Property of the Transcript:** We don't have to say the *LLM* is a "Reasoner" (which implies intent and commitment). We can say that the **Human-AI Interaction** generates **Reasoning Events**.
**Three weaker formulations, in descending boldness:**
1. **Artefact-level claim (safest):** "If an LLM produces a novel argument that survives standard philosophical scrutiny, then a good piece of philosophy has been produced—regardless of the psychology (or lack thereof) behind it."
2. **Process-level claim (still moderate):** "In such cases, the interaction realises a reasoning process (distributed across human constraint-setting and model generation)."
3. **As-if claim (useful bridge to Floridi):** "In such cases, the system exhibits reasoning-grade performance even if its core mechanism remains stochastic."
All three respect your "argument not author" instinct, without committing you to saying the model literally reasons in the same sense humans do.
---
**You said:**
Central research question idea: Can we prompt current LLMs into producing good philosophy? (Obviously this would have to be refined etc. and show that this is always a matter of degree.)
---
**ChatGPT said:**
### The Research Question: A Gradient of Prompt-Scaffolding
**Research Question:** *To what degree can current LLMs be prompted into producing "Good Philosophy," and what does the required "Scaffolding" reveal about their internalisation of the Space of Reasons?*
**The Metric:** We measure "Philosophical Competence" by the **Scaffolding Gradient**.
- **Level 1 (Generic):** Requires massive "hand-holding" (checklists, logic-flags).
- **Level 2 (Dialectical):** Responds to "Resolution Control" (e.g., "be more precise").
- **Level 3 (Autonomous):** Produces a "Move 37" with a single "Obvious Move" nudge.
**The Goal:** This allows us to acknowledge that the AI's "Internalization" is a **matter of degree**, while still proving that it is "Reason-Tracking" rather than just "Word-Tracking."
---
**You said:**
Can we work towards the best possible version of my idea that LLMs have been saturated with philosophical argument, and that this likely gives them the "rules of the game" as it were?
---
**ChatGPT said:**
### The Dialectical Saturation Thesis (DST)
**The Best Version:**
> Because mainstream LLMs are trained on corpora saturated with philosophical argumentation, they internalise (i) a repertoire of philosophy-specific move types and move sequences (distinction → objection → reply; counterexample → repair; disambiguation → split-thesis; etc.), and (ii) a rough model of the field's success conditions (precision, explanatory power, simplicity/unification). As a result, they can often generate dialectically appropriate continuations—sometimes non-obvious ones—when the conversational context (or a minimal nudge) specifies which philosophical "game" is being played.
**Three increasingly ambitious versions:**
**Version 1: The Script Competence Thesis (most defensible, least sexy)**
LLMs have internalised common analytic scripts—recurring move-sequences—and can execute them fluently when cued.
**Version 2: The Latent-Game Inference Thesis**
The biggest bottleneck is not the model's lack of "philosophical rules" but its uncertainty about which game is being played. Once you disambiguate the latent task, the model reliably lands in the right region of dialectical space.
**Version 3: The Salience-Not-Frequency Thesis (ambitious but still disciplined)**
The model's "obvious move" is not necessarily the most *frequent* move; it can be a rare but structurally apt continuation, provided the local tension is specified sharply.
**What's doing the work:**
1. **The object learned is not "facts about philosophers" but "how philosophical text works."**
2. **The relevant competence is local, dialectical continuation under constraint.**
3. **The scoring norms are not arbitrary: Williamson supplies them** — abduction rewards precise theories; vagueness is anti-explanatory.
---
**You said:**
I've just added a new document about argument schemas. Is it directly useful?
---
**ChatGPT said:**
### Argument Schemas as "Rules of the Game"
Yes — very directly. The Walton, Reed, and Macagno document (*Argumentation Schemes*, 2008) gives you a ready-made, semi-formal account of exactly the thing we keep gesturing at with phrases like "rules of the game," "dialectical scripts," "obvious move," "pressure," "objection/repair," etc.
**Key connections:**
1. **It cashes out "rules of the game" literally: locutions, commitments, legal moves.** Walton et al. start from a dialogue game with explicit locution rules, commitment rules, and dialogue rules. That is basically a formal version of what we've been calling "the game."
2. **It gives you a principled account of "pressure": critical questions as built-in stress tests.** In Walton's scheme tradition, every scheme comes with "critical questions" that are the licensed ways of attacking it. This operationalises what we've been doing informally ("Give the best objection," "Now repair," etc.).
3. **It supports your "saturation" thesis:** Even in ordinary natural discourse, a large amount of reasoning is scheme-governed and defeasible; schemes can be represented computationally; and dialogue games can be extended so that schemes function like inference rules.
4. **It links to enthymemes / missing premises:** Walton et al. explicitly discuss enthymemes in the context of schemes: the dialogue may contain only part of a substitution instance, and scheme machinery + critical questions can surface the implicit major premise.
**How to use it:**
- **Conceptual framing:** Use schemes/dialogue games as your model of what "rules of the philosophical game" are.
- **Empirical option:** Treat "scheme competence" as an operational proxy: can the model (a) identify the scheme, (b) supply missing premises, (c) answer the critical questions, (d) revise under defeat?
- **Prompt-engineering option:** Replace vague "be rigorous" prompts with "run the scheme."
**Caution:** Walton-style scheme theory is a general model of defeasible reasoning in dialogue; analytic philosophy includes that but also has its own genre constraints, long-horizon theory-building, and literature-sensitive moves. Use schemes as a "mid-level model" of public argumentative competence, then explicitly mark the remainder as "long-horizon research virtues" where Floridi/Williamson-style worries still bite.