According to a document from *28 January 2026* (as timestamped in the chat export), this “attached chat” is basically a staged attempt to *break* (and then rebuild) a pessimistic Floridi-style diagnosis of LLMs by finding a philosophy-specific escape route - a “sidestep” grounded in (i) *saturation* on philosophical texts and (ii) the thought that much philosophical “verification” is *internal* to the space of reasons rather than empirical. ## Analysis, justification, reasons ### 1) What I’m treating as the “object” to analyse (and why) I’m not treating the chat as a mere summary of Floridi or a set of claims about ML. I’m treating it as a *piece of dialectical engineering* with three aims: - *Aim A (diagnostic):* explain why a sceptical paper (Floridi et al.) feels powerful *and* why your day-to-day experience with LLMs in philosophy doesn’t match its most salient failure mode (“hallucination”). - *Aim B (metaphilosophical):* argue that philosophy is a special domain where “text just is the reasoning” (the *R ≡ T* move), so “pattern completion” can converge with “inferential correctness.” - *Aim C (methodological):* extract a *useful* workflow for actually getting good philosophy out of an LLM (resolution/altitude control, the “obvious move”, generate-critic-select loops), plus an empirical-ish metric (“scaffolding gradient”). So I’m evaluating the chat primarily as: *Does it build a coherent bridge from sceptical theory → lived phenomenology → workable method → plausible research programme?* That’s the right lens because the chat itself explicitly shifts from abstract critique to domain-specific “how it works in the trenches” and then to a proposed research question and metric. ### 2) The “spine” of the chat’s argument The chat’s core structure is: *sceptical anchor* → *philosophy exception* → *mechanism of success/failure* → *novelty frontier*. **(i) Sceptical anchor: Floridi’s “stochastic core / abductive appearance.”** Floridi’s basic claim in the chat’s framing: token predictors can *look* like they’re doing IBE, but (by default) lack an internal truth/verification loop, so “abduction” is at best “appearance.” **(ii) Philosophy exception: internal verification + saturation.** The sidestep says: in philosophy, much “checking” is against *inferential constraints* (validity, consistency, scope, equivocation, counterexamples) rather than sensory world-contact. Therefore, the missing Floridi ingredient (verification) can be supplied *inside* language games. This is explicitly framed as “philosophy as a closed world” or “logic as the board.” Then the saturation thesis strengthens that: because LLMs are trained on loads of analytic philosophy, they learn not just word associations but *position-to-position* transitions (e.g., “Externalism → Twin Earth”), i.e., the “runways” of dialectic. **(iii) Mechanism: the real failure mode is “mushiness”, not hallucination.** You note that the model often doesn’t lie in philosophy; it “hovers” - high altitude, low resolution, genre mismatch. The chat’s explanation is: vagueness is a locally safe equilibrium for next-token systems; philosophy also *rewards* cautious, meta-level framing, so the model imitates that unless forced into commitments. This is where your “obvious move” trick becomes philosophically interesting: it’s treated as a mode-switch that pushes the system out of “survey mode” and into “commitment mode” - releasing the brakes. **(iv) Novelty frontier: Move 37 and the value function problem.** The chat then asks: even if LLMs can do “local continuation”, can they do *deep novelty*? The Move 37 analogy says: novelty requires not just a generator, but *search + value*. So you need a “philosophical value function” (unification, resilience, etc.) and some selection loop. That’s the argument’s spine. Now the interesting part is: which vertebrae are bone, and which are papier-mâché. --- ## 3) What the chat gets genuinely right (strong claims) ### A) The “hallucination vs mushiness” pivot is a serious diagnostic upgrade Floridi’s most vivid empirical symptom is hallucination. Your report is different: the common failure is *not* “falsehood”, but *under-commitment* (semantic fog, too much meta-talk). The chat nails why that’s not an ad hoc excuse: if the system lacks an internal norm that forces checkable commitments, it will rationally drift to outputs that stay “upstream of truth-aptness” unless pushed downstream. That is a clean, mechanism-shaped explanation, not handwaving. This matters because it reframes “LLM brittleness” in philosophy as a *resolution-control problem* rather than a global inability to “reason”. That gives you a tractable research handle. ### B) The “obvious move” is correctly analysed as constraint tightening, not magic The chat’s best methodological point is that “there is an obvious move” is high-bandwidth: it disambiguates the required *answer-shape* and forces the model to instantiate a specific dialectical continuation rather than giving a globally safe survey. In other words, it’s not adding content; it’s setting the *mode*. This is also why it’s philosophically interesting: it suggests what you’re really doing is supplying a local value signal (“advance the dialectic now”), which is a micro-version of the Move 37 “value function” story. ### C) The “internal verification” idea is plausibly true enough to matter The claim is not “philosophy is purely a priori” or “world contact never matters.” The claim is: *many central checks in analytic philosophy are formally/dialectically specifiable* (validity, consistency, equivocation, scope, counterexample resilience). That’s exactly the sort of thing that can be scored without perceptual grounding. Even if you later concede with Williamson that “total evidence” still matters, the internal-check slice is large enough that this is not a gimmick - it’s a *domain difference* with practical consequences. ### D) The “scaffolding gradient” is an actually good research metric The chat proposes measuring philosophical competence by how much prompting “pressure” is needed before the model produces evaluable, specific, dialectically robust output. That’s a *continuous* measure, which fits the reality that competence is not binary. It also operationalises “internalised dialectic” in a way you can actually test. This is one of the best “kicking ideas around” payoffs, because it turns a vibe (“it’s good when pushed”) into a programme. --- ## 4) Where the chat is most vulnerable (and how to patch it) ### A) The strongest claim is also the shakiest: R ≡ T (“text is the reasoning”) The chat’s radical move is: in analytic philosophy, reasoning is about conceptual relations, and the text is the instantiation of those relations; therefore, if the output text is a valid argument, the reasoning has occurred - no “behind the scenes” required. The chat itself flags the pressure point: reasoning is normally thought to involve *error correction*; if the system can’t notice its own mistake without being forced, maybe it’s “drafting” rather than reasoning. **Why this is a real vulnerability:** Even if we grant that philosophy evaluates arguments as public objects, the sceptic can concede “the artefact is good” while denying “the agent reasoned.” That’s not a trivial semantic dispute - it’s your whole target (LLMs *generating good new philosophy*, not merely emitting text that humans can turn into philosophy). **Patch options (and the chat gestures at them):** - *Weaken the identity claim:* Instead of *R ≡ T*, argue *R supervenes on the publicly checkable properties of T within an instituted dialectical practice.* That preserves the “publicity of reason” intuition without over-committing to identity metaphysics. The chat already moves in this direction when it says the reasoning is the interaction (human + AI), not the AI alone. - *Shift from agent to system:* Make the “reasoner” the *human–AI system* (distributed justification loop). This keeps your practical thesis (good philosophy can be produced with minimal prompting) while respecting the “no internal compass” worry. - *Make error-correction public:* Treat “metacognitive check” as a *recursive dialectic* (prompted self-critique, adversarial role-play, consistency probes), not an inner “itch.” This is explicitly suggested in the chat. This preserves what you care about - robust outputs - while removing the easiest sceptical kill shot. ### B) “Inferential role = meaning” is too clean to be universally true A pivotal line is: in analytic philosophy, the meaning of a concept is grounded in its inferential role; therefore an LLM that has mapped the inferential web “has functional semantics.” This *can* be defensible as a working thesis in certain inferentialist traditions, but it’s also where critics will press: do you really want to commit that far? You don’t need the full-strength claim “meaning is purely inferential” to get your paper’s main payoff. **Patch options:** - *Localise it:* “For a large class of *theoretical* philosophical terms (justification, grounding, supervenience…), their operational meaning in the literature is largely captured by inferential roles.” That’s empirically plausible and philosophically safer. - *Distinguish competence from understanding:* You can say the model has *discursive competence* with the term (it can use it correctly within the game) without settling the metaphysics of meaning. The chat already flirts with this by shifting from “inner life” to “outer practice.” ### C) The “telos in syntax” idea is brilliant - but also easy to overstate The optimist punch is that truth-directedness is “baked into the inferential web”; by saturating on the web, the model indirectly ingests the aim; the telos is in the syntax. The sceptical reply (also in the chat) is basically: a compass works because of a field outside the map; if truth is external, mapping talk isn’t enough. The chat’s conditional is important: if “truth” in philosophy is something like coherence + dialectical resilience, the web might be the field. **What’s actually valuable here:** Not the grand metaphysical claim that syntax collapses into semantics, but a more disciplined thesis: *In philosophy, many norms of correctness are publicly encoded in discourse (what counts as an objection, what counts as a non sequitur, what gets treated as ad hoc, etc.). Saturation lets models approximate these norms, but not guarantee truth.* That is both true-ish and useful - and it leaves you room to say “worldly constraints still matter in some areas of philosophy.” ### D) Conservatism / calibration is the real “sting” The chat eventually identifies the key danger: if the model is a high-resolution map of *our* space of reasons, it maps our mistakes too; it may lack a “global reset button” and so be an “ultimate conservative.” This is the place where your paper can either become serious or become hype. If you handle calibration well, the project is compelling. If you hand-wave it, reviewers will (correctly) eat you. **Patch options (concrete, non-handwavy):** - Build calibration into method: require adversarial objections, paraphrase invariance, commitment tracking, and explicit premise extraction as default “verification” steps. (You can present these as *internal* tests.) - Treat “conservatism” as a selectable prior: you can explicitly ask for “least obvious move that still works” and then test it. The chat’s Move 37 framing already pushes here: selection, not generation, is the bottleneck. --- ## 5) The key overlooked gold nugget: saturation on philosophical texts is doing multiple jobs You explicitly complained (in the chat) that earlier summaries missed how central training on philosophy is - and you’re right. The saturation idea isn’t one claim; it’s a cluster: 1. *Dialectical topography*: the model learns common attack/defence trajectories (objection types, repair moves). 2. *Genre priors*: the model learns the “serious philosopher” rhetorical stance - which explains why it defaults to cautious framing, and why your “obvious move” works as a nudge away from that stance. The pdf makes this explicit: top-journal philosophy rewards careful, qualified, scope-aware presentation, which is virtuous in print but becomes a vice in chat. 3. *Implicit value signals*: the literature encodes what counts as elegant, unified, ad hoc, etc. - which is why a “value function” might be extractable or approximable. That’s why saturation matters: it’s simultaneously explaining competence *and* explaining specific failure modes. --- ## 6) Concrete stress tests the chat is implicitly begging you to run Here are stress tests that directly match the chat’s pressure points (and that a referee would recognise as non-fluffy). ### Test set 1: “pattern completion vs rule-following” - *Paraphrase invariance:* restate an argument in multiple syntactically diverse ways; does the model preserve inferential commitments? If it does, you’ve got evidence it isn’t merely latching onto surface templates. - *Adversarial perturbations:* “twist” classic forms (change content words, swap order, insert distractors) and see whether it still tracks validity. This directly targets the “statistical frequency vs logical necessity” worry raised in the chat. ### Test set 2: the “internal verification” thesis - *Explicit-premise extraction (enthymemes):* force the model to output missing premises, then ask for critical questions or counterexamples. This connects nicely to argumentation-scheme tooling (Walton-style) without requiring formal proof assistants. - *Commitment accounting:* keep a running list of premises the model has committed to across turns; penalise equivocation and silent scope shifts. This operationalises “metacognitive check” as public procedure. ### Test set 3: Move 37 / deep novelty - *Generate–critic–select:* explicitly run three roles: generator, hostile critic, repairer. Your “value function” proxies are: clarity, non-question-beggingness, resilience, unification, fruitfulness. The chat already sketches exactly this kind of loop as a method. - *Cross-runway recombination:* take an argument template from one subfield and force a transfer to another (e.g., grounding idioms into decision theory, or semantics/pragmatics into metaphilosophy). A genuine “Move 37” should look weird locally but pay off globally. --- ## 7) Paper-shaping options that fall out of the chat (choose your ambition level) The chat gives you multiple “paper theses.” Here are three that look especially publishable *if* you defend them at the right strength. ### Option 1: The Resolution Management thesis (most defensible) Claim: Floridi’s diagnosis is broadly right about “no built-in justification loop,” but in philosophy the main symptom is *semantic resolution debt*, not hallucination. Therefore, LLM philosophical competence is best measured by how much scaffolding is needed to force checkable commitments (the scaffolding gradient). This is attractive because it’s (a) testable, (b) domain-specific, and (c) doesn’t require controversial metaphysics of meaning. ### Option 2: The Internal verification thesis (moderately ambitious) Claim: A significant subset of philosophical evaluation is internal to the space of reasons; therefore, LLMs can be trained and assessed for philosophical competence without external tools, by scoring dialectical performance (validity, consistency, equivocation detection, counterexample handling). Key move: you explicitly concede that not *all* philosophy is like this - but enough is. ### Option 3: The Performative/public reason thesis (maximally spicy) Claim: In analytic philosophy, the evaluable object is the public argument, not the author’s inner psychological state; hence “lack of itch” is irrelevant to philosophical quality. The “reasoning” is the publicly checkable performance, potentially realised by the human–AI system. This is the “bridge stands” move. If you do this, you must handle calibration and conservatism explicitly, or the paper will be dismissed as fun but unserious. --- ## Final synthesis The chat is valuable because it doesn’t just repeat “LLMs are good brainstormers.” It identifies a more precise (and more philosophically interesting) constellation: - Floridi’s *missing ingredient* is best interpreted as missing *justification pressure*. - In philosophy, the absence of that pressure shows up primarily as *mushiness* (altitude + resolution debt), not constant factual hallucination. - Because LLMs are saturated on philosophical discourse, they often have high *dialectical competence* - but they need either (i) your prompting pressure or (ii) a built-in evaluative loop to convert that competence into checkable commitments. - The most promising “big idea” is not that the model has a ghostly inner reasoner, but that philosophical reasoning is unusually *publicly performative*, making a human–AI system a plausible “unit of agency” for justification. - The honest hard problem is calibration: how to stop a saturated map from confidently leading you over conceptual cliffs. If you build your paper around *that* - especially the “resolution-control” diagnostic plus an explicit evaluative loop (your Move 37 apparatus) - you end up with a thesis that is both ambitious and not silly. Branched from [Branch · Philosophical Analysis of LLMs](https://chatgpt.com/g/g-p-697a425f23048191828a4db7fb6cc730/c/697f7e7d-4d18-8390-8b98-928075302192) Resetting properly. I’m going to do **one** thing here: extract (with quotes) **exactly what the attached conversation says about prompt craft**, and then (separately) offer some *non-committal* ways those pieces could be organised or extended. I’m not treating any of it as “decided”, “settled”, “the right way”, etc. I’m treating it as a pile of candidate moves and hypotheses. ## How I’m reading “prompt craft” in this chat (so you can audit me) In the chat, “prompt craft” is not just “write a nicer prompt”. It’s treated as: - reducing *latent task ambiguity* (“what game are we playing?”) - forcing *commitments* (so outputs become criticisable) - running a *mini verification protocol* inside dialogue - and, sometimes, inducing *selection pressure* towards “tail” continuations (rather than safe surveys) That’s the basis for the extraction below: I’m pulling every place where the chat talks about prompting as (i) constraint/criteria setting, (ii) resolution/altitude control, (iii) protocol design, or (iv) “value function” / selection. ## What the chat says about prompt craft (verbatim or near-verbatim quotes) ### 1) Prompting is framed as constraint engineering, not “ask nicely” The chat explicitly says: - “ *Rock 1 (Well-specified):* This is just **Constraint Engineering**. You provide the ‘Loss Function’ in natural language.” And later it tries to cash out “well-specified” in non-checklist terms: - “A prompt is ‘well-specified’ to the extent that it **pins down the latent task and the success criterion**.” - “So ‘well-specified (tacitly)’ = ‘you’ve reduced uncertainty about what kind of continuation counts as success’.” - “That’s not mysticism. It’s just **constraint tightening**.” It even presents it as an information/entropy reduction story: - “Before your nudge, there are many plausible continuations (multiple genres). After your nudge, the continuation must satisfy a narrower criterion (‘advance the dialectic’). Narrower criterion → fewer acceptable continuations → higher chance the model lands in the region you actually want.” ### 2) The “Obvious Move” is treated as a mode switch that disambiguates the task There are two slightly different formulations of what your “obvious move” does: **(a) “Don’t map; advance the dialectic.”** From the PDF chat transcript: - “Now your ‘obvious move’ trick is interesting because it does pin down the game without you spelling it out. It roughly says: **Don’t map. Don’t hedge. Produce the next move.**” - “That’s a huge constraint. It collapses the space of acceptable continuations from ‘anything helpful-sounding’ to ‘a continuation that advances the dialectic’.” **(b) “Collapse the persona into logic.”** From the other chat file: - It calls this “Methodological Midwifery (The ‘Obvious Move’)” and says it treats the prompter as a “Resolution Controller” who “collapses the model’s ‘Social Persona’ into its ‘Latent Logic.’” - It describes the same move as: “When you say ‘there is an obvious place to go,’ you are not adding new information. You are disambiguating the task… telling the model to ignore the ‘Social Persona’ … and activate the ‘Latent Dialectic’.” It also explicitly characterises your prompting as “interactive constraint tightening” (with an example of you setting a conversational constraint rather than requesting content): - “You didn’t say ‘give me an explanation of X’. You asserted a constraint on the conversation: ‘the hovering metaphor is misleading; you’re over-psychologising; explain why this works without that picture.’ … It’s not magic. It’s a neat way of doing interactive constraint tightening.” ### 3) Prompt craft is repeatedly tied to resolution and altitude rather than “hallucination” This is one of the recurring diagnostic claims: - “you’re not mostly fighting hallucination; you’re fighting **altitude and resolution**.” And it frames vagueness as a default equilibrium you have to break: - “the model… by default, it chooses the ‘foggy’ one to avoid being wrong.” ### 4) The chat gives concrete micro-prompts (word-for-word) This is the most literal “prompt craft” content: a list of prompt moves the chat offers. From the PDF chat transcript: **A. “Moves that increase semantic resolution”** - “ **Crisp it until it breaks.** Ask for a first-pass answer, then immediately: ‘Rewrite that as one claim that could be false. If it can’t be false, it’s too vague; try again.’” - “ **Give me the claim with the quantifiers showing.** Not symbolic logic, just: ‘Does this mean all, most, some, or typically? Say which.’” - “ **Name the sense.** ‘By “X”, do you mean X₁ or X₂? Pick one and commit.’” - “ **Replace the fog words.** Tell it: ‘Delete “in some sense”, “to some extent”, “often”, unless you can specify when.’” **B. “Moves that fix ‘too abstract to be evaluable’ (altitude control)”** - “ **Drop one level.** ‘When would someone who rejects this disagree in a concrete case?’” ### 5) Prompt craft is also described as protocol design (multi-step interaction rules) Later in the PDF transcript, prompt craft is explicitly reframed as imposing a *dialectical method*: - “don’t think of ‘prompting’ as ‘ask nicely’. Think of it as **imposing a dialectical method**.” And then it gives several protocol moves (again: specific): - “ **Commitment ledger.** Force the model to write a numbered list of commitments (definitions + premises + scope claims). Then every subsequent step must cite which commitments it uses and whether it adds any.” - “ **Argument skeleton first, prose last.** Make it output: thesis; premises; inference type; scope; then only afterwards allow rhetorical packaging.” - “ **Objection as adversarial reconstruction.** Require the objection to (i) restate the target view charitably, (ii) identify the single weakest premise, (iii) attack only that, and (iv) show the smallest counterexample.” - “ **Repair under invariants.** If it revises, demand: ‘State exactly what changed and which commitments are preserved.’ Then run an explicit ‘semantic drift audit’: ‘Did the meaning of key terms change?’” ### 6) There’s also a prompt-craft claim about selection pressure (Move 37 framing) This is where prompting is treated as approximating a search/evaluate/select loop: - The PDF says the analogue of AlphaGo’s value network is “a dialectical evaluation loop” and: “your ‘obvious move’ trick is already functioning as a crude selection pressure.” - It then suggests deliberately prompting for less-safe moves: “prompt deliberately for ‘non-salient but defensible’ continuations (e.g., ‘give me the least obvious move that still works, and say what it costs’).” - In the other chat, it similarly says the “Obvious Move” can be used to “override the statistical mean” and push the model towards “the tail.” - And it proposes a “Counter-Intuitive Search” method: “explicitly prompt for ‘Low-Probability but Logically Consistent’ continuations.” ### 7) The chat explicitly warns against one common “prompt craft” move: “just self-correct” It says (in the PDF transcript): - “ ‘just ask it to self-correct’ is not reliable by itself” - and more pointedly: “if you ask a model to critique itself, the critique is not reliably independent from the original error.” So the chat’s line is: if you want “checking”, you add structure (protocols), not a single “please check yourself” request. --- ## A careful synthesis of what these prompt-craft pieces are trying to do Not a conclusion, just a way of seeing the pieces as one tentative picture: 1. **Disambiguate the “game”** The “obvious move” and “don’t map; advance the dialectic” are about telling the model whether you want a survey, a thesis, an objection, a repair, etc. 2. **Force “truth-aptness” and commitments** The “crisp it until it breaks”, “quantifiers”, “name the sense”, and “fog word deletion” prompts are about turning an un-evaluable blob into something that can actually be wrong (and therefore can be argued about). 3. **Prevent silent drift** The “commitment ledger” and “semantic drift audit” are about stopping the model from changing the subject by stealth (changing senses, scope, or definitions midstream). 4. **Simulate a verification loop** The “argument skeleton”, “adversarial reconstruction”, and “repair under invariants” are presented as ways for a skilled prompter to act like the missing feedback loop - inside dialogue. 5. **Add selection pressure for weird-but-defensible continuations** The Move 37 angle adds the thought that “prompting” can approximate search/evaluate/select if you explicitly reward the kind of move you want (“least obvious move that still works”; “low-probability but logically consistent”). --- ## Some non-committal “next-step” ways you could use this (without pretending anything is decided) ### Option A: Treat the prompt craft as a taxonomy You could package the above as four families: - *Task disambiguation* (which “game”) - *Resolution forcing* (truth-aptness / quantifiers / senses) - *Drift prevention* (ledgers / invariants) - *Selection pressure* (least-obvious / tail prompts / adversarial loops) That’s basically just reorganising what’s already said. ### Option B: Treat it as a protocol template (fill-in-the-blanks) Something like: 1. “State the task: survey / thesis / objection / repair.” 2. “State one falsifiable claim.” 3. “Expose quantifiers and senses.” 4. “List commitments.” 5. “Generate the best objection under constraints.” 6. “Revise under invariants + drift audit.” This is basically spelling out the interaction moves the chat already lists. ### Option C: Treat it as an empirical probe The chat hints at a “scaffolding gradient” / “prompt-sensitivity” benchmark - how much intervention is needed before the model stops hovering and does the thing you want. You could turn each micro-prompt into a test: “Does this model respond to *one* ‘quantifiers’ nudge, or do you need five rounds and a ledger?” --- If you want to keep this brutally literal, the quotes above are the “prompt craft” content the attached conversation actually contains. If you want to push beyond literal extraction, the safest next move (consistent with the chat’s own framing) would be to turn those into a compact “protocol prompt” and then torture-test it on a few concrete philosophical mini-cases - but that’s a separate step from “what was said.” Resetting properly: I’m going to treat the “attached chat” as *a transcript of things said* (by you and by an LLM), not as a source of facts about LLMs. Wherever the transcript *claims* “Floridi says X”, I’ll separate: 1. what the transcript *attributes* to Floridi, 2. what Floridi et al. *actually write*, 3. what’s *the chat’s own invention / extrapolation*. The two main artefacts here are the chat transcript and the Floridi paper itself . ## 1) What the transcript is (and what it isn’t) The file “nick - LLMs and Abductive Reasoning” is a record of a conversation where you ask an assistant to help you understand Floridi et al., and then you push the assistant hard with “opponent moves” and engineering-style hypotheticals. So: it’s reliable as evidence of *what was said in that conversation*. It is not, by itself, evidence that any of the assistant’s claims about (say) LLM training, robustness, or the literature are true. ## 2) What the transcript attributes to Floridi vs what Floridi actually says Here are a few central “paper says…” claims that appear in the transcript, with a status check against Floridi et al.’s text. | Claim as the *assistant* frames it in the transcript | Status against Floridi et al. | | --- | --- | | LLMs “generate text based on learned associations rather than performing abductive inferences.” | This is in the abstract (very close wording). | | LLM outputs can appear abductive because training data “encode reasoning structures”; the paper emphasises a “stochastic core” and “abductive appearance”. | Also in the abstract. | | LLMs produce plausible hypotheses “without grounding them directly in truth, semantics, verification, or understanding,” and they “cannot discern truth or verify explanations.” | Again, abstract. | | Reichenbach-style discovery vs justification: LLMs generate candidates but “do not genuinely validate them against reality” (unless augmented); “prior predictive sampling” + lacking an “external feedback loop for posterior evaluation.” | Floridi explicitly discusses Reichenbach’s two-stage picture, says LLMs “perform only the first part,” and uses the “prior predictive sampling / external feedback loop” language. | | LLMs can be misled by fallacies and lack a “metacognitive check” on consistency. | Floridi explicitly says they “lack a metacognitive check” and can generate inconsistent claims. | Now, a key nuance: the transcript repeatedly uses the phrase “truth as an internal objective” as the assistant’s way of packaging Floridi’s point. That phrasing is *the assistant’s gloss*, even when the underlying idea is close to Floridi’s own wording (“aim to model the conditional distribution of tokens… not to evaluate truth”; “lack an external feedback loop…”). So: *content-wise* the gloss tracks passages in Floridi; *phrase-wise* it’s not necessarily Floridi’s own formulation. ## 3) Prompt craft: what is actually said in the transcript Now to the bit you explicitly care about: *prompt craft*. The crucial thing here is that almost all “prompt craft” content is **not Floridi’s**; it is the assistant’s attempt to design interaction protocols that (in the assistant’s view) compensate for what Floridi says is missing (verification/feedback). I’ll quote and label these as *ChatGPT-in-the-transcript proposing prompt moves*, not as established truths. ### 3.1 Prompting as imposing a dialectical method The assistant explicitly reframes prompting like this: “don’t think of ‘prompting’ as ‘ask nicely’. Think of it as imposing a dialectical method.” It then proposes a set of concrete protocol moves (again: proposals, not facts), including: - “Commitment ledger.” - “Argument skeleton first, prose last.” - “Objection as adversarial reconstruction.” - “Repair under invariants” + “semantic drift audit.” - “Paraphrase and invariance testing.” - “Adversarial self-play with role separation… ‘Author’, ‘Opponent’, ‘Referee’.” ### 3.2 Prompting for semantic resolution and “altitude control” Later, the assistant claims you’re “not mostly fighting hallucination; you’re fighting altitude and resolution” and offers specific prompt moves, including: - “Crisp it until it breaks.” - “Give me the claim with the quantifiers showing.” - “Name the sense.” - “Replace the fog words.” - “Drop one level.” And it adds a pairing suggestion that explicitly treats your “obvious move” as something like a throttle, plus a braking step: “There’s an obvious move here. Make it. Now tell me the most likely way that move is subtly wrong.” ### 3.3 “Well-specified” prompts = latent task + success criterion This is a particularly explicit bit of *prompt craft theory* in the transcript. The assistant says: “A prompt is ‘well-specified’ to the extent that it pins down the latent task and the success criterion.” It then lists the “games” it thinks philosophy prompts underdetermine (survey / thesis / premise-to-conclusion / disambiguation / objection / repair) and claims your “obvious move” nudge works because it selects the game: “Don’t map. Don’t hedge. Produce the next move.” And it explicitly characterises this as “constraint tightening,” not content injection: “That’s not mysticism. It’s just constraint tightening.” ### 3.4 Prompting for “Move 37”-style novelty When discussing the AlphaGo analogy, the assistant proposes “seminar-play” versions of prompts, such as: - “There’s a move here that would look mistaken at first glance but would pay off. Make that move.” - “Give me the least obvious continuation that is still defensible.” - “First: give me the orthodox next step. Second: give me a maverick next step… Then compare them…” - “Make the next move, but it must be one that only becomes obviously good after two more moves. Write the three-move sequence.” Again: these are not claims that such prompts *work*; they’re *candidate prompts* the assistant proposes. ## 4) Where the transcript goes beyond Floridi (assistant’s own constructions) A lot of what you pushed the assistant to do is precisely this: go beyond the paper, while trying to stay “close to its framework”. A clean example is the “no tools, but truth-aiming via training” line. The assistant explicitly says its earlier wording could have sounded like a “secret recipe” and then offers three routes (world-supervised tasks, RL from environment feedback, etc.) as ways of building in feedback during training rather than at inference time. The “CEV” brainstorming (Claim–Evidence–Verification) is also an assistant construction: it’s presented as a *blueprint* for engineering the “missing loop” Floridi emphasises, not as something Floridi proposes as a ready-made solution. So if you’re reading the transcript critically: treat these as the assistant’s “opponent moves” and speculative design sketches, not as facts about how LLMs actually are. ## 5) A way to keep the transcript useful without treating it as fact One way to use the transcript without letting it smuggle in “facts” is to treat it as containing three different kinds of material: 1. *Textually grounded reporting* of Floridi’s claims (checkable against the paper). 2. *Conceptual reframings* (“truth as internal objective”, “pressure”, “altitude”) that may be illuminating but are not verbatim Floridi. 3. *Proposed prompt protocols* (commitment ledgers, invariance tests, “obvious move” variants) that are neither Floridi’s nor established empirical results - they’re candidate methods. That separation stops the “LLM said it” → “therefore it’s true” slide, while preserving what’s actually valuable in the artefact: it’s a record of which conceptual levers and prompting levers were being explored, and how they were motivated. If you keep pushing this in the same spirit, the next natural move is to treat each proposed prompt move as a *hypothesis about failure modes* (task ambiguity, drift, non-falsifiability, etc.) and then decide what would count as a fair test of that hypothesis. The transcript already offers the internal vocabulary for that (latent task, success criterion, commitment drift, invariance), without requiring you to believe any of it on authority. ok, I kind of want to start again from the very beginning for the prompt. The root of the idea was to draw on williamson's ideas, as well as bengson and dellsen papers. to develop a prompting strategy to produce novel good philosophy. can we start from first principles as it were, talk about what ideas can be taken from all of these guys and maybe the methodology book too as a way of informing our project here. Also, different question: how can i be sure, once i try to turn this into a prompt, I am conforming to best practices of prompting.? there's plenty of information out there, but i don't know wht to trust and what not to. Starting from scratch, and treating the texts as *sources of candidate constraints* (not as settled marching orders). The project you sketched has two moving parts: 1. a *philosophical* standard for “good” and “novel” philosophical output (that’s where Timothy Williamson, John Bengson, Terence Cuneo, Russ Shafer-Landau, and Finnur Dellsén help); 2. an *engineering* question: how to encode those standards into prompts, and how to tell whether the “prompting best-practice” advice you see online is actually worth trusting. I’ll do both, but I’ll keep them clearly separated. --- ## 1) First principles: what are we trying to get a prompt to do? A prompt is (at minimum) a specification of: - an **aim** (what counts as success) - an **evidence policy** (what can be assumed, what must be flagged as hostage to missing facts) - a **construction method** (how to move from data/cases/constraints to a theory) - an **evaluation method** (how to test, compare, and revise candidates) The texts you mention give you different pieces of that spec. --- ## 2) Candidate ingredients from Williamson ### 2.1 Evidence discipline: “total evidence” as a regulativn explicitly frames abductive philosophy as starting “in principle” from our total evidence, “arguably… the total sum of human knowledge,” with no firewall between armchair methods and science/common sense. That suggests a prompting constraint like: - *Do not treat “what sounds good” as evidence.* - *Treat inconsistency with well-established knowledge as a defeater (or at least a red flag).* - *When you would need extra-empirical facts, mark the point as “empirically hostage” rather than faking it.* You can encode that without pretending the model actually has total evidence by asking it to *explicitly separate* (i) deductions from stated assumptions from (ii) claims that would require external confirmation. ### 2.2 Abduction rewards boldness onecision Williamson is very explicit that abduction rewards “boldly speculative” and “precise” theories, and that vagueness is the fake version of boldness because it evades falsification and explains nothing. Prompt implication (as a candidate design rule): **force falsifiability/precision early**. E.g., require a theory to state what it rules out, what would count against it,s it meaningfully differs from. ### 2.3 Robustness: don’t let one dodgy case crash the whole inquiry He worries about “error-fragility” in methods that treat a single misjudged thought experiment as decisive. Prompt implication: don’t let the model “win” by finding one clever-sounding counterexample. Make it generate *families* of cases, and require stability under small variations (a vtesting). ### 2.4 Simplicity as anti-overfitting (not mysticism) Williamson uses curve-fitting as an analogy: complex theories can overfit noisy data; simplicity can reduce vulnerability to error in the data. Prompt implication: treat simplicity/elegance not as “beauty points” but as a **regulariser** against fitting one’s theory too tightly to one or two beloved cases. ### 2.5 Model-building as a central He treats model-building as a recognisable methodology, especially where systems are complex, and he defines a model (for his purposes) as a “hypothetical example” described precisely enough to be tractable. Prompt implication: have the model build a *tractable toy model* of the target phenomenon, explore consequences, then check fit - instead of only trading intuitions. --- ## 3) Candidate ingredients from Bengson/Cuneo/Shafer-Landau + the methodology book ### 3.1 What counts as “good” output: theoretical understanding They characterise “theoretic properties of a theory: accuracy, being reason-based, robustness (answers many central questions), being illuminating (sometimes genuinely explanatory), being orderly, being coherent (including external fit). A prompt can turn those into an explicit rubric the model must satisfy *and show its work against*. ### 3.2 Method as criteria that move from fine methods as criteria that (i) instruct theory construction and (ii) evaluate theories, with the point being: data underdetermine theory, so method must help you home in on understanding-providing theories. This is already “prompt-shaped”: it naturally becomes a staged instruction set. ### 3.3 The Tri-LeviThey lay out the Tri-Level Method: (i) accommodate and explain the data, (ii) substantiate and integrate the theory’s claims/commitments, (iii) then (as tie-breaker) theoretical virtues. A direct prompt translation is plausible: - *Level 1:* list the data/cases/distinctions and show accommodation + explanation - *Level 2:* defend and explain your own commitments; integrate internally and with adjacent theories - \*Level 3:virtues (simplicity, elegance, etc.) as a tie-breaker Importantly, they also emphasise being **forthcoming** about when you lean on others’ work instead of supplying full substantiation/integration yourself. That maps neatly onto “don’t fabricate citations or empirical premises; mark dependencies.” ### 3.4 Objections as “this theory fails a criterion” They offer a crisp characterisation: a consideratioit gives reason to think the theory performs poorly on one or more criteria - and they map familiar objection-types to criteria (ad hoc → substantiation; contravenes science/common sense → integration; etc.). Prompt implication: have the model generate objections *typed by criterion*, rather than a blob of “possible concerns”. This makes critique more diagnostic and less performative. ### 3.5 “Novelty” via forms of progress: distinctions, refining theories, new possibilities, richer explanatory store They discuss forms olating distinctions, improving existing theories, and expanding the space of serious possibilities, supported by an expanding “explanatory store” (relations like grounding, constitution, supervenience, etc.). Prompt implication: you can demand novelty in specific modes: - “Introduce a distinction whose absence causes confusion here.” - “Give an improved version of the best existing view (minimal change, maximal payoff).” - “Add one new serious option to the space of theories, with explicit costs.” None of that assumes novelty is guaranteed - it just specifies *where* novelty is allowed to appear. --- ## 4) Candidate ingredients from Dellsén: understanding as dependency modelling Dellsén’s core idea: understanding involves grasping a model of the phenomenon’s dependence relations (causal and non-causal, including grounding), with quality determined by **accuracy** and **comprehensiveness**, which can trade off via idealisation/abstraction. This is extremely promptable because it wants an exa *dependency model*. Candidate prompt constraints: - require a dependency map: “X depends on Y; X doesequire both “positive” and “negative” dependencies (what matters and what doesn’t) - require a note on idealisations: “where did you sacrifice accuracy for scope, or scope for accuracy?” This is also a clean way to force philosophical *content* rather than rhetoric: if the model can’t state the dependence claims, it’s probably hovering. --- ## 5) Candidate ingredients from “What is philosophical progress?” (Dellsén et al.) They propose (roughly) that philosophical progress consists in *putting people in a position to increase understanding*, where understanding is better representation of dependence relations; and they explicitly distinguish what *constitutes* progress from what *promotes* it. file9 Two prompt-relevant takeaways: 1. you can treat “progress” as producing artefacts that increase the audience’s capacity to understand (distinctions, arguments, models), not necessarily final verdicts 2. you can separate modules: one module aimed at *understanding outputs* (dependency maps, explanatory organisation), another aimed at *justification/acceptance* (arguments, evidence, integration) That separation matters if you want the prompt to avoid a premature “and therefore this is the one true view” posture. --- ## 6) Putting the pieces together into a prompting strategy (still at the “ingredients” stage) Here’s one way to combine the ingredients without treating it as the final form: ### 6.1 Aims and rubric (Bengson + Dellsén) Tell the model that success is producing a theory that (a) delivers theoretical understanding (six properties) and (b) yields an explicit dependency model with accuracy/comprehensiveness trade-offs declared. ### 6.2 Evidence policy (Williamson + Bengson) Require: - explicit separation between what follows from the stated assumptions and what would require external confirmation (“empirically hostage” markers) - integration checks: flag tensions with well-established results (where relevant) rather than ignoring them ### 6.3 Construction and evaluation loop (Tri-Level + Williamson robustness) Implement an iterative loop: - Level 1: accommodate/explain data - Level 2: substantiate/integrate commitments - Level 3: only then virtues as tie-breaker - robustness tests: vary cases; don’t let one case dominate ### 6.4 Novelty knobs (Bengson + Williamson) Make “novelty” admissible in controlled ways: - produce one new distinction - produce one new serious possibility - produce one toy-model / tractable hypothetical and explore it - produce one “bold but precise” conjecture and list what it rules out That’s a strategy that’s shaped by the texts, but doesn’t assume miracles. --- ## 7) Different question: how to know what to trust about “prompting best practices”? There’s no magic certificate for “I have complied with best practices.” The reliable route is: **trust a narrow set of sources, then empirically test your prompt against failure modes**. ### 7.1 What sources are usually worth higher trust (and why) A practical trust ordering: - Official docs from the major model providers: OpenAI [OpenAI Help Center +1](https://help.openai.com/en/articles/6654000-best-practices-for-prompt-engineering-with-the-openai-api?utm_source=chatgpt.com), Anthropic [platform.claude.com](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-4-best-practices?utm_source=chatgpt.com), Google [Google AI for Developers +1](https://ai.google.dev/gemini-api/docs/prompting-strategies?utm_source=chatgpt.com), Microsoft [Microsoft Learn](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/concepts/prompt-engineering?view=foundry-classic&utm_source=chatgpt.com) Reason: they reflect model-specific behaviour (instruction following quirks, formatting that tends to work, etc.). - Provider cookbooks / engineering guides: e.g., [developers.openai.com](https://developers.openai.com/cookbook/examples/gpt4-1_prompting_guide/?utm_source=chatgpt.com) Reason: they’re closer to “tested patterns” than generic advice. - Everything else (blogs, prompt marketplaces, social media): lower trust by default. Some of it is fine, but it’s noisier and often written to sound confident. ### 7.2 Cross-provider “stable-ish” prompting advice (the overlap) What repeats across the high-trust sources: - **Be explicit and specific** about the task and success criteria [OpenAI Help Center +2 Microsoft Learn +2](https://help.openai.com/en/articles/6654000-best-practices-for-prompt-engineering-with-the-openai-api?utm_source=chatgpt.com) - **Use structure and delimiters** to separate instructions/context/data/output format [Google AI for Developers +1](https://ai.google.dev/gemini-api/docs/prompting-strategies?utm_source=chatgpt.com) - **Define terms/parameters** when ambiguity will matter [Google AI for Developers](https://ai.google.dev/gemini-api/docs/prompting-strategies?utm_source=chatgpt.com) - **Break complex tasks into steps** (or stages) and require intermediate artefacts (lists, rubrics, maps) [OpenAI Help Center +1](https://help.openai.com/en/articles/6654000-best-practices-for-prompt-engineering-with-the-openai-api?utm_source=chatgpt.com) - **Avoid blanket negative constraints** that accidentally block normal reasoning; prefer targeted constraints (this is explicitly noted in Google’s Gemini prompting guidance) [Google Cloud Documentation](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/start/gemini-3-prompting-guide?utm_source=chatgpt.com) ### 7.3 The only way to be “sure”: treat prompt-writing as experiment design If your prompt is meant to produce “novel good philosophy”, you can test it like a methodologist would: - **Ablation tests:** remove one instruction (evidence policy, dependency map requirement, robustness tests) and see what breaks. If nothing breaks, that instruction might be decorative. - **Paraphrase invariance:** restate the same task with different wording; see whether outputs preserve commitments. - **Adversarial cases:** feed it cases that tempt overfitting or vagueness; see if it stays precise (Williamson’s error-fragility worry, operationalised). - **Rubric scoring:** score outputs against the Bengson “theoretical understanding” properties and Dellsén accuracy/comprehensiveness, as explicitly as possible. This is how you avoid “trusting the vibe” of any prompting tip, including ones from respectable sources. --- ## Where this sets up the next step (without jumping ahead) A natural continuation is to turn the “ingredients” above into a *single prompt skeleton* that (i) forces a dependency model output, (ii) enforces Tri-Level evaluation, (iii) includes Williamson-style robustness/precision checks, and (iv) has s (“new distinction”, “new possibility”, “toy model”). That can then be stress-tested with the experiment-style checks above, so the prompt earns its keep rather than being a manifesto. 1\. "1) First principles: what are we trying to get a prompt to do? A prompt is (at minimum) a specification of: an aim (what counts as success)" i think things are actually more subtle than that. For example, an interesting way to get LLMs to generate good novel ideas is the 'make the obvious move' trick. this is my rule of thumb so don't get too excited/overlyfocussed. In this way I am not specifying an aim specifically am I? well, i suppose I am but i still feel that what i am trying to articulate here has something to it. can you help me finish the thought? 2. "an evidence policy (what can be assumed, what must be flagged as hostage to missing facts)" Please note that real-world evidence is not actually that relevant fro a lot of analytic philosophy. of course some work relies on empirical stuff, but often the real world evidence that is relied upon in analytic philosophy is so commonsensical as to be unremarked. now, here I do not what I am alluding to, but do you? –think about recent philosophy papers (look up the most cited analytic philosophy papers of last year or something, you will see very very few are hostage to missing facts? Or do you disagree? 3. in fact the whole of this: "1) First principles: what are we trying to get a prompt to do? A prompt is (at minimum) a specification of: an aim (what counts as success) an evidence policy (what can be assumed, what must be flagged as hostage to missing facts) a construction method (how to move from data/cases/constraints to a theory) an evaluation method (how to test, compare, and revise candidates) The texts you mention give you different pieces of that spec." seems extreley speculative and not actually informed by anyone. find real life RELIABLE sources for this sort of prompt craft and justify your decisions as to who to trust. the date is 2 Feb 2026 Please start the task(s) again completely from scratch, keeping this in mind. And please put MUCH more effort and give this a lot more thought because that was not a good start art all. Starting again from zero, and treating *everything* here as provisional hypotheses plus grounded citations. I’m going to do three things, matching your numbered points: 1. help you finish the thought about why “make the obvious move” feels like it’s doing something *other than* (or at least *not transparently like*) specifying an aim; 2. take your “real-world evidence isn’t that relevant in a lot of analytic philosophy” point seriously, and check it against at least one respectable empirical proxy (not vibes); 3. rebuild “prompt craft” from *reliable* sources (provider docs + research literature), and explain why those sources are worth more trust than the general internet soup. --- ## 1) “Make the obvious move” and the subtlety about aims ### A useful distinction: explicit objectives vs procedural steering One way to cash out your intuition is this: - Sometimes a prompt works by stating a **target property of the final output** (“produce X in format Y”). - Sometimes a prompt works by specifying a **procedure / move-type** that forces the model into a different *mode of continuation*, without describing the final destination in much detail. This second thing is not me inventing a category out of thin air - it shows up in how major providers classify prompts. For example, Google’s Gemini prompting guide explicitly distinguishes input types including “question input”, “task input”, “entity input”, and crucially “completion input” (where you give partial text and ask the model to continue it). That is already a recognition that prompts can be “continue this” rather than “achieve this explicit end state.” [Google AI for Developers](https://ai.google.dev/gemini-api/docs/prompting-strategies) Your “make the obvious move” reads naturally as a **completion / continuation directive** over a *dialectical trajectory*: “given what we’ve said, continue the argument in the way an expert would regard as the next step.” That can feel less like specifying an “aim” because you’re not describing a finished product - you’re selecting a *next action in a practice*. ### The aim is there, but it’s indexical and practice-relative A good way to finish your thought (as a candidate formulation) is: > “I’m not specifying a detailed end state; I’m issuing an indexical directive that relies on shared norms of a practice. ‘Make the obvious move’ means: given the current dialectical position, select the next step that counts as licensed/pressing *by the standards of this discourse*, and commit to it.” That’s a different kind of “aim specification”: it’s not a blueprint, it’s a *deictic* constraint (“from here, go *there* ”), where *there* is fixed by the conversational state plus disciplinary norms. This is also very close in spirit to well-studied prompting effects like “chain-of-thought prompting”: simply eliciting intermediate steps can change performance dramatically without specifying any richer end goal than “show your reasoning.” [arXiv](https://arxiv.org/abs/2201.11903?utm_source=chatgpt.com) Anthropic’s own prompt-engineering overview lists “Let Claude think (chain of thought)” as a general technique category, again treating “procedure prompts” as first-class. [Claude Developer Platform](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/overview) So the subtle point you’re gesturing at can be made pretty cleanly: - Some prompts work by **describing desired outputs**. - Some prompts work by **activating a policy** (a way of proceeding) that is already latent in the model’s training, and which only needs a small “mode switch” cue. Your “obvious move” is plausibly of the second type. ### Why this matters for your project (without pretending anything is settled) If your project is “prompting strategies for novel good philosophy”, this suggests at least one live hypothesis: - For philosophy, *procedural steering prompts* might be more powerful than “rubric prompts”, because they lean on the model’s implicit grasp of what counts as an “appropriate next move” in philosophical writing. That hypothesis is testable, and (importantly) doesn’t require pretending we’ve already defined “success criteria” in some final way. --- ## 2) “Evidence policy” and analytic philosophy: is “hostage to missing facts” rare? Two separate issues get mixed up here, and I think this is what you’re alluding to: 1. **Front-stage evidence** in analytic philosophy (explicitly cited experiments, datasets, measurements) is often minimal outside certain subfields. 2. **Background constraints** are everywhere: ordinary knowledge, Moorean truisms, broadly accepted science, linguistic competence, mathematical facts, etc. They often go unremarked because they’re treated as common ground. So when someone says “total evidence” (Williamson-style) or “hostage to missing facts,” it can sound like they’re demanding a level of empirical tethering that simply isn’t how much analytic philosophy actually runs *most of the time*. ### A decent empirical proxy: which recent papers are most cited within philosophy journals? You suggested looking at “most cited analytic philosophy papers of last year or something.” The “last year” part is genuinely awkward because citation half-lives in philosophy are slow, and 2025 papers often won’t have stabilised citation profiles by 2 Feb 2026. But there is a *credible* proxy in the direction you want: Brian Weatherson ran Web of Science analyses restricted to philosophy journals, asking what papers (published <10 years old) are most cited *at a given time* - a reasonable operationalisation of “what philosophers are talking about.” [Brian Weatherson](https://brian.weatherson.org/quarto/blog/top-ten/top-ten.html) When you look at the 2012–2022 tables, the most-cited-recent papers are overwhelmingly in epistemology/metaphysics/semantics and are not the kind of thing that depends on discovering new empirical facts next week. Here are a few entries from 2020–2022 as listed on Weatherson’s page: Wilson on grounding, Pritchard on epistemology, Plunkett & Sundell on disagreement/semantics, Dasgupta on physicalism, etc. [Brian Weatherson](https://brian.weatherson.org/quarto/blog/top-ten/top-ten.html) That supports *something like* your claim: much highly influential “core” analytic philosophy is not routinely hostage to missing empirical facts. ### But: it’s not zero, and the background “commonsense” base still counts as evidence Even in Weatherson’s lists you see papers that explicitly engage empirical work or experimental philosophy - e.g., Nichols and Knobe on folk intuitions is literally described in cognitive-science terms. [Brian Weatherson](https://brian.weatherson.org/quarto/blog/top-ten/top-ten.html) So I’d put it like this (again: as a candidate articulation, not a verdict): > In a lot of analytic philosophy, the evidential base is dominated by background truisms, conceptual constraints, linguistic competence, and theoretical virtues; front-stage empirical evidence is concentrated in certain subfields. “Hostage to missing facts” is therefore not the *typical* failure mode - but “hostage to tacit presuppositions” absolutely is. And that matters for prompting: the practical “evidence policy” you need is often less about “go fetch new data” and more about *not inventing* empirical claims, plus making tacit presuppositions explicit when they do work. Google’s Gemini 3 prompting guide has a very relevant caution here: broad negative instructions like “do not infer” can backfire by blocking basic reasoning; instead, you should constrain the model to use *provided context* for deductions and explicitly forbid introducing outside information. [Google Cloud Documentation](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/start/gemini-3-prompting-guide) That is basically an engineering version of your point: most of the time you want disciplined reasoning over shared/common ground, not fake “research.” --- ## 3) Prompt craft from reliable sources: who to trust, and why? You’re right to push back on “here’s a four-item spec I made up.” If we want something sturdier, we should build from sources that (a) are close to the systems, and/or (b) are empirically grounded and widely scrutinised. Here’s a defensible trust hierarchy, with reasons and citations. ### Tier 1: Model-provider documentation (primary, model-specific, frequently updated) These aren’t holy writ, but they have two big advantages: they’re written with internal knowledge of model behaviour, and they get updated as models change. - **OpenAI Help: “Best practices for prompt engineering…”** explicitly recommends things like separating instructions from context, being specific about outcome/format, and using examples to enforce output structure. [OpenAI Help Center](https://help.openai.com/en/articles/6654000-best-practices-for-prompt-engineering-with-the-openai-api) - **OpenAI API prompt-engineering guide** explicitly recommends building evals to measure prompt behaviour and pinning model snapshots for consistency. That’s not “prompt influencer lore”; it’s operational guidance for controlling non-deterministic systems. [platform.openai.com](https://platform.openai.com/docs/guides/prompt-engineering) - **Anthropic prompt-engineering overview** is blunt: before “prompt engineering” you should have (1) a clear definition of success criteria, and (2) ways to empirically test against them. [Claude Developer Platform](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/overview) - **Google Gemini prompt design strategies** treats prompt design as iterative and explicitly categorises input types (question/task/entity/completion), which is directly relevant to your “obvious move” being more like completion-steering than goal-specification. [Google AI for Developers](https://ai.google.dev/gemini-api/docs/prompting-strategies) - **Microsoft/Azure OpenAI prompt engineering techniques** gives concrete “best practices” like “Be Specific… restrict the operational space,” and “Give the model an out… respond with ‘not found’ if the answer isn’t present,” which is exactly a practical anti-fabrication policy. [Microsoft Learn](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/concepts/prompt-engineering?view=foundry-classic) Why trust Tier 1 more than random guides? Because the claims are typically phrased as *tested tactics for these systems* and are maintained as living docs (with clear update timestamps, e.g., OpenAI’s help articles updated 22 days ago; Microsoft’s page updated 2025-12-06; Google Cloud updated 2026-01-14). [OpenAI Help Center +2 Microsoft Learn +2](https://help.openai.com/en/articles/6654000-best-practices-for-prompt-engineering-with-the-openai-api) ### Tier 2: Widely cited prompting research (mechanisms and comparative evidence) This is where we get evidence that certain “small prompt moves” reliably change behaviour. - **Chain-of-thought prompting**: demonstrates that eliciting intermediate reasoning steps can substantially improve performance on reasoning tasks. [arXiv](https://arxiv.org/abs/2201.11903?utm_source=chatgpt.com) - **Self-consistency**: samples multiple reasoning paths and selects the most consistent, improving CoT performance - which is directly relevant to your goal of avoiding “safe mush” and instead selecting stable commitments. [arXiv](https://arxiv.org/abs/2203.11171?utm_source=chatgpt.com) - **Tree of Thoughts**: explicitly frames improved performance as coming from exploration over “thoughts” and selection/self-evaluation - conceptually adjacent to your “make the move, then test/repair” workflow. [arXiv](https://arxiv.org/abs/2305.10601?utm_source=chatgpt.com) - **ReAct**: shows the value of interleaving reasoning with action; even if you’re not using tools, the general lesson is that structure in the interaction protocol can matter. [arXiv](https://arxiv.org/abs/2210.03629?utm_source=chatgpt.com) Why trust Tier 2? Because these are not vibes; they’re methods proposed with empirical evaluation and heavy scrutiny (and massive citation counts in the case of CoT). [arXiv +1](https://arxiv.org/abs/2201.11903?utm_source=chatgpt.com) ### Tier 3: Surveys and reviews (useful map, weaker authority) Surveys help you not miss entire families of techniques, but they inherit limitations: they summarise fast-moving literature, and not every included technique is equally robust. - A systematic survey of prompt engineering (Sahoo et al.) explicitly aims to organise techniques and discuss strengths/limitations. [arXiv](https://arxiv.org/abs/2402.07927) - A 2025 open-access review in *Patterns* discusses prompt engineering methods and explicitly notes that methods are evaluated with subjective and objective metrics. [ScienceDirect](https://www.sciencedirect.com/science/article/pii/S2666389925001084) Why trust Tier 3 cautiously? They’re good for orientation and taxonomy, not automatic “do this and it will work.” ### Tier 4: Blogs, “prompt templates”, influencer frameworks (lowest default trust) Some of this is excellent, but as a class it’s unreliable because: - claims are rarely model-version specific, - evaluation is usually anecdotal, - incentives reward confidence and novelty over replicability. If you use them, it’s best as *hypothesis generators*, not authorities. --- ## So how can you be confident you’re following “best practices” for your philosophy prompt? The most reliable answer, consistent with Tier 1 guidance, is: **you don’t trust a practice because it’s called a best practice - you trust it because it survives evals designed around your objective.** This is exactly the stance both Anthropic and OpenAI take in their engineering-facing material: define success criteria, create empirical evaluations, iterate, and use regression testing to avoid “it got worse and we don’t know why.” [Claude Developer Platform +2 platform.openai.com +2](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/overview) ### A minimal eval mindset you can steal for philosophy (without overbuilding) From Anthropic’s evals piece: an eval is “a prompt, a response, and grading logic,” and you run multiple trials because outputs vary. [Anthropic](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) Translate that into your context: - Fix 10–20 micro-tasks that represent “good philosophy” in your sense (e.g., generate a new distinction; generate a novel objection; propose a dependency model; repair under constraints). - For each task, you create a rubric (even if rough) and run repeated trials. - You keep a small “regression set” so when you tweak the prompt, you can see what got better and what broke. That *is* how you get warranted confidence, and it doesn’t depend on trusting the internet. --- ## Bringing this back to your two philosophical pressures (without jumping ahead) 1. Your “obvious move” idea looks like a member of a broader, well-supported category: **procedural steering prompts** (completion/mode prompts) rather than explicit “final output specs.” Provider docs explicitly recognise completion prompts as a type, and research literature shows small procedure cues can matter a lot. [Google AI for Developers +1](https://ai.google.dev/gemini-api/docs/prompting-strategies) 2. Your “commonsense evidence is usually unremarked” point has real bite. And when you look at a citation-based proxy of influential recent work within philosophy journals, the pattern is broadly consistent with your hunch: much highly cited work isn’t waiting on missing empirical facts, though some subfields are. [Brian Weatherson +1](https://brian.weatherson.org/quarto/blog/top-ten/top-ten.html) If you keep those two pressures in view, you get a very plausible project shape: prompt design that (i) steers the model into “make a move” mode, (ii) forces commitments, (iii) runs a lightweight internal test/repair loop, and (iv) is validated by a small, repeatable eval suite rather than by folklore. ### 1) What I meant by “rubric prompts” (and why that label can be misleading) I wasn’t claiming there’s a canonical category called “rubric prompts” in everyone’s docs. I was using it as a shorthand for prompts that **embed an explicit scoring/criterion sheet** for the output. There are (at least) two *different* rubric uses that get mixed together: **A. Rubric for *evaluation* (grading).** This is extremely standard in serious LLM engineering. Anthropic explicitly recommends “detailed, clear rubrics” for grading outputs in evals. [Claude Developer Platform +1](https://platform.claude.com/docs/en/test-and-evaluate/develop-tests?utm_source=chatgpt.com) They also describe using rubrics in real internal testing (system cards) - e.g., “rubrics designed for each individual test case.” [Anthropic +1](https://www.anthropic.com/claude-sonnet-4-5-system-card?utm_source=chatgpt.com) And in their multi-agent research system write-up, they describe judging outputs against rubric dimensions like factual accuracy, citation accuracy, completeness, etc. [Anthropic](https://www.anthropic.com/engineering/multi-agent-research-system?utm_source=chatgpt.com) **B. Rubric for *generation* (writing under criteria).** Here the rubric isn’t just for a judge - it’s part of the instructions, e.g. “your answer will be assessed on X, Y, Z; therefore satisfy X, Y, Z.” There’s also a research-y version of this, where a model *generates rubrics* and then evaluates candidates against them (Chain-of-Rubrics / CoR). For example, RM-R1 explicitly introduces a “chain-of-rubrics mechanism” where the model self-generates rubrics and uses them to evaluate candidate responses. [arXiv +1](https://arxiv.org/pdf/2505.02387?utm_source=chatgpt.com) So, if you want a cleaner terminology (to avoid accidental reification): - **Rubric-as-constraints prompts** (rubric inside the instruction) - **Rubric-as-judge prompts** (rubric used to grade) Your “make the obvious move” trick is not naturally rubric-like in either sense. It’s closer to *procedural steering*. --- ### 2) Developing the idea: when procedural steering is effective, and why analytic philosophy might resemble those cases Here’s a workable hypothesis to explore (not a conclusion): > Some tasks benefit more from prompts that specify *how to proceed* (a procedure, a move-type, a search pattern) than from prompts that specify *what the final product should satisfy* (a rubric/constraints list). Analytic philosophy may be one of those tasks because it is structured as a sequence of licensed moves in a dialectical practice. To develop this, it helps to look at domains where “procedure prompts” have solid evidence behind them, and then ask whether philosophy shares the relevant structure. #### A) Domains where procedural steering has empirical backing **(i) Reasoning problems: “show intermediate steps” (CoT).** Chain-of-thought prompting improves performance on a range of reasoning benchmarks by eliciting intermediate steps. [arXiv +1](https://arxiv.org/abs/2201.11903?utm_source=chatgpt.com) Interpretation: the instruction is not “produce a good answer” (rubric-ish), but “reason *like this*.” **(ii) Search and selection: explore multiple paths, then choose (Self-consistency, Tree-of-Thoughts).** Self-consistency samples multiple reasoning paths and selects the most consistent outcome, improving results over single greedy chains. [arXiv +1](https://arxiv.org/abs/2203.11171?utm_source=chatgpt.com) Tree-of-Thoughts makes the “procedure” even more explicit: generate candidate “thoughts”, evaluate, backtrack, etc. [arXiv +1](https://arxiv.org/abs/2305.10601?utm_source=chatgpt.com) Interpretation: when the space of continuations is large and local plausibility is cheap, procedures that add exploration + selection can help. **(iii) Interactive tasks: interleave reasoning with actions (ReAct).** ReAct explicitly interleaves reasoning traces with actions, which can reduce hallucination/error propagation in some settings. [arXiv +1](https://arxiv.org/abs/2210.03629?utm_source=chatgpt.com) Interpretation: it’s again “do the task by following this kind of trajectory,” not “output must meet these final criteria.” **Caution that matters for your project:** procedural prompting isn’t magic or universal. There’s evidence that CoT gains can be quite problem-class specific and can deteriorate as prompts become less well matched to the task. [NIPS](https://nips.cc/virtual/2024/poster/93898?utm_source=chatgpt.com) So the interesting question becomes: *what makes a domain “procedure-friendly” rather than “rubric-friendly”?* #### B) What those domains have in common (the likely mechanism) A plausible commonality is: there is a **recognisable trajectory** that trained humans follow (solve-by-steps, search-then-select, plan-then-execute). The prompt nudges the model into emitting text that matches that trajectory, which may also better match the latent patterns it has learned. So “procedural steering” tends to work when: - there’s a stable, culturally entrenched *way of proceeding*; - the next correct move is often underdetermined by surface phrasing but constrained by the “game” being played; - quality depends on *path properties* (coherence over steps), not only final-output properties. #### C) Why analytic philosophy might be structurally similar Analytic philosophy (in its mainstream journal style) often has a fairly rigid move grammar: - fix a target claim; - define/clarify terms and distinctions; - give an argument (often modular); - consider a salient objection; - revise/strengthen; - explain payoffs (scope, implications, connections). That’s not “rubric scoring” in the first instance; it’s a *procedure in a practice*. Which makes your “make the obvious move” feel apt: it’s basically “continue the dialectic from here, don’t hover”. In other words, “make the obvious move” looks like a **completion directive over a practice-governed trajectory** (more like CoT/ToT than like a checklist of final answer properties). #### D) A concrete way to compare “procedural” vs “rubric” prompting in philosophy You can make this comparison crisp without committing to grand theory: - **Rubric condition:** “Produce a novel philosophical contribution that will be evaluated on (i) definitional precision, (ii) non-triviality, (iii) objection-handling, (iv) integration, (v) explanatory payoff.” - **Procedural condition:** “From the current dialectical position, make the next obvious move. Then make the next obvious move again. Repeat 3 times, each time tightening commitments.” - **Hybrid condition:** procedural steps *plus* an explicit rubric used only at the end to self-check. Then evaluate with a rubric-as-judge approach (the engineering standard for “is this better?”). [Anthropic +1](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents?utm_source=chatgpt.com) If procedural steering really is doing something distinctive in philosophy, you’d expect: - less “overview fog”; - more commitments earlier; - more naturally placed objections/replies; - more “move-like” novelty (new distinctions, unexpected but licenced transitions), even if not always globally correct. --- ### A small clarification that may help your next step Procedural vs rubric isn’t a binary. Some of the strongest modern prompting techniques are basically **procedures that implement an implicit rubric** (explore candidates, evaluate, select, repair). Self-consistency and Tree-of-Thoughts fit that pattern. [arXiv +1](https://arxiv.org/abs/2203.11171?utm_source=chatgpt.com) So one direction for your project is: treat “make the obvious move” as the *primitive procedural operator*, then decide what kind of lightweight evaluation step (rubric-like) you want to add *after* the move, not before it. That sets you up to explore the interesting comparison you flagged: identify other domains where “next-move” style prompts work, and see whether analytic philosophy shares the same procedural invariants. "2) “Evidence policy” and analytic philosophy: is “hostage to missing facts” rare?" all of what you say here is excellent and I will add it to my draft. "If you keep those two pressures in view, you get a very plausible project shape: prompt design that (i) steers the model into “make a move” mode, (ii) forces commitments, (iii) runs a lightweight internal test/repair loop, and (iv) is validated by a small, repeatable eval suite rather than by folklore." this is you jumping ahead. 3. I am even more confused about rubric than I was. I understand everything else you are saying, is the term important, or a word that maybe you shouldn't have used which has derailed me slightly? Noted on the “jumping ahead”. That “plausible project shape” line was me sketching one possible direction rather than staying inside the slow “we’re still clarifying the moving parts” phase. Let’s bracket it entirely and keep the focus on conceptual clarification, not implied roadmap. On the “rubric” derail: I think this is mostly a terminology mistake on my side. ### What I meant by “rubric” (in plain terms) In current LLM practice, a “rubric” is just a *grading guide*: a list of criteria a judge uses to assess an output (clarity, correctness, completeness, etc.). Anthropic uses the term exactly like that: they describe evaluations that combine unit tests with “an LLM rubric for assessing overall code quality”. [Anthropic](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents?utm_source=chatgpt.com) They also describe judging outputs “against criteria in a rubric” (factual accuracy, citation accuracy, completeness, etc.). [Anthropic](https://www.anthropic.com/engineering/multi-agent-research-system?utm_source=chatgpt.com) Their Claude evaluation docs likewise talk about LLM-based grading and how to make it reliable - the core idea being structured judgement rather than vibes. [Claude Developer Platform](https://platform.claude.com/docs/en/test-and-evaluate/develop-tests?utm_source=chatgpt.com) So, when I said “rubric prompt”, I was trying to gesture at *prompts that explicitly encode criteria* (“this will be assessed on X, Y, Z…”) as opposed to prompts that mostly steer the process (“make the next move”). ### Why it was a bad label here “Rubric prompt” isn’t a standard, settled term in the way “few-shot prompting” or “chain-of-thought” are. What *is* standard is: - “rubric” as a thing used in **evals** (grading), per Anthropic’s engineering and docs. [Anthropic +2 Claude Developer Platform +2](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents?utm_source=chatgpt.com) - and “criteria”/“requirements”/“output format” as things you can include in prompts (provider prompt guides talk like this, not necessarily using the word “rubric”). [OpenAI Platform +1](https://platform.openai.com/docs/guides/prompt-engineering?utm_source=chatgpt.com) So yes: the word was avoidable, and it’s not important enough to keep if it’s slowing you down. ### The distinction that is worth keeping (without the word) What I was aiming at is just this contrast: 1. **Criteria-first prompting**: you tell the model what the answer must satisfy (e.g., “give a definition, a core argument, two objections, a reply, no hedging, explicit scope…”). This is “constraints/criteria” prompting. 2. **Procedure-first prompting**: you tell the model *how to proceed* from where we are (“make the next obvious dialectical move”; then “attack it”; then “repair it”). This is “procedural steering”. Your “make the obvious move” trick is squarely (2). It can work even when you haven’t spelled out an explicit end-product specification, because it’s a directive keyed to shared norms of a practice (dialectical continuation), not a checklist of output properties. ### So: should we drop the term? I think yes, unless you explicitly want it. We can just say: - “criteria prompts” (or “constraints prompts”) - versus “procedural prompts” (or “steering prompts”) and keep “rubric” reserved for the eval context, where reputable sources actually use it. [Anthropic +2 Claude Developer Platform +2](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents?utm_source=chatgpt.com) "The distinction that is worth keeping (without the word) What I was aiming at is just this contrast: Criteria-first prompting: you tell the model what the answer must satisfy (e.g., “give a definition, a core argument, two objections, a reply, no hedging, explicit scope…”). This is “constraints/criteria” prompting. Procedure-first prompting: you tell the model how to proceed from where we are (“make the next obvious dialectical move”; then “attack it”; then “repair it”). This is “procedural steering”. Your “make the obvious move” trick is squarely (2). It can work even when you haven’t spelled out an explicit end-product specification, because it’s a directive keyed to shared norms of a practice (dialectical continuation), not a checklist of output properties." good. this is cool. I would like us to really stress test and develop potentially the ideas you talked about here "2) Developing the idea: when procedural steering is effective, and why analytic philosophy might resemble those cases" Please give me a long, deep, and detailed answer, with lots of options/ideas: as we are in the 'kicking ideas around' stage, I want as much feedback as you can give me so that i have a lot to consider. Alright - let’s really lean into the “stress test” mindset, and keep everything as *candidate hypotheses* plus comparisons, not as a forecast of where the project must go. I’ll do this in three passes: 1. map a bunch of *cases* where procedural steering (in the broad “tell it how to proceed” sense) seems to help, with reliable sources; 2. extract some *candidate common denominators* - why those cases might be procedure-friendly; 3. put analytic philosophy next to those denominators and see where the analogy looks strong, where it looks weak, and what kinds of prompt moves would actually test it. --- ## 1) Where procedural steering has evidence behind it Here “procedural steering” = prompts that primarily constrain *trajectory* (step structure, branching, revision cycles) rather than listing output requirements. ### A) Step-by-step reasoning as a procedure cue The canonical example is *chain-of-thought prompting*: giving or eliciting intermediate steps improves performance on a range of reasoning tasks. [arxiv.org](https://arxiv.org/abs/2201.11903?utm_source=chatgpt.com) What matters for your purposes is not the marketing label “reasoning”, but the mechanism: the prompt constrains the model to generate a certain kind of continuation (a sequence of intermediate steps) rather than a single “answer-shaped” blob. [arxiv.org](https://arxiv.org/abs/2201.11903?utm_source=chatgpt.com) A useful extension is *self-consistency*: sample multiple reasoning paths, then select the most consistent answer. This is procedural steering plus selection pressure. [arxiv.org](https://arxiv.org/abs/2203.11171?utm_source=chatgpt.com) Candidate lesson: some tasks benefit when you force the model to (i) generate *paths*, not just endpoints, and (ii) treat path diversity as a feature rather than noise. [arxiv.org](https://arxiv.org/abs/2203.11171?utm_source=chatgpt.com) ### B) Decomposition procedures: “solve subproblems first” Two strong, well-cited exemplars: Least-to-most prompting: decompose a hard problem into simpler subproblems and solve them sequentially; it’s explicitly motivated as overcoming “easy-to-hard generalisation” failures of simpler CoT prompting. [arxiv.org +1](https://arxiv.org/abs/2205.10625?utm_source=chatgpt.com) Plan-and-solve prompting: explicitly add a “make a plan first, then execute it” stage, aimed at avoiding missing-step and misunderstanding errors in zero-shot CoT. [arxiv.org](https://arxiv.org/abs/2305.04091?utm_source=chatgpt.com) Candidate lesson: when the task has a latent multi-stage structure, the model often needs an explicit procedure that forces it to surface that structure (plan/decompose) before it tries to sprint to the finish. [arxiv.org +1](https://arxiv.org/abs/2305.04091?utm_source=chatgpt.com) ### C) “Step back” / abstraction-first procedures Step-back prompting literally instructs the model to abstract to higher-level principles first, then apply them to the concrete problem, and reports substantial gains on reasoning-intensive tasks. [arxiv.org](https://arxiv.org/abs/2310.06117?utm_source=chatgpt.com) Candidate lesson: sometimes the helpful procedure is not “go forward in more detail”, but “go up a level and pick the governing frame”, then come back down. [arxiv.org](https://arxiv.org/abs/2310.06117?utm_source=chatgpt.com) ### D) Branching + backtracking procedures: make it search Tree of Thoughts (ToT) makes the procedural structure explicit: generate multiple candidate “thoughts”, evaluate them, backtrack/prune, and so on. The paper reports improvements on tasks that require planning/search, including a creative writing task. [proceedings.neurips.cc](https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac703-Paper-Conference.pdf?utm_source=chatgpt.com) Candidate lesson: where the space of plausible continuations is huge and local plausibility is cheap, single-path continuation can get stuck in a locally smooth but globally mediocre groove; branching and evaluation can help. [proceedings.neurips.cc](https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac703-Paper-Conference.pdf?utm_source=chatgpt.com) ### E) Revision loops as a procedure: feedback → refine → feedback Self-Refine is basically “draft, critique, rewrite” as a prompting protocol, and the authors report improvements across multiple tasks without extra training (same model plays generator/critic/refiner). [arxiv.org](https://arxiv.org/abs/2303.17651?utm_source=chatgpt.com) Reflexion is a more agentic variant: use feedback signals, verbalise reflections, store them, and do better in later trials; again, procedural structure matters more than “ask nicely”. [arxiv.org](https://arxiv.org/abs/2303.11366?utm_source=chatgpt.com) Candidate lesson: for many tasks, the first pass is not the main event. The procedure matters because it creates a controlled kind of second-pass pressure. [arxiv.org +1](https://arxiv.org/abs/2303.17651?utm_source=chatgpt.com) ### F) Debate / multi-agent procedures: make it argue with itself There’s also evidence that multi-model or multi-instance debate procedures can improve reasoning/factuality on certain tasks. [arxiv.org](https://arxiv.org/abs/2305.14325?utm_source=chatgpt.com) Candidate lesson: “procedural steering” can include social-structure emulation: you’re not only telling the model *what to output*, but also staging the epistemic situation (two advocates plus a judge). --- ## 2) What seems to make procedural steering work (candidate common denominators) This is the part that matters for your “compare analytic philosophy” move. I’ll keep each as a *hypothesis* you can accept, reject, or modify. ### Hypothesis 1: The task has a stable internal grammar of intermediate states Procedures help when there are recognisable intermediate states that humans use: plan → subproblems → solve → check; or abstract → instantiate; or propose → evaluate → backtrack. That’s exactly what these prompting papers are codifying. [arxiv.org +2 proceedings.neurips.cc +2](https://arxiv.org/abs/2305.04091?utm_source=chatgpt.com) If the task doesn’t have such a grammar (or it’s not stable across contexts), a procedure prompt can just be theatre. ### Hypothesis 2: Local plausibility is cheap, but global coherence is expensive ToT is almost explicit about this: you need “deliberate decision making”, evaluation, and backtracking to make global choices. [proceedings.neurips.cc](https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac703-Paper-Conference.pdf?utm_source=chatgpt.com) Translation: if the model can always produce something that looks locally fine, then “criteria-only” constraints often don’t prevent it from drifting into a high-entropy blob. Procedures force global coordination by making intermediate products inspectable. ### Hypothesis 3: The biggest wins come from reducing a latent ambiguity, not adding more content A lot of these methods can be seen as disambiguators: “don’t answer yet - plan”; “don’t stay concrete - abstract”; “don’t commit to one line - branch”; “don’t stop at first draft - refine”. [arxiv.org +2 arxiv.org +2](https://arxiv.org/abs/2305.04091?utm_source=chatgpt.com) Your “make the obvious move” trick fits this pattern: it doesn’t add information; it selects a *mode of continuation*. ### Hypothesis 4: Selection pressure (even crude) is more important than the first generation Self-consistency and ToT are both selection-heavy: sample diverse paths, then choose. [arxiv.org +1](https://arxiv.org/abs/2203.11171?utm_source=chatgpt.com) If you take that seriously, then what matters is not “can it generate a good line?” but “can you get it into a regime where multiple lines get compared and one is endorsed for reasons you can evaluate?” ### Hypothesis 5: The “reasoning trace” is not automatically trustworthy This matters a lot for philosophy, because philosophy loves beautiful-sounding rationales. There’s solid evidence that chain-of-thought explanations can be *plausible yet unfaithful* (the model’s stated reasoning can diverge from the true causal route to its answer). [arxiv.org +1](https://arxiv.org/abs/2305.04388?utm_source=chatgpt.com) There’s also recent work explicitly focused on measuring faithfulness of CoT (including by “unlearning”), which suggests the community takes this as a live problem. [aclanthology.org](https://aclanthology.org/2025.emnlp-main.504.pdf?utm_source=chatgpt.com) Candidate implication: procedural steering can improve outputs while simultaneously producing persuasive-but-misleading rationalisations. So if you import procedural methods into philosophy, you need a stance like: treat “the trace” as an *argument object* to criticise, not as introspective access to the model’s cognition. That aligns very naturally with analytic philosophy’s attitude anyway: “show your work” is not “trust my soul”, it’s “here is something you can attack”. --- ## 3) Why analytic philosophy might resemble the procedure-friendly cases (and where it might not) Now the fun bit: put analytic philosophy in the same frame and see what follows. ### A) The structural analogy that looks strongest: philosophy has a move grammar Even if you reject any heavy metaphilosophy, journal-style analytic philosophy does have a fairly stable move grammar: state a problem, disambiguate, propose a thesis, argue, consider objections, reply, indicate scope, connect to neighbours. That makes it look closer to tasks where procedural prompting helps: multi-step reasoning, planning, and revision loops. [arxiv.org +2 arxiv.org +2](https://arxiv.org/abs/2305.04091?utm_source=chatgpt.com) So “make the obvious move” can be read as: “continue according to that move grammar; stop hovering.” ### B) The “obvious move” trick has an interesting resemblance to decomposition prompts Least-to-most and plan-and-solve are basically: don’t freewheel; decompose. [arxiv.org +1](https://arxiv.org/abs/2205.10625?utm_source=chatgpt.com) In philosophy, “the obvious move” is often precisely a decomposition move, for example: - “distinguish two senses of X” - “state the argument schema explicitly” - “exhibit the key premise that’s doing the work” - “separate metaphysical from epistemic claims” - “give the best objection and then see what survives” So a hypothesis worth exploring is: your trick works because it triggers a *default decomposition operator* in the learned philosophy-writing policy. That’s attractive because it explains why the trick can work without a detailed final specification. ### C) Philosophy also looks like a “search in conceptual space” problem ToT’s selling point is deliberate search and backtracking. [proceedings.neurips.cc](https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac703-Paper-Conference.pdf?utm_source=chatgpt.com) A lot of philosophical progress (at least locally) is: “try an analysis, hit a counterexample, revise; try a different analysis; unify two distinctions; trade off virtues.” That resembles a branching search over candidate positions more than a single linear computation. So ToT-like procedures may be especially apt as *models* for what you’re trying to do in philosophical prompting - not because philosophy is maths, but because the landscape of “locally plausible” continuations is huge, and global virtues (coherence, scope, fruitfulness) are hard to satisfy in one shot. ### D) The revision-loop analogy is almost too on-the-nose Self-Refine is explicitly “improve initial outputs via iterative feedback and refinement.” [arxiv.org](https://arxiv.org/abs/2303.17651?utm_source=chatgpt.com) Philosophy is practically institutionalised refinement: drafts, objections, revisions, referee reports, seminars, replies. So you can treat “procedural steering in philosophy” as *simulating the epistemic ecology* of philosophical work: not “one answer”, but “draft - adversary - repair - scope check”. The stress test here is to ask: does simulating that ecology produce genuinely better philosophical moves, or just longer, more polished prose? ### E) Where the analogy weakens: “correct answer” is often not available Self-consistency relies on the idea that many reasoning paths converge on a unique correct answer in those benchmarks. [arxiv.org](https://arxiv.org/abs/2203.11171?utm_source=chatgpt.com) In philosophy, there often isn’t a unique convergent answer, and even when there is, it’s not labelled. So self-consistency as “majority vote on answers” may be the wrong analogue; but self-consistency as “diversify paths, then select by a criterion” still makes sense if you replace “correct answer” with “best-supported dialectical position under stated constraints”. That’s not a conclusion - it’s the shape of a test: does path diversity plus criterion-based selection increase philosophical quality, or just increase apparent sophistication? ### F) A lurking danger: procedural steering might amplify clichés “Make the obvious move” risks eliciting the *most conventional* move. In philosophy, that could mean: you get clean dialectical manoeuvres, but you drift towards what’s already common in the literature. So a key stress test for your overall idea is: can you keep the benefits of procedural steering (commitment, structure, non-hovering) while nudging away from “standard seminar moves” into genuinely novel contributions? That’s where you may need procedures that deliberately create controlled deviations - not random weirdness, but “licensed novelty”. --- ## 4) Stress tests you can run on the idea itself (not on any final prompt) Here are several ways to pressure-test “procedural steering is especially effective for analytic philosophy”. Think of these as experimental knobs, not a fixed protocol. ### Test 1: Ambiguity stress - is there even a single “obvious move”? Give the same philosophical setup but craft it so there are at least three plausible “next moves” (define terms vs give counterexample vs propose a distinction). Then see whether “make the obvious move” reliably selects one, and whether the selection is stable across paraphrases. If it’s unstable, that suggests “obviousness” is doing a lot of hidden work (it depends on tacit framing), which you can either embrace (as practice-relative) or treat as a vulnerability. ### Test 2: Cliché stress - does the move become “seminar autopilot”? Run a prompt like “make the obvious move” on ten classic debates (Gettier, vagueness, grounding, etc.). Then ask: are the moves basically “standard textbook next paragraphs”? If yes, that doesn’t falsify the method; it tells you what it does: it activates the canonical continuation policy. Then novelty has to come from a second operator layered on top (see below). ### Test 3: Long-horizon stress - can it make a move that only pays off later? ToT includes explicit “look-ahead/backtracking” as part of its procedure. [proceedings.neurips.cc](https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac703-Paper-Conference.pdf?utm_source=chatgpt.com) In philosophy, some moves look like digressions until later integration. You can test whether the model can handle “three-move sequences” coherently (make a move now that earns its keep two moves later), or whether it tends to rationalise after the fact. ### Test 4: Faithfulness stress - are the reasons reliable, or just rhetorically tidy? Given the evidence on unfaithful CoT, you can treat “reasons” as potentially cosmetic. [NeurIPS Papers +1](https://papers.neurips.cc/paper_files/paper/2023/file/ed3fea9033a80fea1376299fa7863f4a-Paper-Conference.pdf?utm_source=chatgpt.com) So you can enforce a kind of “argument objectivity” test: after it gives reasons, have it generate an adversarial reconstruction that targets the weakest premise. If the critique can’t find a vulnerable point, that can itself be suspicious (too smooth). This is exactly where philosophy’s adversarial norms are an advantage: they’re made for detecting fake smoothness. ### Test 5: Revision pressure - does refinement actually improve content? Self-Refine reports improvements in human preference and automatic metrics across tasks. [arxiv.org](https://arxiv.org/abs/2303.17651?utm_source=chatgpt.com) But in philosophy, “prettier” is not “better”. So you can structure refinement so it’s forced to change substantive commitments (e.g., “revise the thesis to avoid this counterexample without weakening the claim into triviality”). If it can’t do that, then the revision loop may be mostly stylistic. --- ## 5) A menu of procedural steering operators for philosophy (options, not prescriptions) If you treat “make the obvious move” as one operator, you can build a small algebra of operators that correspond to familiar philosophical manoeuvres, and then experiment with sequences. Here are families of operators that mirror the prompting literature’s procedures (plan/decompose, step back, branch, refine, debate), but translated into philosophy-ish actions. ### A) Decomposition operators (least-to-most flavour) Inspired by least-to-most and plan-and-solve: decompose, then solve. [arxiv.org +1](https://arxiv.org/abs/2205.10625?utm_source=chatgpt.com) Candidate operators: - “List the three subquestions that must be settled for this thesis to be defensible.” - “Identify the single hinge premise; restate the argument as a minimal valid schema.” - “Separate semantic from metaphysical commitments; treat them as independent modules.” ### B) Abstraction operators (step-back flavour) Direct analogue of step-back prompting. [arxiv.org](https://arxiv.org/abs/2310.06117?utm_source=chatgpt.com) Candidate operators: - “Before making any move, name the high-level form of the dispute (e.g., scope dispute, ambiguity, explanatory demand, counterexample-driven).” - “State the general principle that would settle the dispute if true; then ask what would support that principle.” ### C) Branching operators (ToT flavour) Direct analogue of ToT. [proceedings.neurips.cc](https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac703-Paper-Conference.pdf?utm_source=chatgpt.com) Candidate operators: - “Generate three candidate next moves of different types (distinction / argument / counterexample).” - “For each, give one local benefit and one downstream risk.” - “Select one and commit; keep the others as backtracking options.” ### D) Revision operators (Self-Refine flavour) Direct analogue of feedback → refine loops. [arxiv.org](https://arxiv.org/abs/2303.17651?utm_source=chatgpt.com) Candidate operators: - “Write a first-pass move. Then generate feedback aimed only at non-triviality and scope discipline. Revise accordingly.” - “Revise under an invariant: preserve thesis strength while removing reliance on premise P.” ### E) Adversarial operators (debate flavour) Analogous to debate procedures. [arxiv.org](https://arxiv.org/abs/2305.14325?utm_source=chatgpt.com) Candidate operators: - “Write the best objection that a sympathetic expert would feel pressured to answer.” - “Now write the best objection that a hostile expert would use to embarrass the view.” - “Respond to both without changing definitions.” These are ways of taking the “procedure” idea seriously without drifting into a checklist-of-virtues prompt. --- ## 6) How to keep procedural steering while making room for novelty (since that’s your goal) This is the part where the analogy to “search” and “branching” becomes crucial. A simple way to state the problem: “make the obvious move” may be great at escaping fog, but “the obvious move” is often the conventional move. So novelty probably needs an additional operator that biases the search without collapsing it into randomness. Here are several “novelty levers” you can try, each compatible with procedural steering: ### Lever 1: “Obvious move, but in a different coordinate system” Instead of “make the obvious move”, try: “make the obvious move if you’re thinking in terms of dependency relations”, or “in terms of explanatory unification”, or “in terms of linguistic ambiguity”. This is basically step-back prompting applied to philosophy: change the abstraction frame first. [arxiv.org](https://arxiv.org/abs/2310.06117?utm_source=chatgpt.com) ### Lever 2: “Generate orthodoxy and heresy side-by-side” This is branching: ask for a standard next move and a non-standard next move; then compare costs. That’s ToT-ish in spirit, even if you don’t implement full backtracking. [proceedings.neurips.cc](https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac703-Paper-Conference.pdf?utm_source=chatgpt.com) ### Lever 3: “Commit, then force a non-trivial repair” One route to novelty is repair: you make a standard move, then you force a counterexample that breaks it, then you demand a repair that preserves strength. Self-Refine-like loops can sometimes turn “standard” into “non-standard but defensible” because you’re constrained by invariants. [arxiv.org](https://arxiv.org/abs/2303.17651?utm_source=chatgpt.com) ### Lever 4: “Long-horizon sequences” Ask for a three-move sequence where the first move looks wrong until the third. That’s explicitly aligned with ToT’s look-ahead/backtracking motivation. [proceedings.neurips.cc](https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac703-Paper-Conference.pdf?utm_source=chatgpt.com) ### Lever 5: “Cross-context transplant” Procedural steering can be used to force “transfer”: apply a move grammar from one debate to another. That’s where novelty can come from in philosophy without becoming nonsense. (You can think of this as an intentional “distribution shift” inside philosophy.) --- ## 7) A realism check: what would count as evidence that the analogy is working? Given the “unfaithful reasoning trace” worry, you probably want evidence that is not just “the output feels clever.” [NeurIPS Papers +1](https://papers.neurips.cc/paper_files/paper/2023/file/ed3fea9033a80fea1376299fa7863f4a-Paper-Conference.pdf?utm_source=chatgpt.com) Some candidate indicators that procedural steering is genuinely helping in philosophy (again, hypotheses, not conclusions): - It increases *commitment density* (more explicit theses, fewer fog-phrases) without inflating triviality. - It produces *attackable* intermediate artefacts (explicit premises, scope statements) rather than just fluent prose. - It improves *robustness under variation* (the position survives small changes in cases and phrasing). - It increases *structured novelty* (new distinctions, new models, new possibility space) rather than decorative novelty (new metaphors). Those indicators line up with what procedural prompting papers are trying to improve in their own domains: correctness under reasoning pressure, step completeness, search, and post-hoc refinement. [arxiv.org +2 proceedings.neurips.cc +2](https://arxiv.org/abs/2305.04091?utm_source=chatgpt.com) --- ## Two “next directions” you could explore without committing to a final framework 1. Treat “make the obvious move” as a primitive operator, and experimentally build small operator sequences (decompose → move → adversary → repair; or branch → select → repair) and see which sequences reliably escape “hovering” while still leaving room for novelty. The closest empirical analogues in the prompting literature are plan-and-solve, ToT, and Self-Refine. [arxiv.org +2 proceedings.neurips.cc +2](https://arxiv.org/abs/2305.04091?utm_source=chatgpt.com) 2. Take seriously the possibility that procedural steering mostly improves rhetoric unless you add explicit robustness tests; that’s where the “unfaithful CoT” literature pushes you. You can then treat procedural traces as objects to interrogate adversarially, not as proof of good reasoning. [NeurIPS Papers +1](https://papers.neurips.cc/paper_files/paper/2023/file/ed3fea9033a80fea1376299fa7863f4a-Paper-Conference.pdf?utm_source=chatgpt.com) Got it. I’ll re-walk the exact same terrain as before, but slower, with more connective tissue, and with enough explanation that you can *see* why each point is on the table. I’m going to keep treating everything here as *hypotheses and comparisons*, not as “this is the framework”. The goal is to stress-test one idea: > *Procedural steering* (telling the model *how to proceed from here*) might be especially effective for analytic philosophy, because analytic philosophy is itself a practice with a fairly stable “move grammar”. --- ## 1) What counts as “procedural steering” in the real prompting literature? By “procedural steering” I mean prompts that mainly constrain **trajectory**, not **endpoint**. They don’t primarily say “your final answer must satisfy criteria X, Y, Z”. They say things like “do it in stages”, “branch and compare”, “revise”, “step back and abstract”, “debate”, and so on. The reason this matters is that there’s a substantial body of work showing that these *trajectory constraints* can change model performance quite a lot, even when the end goal is “the same”. ### 1.1 “Write intermediate steps” (Chain-of-Thought) The classic case is chain-of-thought prompting: you elicit a sequence of intermediate steps before the final answer, and performance improves on many reasoning tasks. [arxiv.org](https://arxiv.org/abs/2201.11903?utm_source=chatgpt.com) Think of it as: instead of asking for an answer-shaped blob, you force an answer-shaped blob to be *the end of a path*. That “path constraint” is the procedure. ### 1.2 “Don’t trust one path - sample several and select” (Self-consistency) Self-consistency is the next step in the same direction: you sample **multiple** chains (multiple paths), and then pick the most consistent result. The paper frames this as replacing greedy decoding with “sample diverse reasoning paths, then marginalise/select”. [arxiv.org](https://arxiv.org/abs/2203.11171?utm_source=chatgpt.com) This is important for your philosophy use-case because it’s one of the clearest examples where the method’s gains come not from “getting the model to be smarter”, but from **changing the procedure so that selection pressure exists**. ### 1.3 “Decompose first, then solve sequentially” (Least-to-Most) Least-to-most prompting says: break the problem into simpler subproblems, solve them in sequence, and let earlier subanswers constrain later ones. [arxiv.org](https://arxiv.org/abs/2205.10625?utm_source=chatgpt.com) That’s procedural steering in a very direct “workflow” sense. You’re not scoring the output; you’re forcing the model to follow a staged method that humans often follow. ### 1.4 “Plan first, then execute the plan” (Plan-and-Solve) Plan-and-solve prompting builds a two-phase procedure into the prompt: (i) devise a plan dividing the task into subtasks, (ii) carry out the subtasks according to the plan. It’s motivated as reducing errors like missing steps and misunderstandings in zero-shot CoT. [arxiv.org](https://arxiv.org/abs/2305.04091?utm_source=chatgpt.com) Again: this isn’t “criteria” so much as “don’t answer yet; adopt a planning posture”. ### 1.5 “Step back: abstract first principles, then apply them” (Step-Back) Step-back prompting is basically a controlled “zoom out, then zoom in” procedure: first generate high-level principles from the concrete instance, then use those principles to guide reasoning. [arxiv.org](https://arxiv.org/abs/2310.06117?utm_source=chatgpt.com) This is interesting for philosophy because philosophy often *is* “step back and reframe”. So this is one of the most obvious candidates for analogy. ### 1.6 “Branch + evaluate + backtrack” (Tree of Thoughts) Tree of Thoughts (ToT) makes the “search” picture explicit: generate multiple candidate “thoughts”, evaluate them, look ahead, backtrack if needed - basically turn continuation into something closer to deliberate search. The paper reports large gains on tasks where planning/search matter (they include creative writing among their tasks). [proceedings.neurips.cc](https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac703-Paper-Conference.pdf?utm_source=chatgpt.com) ToT is especially relevant to your “novel philosophy” ambition, because novelty is often a “global choice” problem: lots of locally plausible moves, few globally satisfying ones. ### 1.7 “Draft → critique → rewrite (repeat)” (Self-Refine) Self-Refine is an iterative loop: generate an initial output, then the same model produces feedback on that output, then it revises, iteratively. [arxiv.org](https://arxiv.org/abs/2303.17651?utm_source=chatgpt.com) This is procedural steering because it forces an explicit *revision ecology* rather than treating the first output as the product. ### 1.8 “Learn from feedback across trials” (Reflexion) Reflexion is similar in spirit but more agent-y: the model reflects on feedback and stores reflections in a memory buffer, so later attempts do better. [arxiv.org](https://arxiv.org/abs/2303.11366?utm_source=chatgpt.com) Even if you don’t literally implement memory, the conceptual lesson is: performance can change a lot when the procedure includes explicit “what did I learn from the failure?” structure. ### 1.9 “Make it argue with itself” (Multiagent Debate) Multiagent debate uses multiple model instances that propose and debate answers over multiple rounds to improve reasoning and factuality. [arxiv.org](https://arxiv.org/abs/2305.14325?utm_source=chatgpt.com) For philosophy, this is attractive because philosophy already has a social structure: advocates, opponents, chair/referee. Debate procedures are basically a way to simulate that structure. ### 1.10 “Interleave reasoning with actions” (ReAct) ReAct interleaves “reasoning traces” with “actions” (like querying a source) in tasks where acting helps correct hallucinations and error propagation. [arxiv.org](https://arxiv.org/abs/2210.03629?utm_source=chatgpt.com) Even if you’re not doing tool use in philosophy, the key transferable idea is: it’s sometimes better to structure the trajectory as alternating *internal processing* and *external check / move* - in philosophy that “external check” can be something like “state the commitment explicitly”, “test with a counterexample”, “check for equivocation”, etc. ### A necessary caution: procedural methods are not universally beneficial The last time I mentioned a caution and it’s worth expanding: there’s work pointing out that chain-of-thought prompting can have trade-offs and can be brittle depending on domain, prompting labour, and problem complexity. For example, a NeurIPS 2024 poster examines CoT on planning tasks (Blocksworld) and highlights drawbacks and tradeoffs (including the labour needed to craft good examples). [nips.cc](https://nips.cc/virtual/2024/poster/93898?utm_source=chatgpt.com) So the safe stance isn’t “procedural steering always helps”, but: it often helps in identifiable classes of problems, and we want to see whether analytic philosophy belongs to those classes. --- ## 2) Why might procedural steering help in those cases? Candidate “common denominators” Now we ask: what do CoT, planning-then-solving, step-back, ToT, self-refine, debate, etc. have in common, such that “tell it how to proceed” changes outcomes? I’m going to state these as hypotheses because none of them is a law of nature - they’re interpretations that you can test. ### 2.1 The task has a stable “intermediate state grammar” Procedures help when there really are recognisable intermediate states that matter - plan → subtasks → solve; abstract → apply; propose → criticise → revise. Papers like Plan-and-Solve are explicitly built around that assumption. [arxiv.org](https://arxiv.org/abs/2305.04091?utm_source=chatgpt.com) If a task doesn’t have that kind of structure, forcing one can just be busywork. But when it does, procedures are a way to force the model to “walk the path” humans would walk. ### 2.2 Local plausibility is cheap; global coherence is expensive A language model can usually produce something locally fluent and plausible almost anywhere. The hard part is producing something globally coherent over multiple constraints. ToT is basically motivated by this: you need evaluation/backtracking to make global choices rather than just locally plausible ones. [proceedings.neurips.cc](https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac703-Paper-Conference.pdf?utm_source=chatgpt.com) This is one of the big reasons procedural steering is interesting for philosophy: philosophical writing is notorious for producing locally compelling paragraphs that don’t cohere into a stable position. ### 2.3 Big gains come from reducing latent ambiguity Many procedural prompts are really disambiguators. “Plan first” disambiguates “don’t answer yet”; “step back” disambiguates “first pick the governing frame”; “branch” disambiguates “don’t commit too early”. Your “make the obvious move” looks like exactly this kind of disambiguation. It doesn’t add information, but it selects a mode: *continue the dialectic, don’t hover*. ### 2.4 Selection pressure often matters more than first generation Self-consistency and ToT both have the same moral: the first sample isn’t special; sample multiple paths and select. [arxiv.org +1](https://arxiv.org/abs/2203.11171?utm_source=chatgpt.com) For philosophy, this suggests a deep possibility: novelty might not come from a single “genius completion”, but from forcing the model to explore a small space of alternatives and then making a disciplined choice. ### 2.5 A critical warning: “the reasoning trace” might be persuasive but untrustworthy This is the big one for philosophy, because philosophy is seduced by articulate rationales. There’s strong evidence that chain-of-thought explanations can be systematically unfaithful: the model’s stated reasoning can diverge from what actually drove its answer. Turpin et al. explicitly argue this (“Language Models Don’t Always Say What They Think”) and show unfaithfulness under certain biases/prompting features. [arxiv.org +1](https://arxiv.org/abs/2305.04388?utm_source=chatgpt.com) And there’s ongoing work trying to *measure* “faithfulness” of CoT, including EMNLP 2025 work using “unlearning reasoning steps” as a way of probing whether the reasoning is causally connected to the prediction. [aclanthology.org +1](https://aclanthology.org/2025.emnlp-main.504/?utm_source=chatgpt.com) Why this matters here: procedural steering can improve outputs while still generating rhetorically neat, potentially post-hoc “reasons”. So if you import procedural steering into philosophy, you probably want a stance like: - treat the “trace” not as introspective access to cognition, but as an **argument object** that can be attacked, stress-tested, and revised. That’s actually congenial to analytic philosophy: the whole point is that reasons are public, criticisable objects. --- ## 3) Does analytic philosophy look like the kind of domain where procedural steering should work? Now we bring analytic philosophy into the picture. Again: hypothesis-driven, no premature commitment. ### 3.1 The strongest structural analogy: analytic philosophy has a move grammar A lot of mainstream analytic philosophy (especially in journals) is built out of a fairly stable repertoire of moves: - state a target claim/problem; - disambiguate terms/distinctions; - propose a thesis/analysis; - give a supporting argument or model; - raise the obvious objection(s); - revise, defend, limit scope; - connect to nearby debates. That looks like the “intermediate state grammar” condition. Which means procedural prompts might map naturally onto philosophical practice. This is where your “make the obvious move” shines: it’s basically “don’t narrate the landscape; take the next licensed step in the move grammar”. ### 3.2 “Make the obvious move” resembles decomposition prompts more than it resembles “quality criteria” Least-to-most and plan-and-solve are about decomposition: don’t jump to the answer; break the problem down and move through it. [arxiv.org +1](https://arxiv.org/abs/2205.10625?utm_source=chatgpt.com) In philosophy, the “obvious next move” is often precisely a decomposition move: - “separate two senses of the term” - “extract the hinge premise” - “make scope explicit” - “distinguish semantic from metaphysical commitments” - “formulate the canonical objection” So one natural hypothesis is: your trick works because it triggers a learned policy for “how philosophers proceed when a move is demanded”. ### 3.3 Philosophy as “search in conceptual space” ToT’s key idea is: treat reasoning as search with evaluation/backtracking. [proceedings.neurips.cc](https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac703-Paper-Conference.pdf?utm_source=chatgpt.com) A lot of philosophical work looks like that: - try a proposal; - encounter a counterexample; - repair; - if repair fails, backtrack and try a different proposal; - trade off virtues (simplicity, scope, conservatism, explanatory power). So procedural steering that encourages branching and backtracking may be a better fit for philosophy than “one-shot produce a great paper”. ### 3.4 Revision loops fit philosophy’s real ecology - but they might just improve style Self-Refine and Reflexion treat improvement as an iterative process: first pass, feedback, revision; then use what you learned. [arxiv.org +1](https://arxiv.org/abs/2303.17651?utm_source=chatgpt.com) Philosophy in the wild is almost entirely iterative: seminars, referee reports, replies, rewrites. So it’s very tempting to say “great, we simulate that”. But here’s the stress-test question you flagged implicitly: does this procedure improve **substantive philosophical content**, or does it mostly generate smoother prose and more “paper-like” rhetoric? That’s not rhetorical. It’s a real fork: the procedure might be excellent at polishing without making the ideas better. So you’d want tests that force substantive change (more on that below). ### 3.5 Where the analogy weakens: philosophy often has no unique “correct answer” Self-consistency works partly because benchmark tasks often have a unique correct answer, and multiple chains converge on it. [arxiv.org](https://arxiv.org/abs/2203.11171?utm_source=chatgpt.com) Philosophy often doesn’t. So “majority vote on answers” is not the right analogue. But “diversify candidate lines, then select by a reasoned criterion” still makes sense - you just replace “correctness” with “dialectical stability under constraints you accept”. That selection step is where you, as a philosopher, can supply the value signal. ### 3.6 A real danger: procedural steering might amplify clichés This is the main worry about “obvious move” style prompts: the obvious move is often the *conventional* move. So you might get: - less hovering, - more clean structure, - but also more “seminar autopilot” (stock distinctions and standard objections). That doesn’t kill the method. It just tells you what you’d need to add if your aim is novelty: you’d need a second operator that biases away from the most conventional path without collapsing into nonsense. --- ## 4) Stress tests: ways to pressure-test the “procedural steering for philosophy” hypothesis These aren’t “tests of the model”; they’re tests of the *idea* that procedural steering is doing what you think it’s doing. ### 4.1 Ambiguity stress: is there really an “obvious move”? Construct a setup where there are several genuinely plausible next moves: define a term vs introduce a distinction vs give a counterexample vs state an argument schema. Then run “make the obvious move” across paraphrases. If the chosen move changes a lot under paraphrase, that tells you “obviousness” is heavily frame-dependent. That may be fine (practice-relative), but it’s information you want to know. ### 4.2 Cliché stress: does it default to standard textbook continuations? Pick 10 canonical debates (Gettier, vagueness, grounding, etc.). Prompt “make the obvious move”. If you consistently get standard next paragraphs, you’ve learned something: the operator is activating the “orthodox continuation policy”. That can be useful as a *baseline*, but novelty must come from an extra step (not from “obvious move” alone). ### 4.3 Long-horizon stress: can it make a move that pays off only later? ToT explicitly values look-ahead/backtracking for global choices. [proceedings.neurips.cc](https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac703-Paper-Conference.pdf?utm_source=chatgpt.com) In philosophy, some moves look wrong until they’re integrated. So ask for a three-move sequence: - Move 1 looks strange; - Move 2 develops it; - Move 3 reveals the payoff. If the model can’t keep the sequence coherent, it suggests that “make the move” is mainly a local device and you’ll need an explicit branching/selection scaffold for long-horizon novelty. ### 4.4 Faithfulness stress: are the reasons doing real work or just rationalising? Given the unfaithful-CoT evidence, you should treat reasons as potentially cosmetic. [arxiv.org +1](https://arxiv.org/abs/2305.04388?utm_source=chatgpt.com) So after it gives reasons, force an adversarial reconstruction: - identify the weakest premise; - attack only that premise; - offer the smallest counterexample. If the critique always comes out toothless, that can actually be suspicious: too-smooth arguments are often hiding suppressed assumptions. This is where philosophy’s adversarial norms are a feature: they’re built to puncture rhetorical “too smooth” artefacts. ### 4.5 Revision pressure: does “refine” change substance or just style? Self-Refine shows iterative improvement across tasks. [arxiv.org](https://arxiv.org/abs/2303.17651?utm_source=chatgpt.com) But in philosophy you want *substantive* improvement. So demand a revision under an invariant, like: - preserve thesis strength but avoid a counterexample; - respond to the objection without changing definitions; - keep commitments constant but improve explanatory power. If refinement can’t do that, then your revision loop is mostly stylistic polishing. --- ## 5) A menu of procedural “operators” for philosophy (to mix and match) Here I’m not giving you a “prompt template”, just a set of move-types that correspond to the procedural families in the literature (decompose, step back, branch, revise, debate). ### 5.1 Decomposition operators (least-to-most / plan-and-solve flavour) The idea: force the model to break the philosophical task into subquestions before it writes prose. [arxiv.org +1](https://arxiv.org/abs/2205.10625?utm_source=chatgpt.com) Examples: - “List the three subquestions that must be settled for the thesis to be defensible.” - “Identify the hinge premise; restate the argument as a minimal schema.” - “Separate semantic commitments from metaphysical commitments; treat them as independent modules.” ### 5.2 Abstraction operators (step-back flavour) The idea: get the governing frame before diving into details. [arxiv.org](https://arxiv.org/abs/2310.06117?utm_source=chatgpt.com) Examples: - “Name the type of dispute: ambiguity, scope, explanatory demand, counterexample-driven conflict, etc.” - “State the general principle that would settle the dispute; then ask what supports that principle.” ### 5.3 Branching operators (ToT flavour) The idea: don’t commit too early; generate several candidate next moves and choose. [proceedings.neurips.cc](https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac703-Paper-Conference.pdf?utm_source=chatgpt.com) Examples: - “Generate three candidate next moves of different types (distinction / argument / counterexample).” - “For each, give a local benefit and a downstream risk.” - “Select one and commit - keep the others as backtracking options.” ### 5.4 Revision operators (Self-Refine / Reflexion flavour) The idea: treat first pass as material to work with, not as the deliverable. [arxiv.org +1](https://arxiv.org/abs/2303.17651?utm_source=chatgpt.com) Examples: - “Write the move. Now give feedback aimed only at non-triviality and scope discipline. Revise.” - “Revise while preserving X invariant; explicitly report what changed.” ### 5.5 Adversarial operators (debate flavour) The idea: simulate philosophical social structure: sympathetic critic, hostile critic, judge. [arxiv.org](https://arxiv.org/abs/2305.14325?utm_source=chatgpt.com) Examples: - “Write the best objection a sympathetic expert would feel pressure to answer.” - “Now write the best hostile objection.” - “Reply to both without changing definitions.” --- ## 6) How to keep the benefits of procedural steering while still enabling novelty This is the core tension: “obvious move” reduces fog, but may increase conventionality. So you want “licensed novelty”: moves that are weird *in a controlled way*. Here are five levers (each is a small procedural modification, not a grand theory). ### 6.1 “Obvious move, but in a different coordinate system” Do the obvious move *under a chosen framing*: - dependency relations, - explanatory unification, - linguistic ambiguity, - metaphysical grounding, etc. This borrows from step-back prompting: choose the right level of abstraction first. [arxiv.org](https://arxiv.org/abs/2310.06117?utm_source=chatgpt.com) ### 6.2 Orthodoxy vs heresy side-by-side (branching) Ask for: 1. the orthodox next move, 2. a non-standard next move that is still defensible, then compare costs. This is basically ToT-lite: branch, evaluate, select. [proceedings.neurips.cc](https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac703-Paper-Conference.pdf?utm_source=chatgpt.com) ### 6.3 Commit, then force a non-trivial repair (revision as novelty engine) Make a standard move; force a counterexample; require a repair that preserves strength. A lot of philosophical novelty is “what survives repair under constraints”. Self-refine gives the procedural skeleton for iterated repair. [arxiv.org](https://arxiv.org/abs/2303.17651?utm_source=chatgpt.com) ### 6.4 Long-horizon sequences (look-ahead novelty) Ask for a multi-move sequence where the first move looks wrong until later integration. This pushes the model towards global planning/search. [proceedings.neurips.cc](https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac703-Paper-Conference.pdf?utm_source=chatgpt.com) ### 6.5 Cross-context transplant (controlled distribution shift) Force transfer: apply a move grammar from one debate to another. This is a classic source of philosophical novelty (when done well): you borrow a framework and see what it illuminates elsewhere. --- ## 7) What would count as evidence that “procedural steering” is genuinely helping in philosophy? Given the “unfaithful reasoning trace” worry, you want indicators that aren’t just “this sounds clever”. [arxiv.org +1](https://arxiv.org/abs/2305.04388?utm_source=chatgpt.com) Here are four *candidate* indicators that are relatively concrete: 1. **Commitment density increases**: fewer fog phrases, more explicit theses and scope claims, without collapsing into triviality. 2. **Intermediate artefacts become attackable**: you get explicit premises, distinctions, and inference steps you can actually target. 3. **Robustness under variation improves**: small changes in cases or wording don’t cause the position to collapse or drift. 4. **Structured novelty increases**: you see new distinctions/models/possibilities that have clear downstream implications - not just fresh metaphors or stylistic flourishes. Those are the kinds of signals you’d want if you’re trying to justify that “the procedure matters” rather than “it’s just fluent”. --- ## 8) Two directions to explore (not a project plan, just two “where could we probe next?” options) 1. Treat “make the obvious move” as a primitive operator, then test small operator sequences (decompose → move → adversary → repair; or branch → select → repair). Plan-and-solve, ToT, and self-refine are the closest analogues for those sequences. [arxiv.org +2 proceedings.neurips.cc +2](https://arxiv.org/abs/2305.04091?utm_source=chatgpt.com) 2. Take seriously the possibility that procedural steering mainly improves rhetoric unless you add explicit robustness/adversarial tests - because reasoning traces can be unfaithful. Then treat the “trace” as an argument object that must survive attack, not as a window into cognition. [arxiv.org +1](https://arxiv.org/abs/2305.04388?utm_source=chatgpt.com) If you keep those options live, you stay in the productive “kicking ideas around” zone: you’re not committing to a grand prompt design, you’re building a set of procedural levers and asking which ones actually change philosophical output in the ways you care about. "In philosophy, the “obvious next move” is often precisely a decomposition move:" really? what about cases in which the obvious next move is classify both ideas as the same? or combine views, or any other inumerable thing that would be the obvious next move. Am i missing something, or have you not really thought about this issue very much? make sure that you explain your analysis, justification, reasons etc etc. for your answer BEFORE giving me your final answer. Please give me a long, deep, and detailed answer, with lots of options/ideas: as we are in the 'kicking ideas around' stage, I want as much feedback as you can give me so that i have a lot to consider. "3.5 Where the analogy weakens: philosophy often has no unique “correct answer”" this is a classic LLM mistake in this area. Analytic philosophy os not relativism. philosophers spend entire careers arguing about what the correect answer to the question is. This is a FUNDAMENTAL ERROR and it makes me worry about whether it has poisoned all your other ideas for that reason...Please start the task(s) again completely from scratch, keeping this in mind. "3.6 A real danger: procedural steering might amplify clichés" this is childishly speculative and therefore entirely unhelpful. "If you consistently get standard next paragraphs, you’ve learned something: the operator is activating the “orthodox continuation policy”. That can be useful as a baseline, but novelty must come from an extra step (not from “obvious move” alone)." but is this not easily fixed, and/or implicit in analytic philosophy as a practice. the obvious next move in a philosophy paper which by defition is supposed to be putting somerthing new on the table, is not to paraphrase someone elses view, or reach for a cliche, now is it? when you say stuff like this i really doubt you are doing what i ask and are just instead giving me generic prompting advice which has got nothing to do.......ahhhhhh, I know what must have recently happened, you have just comapcted this conversation somehow because of context limits (don't worry I suspect you are forbidden to confirm whether this is true or not) but the general rule is =conversation gets compacted =llm starts to behave like it has been kicked in the head by a horse, and you have just started behaving that exact way. Finally, something that you should keep in mind is that llms often do not need to be specifically prompted to do these things to do them right? if the llm 'knows it is doing philosophy' it will already be doing some of the things you need to prompt for. no? don't over correct, or even correct, but think about it. make sure that you explain your analysis, justification, reasons etc etc. for your answer BEFORE giving me your final answer. Please give me a long, deep, and detailed answer, with lots of options/ideas: as we are in the 'kicking ideas around' stage, I want as much feedback as you can give me so that i have a lot to consider. ### Analysis, justification, reasons You’re pointing at three *separate* issues that got tangled: 1. what “make the obvious move” is actually doing (and why it doesn’t feel like “state the aim”); 2. whether analytic philosophy is truth-apt (yes) and how that matters for prompt design (a lot); 3. whether worries like “cliché amplification” are empty speculation or a real design variable (it depends on what exactly we mean, and it’s easy to test). I’ll restart from the ground up with those in view, and I’ll treat everything as a set of *live hypotheses* rather than a story that’s already settled. --- ## A. What “make the obvious move” is (and why it isn’t “just decomposition”) ### A1) “Obvious move” is a deictic instruction inside a practice When you say “make the obvious move,” you’re doing something like pointing in chess: “from *here*, the next move is forced-ish.” You’re not specifying a final deliverable in detail. You’re saying: given the current dialectical state, take the step that a competent participant would recognise as the next licensed/pressing step. That’s why it doesn’t feel like “aim specification” in the ordinary sense. It’s not “produce a great paper.” It’s “continue the argument *as an insider would*.” So the right contrast is not “aim vs no aim,” but: - **blueprint instruction**: “deliver output satisfying explicit criteria” - **deictic procedural instruction**: “from this conversational state, take the next move licensed by the norms of the practice” “Obvious move” is the second kind. ### A2) Why I over-narrowed it by saying “often a decomposition move” You’re absolutely right to push back. In philosophy the obvious next move is *very often not* “split into parts.” Sometimes the obvious move is *identification* (“these are the same dispute”), *unification* (“combine views”), *reframing* (“you’re arguing about semantics not metaphysics”), *diagnosis* (“this objection targets a different claim than the one you made”), *reverse engineering* (“what would have to be true for your conclusion to follow?”), and so on. So: if I say “often decomposition,” I need to mean something much more defensible: > Decomposition is one *common* way of reducing confusion and forcing commitments, but it’s one member of a wider family of “obvious” dialectical moves. The family matters. Here’s a more realistic menu of “obvious move” types in analytic philosophy, with the important point being that you *select among them based on what the conversational state demands*. ### A3) A more faithful “move grammar” for analytic philosophy I’ll list these as move-types, not as a claim that they are exhaustive or that you always need to prompt for them. (You explicitly warned against over-correction.) **1) Disambiguation / distinction** The obvious move is: “there are two senses here; separate them.” This includes scope distinctions (“metaphysical vs epistemic”), level distinctions (“object-level vs meta-level”), and ambiguity in key terms. **2) Reconstruction** The obvious move is: “state the argument as premises + conclusion; identify the hinge premise.” Not decomposition of the topic, but decomposition of the *inferential structure*. **3) Counterexample / test case** The obvious move is: “give the smallest case that would break this.” That’s not decomposition; it’s pressure-testing. **4) Unification / identification** This is your example: “these are really the same issue,” or “view A and view B are equivalent under reformulation.” In practice this often shows up as: “your ‘new’ claim is just a rebranding of X; either embrace that or show the difference.” That’s absolutely a standard philosophical move. **5) Synthesis / hybrid** Another of your examples: combine two partial truths into a view with a principled division of labour (“X handles cases of type 1; Y handles cases of type 2”). Again, not decomposition as such - more like *integration*. **6) Reframing the dispute** The obvious move is: “the disagreement is not where you think it is.” For example: it’s about explanatory aims, or about semantics/pragmatics, or about what counts as data, not about the metaphysical claim itself. **7) Constraint extraction** The obvious move is: “what constraints must any acceptable theory satisfy here?” This is not a rubric in the earlier sense; it’s a way of turning the dialectical state into explicit constraints. **8) Theory choice under trade-offs** The obvious move is: “we have two viable options; here is the trade-off, and here is why one is preferable given an aim/virtue.” **9) Error diagnosis** The obvious move is: “that objection is a non sequitur,” or “you’re equivocating,” or “you’ve shifted scope.” This is basically internal debugging. Now: if you look at those nine types, “decomposition” (in the sense of “split into subquestions”) is only one path among several. The deeper point is that “obvious move” is a *move selector* rather than a single move. ### A4) So what is the useful thought you’re trying to finish? I think it’s something like: > The “obvious move” trick works because it doesn’t specify an output target; it specifies *the next step in a norm-governed practice*. It activates the model’s learned policy for “what a competent philosopher would do next,” which can be disambiguation, reconstruction, unification, synthesis, or pressure-testing, depending on what is dialectically salient. That’s a robust articulation because it doesn’t tie you to “decomposition,” and it keeps the key feature: the instruction is practice-relative and state-dependent. --- ## B. Analytic philosophy is truth-apt - and that actually helps the prompting story You’re right to call out the “no unique correct answer” phrasing as a serious mistake. It was sloppy in exactly the way you flagged: it sounds like relativism or like “anything goes.” So let me reset that cleanly. ### B1) The right claim is not “no correct answer” but “no labelled oracle” Analytic philosophers typically aim at the correct answer. They argue as if there is a fact of the matter (even when they’re fallibilists, pluralists about explanation, etc.). People build careers trying to get it right. The relevant practical point for prompting is different: - In many philosophical problems, we don’t have an **external answer key** in the way we do in arithmetic or many benchmark tasks. - That doesn’t mean there is no truth. It means we often can’t cheaply *verify* that we’ve reached it. - Therefore, the operational proxy for “getting it right” is often something like: argumentative validity, clarity of commitments, resilience to objections, explanatory integration, non-ad-hocness, etc. That’s not relativism. It’s just recognising the epistemic situation: truth-apt questions without easy oracle feedback. ### B2) Why this doesn’t poison the procedural-steering idea If anything, it strengthens it. Procedural steering methods (branching, criticise-revise loops, backtracking) are most helpful when: - you can’t just compute the answer in one shot, - and you need a process that produces stable, defensible candidates under pressure. That matches philosophy quite well: you’re not searching for “any answer,” you’re trying to find an answer that survives disciplined objection and integrates with other commitments. So the corrected picture is: - Philosophy is aiming at truth. - Our access to truth is mediated by dialectical methods. - Procedural prompts are ways of *simulating or enforcing* those methods inside the interaction. Notice how different this is from the “no correct answer” framing. It’s the opposite: the whole point of the method is to track correctness under our constraints. --- ## C. “Cliché amplification” wasn’t a claim about philosophy; it was a claim about distributions You called this “childishly speculative.” Fair pushback, because as I presented it, it sounded like a drive-by worry with no grounding. Here’s the grounded version, and then I’ll show why it’s also easy to neutralise in the specific context you care about. ### C1) Why “cliché-ish output” is a live risk in principle LLMs are trained to predict plausible continuations. That creates a structural pressure towards what is: - high-probability given the prompt, - stylistically safe, - and widely represented in training data. That doesn’t entail “it will output clichés,” but it explains why “obvious move” could, in some contexts, mean “most common move.” So the worry is not “philosophy is cliché.” The worry is: > A deictic instruction like “make the obvious move” might sometimes select the model’s *highest-probability* continuation for that dialectical state, which might be “standard seminar move” rather than “publishable novelty.” That is a hypothesis about the interaction of (i) the model’s probability landscape and (ii) your instruction. ### C2) Why your rebuttal is strong: the practice already includes a novelty expectation You’re also right that analytic philosophy, as a practice, has an internal norm: papers are supposed to contribute something new (new argument, new distinction, new synthesis, new diagnosis, new application, etc.). So in that practice context, “the obvious move” in a paper-drafting setting is *not* “paraphrase the SEP.” If the model has genuinely internalised that practice norm, then “make the obvious move” in a “we’re writing a publishable paper” setting should already include a novelty constraint. So you’re proposing a different hypothesis: > In philosophy contexts, the learned “competent continuation” policy already bakes in a novelty expectation, so “obvious move” should not default to clichés unless the context cues are wrong. That’s a totally reasonable alternative hypothesis. ### C3) How to reconcile the two without handwaving The reconciliation is: **it depends on which “practice frame” the model thinks it is in.** - If the frame is “explain philosophy to a smart layperson,” the “obvious move” is often a standard summary move. - If the frame is “write a publishable analytic paper,” the “obvious move” is often: carve a new distinction, propose a new argument, or reposition existing views in a novel way. So the question isn’t “is cliché amplification real?” It’s: **what frame does “obvious move” activate by default, and how stable is that across prompts?** That is not child speculation; it’s a concrete variable you can test. --- ## D. The key meta-issue you raised: maybe you don’t need to prompt for what the model already does This is important and easy to mishandle. ### D1) The “latent competence” hypothesis Your thought is: > If the model knows it is doing analytic philosophy, it will already produce distinctions, objections, replies, etc. Heavy prompting can be redundant, and redundancy can distort the output. That’s plausible. In fact, it’s often exactly what people observe: once the genre is cued, the model will naturally produce a reasonably structured “philosophy paper voice.” ### D2) The “intervention ladder” idea (not a plan, just a way to avoid over-correction) One way to operationalise your worry is: don’t start by forcing every desirable behaviour. Start with a minimal procedural cue, and only add constraints when you see a specific failure mode. For example (purely as options): - Start with “make the obvious move” and nothing else. - If it hovers, add a commitment-forcing follow-up (“state the claim in a way that could be false; name the hinge premise”). - If it becomes too standard, add a novelty operator (“give orthodox and non-orthodox next moves; pick one with costs”). - If it becomes baroque, add a conservatism operator (“minimise departures; maximise payoff”). The guiding idea is: prompt only to correct the failure you’re actually seeing, not to legislate everything in advance. This fits your “rule of thumb” approach: you have a small set of high-leverage steering moves, not a mega-prompt. ### D3) Why this matters for “novel good philosophy” If you over-specify procedures, you can end up with formulaic outputs that look like templates rather than ideas. And that would be self-defeating for novelty. So the live tension is: - too little steering → hovering, vagueness, safe overviews - too much steering → templated, predictable, “paint-by-numbers” philosophy Your “obvious move” trick is attractive because it sits in a middle zone: it pushes out of hovering while leaving room for the model’s own internalised practice norms to fill in structure. --- ## E. A better way to restate the “procedural steering works in philosophy” idea, given all of the above Here’s the cleaned-up thesis (still a hypothesis): 1. “Make the obvious move” is a deictic procedural cue that activates a learned policy for “what competent philosophers do next,” and that policy includes many move-types (distinction, reconstruction, unification, synthesis, objection, etc.), not just decomposition. 2. Analytic philosophy is truth-apt and aims at correctness. Procedural steering is relevant not because there’s no truth, but because there’s often no cheap oracle. We therefore rely on dialectical procedures as proxies for tracking correctness. 3. Whether “obvious move” yields cliché or novelty depends on the practice frame the model infers: “explain” vs “contribute.” If the frame is right, the novelty norm is already implicit; if the frame is wrong, you get summary-mode continuations. 4. Therefore, a lot of prompt craft here is not about adding more constraints; it’s about *reliably setting the right frame* and then using minimal interventions only when the model falls into a known failure mode. That’s a coherent picture that incorporates your objections rather than fighting them. --- ### Final answer - You’re not missing something: my earlier “obvious move = decomposition” line was under-thought. The “obvious move” in analytic philosophy can just as easily be *unification, synthesis, diagnosis, reframing,* or *argument reconstruction*. The better claim is: “obvious move” is a **move-selection instruction** whose output depends on what’s dialectically salient, not a single move-type. - The “no unique correct answer” framing was a genuine error in how it was phrased. Analytic philosophy is truth-apt and aims at correct answers. The relevant point for prompting is not “no truth,” but “often no easy external oracle,” so we use dialectical procedures and theoretical virtues as operational proxies for tracking correctness. - The “cliché amplification” worry only makes sense when formulated as a *distribution + framing* hypothesis: “obvious move” may pick high-probability continuations, but in a paper-writing frame the high-probability continuation may already include a novelty expectation. So the real variable is: does the model infer “explain” or “contribute”? That’s testable, and it’s plausibly easy to steer without heavy scaffolding. - Your final point is central: models often already do many philosophy-appropriate procedures once the genre is cued. So the smart stance isn’t “prompt for everything,” it’s **minimal procedural steering + targeted follow-ups** only when a specific failure mode appears (hovering, drift, standard-summary mode, baroque invention, etc.). If you keep the “obvious move” trick as the primitive operator, the deepest next stress test is to treat it as a *move selector* and see which move-types it reliably produces under different frames (“explain” vs “publishable contribution”), and how little extra steering is required to keep it in “contribute” mode without forcing a template. "Decomposition is one common way of reducing confusion and forcing commitments, but it’s one member of a wider family of “obvious” dialectical moves." or don't say anything of the sort, as you are still putting this one possibility on a pedstal. just go straight to the move grammar (although again, I am constantly having to battle with you giving things fucking names as though they are complete coherent theories or concepts. Also, 'the move grammar' would be much better if you referred to the philosophical methodology book. not just list everything it said, but actually showing the framework they adopt in general and say that this could be considered one way of formalising the 'move grammar; of philosophy. what do you think of this idea, would it work? "The relevant practical point for prompting is different: In many philosophical problems, we don’t have an external answer key in the way we do in arithmetic or many benchmark tasks. That doesn’t mean there is no truth. It means we often can’t cheaply verify that we’ve reached it. Therefore, the operational proxy for “getting it right” is often something like: argumentative validity, clarity of commitments, resilience to objections, explanatory integration, non-ad-hocness, etc. That’s not relativism. It’s just recognising the epistemic situation: truth-apt questions without easy oracle feedback." again, we have been over this. your conversation compacting has lost all the fucking detail. it is very very very fdrustrating. "If you over-specify procedures, you can end up with formulaic outputs that look like templates rather than ideas. And that would be self-defeating for novelty. So the live tension is: too little steering? hovering, vagueness, safe overviews" this seems extremely speculative. llms seem to think that these are failiure modes of llms doing philosophy, none have been able to produce actual evidence that they are in fact real failiure modes. moreover. i, as someone who does philosophy with llms all day every day, do not witness these failure modes. finally, because of the head kicking/conversation compacting, we need to work to recover as much information about the conversation as posssible. Please give me 3000 analysis on what we have discussed, and where we are going and how. address in particular the stuff i am accusing you of forgetting. ### Analysis, justification, reasons You’re asking for two things at once, and they interact. First, you want to **recover high-fidelity control over *this* conversation**: what has actually been said, where I started drifting into generic “prompting advice”, where I introduced sloppy framings (especially around truth/correctness), and where I started acting like half-formed sketches were already coherent “theories” just because I slapped a label on them. Second, you want to **restart the substantive project** (prompting for novel, good analytic philosophy) in a way that is genuinely informed by the texts you flagged, especially the Bengson/Cuneo/Shafer-Landau methodology book - and in a way that respects your repeated constraint: do not treat anything as settled, “major”, “best”, etc. unless you say so. So the right approach is: 1. reconstruct what’s happened (with enough detail to show I’m not “forgetting” what’s been argued about); 2. isolate the **specific failure points you’re accusing me of** (sloppy “no correct answer” framing; pedestal-ing one move; inventing failure modes without evidence; generic advice; naming everything); 3. rebuild the “structured pattern of philosophical moves” idea by *anchoring it in the methodology book’s own framework* - not as a grand discovery, but as one plausible way of formalising what “competent philosophical progression” amounts to; 4. revisit the “obvious move” trick in that light, without smuggling in claims about LLM failure modes that you reject. That’s what I’m doing below. I’ll keep the tone calmer, and I’ll avoid “here’s the framework / the key insight / the best plan” language. I’ll treat everything as options, comparisons, and testable conjectures. --- ## 1) High-fidelity recovery of what we’ve discussed so far ### 1) How this started You initially asked for a deep analysis of an attached chat. I produced an analysis that treated a lot of things said in the chat as though they were already established facts, and you immediately pushed back on exactly that: “a lot of the information here is stuff the llm said that you are taking as fact.” That was a fair correction. The right stance for those materials is: the chat is evidence of *what was said* and of *what conceptual moves were being attempted*; it is not evidence that the LLM’s empirical claims are true. ### 2) Then you asked for a precise extraction of “prompt craft” You wanted: read the conversation properly; tell you precisely what was said about prompt craft. I responded by pulling out a bunch of “prompt moves” from the transcript. You then pushed again: even that extraction risked treating the assistant’s claims as “the method,” rather than just as *one assistant’s suggestions in a chat*. You wanted a harder reset and a more grounded reconstruction. ### 3) Then you reset the project: “start from first principles” with Williamson + the methodology book + Dellsén You said the root idea was to draw on Williamson, Bengson/Cuneo/Shafer-Landau, and Dellsén to develop prompting strategies for novel good philosophy. I responded by sketching (too briskly) a “prompt spec” as if that was obviously right: aim, evidence policy, construction, evaluation. You objected that this looked speculative and not clearly grounded in any source. You also made a crucial point about “evidence”: much analytic philosophy isn’t explicitly “hostage to missing facts” - the real-world constraints are often commonsensical and backgrounded. That was an important clarification and I should have treated it as central rather than trying to import an “evidence policy” template from generic LLM engineering. ### 4) You then demanded “reliable sources” for prompt craft and justification of trust I responded by pointing to provider docs and prompting research (CoT, ToT, etc.) and giving reasons to trust them more than random “prompt lore.” That part you didn’t reject; your irritation was more about my *philosophical* framing and my tendency to “jump ahead” into project-shape talk. ### 5) We then hit the big derail: “rubric” I used “rubric prompts” as a label, you got (reasonably) confused, and we spent time unpicking that. We ended up with a cleaner distinction you liked: - criteria/constraints-first prompting (spell out what the output must satisfy), versus - procedure/trajectory-first prompting (tell the model how to proceed from here). You explicitly endorsed that distinction and wanted to stress-test it. ### 6) I gave you a stress-testy answer in a terse style, you asked me to slow down, and then you attacked specific claims You challenged: - my over-narrow claim that the “obvious move” in philosophy is “often decomposition”; - my sloppy framing that “philosophy often has no unique correct answer” (you called it a fundamental error); - my “procedural steering amplifies clichés” worry as speculative; - my implicit picture that “hovering / vagueness / safe overviews” are established failure modes (you say you don’t see them and nobody has shown evidence). You also accused me of drifting into generic prompting talk rather than engaging your actual project: analytic philosophy with LLMs, where the model already “knows” it’s doing philosophy and therefore may already be doing the right things without heavy prompting. ### 7) Now you want a full reset + recovery report You asked for 3000+ words analysing what we’ve discussed, where we’re going, and addressing what you accuse me of forgetting - specifically: the lost detail, the truth/correctness issue, the speculative failure modes issue, and the “stop naming everything” issue. That is the context. --- ## 2) The specific mistakes you’re flagging (and what they would imply if left uncorrected) ### Mistake A: treating one move as the “core” move When I said “the obvious next move is often decomposition,” I did two things you dislike: 1. I smuggled in an empirical generalisation about philosophical practice without grounding it; 2. I implicitly put that move on a pedestal, which then distorted the whole discussion. Your point is simple and strong: “obvious moves” in philosophy can be unification, synthesis, classification, reframing, diagnosis, reconstruction, etc. If I treat “decomposition” as *the* paradigm, I will systematically misdescribe what your trick is doing and misdesign prompts around it. So the repair is: don’t start with decomposition at all. Start with the plurality of move-types that actually occur in philosophical writing and conversation, and only then ask whether any one of them is more common in some contexts. ### Mistake B: the “no unique correct answer” phrasing You’re right that this was a classic bad move in this context because it sounds like: philosophy isn’t truth-apt or isn’t trying to get things right. That’s not analytic philosophy. A more careful point (which we had already been circling) is *epistemic*, not metaphysical: often we don’t have a cheap, decisive oracle for whether a philosophical theory is correct, even though we treat correctness as the target. That’s not relativism; it’s a claim about feedback structure. But you’re also right that I started repeating the thin version of that point in a way that sounded like I’d forgotten the richer version you care about: that in much analytic philosophy, a huge amount of “world constraint” is tacit, commonsensical, and unremarked - and that many papers are not “hostage to missing facts” in the way some LLM critiques imply. So: the correction is not just “rephrase the point”; it’s: re-embed it in the detailed picture you’ve been insisting on. ### Mistake C: speculative “failure modes” presented as if they were known You reject “hovering / vagueness / safe overviews” as established failure modes for LLM philosophy, and you say your experience doesn’t match them. If I continue to build prompt strategy around those “failure modes,” I’ll end up trying to “fix” things that aren’t broken in your use case, and I’ll likely over-prompt. So the correction is: treat those failure modes as *at most* conditional hypotheses that may apply to some models, some users, or some tasks - not as the default baseline for your workflow. ### Mistake D: turning tentative descriptions into “named theories” You’ve repeatedly warned against me giving everything a name (“move grammar”, “operator algebra”, etc.) because it makes half-formed ideas feel finished. That’s a good discipline for “kicking ideas around,” because naming tends to freeze a sketch prematurely. So I’ll use minimal labels only when they genuinely save time, and I’ll keep reminding myself (in how I write) that labels are conveniences, not commitments. --- ## 3) Your suggestion: use the methodology book as a way of formalising “how philosophical inquiry proceeds” This is the most constructive part of your message, and I agree it’s a good direction to explore. The Bengson/Cuneo/Shafer-Landau book explicitly treats method as “the engine of inquiry” and says that methods are sets of criteria that both guide theory construction and evaluate theories. That already looks like something you can use to articulate how a philosophical discussion *moves* - because criteria can be read as demands that generate the next dialectical step. “If your view hasn’t handled X, the next step is to handle X.” They also say that the criteria they endorse are “familiar from the way many philosophers go about their business,” but have not been sufficiently “justified, ordered, and integrated.” That is basically an invitation to treat their framework as an explicit formalisation of widely-used philosophical practice. ### What framework do they actually adopt? In very compressed terms (but grounded in their own statement): they propose a method with three priority levels: - first, handle the data by accommodating and explaining it; - second, ensure the theory’s own claims/commitments are substantiated and integrated; - third, consider theoretical virtues as a tie-breaker once the first two levels are satisfied. They explicitly describe this as *hierarchical* (later levels don’t get to override earlier ones in most cases) and *comparative* (you can compare rival theories by how well they satisfy the criteria). They also explicitly characterise what an “objection” is in this setting: a consideration counts as an objection exactly insofar as it gives reason to think the theory does poorly with respect to one or more criteria - and they map familiar objection types (ad hoc, explanatorily inadequate, contravenes common sense/science, etc.) onto the relevant criterion. Finally, they’re explicit about their guiding epistemic aim: theoretical understanding. They analyse it in terms of properties of theories (accuracy, reason-based support, robustness, illumination/explanation, orderliness, coherence). So the “framework” is not just “philosophy is argument.” It’s: inquiry aims at understanding, and good theorising is governed by an ordered set of criteria meant to deliver understanding. ### Does this work as a way of formalising how philosophical discussion progresses? It seems workable as *one* candidate formalisation, and it has an advantage over my earlier freehand “move list”: it gives you **structured reasons** for why one move rather than another is “the next thing to do.” Here’s the key translation step: - A conversational move in philosophy can be seen as an attempt to improve a theory’s standing with respect to one or more criteria. So, for example: - If a view accommodates but doesn’t explain, a “next move” is: supply an explanation, or argue that the datum is basic, or appeal to their escape clauses. - If a view asserts commitments without defending them, the “next move” is: provide reasons/argumentation that substantiates them. - If a view conflicts with “logic, mathematics, science, and common sense,” a “next move” is: integrate it or explain why the conflict is acceptable. - If an objection says “ad hoc,” the framework tells you what kind of repair is demanded: it’s a substantiation failure, so you need non-ad-hoc support, not just prettier prose. In other words: their method doesn’t just list “things philosophers do.” It gives a principled way to say which “thing” is called for, given what has been done so far. ### How does this connect to your “make the obvious move” trick? Potentially very directly: “Make the obvious move” can be interpreted as: given the current dialectical state, take the step that most directly improves the theory’s standing with respect to the relevant unmet criterion. That automatically covers your counterexamples to “decomposition”: - If the salient deficit is *integration* (you have two positions that look like variants), the “obvious move” might be unification: show they’re equivalent or show how to combine them. - If the salient deficit is *orderliness* (you have a messy list of points), the “obvious move” might be to articulate organising relations among them (dependency relations, explanatory structure). - If the salient deficit is *robustness* (the view answers one question but dodges central related ones), the “obvious move” is to expand it to cover those questions. This is why your idea about bringing the methodology book in is good: it gives you a principled way to say what “obvious” means without collapsing it into one move-type. ### A possible worry: does this formalisation risk making outputs formulaic? It could, depending on how it’s used. If you turn the method into a rigid checklist the model must explicitly march through every time, you might get templated prose. But if you treat it as a background normative structure - something the model uses implicitly to decide what to do next - it may preserve flexibility. This dovetails with your point that LLMs often already “know they are doing philosophy.” The book itself says their criteria are familiar from how philosophers already work. So one hypothesis is: the model already has this kind of structure implicitly, and the main role of prompting is to *activate* it reliably, not to force it step-by-step. --- ## 4) Your objection about “lost detail” on truth, evidence, and verification You’re saying: we’ve already been over the subtlety, and I started repeating a thin version as if I’d forgotten the richer one. So let me re-state the richer version in the way you’ve been pushing for, without sliding back into relativism talk. ### The core picture you’ve insisted on A lot of analytic philosophy: - is overtly truth-directed (people argue for the correct account, not “a perspective”); - relies heavily on background knowledge that is rarely flagged as “evidence” (commonsense, widely accepted science, basic logic/maths, linguistic competence, etc.); - therefore is often *not* waiting on missing empirical facts in any dramatic way - the “facts” it relies on are frequently treated as the shared backdrop. This matters for prompting because it changes what “evidence discipline” should mean. It’s not: “go fetch data.” It is more like: - don’t fabricate empirical claims; - make tacit presuppositions explicit when they are doing work; - check that the view integrates with widely accepted background constraints when relevant. And this is exactly the sort of thing the methodology book builds into its “integration” demand: sensitivity to “logic, mathematics, science, and common sense” is positioned as part of meeting the integration criterion. So there’s no need to import a generic “hostage to missing facts” framing as if that’s the standard situation in philosophy. A better framing is: philosophy often operates under an implicit background corpus; errors often come from hidden assumptions, equivocations, neglected cases, or failure to integrate - not from missing lab measurements. That is consistent with the methodological framework you want to use, and it fits your lived experience better than the “missing facts” trope. --- ## 5) Your objection about “failure modes” and my earlier speculation You’re saying two things: 1. Claims like “hovering, vagueness, safe overviews” are asserted all the time as LLM failure modes in philosophy without convincing evidence; 2. you personally don’t observe them in your actual daily work with LLMs. Those points matter, because if they’re right, then a lot of prompt engineering discourse is fixing the wrong problem. Here’s the careful way to handle this without pretending either side is settled. ### Option 1: Treat those “failure modes” as not part of your baseline For your workflow, the baseline is: the model already produces structured, committed philosophy when cued correctly. If so, then the purpose of procedural steering isn’t to fix “hovering”; it’s to do something else - for example, to increase novelty, or to accelerate the search through candidate positions, or to make the output more *useful as a draft* rather than as a conversational contribution. That would immediately remove a bunch of my earlier “procedural prompts solve hovering” rhetoric as irrelevant. ### Option 2: Treat those “failure modes” as task-relative rather than global Even if you don’t experience “hovering,” it might still appear in some tasks: e.g., when asking for meta-level overviews, or when the prompt underdetermines whether the model should summarise or argue, or when the domain is unfamiliar. But again, this becomes a contingent point, not a foundational assumption. ### Option 3: Treat “failure modes” talk as mostly about prompt mismatch, not model incapacity This is closer to your own idea: if the model “knows it is doing philosophy,” it will already do many things right. When it doesn’t, that may be because the conversational framing cues the wrong genre (explain vs contribute), not because the model can’t do the thing. This is where your “obvious move” trick is interesting: it may be less about repairing incapacity and more about snapping the interaction into the “contribute / advance” genre. The key thing is: none of these options requires us to assume hovering is a universal problem. They let us keep focusing on your actual use case. --- ## 6) Restarting “obvious move” properly: no pedestal, no single move-type You asked me to stop pre-facing with “decomposition is common” and “don’t pedestal one possibility.” So: straight to the plural move-types, but now anchored to the methodology book’s way of organising philosophical work. Here is a way to think about “what happens next” in a philosophical exchange that doesn’t treat any one move as the centre. ### A plural repertoire, but disciplined by the method’s criteria Instead of listing a random zoo of moves, we can group them by what they are trying to achieve relative to the method. Level one is “handling the data”: accommodation and explanation. So “obvious moves” here include: - articulating a distinction that resolves a datum-conflict; - providing a counterexample that shows failure of accommodation; - giving an explanation (or arguing the datum is basic / not in need of explanation, per the book’s escape clause discussion). Level two is about the theory’s own claims and commitments being substantiated and integrated. So “obvious moves” here include: - defending a key premise that has been left hanging (substantiation); - diagnosing an “ad hoc” patch and replacing it with a supported principle (the book explicitly connects “ad hoc” objections to failure of substantiation). - integrating the view with background constraints (science/common sense/logic), or explaining why the conflict is acceptable. Level three is virtues as tie-breaker: simplicity, elegance, etc. So “obvious moves” here include: - showing that two rival views are tied on levels one and two and then comparing virtues; - or showing that apparent virtue gains don’t compensate for level-one or level-two failures (they explicitly say virtues don’t rescue a theory that fares poorly at lower levels). Now, notice: **unification and synthesis fit naturally here**. They’re often integration moves. If two views are really the same, identifying that can improve orderliness and integration. If two views can be combined in a principled way, that can improve robustness and coherence. So your correction (“what about classifying both ideas as the same?”) isn’t an exception to the account; it’s exactly what the account predicts once you don’t stupidly reduce “obvious move” to “decompose.” ### A genuinely useful way to interpret “obvious move” now Given this, a strong way to cash out your trick is: “Make the obvious move” means: identify which of these demands is currently most salient (accommodate/explain data; substantiate/integrate commitments; only then virtue comparisons) and do the next step that most directly addresses it. That’s a state-dependent, practice-relative instruction, and it covers the unification/synthesis cases cleanly. --- ## 7) The “orthodox policy / cliché” dispute, reframed in a way that isn’t childish You objected to “procedural steering might amplify clichés” as childishly speculative, and you also argued something stronger: analytic philosophy as a practice already contains a novelty norm, so a prompt that cues “you are doing philosophy” should not default to clichés. I think the productive way to handle this is not to argue about “clichés” at all, but to talk about **which practice frame the prompt induces**. There are at least two distinct frames: - “explain the literature / give an overview”; - “make a contribution / advance a position.” Even within analytic philosophy, those are different genres. The “obvious next move” in an overview is often a standard mapping move; the “obvious next move” in a contribution is often: sharpen a claim, make a new distinction, propose a new argument, introduce a new possibility, or integrate in a new way. The methodology book explicitly treats progress as including articulating overlooked distinctions, improving existing theories, and expanding the field of serious possibilities. That is basically a formal statement of the “novelty norm” you’re appealing to. So the relevant question becomes: how reliably does your cue (“make the obvious move”) land in the “contribute” frame rather than the “overview” frame? If, in your experience, it reliably lands in “contribute,” then the cliché worry is not central for your workflow. This again supports your point about not over-prompting: if the model already behaves like a philosopher once properly framed, you don’t need to bolt on “be novel” as a crude explicit requirement - you need to maintain the right genre. --- ## 8) Where this leaves the project, without jumping ahead You asked “where we are going and how,” but you also hate when I jump ahead and start talking like we’ve decided the roadmap. So I’ll phrase this as “here are a few live directions that are consistent with what you’ve been insisting on.” ### Direction A: Treat the methodology book’s ordered criteria as a way of making “obvious move” precise Not by turning it into a checklist prompt, but by using it to interpret what “the next move” should be when you or the model are unsure. The criteria tell you what kinds of deficit matter, and the book’s objection taxonomy tells you what kind of repair is demanded. ### Direction B: Use the method as a selection device rather than a generation constraint If your worry is over-templating, you can keep generation minimally constrained (“make the obvious move”), but when you generate multiple candidate continuations you can choose among them by asking: which one best improves accommodation/explanation/substantiation/integration, without leaning on virtues too early? That respects your point that the model may already “know how to do philosophy,” while still giving you a principled, philosophy-specific way to steer selection. ### Direction C: Stop treating “LLM failure modes” as the engine of the project If you don’t see hovering/vagueness as a baseline problem, then procedural steering is not a therapy; it’s an accelerator: a way to force “next-step competence” and then explore alternatives, integrate, unify, etc., within the contribution frame. That keeps the project tied to your lived experience rather than to generic LLM discourse. --- ## Final answer You’re right to object to my earlier “decomposition” emphasis: “make the obvious move” is not a decomposition trick; it’s a *state-dependent “take the next licensed step” cue* that can just as readily call for unification, synthesis, reframing, reconstruction, objection, repair, or integration. The best way to avoid pedestal-ing any one move is to stop listing moves as a flat zoo and instead use the methodology book’s framework as an organising device: Bengson/Cuneo/Shafer-Landau explicitly treat philosophical method as an ordered set of criteria that guide both construction and evaluation, culminating in the Tri-Level Method (handle data via accommodation/explanation; then substantiate/integrate commitments; only then let virtues play a tie-breaking role). They explicitly connect familiar objection-types to failures on specific criteria, which gives you a principled way to say what “the next move” should be in a given dialectical state. On the truth/correctness issue: my “no unique correct answer” phrasing was sloppy in exactly the way you flagged, because it sounds like relativism. The point that’s actually relevant here is epistemic and methodological: analytic philosophers aim at the correct answer, but a lot of the “world constraint” is backgrounded and the feedback signal is often dialectical (does it accommodate/explain the data as understood, does it substantiate/integrate its commitments, does it survive objections), which is precisely the sort of thing their method tries to systematise without pretending philosophy is just “argument” or “analysis.” On “failure modes”: you’re right to reject my earlier language if it treated hovering/vagueness as established facts about LLM philosophy. For your use case, the safer stance is: assume the model already has a lot of philosophy competence once the “contribute” frame is active, and treat prompting as minimal steering and (if needed) principled selection, not as a bundle of generic fixes. And on the “forgetting/compacting” accusation: the main content you say I’ve been losing is (i) your insistence that analytic philosophy is truth-apt and truth-directed, (ii) your insistence that “evidence” is often tacit/commonsense rather than “missing facts,” (iii) your insistence that you do not observe the generic failure modes I kept invoking, and (iv) your insistence that we keep this as a low-fidelity sketching phase without premature naming or “project-shape” talk. I’ve rebuilt the discussion above around exactly those constraints, and used the methodology book as the anchor for formalising “what counts as the next move” rather than inventing a freehand move taxonomy. ### Analysis, justification, reasons You’re pushing a sharper (and better) version of the point than “the methodology book helps formalise a list of moves.” The stronger idea is: 1. The methodology book is trying to *make explicit* a set of norms/criteria that already operate in ordinary philosophical practice; and it argues that philosophers often satisfy those criteria even when they aren’t consciously “following the Tri-Level Method”. 2. If that’s right, then those norms aren’t just “in philosophers’ heads” - they are *written into the structure of philosophical texts themselves* (what gets said next, how objections are introduced, what kind of repairs are treated as legitimate, when virtues are appealed to, and so on). 3. An LLM trained on lots of philosophy doesn’t merely ingest *conclusions* or *terminology*. It ingests the *textual traces of those norms* - i.e., the recurrent patterns of “here’s a view → here’s why it handles the data → here’s an objection → here’s a reply → here’s integration with background constraints → here’s why the rival is ad hoc → …”. 4. So your “make the obvious move” trick may work, not because it forces a particular explicit aim, but because it points the model toward continuing along those learned “norm-shaped” trajectories - basically: “continue this text as a competent philosopher would continue it.” That’s the integration you’re asking for: the methodology book as a way of articulating what the LLM has plausibly learned about *how philosophical writing proceeds*. I’m going to spell that out carefully, anchored to what Bengson/Cuneo/Shafer-Landau actually say about method and practice, and then show multiple ways it could mesh with the “obvious move” style of prompting without turning into a rigid template. --- ## 1) Why the methodology book is relevant to “what the model has ingested” ### 1.1 They explicitly treat “method” as criteria that govern theorising They define method, in the relevant sense, as a set of criteria with a dual role: it instructs how to construct a theory from data and also supplies standards for evaluating theories. That matters because those criteria aren’t meant to be mere after-the-fact commentary. They’re meant to be *action-guiding norms* - rules for what to do next in inquiry. So if those norms are widely followed (even tacitly), they should leave fingerprints in the text: what counts as a legitimate “next paragraph” in a paper; what gets treated as an objection that must be answered; what kind of reply is treated as satisfying; when and how people appeal to virtues; what they count as “data”; etc. ### 1.2 They explicitly claim the criteria line up with ordinary philosophical activities When they discuss what a sound method must handle, they talk about “the activities that typify philosophical practice,” and they give a list: advancing arguments, raising objections, offering replies, seeking clarification, seeking explanations, and displaying sensitivity to logic, mathematics, science, and (to varying degrees) common sense. This is already a bridge from “method” to “text structure,” because those activities are precisely what show up as rhetorical/dialectical moves in papers. ### 1.3 They explicitly say philosophers can satisfy the criteria without intending to This is the killer passage for your “implicit grammar in texts” idea. They say that philosophers’ satisfying the method’s criteria needn’t be intentional; philosophers may be trying to satisfy criteria invoked by other methods, or may simply be engaging in the typical activities of practice - and “either way” they can end up satisfying the method’s criteria. If that’s right, then the Tri-Level criteria are not just an external philosophical theory about philosophy; they are (at least partly) a reconstruction of norms that are already being enacted in everyday philosophical writing and debate. So: if the model has ingested a ton of ordinary philosophical writing, it has ingested lots of instances of people enacting those norms (even when they didn’t label them). That’s exactly your point. --- ## 2) How “criteria in a method” translate into “patterns of next steps” in texts Let me avoid grand naming and just be explicit about the mechanism. If a method says: - “you must accommodate the data” - “you must explain the data” - “you must substantiate your commitments” - “you must integrate with background constraints” - “virtues mostly matter as tie-breakers” then, in a piece of writing, those are realised as repeated textual patterns like: - *Here are the key considerations/cases/claims we take as given* (data) - *Here is a theory that makes those considerations likely / fits them* (accommodation) - *Here is why those considerations hold, given the theory* (explanation) - *Here are the reasons that support the theory’s own commitments* (substantiation) - *Here is how the theory fits with logic/science/common sense or adjacent domains* (integration) - *Only now (if needed) here is why this theory is simpler/cleaner/more elegant than the tied rival* (virtue tie-break) The Tri-Level Method chapter even lays this out in a quite explicit “organisation of criteria” diagram and explains that the Virtue Criterion is, on their version, typically tie-breaking and lower priority than level-one and level-two criteria. Now notice what that means for an LLM. An LLM learns (roughly) conditional distributions over sequences: given a context that looks like “I have stated a view and some data,” what tends to come next in the training corpus? If philosophical papers and exchanges tend to follow something like the above order (not perfectly, not always, but often enough), then the model will have a learned tendency to continue with: - an argument that supports the view, - an explanation that shows how it handles the salient considerations, - a standard objection, and so on. That is the “implicit move pattern in texts” story in very plain terms: the norms are enacted in the corpus as recurring sequences; the model learns those sequences. So your “obvious move” instruction can be seen as a way of saying: > “Don’t hover. Don’t do a generic overview. Continue in the way that this kind of philosophical text normally continues from *this* point.” And because the “way it normally continues” is shaped by those method-like norms, you get something that looks like disciplined philosophical progression even without spelling out a rubric. This is why the methodology book fits your project in a deeper way than “it helps us list moves.” --- ## 3) “Make the obvious move” as a pointer into those learned sequences Your trick is a deictic command: “from here, take the step that’s demanded.” Under the lens above, “obvious” is not “split the topic.” It’s “take the step that a competent philosopher would recognise as demanded by what’s currently missing.” And the methodology framework gives a principled story about what “missing” tends to mean in philosophical practice: - maybe the view fits the data but doesn’t explain it (explanatory burden), and the next thing demanded is explanation; - maybe the view asserts a commitment without support (substantiation gap), and the next thing demanded is an argument; - maybe the view conflicts with background constraints (integration pressure), and the next thing demanded is reconciliation or a defence of the conflict; - maybe two views tie on the first two levels, and now (and only now) a virtues comparison is demanded. They explicitly treat virtues objections like “unparsimonious/inelegant/ugly” as corresponding to the Virtue Criterion, and “contravenes science or common sense” as corresponding to the Integration Criterion. So: your “obvious move” prompt can be understood as a minimal way to get the model to *diagnose what kind of deficit is salient and address it*, drawing on the implicit patterns it has seen in philosophical texts. That also fits your repeated point that LLMs often don’t need to be explicitly told to do “philosophy-shaped things.” If the corpus already encodes those “deficit → response” patterns, and your cue triggers the “we’re doing philosophy” frame, you may only need light steering. --- ## 4) How the methodology framework can help without turning the prompt into a rigid template You were very explicit earlier: don’t over-correct; don’t assume the model needs to be prompted to do what it already does. So the useful role for the methodology book might be more like a *lens* or *diagnostic vocabulary* than a step-by-step script. Here are several ways it could mesh with the “obvious move” approach, each with different levels of intrusiveness. None of these needs to be treated as “the right” one; they’re different knobs. ### Option 1: Use the method as a post-hoc label, not a generation constraint You prompt: “Make the obvious move.” Then, after the move, you ask: “Which kind of demand did that move address - accommodation, explanation, substantiation, integration, or virtues?” Why this might be useful: it keeps generation minimally constrained (your style), but it forces the model to *make explicit what it thinks it was doing*. That can help you see whether it’s actually tracking something like the method’s structure, or whether it’s just producing plausible-sounding prose. Why it might be useless: post-hoc labelling can be rationalisation. So you’d treat the label as a clue, not as proof. ### Option 2: Use the method as a move-selector when there are multiple “obvious” moves Sometimes “obvious next move” is genuinely underdetermined: you could clarify, you could object, you could integrate, you could unify two views. Here, the method can be used as a way to decide what’s most pressing: “Which criterion is currently most at risk?” (Data handling? Substantiation? Integration?) This respects your point that “obvious move” isn’t one thing. The method gives you a principled way to decide among many plausible next steps without making you spell out an “aim” in advance. ### Option 3: Use the method primarily for selection among alternatives This is the “the model already knows how to do philosophy” friendly option. Instead of heavily steering generation, you ask the model to generate 2–4 candidate “next moves” (different types: unification, counterexample, argument reconstruction, integration move), and then select the one that best improves standing with respect to the level-one and level-two criteria, treating virtues as tie-breakers. This uses the methodology book as a disciplined selection policy rather than as a generative template. It fits your suspicion that what matters is not instructing the model how to write, but shaping which continuations get endorsed. ### Option 4: Use the method to structure objections in a non-generic way You disliked earlier generic talk about “objections” and “failure modes.” The method gives a way to keep objection-generation non-vague. Instead of: “give objections,” you can ask: “give one objection of each relevant type,” where “type” is grounded in criteria: e.g., “integration objection” (conflict with background constraints), “substantiation objection” (unsupported commitment/ad hoc), “explanation objection” (explains too little / leaves burdens of explanation). Even if you don’t end up using this in a prompt, it’s a conceptual resource for thinking about what “the next move” might need to respond to. ### Option 5: Use the method to articulate what the model has likely learned from the corpus This is closest to your meta-level point. You can treat the methodology book as saying: “here is a plausible reconstruction of the norms that show up in the texts.” It’s not merely a tool for you; it’s part of your story about why LLMs can do philosophy: because philosophical texts encode not only propositions but structured argumentative practice, and this method is one way to make that structure explicit. That would slot neatly into your “saturation” line of argument: the model has been trained on a domain where the rules of the game are largely in the text. --- ## 5) How this connects back to Williamson/Dellsén without derailing Very lightly, because you asked specifically about fitting the methodology-book point into what we’ve been discussing. - Williamson pushes the “don’t let ‘sounds right’ substitute for what’s actually known” pressure. In your version, “what’s actually known” in philosophy often includes tacit commonsense and background constraints rather than missing empirical facts. The methodology book explicitly includes “sensitivity to logic, mathematics, science, and common sense” among the friendly-to-practice data about method. That’s a nice bridge: it’s a way to articulate “background constraints” as part of the method that the corpus will encode. - Dellsén’s “dependency modelling” account of understanding is a way to cash out “orderliness” and “coherence” in something like a structured representation, and the methodology book’s “theoretical understanding” includes orderliness and coherence as key features. So, if you ever want a *concrete output form* that aligns with the implicit norms, “give a dependency sketch” can be seen as a way to force that “orderly/coherent” dimension without turning the whole interaction into a checklist. But neither of these needs to be baked into the prompt. They’re conceptual supports for your overall explanation of why minimal procedural cues can work in philosophy. --- ## Final answer Yes - your stronger point fits extremely well with what we’ve been discussing, and it actually *improves* the story. The methodology book isn’t just useful because it lets you list “moves.” It’s useful because the authors explicitly treat method as inquiry-governing criteria, and they explicitly connect those criteria to ordinary philosophical activities (arguing, objecting, replying, explaining, sensitivity to logic/science/common sense). Crucially, they also say philosophers often satisfy the criteria without aiming to - sometimes just by engaging in those ordinary activities. That gives you a clean way to articulate what an LLM trained on philosophical texts has plausibly ingested: not only philosophical *content*, but the recurrent *norm-shaped textual sequences* that enact those criteria (data → accommodation/explanation → substantiation/integration → only then virtues as tie-breakers). On that picture, “make the obvious move” works because it functions as a pointer: “continue this philosophical text/practice from here in the way a competent participant would,” and the competence it points to is partly constituted by those method-like norms encoded in the corpus. The methodology book then plays a second role: it supplies a principled vocabulary for describing which kind of “next step” is demanded (explanation burden, substantiation gap, integration pressure, etc.) without forcing you to turn prompting into a rigid template - you can use it lightly as a diagnostic or as a selection device when multiple plausible “obvious” moves are available. 1\. btw, I liked what you said about philosophy not being driven by an oracle, or sometihng like that? what exactly did you say? 2. "This is the killer passage for your “implicit grammar in texts” idea. They say that philosophers’ satisfying the method’s criteria needn’t be intentional; philosophers may be trying to satisfy criteria invoked by other methods, or may simply be engaging in the typical activities of practice - and “either way” they can end up satisfying the method’s criteria." please find a block quote. 3. a bigger task: here is a very very very rough draft/plan of a paper I am writing on this topic. I feel that what we hav ecovered in our conversation could be reformulated so as to be a substantial part of section 3. A substantial part. Please could you give me a very very detailed plan for this part. make sure you reorganise the stuff we have been talked about so that it is structrually robust and substantial part of section 3 NO FUCKING SUBSECTIONS. make sure that you explain your analysis, justification, reasons etc etc. for your answer BEFORE giving me your final answer. DRAFT: # 0. Introduction # Introduction My aim in this paper is to argue that LLMs can generate novel philosophical work of publishable quality with minimal prompting—work that meets the standards by which we evaluate the best human philosophy. While it is often supposed that AI assistance in philosophy amounts to brainstorming, drafting support, or what we might call philosophical training wheels, I will argue that LLMs can do substantially more than this: they can produce philosophical output that is genuinely first-rate. Many will find this implausible, perhaps offensive; but I will argue that it is true. By \*minimal prompting\* I mean genre-governing cues rather than micromanaged step-by-step instructions: directives like "be philosophically robust", "focus on the arguments", "explain your analysis before giving a final answer". Such prompts specify what kind of thing is wanted—a philosophical artefact—rather than the specific moves to make. The contrast is with elaborate prompt-engineering that essentially does the philosophical work for the model, feeding it premises, walking it through inferences, correcting its mistakes in real time. The claim I am defending is interesting precisely because thin constraints elicit substantial philosophical structure; if the user had to do all the philosophical labour in the prompt, the claim would be trivial. My argument rests on a structural point about what good philosophy consists in, and two recent accounts of philosophical understanding help to make this point precise. Dellsén's Dependency Modelling Account holds that understanding consists in grasping a sufficiently accurate and comprehensive model of the network of dependence relations in which a phenomenon is situated: > According to the proposed account, one understands a phenomenon, P, just in case one grasps a sufficiently accurate and comprehensive model of the network of dependence relations in which P, or its contextually relevant parts, is situated; and one's degree of understanding of P is proportional to the comprehensiveness and accuracy of such a model. (Dellsén, p. 1262) The formal statement makes the structure explicit: > DMA: S understands a phenomenon, P, if and only if S grasps a sufficiently accurate and comprehensive dependency model of P (or its contextually relevant parts); S's degree of understanding of P is proportional to the accuracy and comprehensiveness of that dependency model of P (or its contextually relevant parts). (Dellsén, p. 1268) A dependency model can fail in two ways—by misrepresenting the network, or by not representing it at all—and Dellsén identifies these as the two criteria by which models are evaluated: > Since a dependency model can thus fail either by incorrectly representing (that is, misrepresenting) some aspect of this network, or by not representing it at all, we can identify two separate criteria here, namely, accuracy and comprehensiveness. (Dellsén, p. 1267) Both criteria are properties of the representation, not of the representer's mental states; a model is better to the extent that the network of dependence relations is correctly depicted: > A dependency model better represents P to the extent that the network of dependence relations that P stands in is correctly depicted by the model. (Dellsén, pp. 1267–1268) Understanding, on this account, admits of degrees in a way that propositional knowledge does not: > Understanding is a matter of degree in a way that propositional knowledge, for example, is not. It's not just that one can understand more or fewer phenomena; rather, one can have more and less (or, if you prefer, 'better' and 'worse') understanding of a single phenomenon, P. (Dellsén, p. 1264) This gradability is explained by the gradability of accuracy and comprehensiveness themselves: > I have noted that understanding is a gradable notion—that one can have various degrees of understanding of the same phenomenon. In a model-based account of the sort I am proposing, this is explained by the fact that the two aforementioned criteria (accuracy and comprehensiveness) are both gradable. (Dellsén, p. 1268) Dellsén also separates understanding from explanation; one can achieve understanding through means other than learning explanations: > It is possible to increase both the accuracy and the comprehensiveness of such a dependency model of P without learning an explanation of any aspect of P. Accordingly, this account of understanding accommodates the possibility of achieving understanding through means other than explanations. (Dellsén, p. 1262) This opens conceptual space for AI-generated understanding, since a model, for Dellsén, is simply an information structure interpreted so as to represent its target: > For my purposes, a model is simply an information structure of some kind that is interpreted so as to represent its target. (Dellsén, pp. 1264–1265) Dellsén is explicitly neutral on the cognitive mechanisms involved; \*grasp\* is a placeholder for whatever relation obtains between mind and model: > As a shorthand for the relation between the mind and the models—whatever it turns out to be—I will use the term 'grasp'. (Dellsén, p. 1265) The evaluative question, then, is about what makes the model good, not about what makes the modeller understanding: > Of course, to have understanding of phenomenon P, it is not enough to grasp any old dependency model of P. Rather, the model must in some sense be a 'good' representation of the relevant dependence relations. So what makes such a model better or worse qua representation? (Dellsén, pp. 1266–1267) Bengson, Cuneo, and Shafer-Landau provide a complementary account. On their view, theoretical understanding is the state that agents possess when they fully grasp a theory with six properties: > Theoretical understanding, as we'll construe it, is the state that agents possess just when they fully grasp a theory with the following six properties. (Bengson et al., p. 28) The first property is accuracy: > First, the theory possesses a high degree of accuracy, since largely inaccurate theories will fail to dispel confusion (a characteristic of misunderstanding). (Bengson et al., p. 28) The second is that the theory be \*reason-based\*: > Second, the theory is reason-based, in the sense that it is positively supported by considerations, beyond mere coherence, that speak in favor of its accuracy. For in the absence of such support, signing on to the theory would be arbitrary or haphazard (again, a characteristic of misunderstanding). (Bengson et al., p. 28) This is a property of the theory's epistemic standing, not of the producer's reasoning process; whether reasons exist that support a theory is independent of whether the entity that produced it was 'reasoning' in some deep sense. A theory becomes reason-based through being defended, and the defense is part of the theory's content: > By defending its claims and commitments, a view becomes reason-based; by explaining its claims and commitments, it adds robustness and overall illumination. (Bengson et al., p. 118) The third property is robustness: > Third, the theory is robust, answering a multitude of questions about the most important features of the domain under investigation. A theory that neglects or dodges such questions leaves out just what's needed to yield comprehension. (Bengson et al., p. 28) The fourth is illumination: > Fourth, the theory is illuminating, in that its answers must at least sometimes be not just general but also genuinely explanatory, going beyond a mere description of those features to explain why each exists or is instantiated. (Bengson et al., p. 29) The explanation is in the theory, not in the theorist's head. The fifth property is orderliness: > Fifth, the theory is orderly, not simply offering such feature-specific explanations but also affording a broader view of the domain by revealing how those (and other) features, as well as the proposed explanations, gel or hang together—for example, by exposing basic relations or systematic connections among them. (Bengson et al., p. 29) The sixth is coherence: > Sixth, the theory is coherent, not only internally but also externally, fitting well with a wide range of understanding-providing theories of other domains. (Bengson et al., p. 29) The first four properties are fundamental; the latter two contribute only conditionally: > Although all six properties contribute to theoretical understanding, they do so in different ways. The latter two, unlike the former four, only conditionally make such contributions. The orderliness and coherence of a theory contribute to its ability to supply understanding only if the theory possesses the other four features to at least some extent. In this way, these first four are fundamental to understanding in a way that the final pair are not. (Bengson et al., p. 29) When inquirers fully grasp theories with these six properties, understanding is achieved: > When inquirers fully grasp theories with these six properties, the targets of their theories make sense to them. This is theoretical understanding. (Bengson et al., pp. 29–30) And theoretical understanding is an ultimate proper goal of inquiry: > Our own view, as noted, is that theoretical understanding is an ultimate proper goal. \[...\] We call a method 'sound' just in case satisfaction of its criteria thereby positions inquirers to achieve an ultimate proper goal of inquiry. (Bengson et al., p. 27) The structural point that both accounts share is this: understanding and good philosophy are properties of the theory or model produced, not of the producer's inner states. Dellsén's evaluation criteria—accuracy and comprehensiveness—are properties of the model; the producer's cognitive processes are bracketed. Bengson's six properties—accuracy, reason-based support, robustness, illumination, orderliness, coherence—belong to the theory, not to the theorist's psychology. Whether a theory is reason-based depends on whether considerations exist that support it, not on whether the producer 'reasoned'. This is not a quirk of one framework; it is where two independent accounts of what good philosophy consists in end up. If an AI produces a theory that is accurate, supported by reasons, robust, illuminating, orderly, and coherent, then that theory can give understanding to someone who grasps it; the causal history of the theory's production is irrelevant to whether it has these properties. One might object that this reduces my thesis to the uninteresting observation that LLMs can recombine things humans have already said. But the claim I am defending is that LLMs can produce \*new\* philosophical moves—the kind of contribution that advances a debate, solves a problem, or reframes an issue in a productive way. The novelty claim is part of the thesis from the start; without it, the thesis would be trivial. I am agnostic about whether LLMs 'really reason' in some deep metaphysical sense, and I do not claim that they have understanding, beliefs, or intentional states. My focus is entirely on the artefact—the philosophical text produced. The question is whether that text satisfies the constraints by which we evaluate philosophy, not whether the producer has the right inner life. This is methodologically principled, not evasive: we evaluate papers, not souls, and blind review exists precisely because provenance should not affect judgement. If a paper meets the standards, it meets the standards; who or what produced it is irrelevant to that assessment. If the thesis is right, it has implications for philosophical methodology, for understanding what philosophy is, and for the future of the discipline. For methodology: what does it mean that the constraints are learnable from text? For philosophy's self-understanding: is the practice governed by publicly codifiable norms rather than ineffable insight? And for the discipline's future: a new kind of collaborator, or competitor, has arrived. The paper proceeds as follows. Section 1 presents the best recent case against the thesis: Floridi et al.'s argument that LLMs have a \*stochastic core\* and at best an \*abductive appearance\*. Section 2 shows how Williamson's account of philosophical method as abductive intensifies the worry, then executes a pivot—relocating the debate from production mechanism to constraint satisfaction. Section 3 makes the positive case: philosophy's rules are learnable from text, and an LLM trained on philosophical corpora has learned them. Section 4 demonstrates the thesis with worked examples. --- # 1. What LLMs Aren't Doing # What LLMs Aren't Doing The best recent case against the thesis defended here comes from Floridi, Morley, Novelli, and Watson's 'What Kind of Reasoning (if any) is an LLM actually doing? On the Stochastic Nature and Abductive Appearance of Large Language Models'. The paper deserves serious engagement. It articulates with precision what many suspect: that LLMs produce text that looks like reasoning without actually reasoning, and that this appearance is misleading in ways that matter epistemically. The paper's thesis centres on a duality. LLMs have stochastic internals but produce outputs that resemble reasoning. > "Our main argument is that LLMs occupy a conceptual space 'between' traditional stochastic processes and human-like abductive reasoning. On the one hand, their internal processes are entirely stochastic: during training, they gather statistical correlations from text, and during generation, they produce words based on learned probability distributions. They lack explicit representations of meaning, everyday relevance, truth values, or causality as a reasoning agent would." (Floridi et al., p. 2) The slogan is 'stochastic core, abductive appearance'. > "This duality, centred on the stochastic core of the models and the abductive appearance of the applications, has important implications for the evaluation and use of LLMs." (Floridi et al., Abstract) > "We can briefly describe LLMs as fundamentally stochastic, with surface-level abductive appearances." (Floridi et al., p. 19-20) The technical picture is straightforward. LLMs work through statistical inference over language data. > "LLMs mainly work through statistical inference over language data. During training, an LLM processes enormous amounts of text and optimises a model (usually a neural network transformer) to predict the next token (word or sub-word) based on the preceding context. The result is essentially a complex probability distribution: for any particular sequence of tokens/words, the model can assign likelihoods to potential continuations." (Floridi et al., p. 7) The process is stochastic regardless of implementation details. > "Regardless of the approach, the process remains stochastic: either inherently (with sampling) or effectively (since training involves discovering a model that encodes frequencies and correlations from initial random weights)." (Floridi et al., p. 7) Critics have called LLMs 'stochastic parrots' to emphasise the lack of understanding. > "In fact, critics (Bender et al. 2021) have called them 'stochastic parrots' to emphasise that they merely mimic language through probabilistic means, without any understanding or reasoning." (Floridi et al., p. 2) The 'stochastic core' is not a dismissal but a precise characterisation: LLMs are probability-distribution samplers over tokens. The question is what this entails about their outputs. The outputs appear to share a phenomenological similarity to human reasoning. Floridi is clear that this appearance is not accidental. > "Conversely, their outputs appear to share a phenomenological similarity to human reasoning. This effect is deliberately achieved through interface design, which encourages users to interpret outputs as explanations, commonsense reasoning, or analogies, but it also relates to the abductive patterns present in the data used to train the models. The result is a compelling illusion of genuine and structured inferential reasoning." (Floridi et al., p. 2-3) > "When their output exhibits an apparent abductive quality – often reinforced by interface design – this effect is due to the model's training on human-generated texts that encode reasoning structures." (Floridi et al., Abstract) Abductive reasoning varies in strength. Weak abduction involves hypothesis generation without strong commitment; strong abduction involves inferring the most probable explanation among alternatives. > "Abductive reasoning varies in strength. Sometimes, a distinction is made (Calzavarini & Cevolani 2022) between weak abduction—hypothesis generation without strong commitment—and strong abduction—inferring the most probable or best hypothesis. Weak abduction involves constructing a plausible story from the facts. Strong abduction entails choosing the best explanation among alternatives, which aligns more closely with IBE proper and may require comparative judgment or additional evidence." (Floridi et al., p. 4) LLMs seem to perform at least weak abduction. > "LLMs today seem to perform at least weak abduction: when presented with a scenario or riddle, they often generate a plausible explanation for it. They can even seem to carry out a form of strong abduction when all candidate hypotheses are explicitly provided, by selecting the most suitable one." (Floridi et al., p. 4) Floridi does not deny that outputs look abductive. He denies that the internal process is abductive. The distinction between appearance and mechanism is load-bearing for his argument. LLMs cannot use truth as a filter because their objective function does not reference truth. > "They \[LLMs\] generate text based on learned associations rather than performing abductive inferences... without grounding them directly in truth, semantics, verification, or understanding, and without any abductive reasoning." (Floridi et al., Abstract) > "This process lacks explicit logical rules, deliberate hypothesis testing, or reference to an external world model. It is driven solely by data correlations." (Floridi et al., p. 8) LLMs do not possess an inherent concept of truth or verification beyond what their training data provides. > "They also do not possess an inherent concept of truth or verification beyond what their training data provides. The 'stochastic parrots' metaphor highlights two limitations: (a) LLMs are limited by their training data; they can remix, rephrase, and build on the data, and can be creative, but if the data contain factual gaps or biases, so will the model; and (b) LLMs do not know whether they are right/correct or wrong/incorrect." (Floridi et al., p. 9) A consequence is that LLMs cannot lie in the ordinary sense. > "A consequence of (b) is that they cannot lie in the ordinary sense... one must have an understanding of what counts as true or false." (Floridi et al., p. 9) This is the epistemic heart of Floridi's critique. The objective function is next-token probability, not truth-tracking. Hallucinations follow directly from this architecture. Hallucinations reveal the underlying architecture. They are not bugs but features of a system optimised for plausibility rather than truth. > "An illustrative example is the phenomenon of AI 'hallucinations,' in which an LLM invents a non-existent source or confidently offers a fabricated statement or explanation. For instance, when asked about a historical figure's cause of death, the LLM might create a plausible narrative if it cannot recall the fact, because providing any answer with an authoritative tone is statistically more likely than stating 'I don't know' (especially if training data rarely include the AI saying it does not know)." (Floridi et al., p. 9) The abductive style of LLM outputs is a double-edged sword. > "This tendency shows that the abductive style of LLM outputs is a double-edged sword: the model proposes an explanation or answer because that is what fluent, human-like responders do, and because models are trained to be 'helpful', projecting certainty so as not to undermine their perceived credibility." (Floridi et al., p. 9) > "The result can be a convincing but entirely incorrect answer, essentially a confabulation." (Floridi et al., p. 9-10) Floridi coins the term 'over-abduction' to describe this tendency. > "In terms of IBE, it is as if the model always chooses an explanation, even when none is justified—it cannot 'resist' explaining because generating a plausible and preferable continuation is its task. This could be termed over-abduction: a human reasoner might say, 'I'm not sure; more information is needed', while the LLM often makes a guess regardless." (Floridi et al., p. 12) The model is structurally incapable of epistemic humility because silence is not a probable continuation. Floridi invokes Reichenbach's distinction between discovery and justification to locate what LLMs can and cannot do. > "Reichenbach (1938) and subsequent philosophers of science described inference as comprising two parts: the context of discovery, where abduction or IBE generates hypotheses; and the context of justification, where we test those hypotheses, almost always via statistical inference. This two-stage model is simple but effective: abduction provides the candidate, and induction assesses it." (Floridi et al., p. 6) LLMs perform only the first part. > "Interestingly, LLMs seem to perform only the first part. They generate candidates (explanations, answers) but do not genuinely validate them against reality (unless they are specifically augmented by other systems, which only proves the point). They aim to model the conditional distribution of tokens in text, not to evaluate truth." (Floridi et al., p. 6-7) In statistical terms, LLMs perform prior predictive sampling but lack an external feedback loop for posterior evaluation. > "In statistical terms, LLMs perform prior predictive sampling but lack an external feedback loop for posterior evaluation. We will revisit this limitation when discussing their tendency to hallucinate plausible but false information." (Floridi et al., p. 7) Prior predictive sampling means generating from learned distributions. Posterior evaluation means updating beliefs given evidence. LLMs do the former without the latter. Floridi's picture of genuine abductive reasoning includes truth-directedness and verification. > "Abduction thus contrasts with deduction (which reasons forward, in this case from cause to effect with certainty) and with induction (which generalises from many wet-lawn observations to a potentially probabilistic rule)." (Floridi et al., p. 3) > "IBE can be understood as a form of abduction that adds a comparative evaluation step: multiple candidates are generated, then weighed by criteria such as simplicity, coherence with background knowledge, scope of explanation, and so on. The 'best' explanation is then inferred as the most likely to be true." (Floridi et al., p. 3) Both abduction and IBE are defeasible. > "Both abduction and IBE are defeasible types of inference: their conclusions can be wrong, even if the reasoning appears sensible, because new evidence or information can defeat or invalidate them." (Floridi et al., p. 4) What LLMs lack is clear. > "What is clear is that LLMs lack specific abilities that human reasoners have. They do not understand the text they generate in the way humans assign meaning; they lack grounded semantics connecting words to the physical world or perceptual experiences." (Floridi et al., p. 9) Floridi's implicit standard for 'real' abduction includes truth-directedness, verification capacity, defeasibility-awareness, and grounded semantics. The symbol-grounding problem is central to Floridi's critique. > "They do not understand the text they generate in the way humans assign meaning; they lack grounded semantics connecting words to the physical world or perceptual experiences (Harnad 1990, Harnad 2024)." (Floridi et al., p. 9) > "Unlike a human expert, current LLMs do not have direct perceptual or embodied access to the world, nor do they possess conscious mental states. They can simulate expressions of positionality and uncertainty in language, but these are not grounded in lived experience; their reliability depends entirely on training, calibration, and system design, rather than on human-like understanding." (Floridi et al., p. 9) Any connection to real-world evidence must be deliberately engineered. > "Any connection to real-world evidence must be deliberately engineered, as in retrieval-augmented systems, which provide external access to information rather than embodied grounding." (Floridi et al., p. 9) The smoke/fire example illustrates the difference between human inference and LLM output. > "We see smoke and infer the presence of fire because we know that fire typically causes smoke. Large Language Models (LLMs) see the word 'smoke' and often output 'fire' because, in their training data, these words frequently co-occur as cause and effect. The difference is subtle: humans infer real fire in the world, while LLMs predict 'fire' within sentences." (Floridi et al., p. 16-17) > "However, if asked, 'There is smoke. What is a possible cause?', it will answer 'Fire' in a causal sense, not just to complete a sentence, because it has learned that causal relation as a linguistic association." (Floridi et al., p. 17) Floridi acknowledges that the outputs can be identical. The difference is in what they are about. Humans' inference is world-directed; the model's is text-distribution-directed. The LLM's extensive training on language has endowed it with a vast repository of commonsense causal knowledge. > "Essentially, the LLM's extensive training on language has endowed it with a vast repository of commonsense causal knowledge, although not explicitly structured. It 'knows' that slippery floors cause falls, that not eating causes hunger, that polls predict elections, and so on—because it has processed countless expressions of these relations." (Floridi et al., p. 17) What it lacks is experiential or embodied grounding. > "What it lacks, however, is an experiential or embodied grounding of that knowledge. It does not have sensorimotor verification, such as pushing a cup off a table and causing it to fall. But in language, it has likely encountered 'the cup fell off the table after being pushed,' which associates 'push' with 'fall.'" (Floridi et al., p. 17) Floridi grants that linguistic associations encode causal structure. The question is whether philosophy requires more than linguistic and inferential structure. The 'phenomenology of plausibility' explains why users are fooled. > "When users interact with an LLM-based AI, such as a chatbot or assistant, they often perceive the AI's responses as if they were created by an intelligent mind reasoning through the question. The AI's answer 'makes sense': it addresses the question with relevant points, sometimes even providing justification or analogies. This phenomenology of plausibility can be pretty compelling. It explains why people have attributed understanding and even sentience or consciousness to advanced chatbots." (Floridi et al., p. 10) What underpins this phenomenology is the LLM's training on human language. > "What underpins this phenomenology? In large part, it is because the LLM's training on human language enables it to mimic how humans communicate explanations and reasons." (Floridi et al., p. 10) > "From the user's perspective, it genuinely feels as if the model has reasoned to that answer." (Floridi et al., p. 10) Interface design amplifies this effect. > "This effect is deliberately achieved through interface design, which encourages users to interpret outputs as explanations, commonsense reasoning, or analogies." (Floridi et al., p. 2) Sycophancy compounds the problem. > "This problem is compounded by the recently observed behaviour of 'sycophancy'... the tendency of LLMs to generate outputs that prioritise alignment with user beliefs or preferences over factual accuracy. Because users frequently prefer convincingly-written sycophantic responses to correct responses, they are not minded to 'fact-check' if the output supports their explanation." (Floridi et al., p. 12) Interface design is a real confounder. But the argument to be developed concerns the text itself, evaluated by competent readers, not naive users. The crucial concession in Floridi's paper concerns what LLMs have learned from their training data. Human-written text often results from IBE. > "Human-written text in its training data often results from IBE. For example, many Wikipedia articles, Q&A forums, or scientific papers present evidence and then offer an explanation or conclusion. The model has absorbed these patterns. Therefore, when prompted to explain something, it generates a response that not only states a fact but often justifies it, following a structure like 'We observe X; a plausible explanation is Y, because...'. It is likely to include causal connectives ('because', 'thus', 'therefore') and explicit reasoning steps, because that is how explanations are typically structured in the training data." (Floridi et al., p. 10) Abductive structure can be learned as a linguistic template without abductive commitment to truth. > "LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations." (Floridi et al., p. 8) Floridi admits the structure is learned. The move to be made in what follows: in philosophy, those structures are not decorative. They are constitutive. Floridi acknowledges what LLMs can do. They can help generate hypotheses and support human reasoning. > "They can help generate hypotheses and support human reasoning, but their outputs must be critically examined because they cannot discern truth or verify explanations." (Floridi et al., Abstract) > "A notable application is assisting human judgment: LLMs can generate hypotheses that a person might not have considered, effectively broadening the scope of abductive search." (Floridi et al., p. 12) > "In this way, the LLM functions as an abduction generator, supporting the human reasoner during the discovery phase." (Floridi et al., p. 12) > "Experiments in human–AI collaborative reasoning (Zhou et al., 2024) indicate that LLMs can provide creative inputs or initial explanatory hypotheses, albeit often mixed with irrelevant suggestions. In essence, LLMs function like brainstorming assistants that toss out ideas without filtering for quality." (Floridi et al., p. 12) LLMs perform remarkably well on abductive tasks. > "LLMs perform remarkably well, often at a near-human level \[on the Abductive Natural Language Inference challenge\]. Such findings already suggest that LLMs, despite lacking explicit reasoning, recognise patterns that align with human explanatory preferences." (Floridi et al., p. 4-5) Floridi grants that LLMs can generate good hypotheses. The question becomes: if the hypothesis satisfies philosophical standards when evaluated, what more is required? The key passages for the redirect concern what training data encodes. > "When their output exhibits an apparent abductive quality – often reinforced by interface design – this effect is due to the model's training on human-generated texts that encode reasoning structures." (Floridi et al., Abstract) > "LLMs have effectively absorbed patterns of human abductive reasoning as expressed in writing." (Floridi et al., p. 8) > "Through exposure to billions of words, an LLM acquires a broad range of information about the world. It 'knows', in a statistical sense, many facts, relationships, and even commonsense truths, simply because these are reflected in language use. It also learns common patterns of explanation and argument, such as how 'because' often introduces an explanation, and that scientific questions are answered with specific explanatory forms." (Floridi et al., p. 8) > "Philosophy of mind and cognitive science might see this as outsourcing part of the cognitive labour involved in hypothesis generation to an artificial agent. This artificial agent achieves this through stochastic pattern matching over the corpus of human culture." (Floridi et al., p. 12) Floridi thinks 'reasoning structures' are merely formal. The argument to follow is that in philosophy, the formal structures are the discipline's method. Floridi acknowledges limits to his analysis. > "We have treated 'LLMs' somewhat generally, focusing mainly on the latest large models as of 2025, with the GPT series as a reference point. Smaller or less trained models might not exhibit the abductive illusion as strongly; their outputs can be clearly incorrect." (Floridi et al., p. 18) > "We remain agnostic about future possible systems." (Floridi et al., p. 18) > "Furthermore, our epistemological level of abstraction (stochastic versus abductive) may not capture all nuances." (Floridi et al., p. 18) The remarkable passage is Floridi's own question about whether process matters. > "In particular, if an AI can generate the same explanatory hypothesis a human would, does it matter that the process was different? From an epistemological standpoint, perhaps yes—justification is significant—but regarding the content of the hypothesis and our interpretation of it, maybe not." (Floridi et al., p. 12-13) Floridi raises the question but does not develop it. He gestures at 'justification' mattering without saying why. The argument to come is that for philosophy, justification is artefact-level, not producer-level. The answer's convincingness stems from alignment with known mechanisms and training data. > "The answer's convincingness stems from its alignment with known causal mechanisms (batteries and temperature) and from the knowledge it acquired from its training data. Because the explanation is consistent with common sense, users tend to accept it as reasonable. Essentially, the LLM manages to produce the same explanation a human reasoner would likely choose." (Floridi et al., p. 11) LLMs can falter in less common situations. > "However, in less common situations, LLMs can falter or produce a confident-sounding explanation that is subtly incorrect... Unlike a human doctor, who carefully weighs evidence (or at least can and should), the LLM 'does not know what it does not know'—it has no awareness of its own ignorance—nor does it necessarily detect subtle inconsistencies." (Floridi et al., p. 11) The danger of conflation is real. > "On the other hand, the dangers are clear: if one conflates the surface with the core—if one assumes the LLMs genuinely 'know what they are talking about'—one can be misled." (Floridi et al., p. 20) > "An explanation can be coherent and convincing (even optimal by IBE criteria) and yet still false. LLMs lack an epistemic compass to navigate that distinction." (Floridi et al., p. 20) Floridi is clear about what he is not claiming. > "We are not claiming that LLMs hold literal beliefs or follow Peirce's method of hypothesis internally. Instead, we argue that the output structure of LLMs often resembles that of an abductive reasoning process... and this resemblance is not random but systematic, resulting from training on human explanations." (Floridi et al., p. 15) The resemblance is systematic because LLMs have been trained on human explanations. This is the thread to pick up. --- # 2. Abduction and Philosophy # Abduction and Philosophy Williamson's account of philosophical method appears to intensify the problem. If philosophy is essentially abductive, and LLMs cannot do abduction in any genuine sense, then LLMs cannot do philosophy. The argument would be quick and decisive. Williamson proposes that philosophy should use a broadly abductive methodology. > "I propose that philosophy should use a broadly abductive methodology. Indeed, to some extent it already does so. I propose that it should do so in a bolder, more systematic, more self-aware way." (Williamson, p. 356) This is not a peripheral claim. Williamson finds it surprising that anyone would think philosophy could proceed without inference to the best explanation. > "My dominant reaction was, and to some degree still is, surprise at the idea that philosophy could or should get by without something like inference to the best explanation." (Williamson, p. 351) Even systematic philosophy of language requires abduction. > "I still favor inference to the best explanation and an abductive methodology in philosophy (Williamson 2013a: 423–9). Indeed, it is hard to see how the kind of positive, systematic, general theory that Dummett sought in the philosophy of language could be established by any other means." (Williamson, p. 335) David Lewis's modal realism serves as the paradigm case. Lewis postulates possible worlds because they follow from his modal realism, which he regards as the best theory in respect of simplicity, strength, elegance, and explanatory power. > "Lewis postulates them because they follow from his modal realism, which he regards as the best theory of possibility, necessity, and related phenomena, in respect of simplicity, strength, elegance, and explanatory power: to use C. S. Peirce's term broadly, Lewis's argument for modal realism is abductive." (Williamson, p. 314) > "We can take Lewis's modal realism as a case study for the resurgence of speculative metaphysics in contemporary analytic philosophy." (Williamson, p. 314) Lewis explicitly moved toward abductive justification over time. > "By the time he wrote what became the canonical case for modal realism, his book On the Plurality of Worlds (Lewis 1986b), based largely on his 1984 John Locke lectures at Oxford, Lewis's perspective had changed. He talks much less about linguistic matters, and much more about the abductive advantages of modal realism as a theoretical framework for explaining a variety of phenomena, many of them non-linguistic." (Williamson, p. 318) Theoretical virtues are the currency of philosophical evaluation. > "It should be elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated. It should be informative and general. In brief, it should combine simplicity with strength." (Williamson, p. 354) > "Abduction also rewards virtues such as simplicity, elegance, generality, and unificatory power, which all tend to make for bold theories." (Williamson, p. 366) The comparative structure of abductive evaluation is explicit. > "A theory T is a better potential explanation of evidence E than a theory T\* if and only if T would explain E if T were true better than T\* would explain E if T\* were true – in brief, T would explain E better than T\* would." (Williamson, p. 354) > "When a theory T scores highly enough as a potential explanation of our evidence E, and better than its rivals, we may infer T from E by inference to the best explanation." (Williamson, p. 354) Theoretical virtues do real work. They discriminate between empirically equivalent theories. > "When two theories make the same observable predictions, inference to the best explanation may still be able to select one over the other because the former is simpler and less ad hoc." (Williamson, p. 355) Williamson extends the Forster-Sober account of simplicity to philosophy. Simplicity is not merely an aesthetic preference; it protects against over-fitting. > "Malcolm Forster and Elliott Sober (1994) have made a strong case that at least part of the story concerns the problem of 'over-fitting' in natural science. I will suggest that their account has a significant moral for the role of simplicity in philosophy." (Williamson, p. 367) > "Forster and Sober's rationale for the criterion of simplicity can be extended to philosophy, even though quantitative data are not involved. For something very like the problem of over-fitting occurs in philosophy too." (Williamson, p. 368) The restriction helps avoid mistaking noise for signal. > "The restriction helps us avoid mistaking noise for signal, which we do if we fit the current data too closely. This account of the role of simplicity and similar aesthetic criteria in abductive methodology is consistent with a fully realist, non-pragmatist understanding of science." (Williamson, p. 368) This is realist, not pragmatist. Simplicity is truth-conducive because it prevents over-fitting. > "By giving weight to simplicity and elegance as a counterbalance to evidential fit, an abductive methodology avoids error-fragility in both the experimental sciences and philosophy." (Williamson, p. 370) Simpler theories tend to be predictively more accurate. > "Although it typically leads to equations that fit present data slightly less well, they tend to be predictively more accurate, that is, to fit future data better. The reason is that they are less vulnerable to distortion by errors in the data." (Williamson, p. 368) > "scientific experience shows that doing so leads to the problem of over-fitting, where such equations tend to be predictively inaccurate: although they fit present data well, they fit future data badly." (Williamson, p. 367–368) Abductive methodology rewards boldness. Bolder theories are riskier but stronger. > "Contrary to some stereotypes of analytic philosophy, abduction rewards boldly speculative theories. Bolder theories are riskier but stronger, in other words more informative; they entail more and so tend to have more explanatory potential, but are easier to falsify." (Williamson, p. 366) Precision is a virtue because it enables falsification. Vagueness masquerades as boldness. > "For the same reason, abduction rewards precise theories. Many wildly unclear, obscure, and vague theories have the air of setting off into the unknown, and so look bold, when really they are the opposite. Since it is quite unclear what they are supposed to entail, they avoid the risk of falsification, but by the same token they give up the hope of explaining anything." (Williamson, p. 366) > "As already emphasized, an abductive methodology will not infrequently lead us to false theories. But the clearer those theories are, other things equal, the better able we are to discover their falsity, and so learn from our mistakes. This is another advantage of precise theories over vague ones, which are much harder to falsify." (Williamson, p. 366) Vague theories rank low on the abductive scale. > "Many wildly unclear, obscure, and vague theories have the air of setting off into the unknown, and so look bold, when really they are the opposite." (Williamson, p. 366) > "Such theories rank low on the abductive scale." (Williamson, p. 366) If a theory does well by abductive criteria, that is reason to take it to be true. > "If a theory does well by abductive criteria, that is reason to take it to be coherently meaningful as well as true." (Williamson, p. 352) We rank theories as potential explanations before knowing whether they are true. > "We can rank theories (or hypotheses) as potential explanations of our evidence. The point of the qualifier 'potential' is that a false theory is not the actual explanation of the data; in that sense, it does not really explain them. But we need to rank theories as potential explanations before knowing whether they are true, in order then to use the ranking to guide our judgments as to which theory is true." (Williamson, p. 353–354) At a bare minimum, a theory must be consistent with the evidence. > "At a bare minimum, T must be consistent with E. In brief, the closer T comes to entailing E, the better (ceteris paribus)." (Williamson, p. 354) A philosophical theory must cohere with the total evidence base. > "From what evidence base should we start when applying abduction to the construction and selection of philosophical theories? As always, the answer is in principle: our total evidence. That is arguably no less than the total sum of human knowledge. It includes whatever knowledge the natural and social sciences, philosophy, and common sense have already gained. None of our knowledge is irrelevant in principle to philosophy, for any philosophical theory inconsistent with any of it is false (since what is known is true)." (Williamson, p. 356–357) Deduction still plays a major role within abductivist inquiry. > "Deduction still plays a major role within abductivist inquiry, since deducing consequences from a theory (usually with some auxiliary hypotheses) is integral to the explanatory enterprise. More generally, abduction values theories of great deductive strength (at least, when they are consistent with the evidence)." (Williamson, p. 365) The problem with purely deductive methodology is that it leads to stalemate. > "All too often, if the argument is deductively valid, opponents simply reject one of those informative universal premises as 'question-begging.' One can try deducing the rejected premise from further informative universal premises, but that way an infinite regress looms." (Williamson, p. 364) Abduction bypasses deductive deadlocks. > "An abductive methodology bypasses deductive deadlocks, by encouraging both the accumulation of more evidence of various kinds and the development of better explanations of that evidence (which may simply bring it under illuminating generalizations)." (Williamson, p. 365) Robustness is a methodological desideratum. The method should tolerate error in the input. > "Since we cannot keep our premises completely free of error, we need robust methods of theory choice that do not crash every time an error enters." (Williamson, p. 369–370) > "Joshua Alexander and Jonathan Weinberg (2014) have argued that the method of thought experiment is error-fragile, in the sense that it tends to multiply the effect on the output theories of any errors in the supposed input evidence." (Williamson, p. 369) The standards of analytic philosophy are publicly articulable and applicable by contemporary standards. > "Thus, simply using the methods of analytic philosophy critically, by contemporary standards, takes some sophistication in both semantics and pragmatics, irrespective of the subject matter under philosophical investigation. That is a robust legacy from analytic philosophy of language for all philosophy." (Williamson, p. 346) All of this seems to vindicate Floridi's worry. If philosophy proceeds by abduction, and LLMs cannot do abduction, then LLMs cannot do philosophy. But this argument moves too fast. It assumes that the question 'can LLMs do abduction?' is the right question to ask. It is not. The pivot is this: look at what Williamson says abductive competence consists in. Weighing theoretical virtues. Preferring simpler theories. Avoiding over-fitting. Seeking integration with other commitments. These are all features of the theory. They are publicly articulable. You can check whether a paper exhibits them by reading the paper. Theoretical virtues are intrinsic to the theory. > "Apart from its relation to E, the more T has the intrinsic virtues of a good theory, the better (ceteris paribus). It should be elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated." (Williamson, p. 354) The word 'intrinsic' is key. Theoretical virtues are features of the theory itself, not of the theorist. They are properties of the artefact, not of the producer's inner states. Theories rank on the abductive scale. > "Such theories rank low on the abductive scale." (Williamson, p. 366) The ranking is of theories, not of theorists. The criteria are publicly checkable even if their ultimate justification is unclear. > "The foregoing sketch leaves it far from clear what makes abduction such a good method. For instance, why should aesthetic criteria such as elegance contribute to the pursuit of truth? Nevertheless, the central role of abduction in the success of the natural sciences provides good reason to think that it is a good method, even though we do not fully understand why." (Williamson, p. 356) We can use the method without fully understanding why it works. The criteria are practically applicable. The criteria have determinate content. They license some moves and not others. > "Of course, abductive criteria of simplicity and elegance do not license one simply and elegantly to ignore recalcitrant data. Rather, they encourage a more critical attitude to the data." (Williamson, p. 368) The standards are publicly applicable to the philosophical community's output. Williamson diagnoses the community, not individual minds. > "Strikingly, the philosophical community showed very little aversion to the multiplication of complication. A firmer preference for simplicity and elegance would have warned the community that something was going wrong. Indications of over-fitting remain quite widespread in analytic philosophy." (Williamson, p. 369) This is a diagnosis of papers, not minds. Williamson's anti-over-fitting point gets a new application. A paper that commits to the simplest view satisfying its explanatory target, and explicitly rejects complexity-adding repairs unless they bring compensating gain, is exhibiting the robustness strategy Williamson defends. This can be assessed from the text. The question is whether the text exhibits simplicity, not whether the producer experienced simplicity-preferring cognitive states. The real question is not 'can LLMs do abduction internally?' That is a question about mechanism we may never answer. The real question is 'can LLM outputs instantiate the constraint structure that distinguishes good philosophy from persuasive dialectic?' That question is answerable. And the answer, as the following sections will show, is yes. This relocates the debate. Any anti-LLM argument now has to point to specific text-internal failures — equivocations, illicit premises, ad hoc repairs, question-begging moves. Gesturing at production mechanism is no longer sufficient. Saying 'but it's just statistics' is not a philosophical objection. Show me the flaw in the paper, or accept that the paper is good. Taking Williamson seriously means looking at what the standards are and where they apply. The standards are features of theories. The evaluation is of artefacts. The production mechanism drops out. --- # 3. Learning the Game # Learning the Game LLMs trained on philosophical corpora have internalised the constraint structure. Not as explicit rules they can articulate, but as practice-patterns — the way a native speaker learns grammar without learning rules, the way a chess player learns positional intuitions without learning explicit algorithms. The corpus of philosophical work encodes the rules. LLMs have ingested that corpus. They can play the game. Argumentation schemes are structures of inference that represent common types of arguments. > "Argumentation schemes are forms of argument (structures of inference) that represent structures of common types of arguments used in everyday discourse, as well as in special contexts like those of legal argumentation and scientific argumentation." (Walton et al., p. 1) They include deductive and inductive forms but also a third category — defeasible, presumptive, or abductive. > "They include the deductive and inductive forms of argument that we are already so familiar with in logic. However, they also represent forms of argument that are neither deductive nor inductive, but that fall into a third category, sometimes called defeasible, presumptive, or abductive." (Walton et al., p. 1) This third category is where philosophical argumentation lives. Arguments in this space may not be very strong by themselves, but may be strong enough to warrant rational acceptance of their conclusion. > "Such an argument may not be very strong by itself, but may be strong enough to provide evidence to warrant rational acceptance of its conclusion, given that its premises are acceptable." (Walton et al., p. 1) This captures the epistemic standard for philosophy: not proof, but warranted acceptance. The schemes encode what counts as warrant. Argumentation schemes are codifiable. They have been codified. > "The most useful and widely used tool so far developed in argumentation theory is the set of argumentation schemes." (Walton et al., p. 1) > "This book provides a systematic analysis of many common argumentation schemes and a compendium of ninety-six schemes." (Walton et al., book description, p. i) Ninety-six schemes with explicit structure. The rules exist. Each scheme comes with built-in critical questions. > "Each argument of this type is presented as providing only a defeasible support for its conclusion, subject to critical questioning in a context of dialogue. Matching each argumentation scheme is an appropriate set of critical questions." (Walton et al., p. 3) The method of evaluation is explicit. > "The method of evaluation of an argument fitting a scheme is that once the argument is put forward by a proponent, it may be defeated if the respondent asks an appropriate critical question that is not answered by the proponent." (Walton et al., p. 3) The dialectical structure is: move, critical question, response. This is the game, and its rules are stated. Consider the appeal to expert opinion. It has six critical questions. > "1. Expertise Question: How credible (knowledgeable) is E as an expert source? > 2. Field Question: Is E an expert in the field that A is in? > 3. Opinion Question: What did E assert that implies A? > 4. Trustworthiness Question: Is E personally reliable as a source — for example, is E biased? > 5. Consistency Question: Is A consistent with what other experts assert? > 6. Backup Evidence Question: Is E's assertion based on evidence?" (Walton et al., p. 32-33) These critical questions are the rules for when an argument from authority succeeds or fails. A model that has seen enough philosophical argumentation has encountered the pattern: authority-appeal, challenge on these dimensions, response. The structure is in the text. Critical questions specify the conditions of defeat. > "The critical questions matching a scheme can be seen as representing additional relevant factors that might cause an argument to default." (Walton et al., p. 38) When a proponent presents a cogent presumptive argument, the respondent must respond appropriately. > "When a proponent presents a cogent presumptive argument that fits one of the schemes to the respondent, he needs to respond to it in some appropriate way. The obvious way is for him to either accept the argument or ask one of the set of appropriate critical questions." (Walton et al., p. 36) This specifies the dialogue rules. A scheme-fitting argument creates a commitment unless a critical question is raised. The rules are explicit. The study of schemes aspires to taxonomic systematisation. > "The goal of the study of schemes is to develop an integrated theory of argumentation schemes that will have at its core an ontology of schemes organized in a tree structure, with top nodes representing the most general types of schemes, and lower nodes representing increasingly more specific types." (Walton et al., p. 347-348) These are not ad hoc patterns. They have hierarchical structure that can be learned. Schemes serve as tools for enthymeme completion. The problem of attribution is the problem of interpreting a claim based on a text of discourse. > "The problem of attribution is one of interpreting a claim supposedly made, based on a quotation, or given text of discourse, that records what the arguer actually said or wrote." (Walton et al., p. 213) This is the problem an LLM must solve when generating philosophical text. What premises are in play? What must be supplied? The machine should give preference to missing premise candidates that are true, represent common knowledge, or are at least plausible in context. > "In addition to generating an argument that is structurally correct by some standard of inference, the machine should give preference to missing premise (or conclusion) candidates that are true, that represent common knowledge, or that are at least plausible, in context." (Walton et al., p. 216) Walton is describing an enthymeme machine — a system that fills in implicit premises. This is what LLMs do when they complete a philosophical argument. Scheme identification enables premise completion. > "Araucaria is equipped with a set of argumentation schemes. When a user constructs an argument diagram, she can identify the scheme that fits a given set of premises and conclusion that she has identified as an argument in a given text. Araucaria can then fit the scheme to the specified parts of the argument, and identify the missing premises required by that scheme." (Walton et al., p. 215) Scheme identification, premise completion. This is the inference pattern LLMs have learned. Implicit premises are based on common knowledge. > "According to Govier (1992, p. 120), an implicit premise in an argument is based on common knowledge if it states something known by virtually everyone, depending on audience, context, time, and place." (Walton et al., p. 208) Philosophical corpora encode common knowledge for philosophy — the shared premises and background assumptions of the discipline. This tacit knowledge is notoriously difficult to capture explicitly. > "In the literature on planning in AI (Carberry, 1990), these assumptions would be classified as domain-dependent knowledge, and they are notoriously difficult to capture in a principled way. But they are not based on specialized expert knowledge. They represent common knowledge about the way things can normally be expected to work in a typical situation known to both sides in a conversation." (Walton et al., p. 209-210) Difficult to capture explicitly, but it can be absorbed from exposure to enough discourse. This is the LLM's advantage. Schemes are text-recognisable. A theory of argumentation schemes should be rich enough to cover a large proportion of naturally occurring argument. > "Thus a theory of argumentation schemes should be... rich and sufficiently exhaustive to cover a large proportion of naturally occurring argument." (Walton et al., p. 39) The patterns exist in text and can be extracted from it. Schemes must be formalised so that a coder can recognise a particular argument as fitting a scheme. > "In order to be useful in logic, artificial intelligence, and related scientific fields, schemes must be formalized, meaning that they have to be codified in some precise way so that the coder (whether machine or human) can recognize a particular argument as fitting a scheme and then use it to derive conclusions from the given set of premises based on that identification." (Walton et al., p. 364) Schemes are recognisable from text. Recognition enables inference. The pattern-matching is the skill. Once an argument is recognised as fitting a scheme, it can be reconstructed. > "Once an argument is recognized as fitting a scheme, an argument markup, utilizing an argument diagram, can reconstruct the argument in a given case using the scheme as a template or pattern on which to frame the reconstruction." (Walton et al., p. 364) Scheme-fitting enables reconstruction. The same process by which an LLM generates well-formed arguments. Scheme competence is teachable from the ground up. > "This volume surveys all aspects of argumentation schemes from the ground up, taking the reader from the elementary exposition of the first chapter to the latest state of the art in the research efforts to formalize and classify the schemes as outlined in the last three chapters." (Walton et al., p. 6) It can be learned through systematic exposure. Even the procedural question of when to stop pressing is part of the scheme structure. > "Presumptive schemes are defeasible. They are not deductively valid. The question, then, is how long the process of asking critical questions can continue before the argument must finally be accepted as binding the respondent to acceptance of the conclusion, if he has accepted the premises." (Walton et al., p. 30-31) This is dialectical judgement, and it too is learnable from seeing how arguments terminate. Defeasible reasoning is the core of philosophical argumentation. A defeasible argument is one in which the conclusion can be accepted tentatively but may need to be retracted as new evidence comes in. > "A defeasible argument is one in which the conclusion can be accepted tentatively in relation to the evidence known so far in a case, but may need to be retracted as new evidence comes in." (Walton et al., p. 2) This is how philosophical positions work. Held tentatively, revisable under new arguments. Presumptive arguments are necessary but dangerous. > "To use a phrase from Anderson, Schum, and Twining (2005, p. 262), such presumptive arguments are necessary but dangerous. We need to use them as heuristics that provide rational grounds for accepting a conclusion tentatively even if it has not been conclusively proved, but we have to remain open-minded when we use such arguments, because they are fallible and inherently subject to default." (Walton et al., p. 2) Necessary but dangerous captures the epistemic situation of philosophical argument. LLMs that have learned these patterns have learned the core mode of philosophical inference. The recognition of the importance of defeasible argumentation has led to a paradigm shift. > "The recognition of the importance and legitimacy of defeasible argumentation has led to a recent paradigm shift in logic, artificial intelligence, and cognitive science. Common forms of defeasible arguments were long categorized as fallacious in logic textbooks. It is been only recently that, as these informal fallacies have been studied more intensively, more and more instances have been recognized where the forms of argument underlying them are reasonable, but inherently defeasible." (Walton et al., p. 2) Philosophy has always operated in this space. The schemes codify it. Schemes create dialectical obligations. When an argument fitting one of the schemes is put forward, it shifts the burden of proof. > "One way to approach the problem of showing how presumptive argumentation schemes can be binding on a respondent is to utilize the notion of burden of proof. According to this approach, when an argument fitting one of the schemes is put forward in a given case, it shifts the burden of proof to the side of the respondent in a dialogue." (Walton et al., p. 35) Making a move puts pressure on the interlocutor. This is the game-like structure. Walton envisions computational systems doing what LLMs do. > "According to this approach, the system can use argumentation schemes and critical questions to identify nonexplicit premises in an incomplete argument identified by a user in a given text of discourse." (Walton et al., p. 216) Using learned scheme structure to complete arguments. The study of schemes has a long history going back to Aristotle's topics. > "As we show in Chapter 8, the study of schemes has a long history going back to Aristotle's topics — common types of argument, often called commonplaces, that Aristotle saw as fundamental building blocks in a branch of logic he called dialectic." (Walton et al., p. 5) Philosophical method has always been reflexive. The study of argument is itself argument. The rules are discussable within the practice. What makes an argument strong is structural correctness according to accepted standards. > "What makes an argument strong? The reply suggested earlier is that the critic needs to put in the assumptions required to make an argument come out valid, or at least structurally correct according to whatever the accepted standards are." (Walton et al., p. 217) If the standards are satisfied, the argument is strong. There is no hidden further fact. Cogency is defined by scheme-fitting plus requirement-satisfaction. > "A presumptive argumentation scheme imposes a relation of conveyance on the respondent such that if he accepts the premises, and if the scheme is applicable, and if all the requirements of the scheme are met, the conclusion is conveyed to him by these factors. That doesn't mean he has to accept it, or cannot now present reasons for doubting it; but it does mean that he has now been given a cogent reason for accepting it, and he has to deal with that." (Walton et al., p. 36) This is the public, checkable standard that makes philosophical evaluation possible. Philosophy is peculiar in a way that strengthens the case. Unlike biology, which is grounded in cells, or physics, which is grounded in particles, philosophy is grounded in the space of reasons itself. The objects of philosophical study are logical and inferential relations, not external entities. In philosophy, the map is the land. Floridi's complaint that LLMs lack world-contact loses much of its force when the world in question is the system of reasons the model has internalised. Philosophy's verification is largely internal: validity, consistency, dialectical robustness. The symbol-grounding problem that plagues LLMs in empirical domains is significantly weakened here. For competent philosophical readers, 'looks like good philosophy' in the strong sense just means 'the standards are satisfied in the text'. Experts are not tracking genre markers — signposting, objections-and-replies sections, theoretical-virtue talk. They are checking whether conclusions are earned, costs paid, equivocations avoided. When the standards are satisfied, appearance is reality. The paper is good. When they are not satisfied, the output does not even register as philosophy. It gets filtered into 'not philosophy' before it can count as the feared middle category. This matches reported experience. Working with LLMs, one either gets genuinely good philosophy or obvious sludge, rarely 'plausible but bad'. The feared 'abductive appearance without substance' category is rare or empty because the appearance competent readers track is the substance. After this, Floridi-style 'abductive appearance' dismissals become question-begging. One cannot dismiss an LLM output as only appearing to do philosophy if the text satisfies the very standards by which good philosophy is recognised. Saying 'it only appears valid' when the inference has been checked and is valid is a verbal trick, not a philosophical objection. The burden shifts to the sceptic. Name the specific failure, or accept that the artefact is good. --- # 4. How to Generate Philosophy with AI # How to Generate Philosophy with AI \*Burden\*: Show the thesis in action with worked examples. This section makes it vivid. You need at least one case where: - The prompt is minimal (genre-cueing, not micromanaged) - The output exhibits genuine philosophical structure: hinge identification, cost-accounting, alternative-theory comparison, sensitivity to objections - You can evaluate it against the standards and show it passes The reader should be able to \*see\* what you mean by constraint-satisfaction, not just take your word for it. You might also include a stress-test case — something that exposes where failure IS identifiable text-internally. The pseudo-robustness example (the semantics-reduces-to-physics prompt) could serve: you show that \*when\* standards are violated (equivocation, bait-and-switch), the violations are identifiable from the text. This supports your claim that evaluation is artefact-level: you don't need to know it was an LLM to see the flaw. This section comes last (before conclusion) because it's evidence, not argument. You want the reader to have the framework before seeing the examples. --- # 5. Conclusion # Conclusion \*Burden\*: Restate the thesis, sum up the argument, gesture at implications. \*\*Restate\*\*: LLMs can produce novel, first-rate philosophy with minimal prompting. The question isn't "do they really reason?" but "do their outputs satisfy the constraint structure of good philosophy?" The answer is yes — often enough to matter. \*\*The argument in brief\*\*: Floridi's "abductive appearance" critique and Williamson's centrality-of-abduction picture seem to block LLM philosophy. But philosophical evaluation is artefact-level: we assess texts, not producers. The relevant standards — precision, cost-accounting, non-ad hocness, defeater-sensitivity, fair treatment of rivals — are text-internal and publicly codifiable. LLMs trained on philosophical corpora have internalised these standards as practice-patterns. Minimal prompting cues the relevant genre; the latent dialectical structure does the rest. Philosophy's peculiar self-grounding nature (the map IS the land) and the collapse of appearance/reality for competent readers further support the thesis. \*\*Implications\*\*: \*For philosophical methodology\*: What does it mean that the standards are learnable from text? Perhaps philosophy's "rules of the game" are more public and codifiable than we assumed. Perhaps philosophical competence is more like fluency than genius. \*For understanding what philosophy is\*: The thesis suggests that philosophy is a practice governed by publicly accessible norms — not ineffable insight, not special cognitive access, but skill with reasons as they appear in text. \*For the discipline's future\*: A new kind of collaborator (or competitor) has arrived. How should philosophers respond? The answer isn't panic or dismissal; it's recognition that the standards remain the same. Good philosophy is good philosophy. Evaluate the work. ### Analysis, justification, reasons You’re asking for three deliverables that fit together: 1. the *exact* “no oracle / no answer key” formulation you liked, so you can reuse it in the paper without accidental drift; 2. a block quote from Bengson/Cuneo/Shafer-Landau that supports the “implicit norms in texts” idea (specifically: satisfying the method’s criteria needn’t be intentional); 3. a very detailed plan for a substantial chunk of Section 3 (“Learning the Game”), reorganising what we’ve uncovered in our conversation into something that can actually sit inside your draft as real argumentative machinery. Two constraints I’m enforcing while doing this: - I’m not going to turn tentative ideas into “finished theories” by giving them extra names. I’ll use your wording where possible (e.g., “obvious move”), and I’ll describe structures rather than baptise them. - For the requested book quotation: I can give you a short verbatim excerpt, but I can’t paste long stretches of the book’s text. So I’ll give one tight quote that contains the key claim, and then I’ll paraphrase the immediately surrounding sentences so you still get the full evidential support. For the plan: you asked “NO SUBSECTIONS”, so I’ll structure it as a sequence of paragraphs/moves that would make up a large continuous part of Section 3, without proposing subheadings like “3.1 / 3.2”. --- ## Final answer ## 1) The “no oracle” point - what I said Here is the core formulation I used (verbatim from what I wrote earlier in this chat): > “The right claim is not ‘no correct answer’ but ‘no labelled oracle’. In many philosophical problems, we don’t have an external answer key in the way we do in arithmetic or many benchmark tasks. That doesn’t mean there is no truth. It means we often can’t cheaply verify that we’ve reached it.” And here is the continuation I paired with it (also verbatim), which you may or may not want to keep attached: > “Therefore, the operational proxy for ‘getting it right’ is often something like: argumentative validity, clarity of commitments, resilience to objections, explanatory integration, non-ad-hocness, etc.” If you want this to be maximally aligned with your broader thesis, the safest use is: keep the “no labelled oracle” bit to characterise *feedback structure*, and then immediately reaffirm that philosophy is truth-directed (your earlier complaint was about any slide into relativism). --- ## 2) The methodology book passage - block quote This is the sentence that contains the “needn’t be intentional” claim, from Bengson, Cuneo, and Shafer-Landau, *Philosophical Methodology: From Data to Theory*, Chapter 6 (“Philosophical Progress”), section 2.5, p. 162. > “To be clear, our claim is not that the construction of views that satisfy the method’s criteria has always been intentional.” Immediately after that sentence (still on p. 162), they continue by offering two routes to non-intentional satisfaction: philosophers may have aimed to satisfy criteria from other methods, or they may simply have been engaging in standard philosophical activities (arguing, objecting, replying, explaining); either way, they end up satisfying the Tri-Level Method’s criteria. That’s exactly the leverage you wanted: the Tri-Level structure is not just a philosopher’s after-the-fact reconstruction; the authors explicitly treat it as something that can be realised by ordinary philosophical practice without being self-consciously followed. --- ## 3) A detailed plan for a substantial part of Section 3 ### Analysis of what Section 3 needs to do (in your draft’s dialectic) Section 3 is where your paper stops being primarily defensive (Floridi/Williamson as the strongest sceptical frame) and starts doing the positive work: explaining why minimal prompting can yield genuinely high-level philosophical output. Your draft already has a strong “shape” for Section 3: the claim that philosophical corpora *encode the rules of the game*, and that LLMs trained on them can “play” - plus a move to argumentation schemes (Walton et al.) as explicit codifications. The conversation added something that can make this much sturdier: - Bengson/Cuneo/Shafer-Landau provide a *method-level* story about what competent philosophical theorising is like (criteria, ordering, and a conception of objections as targeting criteria), and - crucially - they explicitly say philosophers can satisfy these criteria without aiming to. That supports your claim that the norms are *in the texts*, not merely in philosophers’ private intentions. - You have a nice way to connect that to minimal prompting: the “obvious move” cue isn’t a checklist; it’s a practice-indexical directive that tells the model to continue the dialectical sequence in the way the corpus has taught it competent philosophers proceed. So the plan below builds Section 3 around that pairing: (i) a method-level account of implicit norms in philosophical writing, (ii) a mechanism for why a minimal cue can reliably activate those norms, (iii) an explicit bridge to Walton’s schemes as a complementary codification at a different level of granularity. ### The plan (paragraph-by-paragraph, no subsections) 1. Start by restating the specific target of Section 3 in one paragraph: the issue is not whether the model has inner “reasoning states” but whether it has absorbed and can enact the publicly checkable norms by which philosophical texts are produced and assessed. Tie this to your introduction’s “minimal prompting” definition: you want genre-governing cues that trigger competence, not micromanagement. (This paragraph is doing rhetorical positioning: the reader needs to know what counts as success for this section.) 2. Immediately introduce the methodological book as a way of making precise what “competent philosophical progression” amounts to, without claiming it is the only account. Use their own framing: method is “the engine of inquiry”, and methods are sets of criteria that guide both theory construction and evaluation. The point here is not to teach the whole Tri-Level Method; it is to give the reader a concrete handle on what “norms of philosophical writing” look like when made explicit. 3. In the next paragraph, give the reader the minimal content of their model of inquiry: philosophical inquiry has (at least) data + method-of-theorising as major components, with method functioning as criteria for construction/evaluation, and the aim being theoretical understanding. You can cite the book’s own statement of its guiding questions and structure. This sets up the idea that philosophical writing has an implicit “workflow” because inquiry has an organised structure. 4. Now bring in the key bridge to your project: quote and then explain the “killer” point you asked about - that satisfying the method’s criteria needn’t be intentional, because philosophers may simply be engaging in ordinary practice activities and thereby end up satisfying the criteria. After stating that point, make explicit what you want it to licence: if philosophers can satisfy these criteria without self-conscious adherence, then philosophical texts will tend to *instantiate* the criteria as patterns of exposition and dialectical response (what gets done next, what counts as an objection, what counts as a repair). That’s the link from “method” to “corpus structure”. 5. Next paragraph: articulate the “implicit in the texts” thesis carefully, without overclaiming. Something like: philosophical corpora don’t just contain conclusions; they contain recurring patterns of how philosophers move from a dialectical state to its demanded next step (defence of commitments, explanation of data, integration with background constraints, and so on). The methodology book helps you specify what sorts of “demands” commonly drive that progression (accommodation/explanation, substantiation/integration, virtues only later). This is where you cash out your “grammar” idea without naming it: the criteria create typical “next things to do”, and those “next things” show up in texts. 6. Now slot in the “no labelled oracle” point in a way that doesn’t antagonise your reader: philosophy is truth-directed, but often lacks cheap external answer keys; therefore, the practice relies heavily on public, text-assessable constraints (validity, explanatory fit, integration, non-ad-hocness, etc.) as the way to track truth under conditions of limited direct verification. This is where the phrase you liked earns its keep. The job of this paragraph is to explain why text-internal norms matter so much: they’re not substitutes for truth, they’re the discipline’s way of pursuing it. 7. Next paragraph: connect the previous two. If the practice relies on public criteria and the texts instantiate them as recurring patterns, then a model trained on that text is positioned to learn the patterns - not necessarily as explicit rules, but as reliable expectations about what comes next in philosophical writing. This is your “learn the game” claim, now backed by the methodology book’s own insistence that the criteria are familiar from practice and can be satisfied without intending to. 8. Now introduce your “obvious move” as the minimal prompt that exploits exactly that: it does not feed premises or walk the model through a proof; it functions like a deictic instruction (“from here, do what’s demanded”). Your explanation here should explicitly include the correction you insisted on earlier: “obvious move” is not a single kind of move (not just distinctions/decomposition); depending on what is currently missing, the next demanded step could be unification, synthesis, reconstruction, handling an objection, strengthening an explanation, integrating with background constraints, etc. The methodology book helps here too: because it characterises objections as targeting deficits with respect to the criteria, it implicitly characterises what kinds of repairs count as “the next thing to do.” 9. At this point you can bring in a second methodological reinforcement from Bengson et al. that also fits your draft: their emphasis on shared frameworks that underwrite disagreements, including distinctions, inventories of similarities/differences, necessary-condition claims, records of dead ends, rosters of open possibilities, and so on. This supports your earlier claim (in our conversation) that “real-world evidence” in analytic philosophy is often backgrounded and unremarked: the background is precisely these shared frameworks, and they’re heavily textual. It also supports your stronger “saturation” thought: the model has been trained not just on controversial theses but on the shared background scaffold that makes serious philosophical disagreement possible. 10. Now transition to Walton et al. as a complementary codification at a different scale. The way to frame the transition so it doesn’t feel like a non sequitur is: Bengson et al. give you method-level criteria for constructing and appraising theories; Walton gives you mid-level or micro-level regularities of argument forms and critical questions that structure dialogue. (You already drafted this, so the work is to integrate, not replace.) This paragraph’s job is to justify why both belong in Section 3: one makes explicit the “theory appraisal” norms; the other makes explicit the “move/response” norms. 11. Next paragraph: use Walton’s concept of argumentation schemes and critical questions to reinforce your “learnable from text” thesis: schemes are explicitly described as common inferential structures used in ordinary and specialised contexts, and they come with matched critical questions; this makes the dialectical structure “move → challenge → response” explicit. Now connect back to LLM training: the corpus contains countless instances of scheme-like patterns being enacted, so a token predictor trained on philosophical corpora plausibly learns a distribution over such patterns. 12. Then make one paragraph doing the “two-level” synthesis: Bengson et al. give you criteria that govern what counts as an adequate theory (accommodation, explanation, substantiation, integration; virtues as tie-breakers); Walton gives you tools for modelling how arguments are advanced and defeated in dialogue. Taken together, they give you a principled way to say what it is for a model to have internalised philosophical competence as expressed in texts: it can (i) generate theory-shaped contributions that satisfy the criteria and (ii) navigate scheme-like challenge/response dynamics. Keep the tone careful: you are not claiming this proves the model “reasons”; you are claiming it explains why minimal cues can trigger robust philosophical behaviour. 13. Finally, close this part of Section 3 by turning the screw on Floridi’s “abductive appearance” line in a controlled way. The key move is: if the norms that constitute competent philosophical practice are public, textually manifested, and assessable by competent readers, then an output that satisfies them is not “mere appearance” in the relevant evaluative sense - it is an instance of the artefact meeting the discipline’s standards. You can say explicitly: this doesn’t settle metaphysical questions about mental states; it relocates the argumentative burden to the level that matters for peer review - show the specific textual failure (equivocation, ad hoc patch, unmet explanatory burden, integration failure), or accept that the artefact meets the standards. That is a substantial chunk of Section 3 because it gives the reader a clear chain: methodological criteria → enacted in texts even without intending to → therefore learnable from texts → minimal prompts can point the model at continuing those patterns → scheme-level codification reinforces the “learnable game” story → this reframes what “abductive appearance” can reasonably mean for philosophical evaluation. ### Two placement variants (in case you want different emphasis) Variant A: Put Bengson et al. before Walton (as above). This makes the section feel like: first the macro norms of inquiry, then the micro patterns of argumentation. Variant B: Start with Walton (as you currently do), then bring Bengson et al. as the “higher-level glue” that shows how local argument patterns are embedded in inquiry that aims at understanding and is governed by criteria. This variant makes the section feel more “from dialogue mechanics up” rather than “from methodology down”. Both variants preserve your core ambition: Section 3 isn’t “LLMs are good at writing”; it’s “philosophical competence is publicly codified in texts in ways that a language model can internalise, and minimal cues can activate that competence in contribution mode rather than overview mode.” "Now slot in the “no labelled oracle” point in a way that doesn’t antagonise your reader: philosophy is truth-directed, but often lacks cheap external answer keys; therefore, the practice relies heavily on public, text-assessable constraints (validity, explanatory fit, integration, non-ad-hocness, etc.) as the way to track truth under conditions of limited direct verification. This is where the phrase you liked earns its keep." a small point: maybe you could compare it to code. code does have a clear yes it worked no it doesnt. this also allows us to point out the huge improvmenets in codeing that llms do now. could you write out the whole plan again verbatim apart from this addition. i want to give it to another llm 1. Start by restating the specific target of Section 3 in one paragraph: the issue is not whether the model has inner “reasoning states” but whether it has absorbed and can enact the publicly checkable norms by which philosophical texts are produced and assessed. Tie this to your introduction’s “minimal prompting” definition: you want genre-governing cues that trigger competence, not micromanagement. (This paragraph is doing rhetorical positioning: the reader needs to know what counts as success for this section.) 2. Immediately introduce the methodological book as a way of making precise what “competent philosophical progression” amounts to, without claiming it is the only account. Use their own framing: method is “the engine of inquiry”, and methods are sets of criteria that guide both theory construction and evaluation. The point here is not to teach the whole Tri-Level Method; it is to give the reader a concrete handle on what “norms of philosophical writing” look like when made explicit. 3. In the next paragraph, give the reader the minimal content of their model of inquiry: philosophical inquiry has (at least) data + method-of-theorising as major components, with method functioning as criteria for construction/evaluation, and the aim being theoretical understanding. You can cite the book’s own statement of its guiding questions and structure. This sets up the idea that philosophical writing has an implicit “workflow” because inquiry has an organised structure. 4. Now bring in the key bridge to your project: quote and then explain the “killer” point - that satisfying the method’s criteria needn’t be intentional, because philosophers may simply be engaging in ordinary practice activities and thereby end up satisfying the criteria. After stating that point, make explicit what you want it to licence: if philosophers can satisfy these criteria without self-conscious adherence, then philosophical texts will tend to instantiate the criteria as patterns of exposition and dialectical response (what gets done next, what counts as an objection, what counts as a repair). That’s the link from “method” to “corpus structure”. 5. Next paragraph: articulate the “implicit in the texts” thesis carefully, without overclaiming. Something like: philosophical corpora don’t just contain conclusions; they contain recurring patterns of how philosophers move from a dialectical state to its demanded next step (defence of commitments, explanation of data, integration with background constraints, and so on). The methodology book helps you specify what sorts of “demands” commonly drive that progression (accommodation/explanation, substantiation/integration, virtues only later). This is where you cash out your “grammar” idea without naming it: the criteria create typical “next things to do”, and those “next things” show up in texts. 6. Now slot in the “no labelled oracle” point in a way that doesn’t antagonise your reader: philosophy is truth-directed, but often lacks cheap external answer keys; therefore, the practice relies heavily on public, text-assessable constraints (validity, explanatory fit, integration, non-ad-hocness, etc.) as the way to track truth under conditions of limited direct verification. Add a compact comparison to code: in programming, you often do have a relatively crisp “oracle” (the code compiles, runs, and/or passes tests), which makes both evaluation and iterative improvement straightforward. Use that contrast to highlight why LLM progress has been so visible in coding: success signals are clearer, feedback loops are tighter, and outputs are easy to score. Then pivot back: philosophy lacks that kind of immediate runtime verdict, so the discipline’s public constraints and dialectical procedures play an especially central role in tracking truth. 7. Next paragraph: connect the previous two. If the practice relies on public criteria and the texts instantiate them as recurring patterns, then a model trained on that text is positioned to learn the patterns - not necessarily as explicit rules, but as reliable expectations about what comes next in philosophical writing. This is your “learn the game” claim, now backed by the methodology book’s own insistence that the criteria are familiar from practice and can be satisfied without intending to. 8. Now introduce your “obvious move” as the minimal prompt that exploits exactly that: it does not feed premises or walk the model through a proof; it functions like a deictic instruction (“from here, do what’s demanded”). Your explanation here should explicitly include the correction you insisted on earlier: “obvious move” is not a single kind of move (not just distinctions/decomposition); depending on what is currently missing, the next demanded step could be unification, synthesis, reconstruction, handling an objection, strengthening an explanation, integrating with background constraints, etc. The methodology book helps here too: because it characterises objections as targeting deficits with respect to the criteria, it implicitly characterises what kinds of repairs count as “the next thing to do.” 9. At this point you can bring in a second methodological reinforcement from Bengson et al. that also fits your draft: their emphasis on shared frameworks that underwrite disagreements, including distinctions, inventories of similarities/differences, necessary-condition claims, records of dead ends, rosters of open possibilities, and so on. This supports your earlier claim that “real-world evidence” in analytic philosophy is often backgrounded and unremarked: the background is precisely these shared frameworks, and they’re heavily textual. It also supports your stronger “saturation” thought: the model has been trained not just on controversial theses but on the shared background scaffold that makes serious philosophical disagreement possible. 10. Now transition to Walton et al. as a complementary codification at a different scale. The way to frame the transition so it doesn’t feel like a non sequitur is: Bengson et al. give you method-level criteria for constructing and appraising theories; Walton gives you mid-level or micro-level regularities of argument forms and critical questions that structure dialogue. (You already drafted this, so the work is to integrate, not replace.) This paragraph’s job is to justify why both belong in Section 3: one makes explicit the “theory appraisal” norms; the other makes explicit the “move/response” norms. 11. Next paragraph: use Walton’s concept of argumentation schemes and critical questions to reinforce your “learnable from text” thesis: schemes are explicitly described as common inferential structures used in ordinary and specialised contexts, and they come with matched critical questions; this makes the dialectical structure “move → challenge → response” explicit. Now connect back to LLM training: the corpus contains countless instances of scheme-like patterns being enacted, so a token predictor trained on philosophical corpora plausibly learns a distribution over such patterns. 12. Then make one paragraph doing the “two-level” synthesis: Bengson et al. give you criteria that govern what counts as an adequate theory (accommodation, explanation, substantiation, integration; virtues as tie-breakers); Walton gives you tools for modelling how arguments are advanced and defeated in dialogue. Taken together, they give you a principled way to say what it is for a model to have internalised philosophical competence as expressed in texts: it can (i) generate theory-shaped contributions that satisfy the criteria and (ii) navigate scheme-like challenge/response dynamics. Keep the tone careful: you are not claiming this proves the model “reasons”; you are claiming it explains why minimal cues can trigger robust philosophical behaviour. 13. Finally, close this part of Section 3 by turning the screw on Floridi’s “abductive appearance” line in a controlled way. The key move is: if the norms that constitute competent philosophical practice are public, textually manifested, and assessable by competent readers, then an output that satisfies them is not “mere appearance” in the relevant evaluative sense - it is an instance of the artefact meeting the discipline’s standards. You can say explicitly: this doesn’t settle metaphysical questions about mental states; it relocates the argumentative burden to the level that matters for peer review - show the specific textual failure (equivocation, ad hoc patch, unmet explanatory burden, integration failure), or accept that the artefact meets the standards. Variant A: Put Bengson et al. before Walton (as above). This makes the section feel like: first the macro norms of inquiry, then the micro patterns of argumentation. Variant B: Start with Walton (as you currently do), then bring Bengson et al. as the “higher-level glue” that shows how local argument patterns are embedded in inquiry that aims at understanding and is governed by criteria. This variant makes the section feel more “from dialogue mechanics up” rather than “from methodology down”. 1. Start by restating the specific target of Section 3 in one paragraph: the issue is not whether the model has inner “reasoning states” but whether it has absorbed and can enact the publicly checkable norms by which philosophical texts are produced and assessed. Tie this to your introduction’s “minimal prompting” definition: you want genre-governing cues that trigger competence, not micromanagement. (This paragraph is doing rhetorical positioning: the reader needs to know what counts as success for this section.) 2. Immediately introduce the methodological book as a way of making precise what “competent philosophical progression” amounts to, without claiming it is the only account. Use their own framing: method is “the engine of inquiry”, and methods are sets of criteria that guide both theory construction and evaluation. The point here is not to teach the whole Tri-Level Method; it is to give the reader a concrete handle on what “norms of philosophical writing” look like when made explicit. 3. In the next paragraph, give the reader the minimal content of their model of inquiry: philosophical inquiry has (at least) data + method-of-theorising as major components, with method functioning as criteria for construction/evaluation, and the aim being theoretical understanding. You can cite the book’s own statement of its guiding questions and structure. This sets up the idea that philosophical writing has an implicit “workflow” because inquiry has an organised structure. 4. Now bring in the key bridge to your project: quote and then explain the “killer” point - that satisfying the method’s criteria needn’t be intentional, because philosophers may simply be engaging in ordinary practice activities and thereby end up satisfying the criteria. After stating that point, make explicit what you want it to licence: if philosophers can satisfy these criteria without self-conscious adherence, then philosophical texts will tend to instantiate the criteria as patterns of exposition and dialectical response (what gets done next, what counts as an objection, what counts as a repair). That’s the link from “method” to “corpus structure”. 5. Next paragraph: articulate the “implicit in the texts” thesis carefully, without overclaiming. Something like: philosophical corpora don’t just contain conclusions; they contain recurring patterns of how philosophers move from a dialectical state to its demanded next step (defence of commitments, explanation of data, integration with background constraints, and so on). The methodology book helps you specify what sorts of “demands” commonly drive that progression (accommodation/explanation, substantiation/integration, virtues only later). This is where you cash out your “grammar” idea without naming it: the criteria create typical “next things to do”, and those “next things” show up in texts. 6. Now slot in the “no labelled oracle” point in a way that doesn’t antagonise your reader: philosophy is truth-directed, but often lacks cheap external answer keys; therefore, the practice relies heavily on public, text-assessable constraints (validity, explanatory fit, integration, non-ad-hocness, etc.) as the way to track truth under conditions of limited direct verification. Add a compact comparison to code: in programming, you often do have a relatively crisp “oracle” (the code compiles, runs, and/or passes tests), which makes both evaluation and iterative improvement straightforward. Use that contrast to highlight why LLM progress has been so visible in coding: success signals are clearer, feedback loops are tighter, and outputs are easy to score. Then pivot back: philosophy lacks that kind of immediate runtime verdict, so the discipline’s public constraints and dialectical procedures play an especially central role in tracking truth. 7. Next paragraph: connect the previous two. If the practice relies on public criteria and the texts instantiate them as recurring patterns, then a model trained on that text is positioned to learn the patterns - not necessarily as explicit rules, but as reliable expectations about what comes next in philosophical writing. This is your “learn the game” claim, now backed by the methodology book’s own insistence that the criteria are familiar from practice and can be satisfied without intending to. 8. Now introduce your “obvious move” as the minimal prompt that exploits exactly that: it does not feed premises or walk the model through a proof; it functions like a deictic instruction (“from here, do what’s demanded”). Your explanation here should explicitly include the correction you insisted on earlier: “obvious move” is not a single kind of move (not just distinctions/decomposition); depending on what is currently missing, the next demanded step could be unification, synthesis, reconstruction, handling an objection, strengthening an explanation, integrating with background constraints, etc. The methodology book helps here too: because it characterises objections as targeting deficits with respect to the criteria, it implicitly characterises what kinds of repairs count as “the next thing to do.” 9. At this point you can bring in a second methodological reinforcement from Bengson et al. that also fits your draft: their emphasis on shared frameworks that underwrite disagreements, including distinctions, inventories of similarities/differences, necessary-condition claims, records of dead ends, rosters of open possibilities, and so on. This supports your earlier claim that “real-world evidence” in analytic philosophy is often backgrounded and unremarked: the background is precisely these shared frameworks, and they’re heavily textual. It also supports your stronger “saturation” thought: the model has been trained not just on controversial theses but on the shared background scaffold that makes serious philosophical disagreement possible. 10. Now transition to Walton et al. as a complementary codification at a different scale. The way to frame the transition so it doesn’t feel like a non sequitur is: Bengson et al. give you method-level criteria for constructing and appraising theories; Walton gives you mid-level or micro-level regularities of argument forms and critical questions that structure dialogue. (You already drafted this, so the work is to integrate, not replace.) This paragraph’s job is to justify why both belong in Section 3: one makes explicit the “theory appraisal” norms; the other makes explicit the “move/response” norms. 11. Next paragraph: use Walton’s concept of argumentation schemes and critical questions to reinforce your “learnable from text” thesis: schemes are explicitly described as common inferential structures used in ordinary and specialised contexts, and they come with matched critical questions; this makes the dialectical structure “move → challenge → response” explicit. Now connect back to LLM training: the corpus contains countless instances of scheme-like patterns being enacted, so a token predictor trained on philosophical corpora plausibly learns a distribution over such patterns. 12. Then make one paragraph doing the “two-level” synthesis: Bengson et al. give you criteria that govern what counts as an adequate theory (accommodation, explanation, substantiation, integration; virtues as tie-breakers); Walton gives you tools for modelling how arguments are advanced and defeated in dialogue. Taken together, they give you a principled way to say what it is for a model to have internalised philosophical competence as expressed in texts: it can (i) generate theory-shaped contributions that satisfy the criteria and (ii) navigate scheme-like challenge/response dynamics. Keep the tone careful: you are not claiming this proves the model “reasons”; you are claiming it explains why minimal cues can trigger robust philosophical behaviour. 13. Finally, close this part of Section 3 by turning the screw on Floridi’s “abductive appearance” line in a controlled way. The key move is: if the norms that constitute competent philosophical practice are public, textually manifested, and assessable by competent readers, then an output that satisfies them is not “mere appearance” in the relevant evaluative sense - it is an instance of the artefact meeting the discipline’s standards. You can say explicitly: this doesn’t settle metaphysical questions about mental states; it relocates the argumentative burden to the level that matters for peer review - show the specific textual failure (equivocation, ad hoc patch, unmet explanatory burden, integration failure), or accept that the artefact meets the standards.