# opus[1m]
## Skills Used
- epistemic-discipline
- source-work
## Retrieval Notes
- Session id: `0606563d-5f0c-4a1f-9d9b-6484ce38e15b`
- Last activity: `2026-03-18T09:36:24.047Z`
## My Notes
<!-- Add your notes here. This section is preserved across syncs. -->
## Conversation
### User
opus[1m]
### User
/draft-audit
"Generating Philosophy - Text-Internal Evaluation" I want to work with you on this project. Basically I've been looking at section 01 one and two and you will see from my inline comments I am very very very unhappy with these things. I don't want you to make any edits yet. I just want to talk to you about this stuff in the chat and see if we can sort it out. This is as you can see quite a hard task. I've been working on this for ages with you and we just can't get these sections correct and there's a multitude of problems as you can see from all my comments. So yeah. You need to really knuckle down for this one, please.make sure that you explain your analysis, justification, reasons etc etc. for your answer BEFORE giving me your final answer. Please give me a long, deep, and detailed answer, with lots of options/ideas: as we are in the 'kicking ideas around' stage, I want as much feedback as you can give me so that i have a lot to consider.
### Assistant
Let me load up the project context and find those manuscript sections.
### Assistant
Now let me read the two sections you're unhappy with, plus the session file for current context.
### Assistant
Now let me read the Pigliucci paper Nick specifically flagged in his comments, and the paragraph-level overview for additional context.
### Assistant
I've now read both sections with your comments, the introduction, the session file, and the Pigliucci paper in full. Let me work through this properly.
Right. Let me start by just cataloguing what's happening in these sections, because the problems are layered — some are sentence-level (voice), some are paragraph-level (clarity), and some are section-level (structure). I need to separate these or they'll blur together.
## Voice problems — the "not how i write" markers
There's a pattern to the voice failures, and it's worth naming it precisely. The recurring issue is short, flat, declarative metacommentary sentences that announce what the text is about to do or has just done, rather than doing it. Examples:
- "is not like this" (Section 1, ¶1)
- "The arguments are the contribution." (Section 1, ¶1)
- "Philosophical evaluation has this character." (Section 1, ¶5)
- "The phrase marks an absence." (Section 2, ¶1)
- "This is more than a passing acknowledgement." (Section 2, ¶4)
- "An analogy may clarify the relationship." (Section 2, ¶5)
- "The same applies to philosophy." (Section 1, ¶8)
These are all the same kind of sentence: they're signposts that an LLM writes to tell the reader "here's where we are in the argument." They're empty — they don't carry philosophical content. They read like a student essay. Nick's analytic voice doesn't do this; his prose carries its own momentum through the progression of ideas, not through metacommentary labels.
Hmm... but wait, the voice problem goes deeper than just signpost sentences. There's also the triplicate example issue, which Nick flags with increasing fury throughout. Let me count:
- "identifies a hypothesis (the battery), provides a reason (cold weather reduces battery efficiency), and presents the explanation with the connectives and qualifications" — Section 2, ¶1
- "state the objection, make a concessive move, identify a flaw" — Section 2, ¶2
- "weigh evidence, or select among hypotheses" — Section 2, ¶4
- "elegance, coherence, illumination of subject matter" — Section 1, closing
- "handled objections, engaged with existing work, and contributed something non-trivial" — Section 1, ¶10
- "move and countermove that constitute the discipline's evaluative practice" — Section 1, ¶9
- "argument from analogy, from consequences, from expert opinion" — Section 1, ¶9
This is a deep LLM pattern — the compulsive need to give three instances of everything. It's not that three examples is always wrong; sometimes you need three. But here they're functioning as filler, as a way of gesturing at a category rather than explaining it. The triplets make the writing feel incantatory rather than argumentative. Each triplet is a place where the text should instead be doing real work — explaining what it means, giving a single well-chosen example with enough context, or just making the point directly.
So the voice problem has two layers: (1) empty signpost sentences, (2) triplicate filler examples. Both are LLM habits rather than writing habits.
## The listicle problem — Section 1's structural failure
Now, the bigger structural issue. Nick's comment on the second half of Section 1 is devastating and correct: "the second half of this section has just turned into fucking listicle paragraphs where you give me a disconnected fact in each paragraph."
Let me trace how this happens. The section starts well enough — Watson/Crick vs Quine establishes that philosophy lives in its arguments, not in prior facts reported. Then Lipton's likeliest/loveliest distinction is introduced. So far there's a thread: what makes a contribution in philosophy, and how do we evaluate it?
But then the thread breaks. After the Lipton paragraph (which Nick rightly says is "a load of shit jammed together"), we get:
- Semmelweis (¶4) — dropped in without setup
- "Philosophical evaluation has this character" (¶5) — a one-paragraph non-explanation
- Williamson on overfitting (¶6) — disconnected, purpose unclear
- Deep Blue (¶7) — reader has no idea why
- Dellsén on progress (¶8) — another disconnected theorist
- "The same applies to philosophy" (¶9) — another thin transition
- Bengson et al. on criteria (¶10) — yet another theorist introduced and dropped
- Walton on argumentation schemes (same ¶10) — and another
- The philosophical corpus paragraph (¶11) — tries to synthesise but reader is lost
- Self-evidencing explanations (¶12) — another theorist dropped in
The problem is clear: each paragraph introduces a source and summarizes what they say, but doesn't connect the source to a developing argument. The reader is being told "here's what Lipton says, here's what Williamson says, here's what Gaut says, here's what Dellsén says, here's what Bengson says, here's what Walton says" without understanding what question each source is answering and why it matters at that point in the argument.
What should this section actually be doing? Let me think...
The section needs to establish: philosophical evaluation concerns properties of texts — properties assessable by reading the arguments. The question is what those properties are, and why their being text-internal matters for the LLM question.
So the thread should probably be something like:
1. Philosophy lives in its arguments (Watson/Crick vs Quine — this works)
2. But what makes arguments good or bad? (This is the question that needs to be explicitly posed)
3. Various people have answers to this — but instead of listing them, the text needs to ARGUE for a particular picture of philosophical quality
4. That picture needs to be: quality is assessable from the text itself, and the corpus is filtered for it
The problem is that the section currently tries to establish (3) by name-dropping sources rather than by building the case. Each source gets its own paragraph rather than being woven into a single developing argument.
Hmm, but I should think about this differently. Maybe the issue isn't "too many sources" but "each source gets its own paragraph." What if the sources were integrated — if Lipton, Williamson, and Bengson were all contributing to the same paragraph or two, each adding one piece to a single picture? That would avoid the listicle feel.
Actually, wait. Let me reconsider the whole section structure. There might be a more radical problem: the section is trying to do too many things at once. It's simultaneously trying to:
(a) Establish that philosophy is in the text, not in the discoverer
(b) Introduce the likeliest/loveliest distinction from Lipton
(c) Introduce Williamson's theoretical virtues
(d) Introduce the idea that the corpus is filtered for quality
(e) Introduce argumentation schemes
(f) Introduce self-evidencing explanation
(g) Make the process-independence point (Gaut, blind review)
(h) Make the progress point (Dellsén)
That's too many things for one section. Some of these belong together naturally — (a) and (g) are about the same thing (what matters is the text, not who made it). (b) and (c) are about what makes philosophy good. (d) and (e) are about how those qualities are visible in the corpus.
But (f) — self-evidencing explanation — is actually a separate and quite powerful point that might work better if it got more room to breathe, or was moved elsewhere. And (h) — Dellsén on progress — might actually belong in the introduction rather than here.
Let me now think about the Lipton paragraph specifically, because Nick calls it "a load of shit jammed together" and I think diagnosing what's wrong there reveals something about the whole section.
The paragraph currently does this:
1. Introduces likeliest vs loveliest (by restating the block quote — Nick's point about "you're just restating exactly what was in the block quotes")
2. Says they come apart
3. Gives tiny example of lovely without probable
4. Connects to philosophy lacking empirical tests
5. Lists Williamson's virtues
6. Claims these are loveliness-properties
7. Concludes philosophical evaluation is IBE for loveliness
That's too much for one paragraph and none of it gets adequate treatment. The likeliest/loveliest distinction is genuinely important for the paper — it's doing real work later when talking about LLM outputs. But it's introduced too fast, explained too thinly, and connected to Williamson before the reader has properly understood it.
Nick's point about "you're just restating exactly what was in the block quote" is exactly right. After a block quote, the text should BUILD ON what the quote says — explain its significance, draw out an implication, show why it matters for the present argument. It shouldn't paraphrase the same content back.
Then immediately after — Semmelweis. The idea is to give a concrete example of the likeliest/loveliest distinction. Actually, this is a good idea in principle — Semmelweis is Lipton's own example and it makes the distinction vivid. But it arrives without setup. The reader has no reason to expect a historical medical example. The transition needs to explicitly say: "here's what this looks like in practice" or something that connects the abstract distinction to the concrete case.
OK, so there are things to preserve (Watson/Crick opening, the Lipton distinction as a framework, some of the synthesis about what philosophical quality consists in) and things that need to go or be radically reworked (the listicle structure, the disconnected paragraphs, the signpost sentences, the triplicate examples).
## Section 2's structural failure
Nick stopped reading Section 2 at line 18: "I stopped reading here because it's.. the structure here is just a mess."
Let me trace what happens. The section opens with Floridi's zeroth-order abduction argument — this exposition is actually pretty good, though it has some voice issues. The block quote from Floridi is well chosen. The Lipton generation/selection distinction is brought in effectively.
Then comes the "worry this raises for philosophy" paragraph — again, OK content but with the triplicate example problem and some "not how i write" markers.
Then: "We grant this characterisation at the level of mechanism." This is the argumentative turn — the paper concedes Floridi's mechanism point but disputes what follows. This is the right move. But then: "Floridi et al. themselves provide the materials for a response." Nick calls this "a very abrupt and disorientating change switch turnaround." And he's right — the text goes from "Floridi is right about mechanism" to "but Floridi's own words undermine their conclusion" in the space of a sentence, without properly staging the redirect.
The problem here is that the turn needs more room to breathe. The concession needs to land. And then the redirect needs to be explicitly framed: "the question isn't whether Floridi is right about the mechanism, but what the mechanism operates over."
After the turn, we get the convergence claim paragraph: "Statistical plausibility, within this corpus, converges with philosophical quality." Nick says the content is good but the explanation is "a fucking disaster" — problems at both paragraph level and section level with structure and ordering of information.
I think the issue is that the convergence claim is the single most important thesis of the paper and it's being stated rather than built up to. The reader needs to be walked through the reasoning: (1) statistical plausibility depends on training data, (2) the training data is a filtered philosophical corpus, (3) what was filtered for is philosophical quality, (4) therefore plausibility in this distribution approximates philosophical quality. Each step needs its own moment. Instead, the whole thing is compressed into one dense paragraph that a reader unfamiliar with the argument would find impenetrable.
Then the grammar analogy — "An analogy may clarify the relationship" — ugh. But the analogy itself is quite good: children acquire grammatical competence through exposure to grammatical speech without learning grammar rules. The disanalogy (children become genuine speakers, LLMs don't) is honestly noted. This is potentially useful material, but it needs to not be introduced with a signpost sentence.
After line 18 (where Nick stopped) — the borrowed calibration argument, Model A/B, Floridi's hedge, levels of description, latent vs expressed, empirical questions. I don't have Nick's comments on these (since he stopped reading), but based on the patterns, I'd expect similar problems: dense content, poor paragraphing, disconnected transitions.
## The Pigliucci tension
Now — Nick's comment on "there is nothing prior that it reports" in Section 1. This is a genuinely important philosophical point, not just a writing issue.
The current text says: "Not whether a text accurately reports something prior to it, since there is nothing prior that it reports."
Nick flags this: "is this really the right way to put things? Especially given the work from the Italian guy that we draw on about evoking conceptual landscapes."
Having now read Pigliucci's chapter, I can see the tension clearly. Pigliucci argues that philosophy is "empirically informed evoking" — its starting points are empirical data about the world. He writes:
> "the basic parameters that philosophers use as their inputs, the starting points of their philosophizing, their equivalent of axioms in mathematics and assumptions in logic (or rules in chess) are *empirical* data about the world."
And:
> "philosophy attempts to clarify things, or to analyze in order to bring about understanding, not really to discover new facts, but rather to evoke rational conclusions arising from certain ways of looking at a given problem or set of facts."
So for Pigliucci, philosophy IS connected to the world — it takes empirical starting points and explores the conceptual landscapes they evoke. The "evoked truths" have rigid properties once evoked. Philosophy isn't pure invention, but it's also not discovery (of pre-existing facts). It's a third thing.
The current formulation — "there is nothing prior that it reports" — is too strong. It suggests philosophy is entirely untethered from prior reality. But on Pigliucci's account, there IS something prior (the empirical constraints, the "axioms" drawn from experience), and the philosophical contribution consists in the evocation of conceptual landscapes from those starting points. The contribution is not a report of the prior reality, but it is constrained by it.
This matters for the paper because Zahavy's objection turns precisely on this point — that LLMs lack contact with the world, and therefore can't do the "empirically informed" part of "empirically informed evoking." If Section 1 says "there is nothing prior," it hands Zahavy the win too easily. The paper needs a more nuanced formulation: philosophy's contributions are not reports of prior reality, but they are constrained by it — and the constraints are themselves propositionally encoded in the texts that transmit them.
Actually, I think this tension goes even deeper. The Watson/Crick vs Quine opening is built on a dichotomy: science discovers pre-existing facts, philosophy doesn't. But Pigliucci's framework dissolves this dichotomy with the "evoked" category — things that don't have prior existence but DO have rigid properties once evoked. Philosophy doesn't discover facts AND doesn't invent arbitrarily — it evokes constrained conceptual landscapes. The current opening might need to acknowledge this rather than relying on a simple discovery/non-discovery binary.
Hmm, but I need to be careful here. The opening's purpose is to establish that philosophical evaluation concerns properties of the text itself. You don't need the strong claim "there is nothing prior" to make this point. You could instead say: "Whatever empirical realities philosophy engages with, the philosophical contribution lies in the arguments made about them — in the distinctions drawn, the objections anticipated, the coherence of the position constructed. And these are properties of the text." That preserves the text-internal evaluation thesis without overcommitting on the metaphysics of what philosophy is about.
## What about the "et al." issue?
Section 2 opens with "Floridi, Nobre, and Taddeo (2024)" — Nick wants "Floridi et al. (2024)" for papers with more than two authors. This is a simple formatting rule. The rest of Section 2 correctly uses "Floridi et al." — it's just the first occurrence that gets it wrong.
Similarly in the introduction: "Floridi, Nobre, and Taddeo (2024)" should be "Floridi et al. (2024)."
## Attribution problems
Several comments flag unclear attribution:
1. In Section 1, ¶8 (Dellsén): "Suppose a scientist publishes an important finding and then dies..." — Nick asks whether this example is from Dellsén and says it's academically dishonest not to make that clear. Looking at this, I think the text IS presenting a Dellsén example but doesn't explicitly say so. It should say "Dellsén and colleagues illustrate this with a thought experiment: suppose a scientist..."
2. In Section 1, ¶10 (Bengson): "A paper that scores well on these criteria handles objections in recognisable ways" — Nick asks "is Bengson saying this? Or is this supposed to be me?" The sentence follows the Bengson criteria list and could be read either way. It needs to be clear: either "On this account, a paper that..." (attributing to Bengson) or "We take this to mean that a paper..." (marking it as interpretation).
3. In the Introduction, ¶5: "On Nietzsche's account (as Sorgner reads it), philosophers are creators of values, expressing drives and psychophysiology that LLMs lack." Nick is furious: "neither Nietzsche, nor Sorgner write about LLMs. PLEASE STOP BEING INNACCURATE ABOUT WHAT WRITERS ARE WRITING ABOUT." The sentence implies Sorgner was addressing LLMs when his reading of Nietzsche says nothing about LLMs. The text should say something like: "On Nietzsche's account (as Sorgner reads it), philosophers are creators of values, expressing drives and psychophysiology. An LLM, on such a view, would lack the requisite conditions entirely." That keeps the inference about LLMs as the paper's own, not Sorgner's.
## Structural options for Section 1
OK, so what do we actually DO about Section 1? Let me think about different approaches.
Option A: Keep the current structure but fix the problems within it.
- Rewrite voice issues sentence by sentence
- Add proper transitions between theorist paragraphs
- Remove triplicate examples
- Integrate sources rather than listing them
This is the least disruptive approach but I worry it won't solve the listicle problem. You can't fix a listicle by adding transitions between list items — the items need to be woven into a single argument.
Option B: Reorganize around questions rather than sources.
Instead of: "Lipton says X, Williamson says Y, Bengson says Z, Dellsén says W"
Do: "What makes philosophy good? [Answer drawing on multiple sources] How is this visible in texts? [Answer drawing on multiple sources] Why does it matter who produces the text? [Answer drawing on multiple sources]"
This would solve the listicle problem by organizing around the argument's questions rather than around individual theorists. Each question would draw on whichever sources are relevant, so a single paragraph might cite Lipton, Williamson, and Bengson together rather than giving each their own paragraph.
Option C: Thin the sources dramatically.
The section might just be trying to do too much. What if it used Lipton as the primary framework (likeliest/loveliest) and Williamson as the primary articulation of what philosophical virtues are, and relegated Bengson, Walton, Dellsén, and Gaut to footnotes? The main argument would be: (1) philosophy lives in arguments, not discoveries; (2) the relevant evaluative standard is loveliness not likeliness (Lipton); (3) loveliness consists in elegance, unity, non-ad-hocness (Williamson); (4) these properties are text-internal; (5) therefore what matters is whether a text exhibits them, not who produced it. That's a clean five-step argument that can be made in maybe seven paragraphs rather than fourteen.
Option D: Split the section.
Maybe this is actually two sections: one that establishes "philosophy lives in its arguments" (Watson/Crick, Quine, the Pigliucci evocation point, blind review) and one that establishes "what philosophical quality consists in" (Lipton, Williamson, Bengson, Walton). The current section tries to do both and neither gets adequate treatment.
I'm actually drawn to a combination of B and C. The section needs to be organized around questions (not sources) and it probably needs fewer sources in the main text (with others in footnotes). The footnote structure already has some of this — [^criteria] already collects the different theorists' vocabulary.
Option E: The "one good example" approach.
Instead of the Semmelweis example being one of many items in a list, make it the section's anchor. Introduce the likeliest/loveliest distinction properly, illustrate it with Semmelweis (properly set up), then immediately apply it to philosophy — showing that when we evaluate a philosophical argument, we're assessing loveliness (whether it would illuminate if correct), not likeliness (whether we can empirically confirm it). Then connect this to the "text-internal" thesis: if what we evaluate is loveliness, and loveliness is assessable by reading the argument, then evaluation is text-internal. This gives the section a single spine rather than a series of disconnected points.
Actually, I think there's something promising in combining E with C. The section would: (1) Watson/Crick vs Quine opening (what philosophy IS), (2) Lipton's distinction as the framework (what makes philosophy GOOD), (3) Semmelweis as the illustration (what this LOOKS LIKE), (4) application to philosophy — philosophical evaluation is loveliness evaluation — drawing on Williamson's specific articulation of the virtues, (5) the text-internal thesis: loveliness is assessable from the text, so the producer doesn't matter. Bengson, Walton, Dellsén, Gaut, self-evidencing explanation — these all go to footnotes or get absorbed into the main argument where they naturally fit (Gaut supports point 5, Dellsén provides background on progress, etc.).
## Structural options for Section 2
Section 2's problems are somewhat different. The content is better but the ordering and staging are wrong. The argumentative turn — from "Floridi is right about mechanism" to "but what the mechanism operates over matters" — is too abrupt. And the convergence thesis is stated rather than built up to.
Let me think about what the section actually needs to do:
1. Present Floridi's argument faithfully (zeroth-order abduction, no real reasoning)
2. Acknowledge the force of the argument at the mechanism level
3. Redirect: the question isn't the mechanism but the data
4. Build the convergence thesis: filtered corpus → statistical plausibility ≈ philosophical quality
5. Handle the "but it's just statistics" objection (levels of description)
6. Handle the "but it hasn't earned its standards" objection (borrowed calibration)
7. Clarify what the model has learned (Model A/B)
The current text does all of these but in the wrong order and with bad staging. The most important move — the convergence thesis — gets one compressed paragraph instead of being built up step by step.
One option: slow down the turn. After granting the mechanism point, explicitly pose the question that drives the rest: "But what does it mean for a stochastic process to be plausible, and plausible relative to what?" Then spend two or three paragraphs building: (a) statistical plausibility is relative to training data, (b) the training data for philosophy is a filtered corpus, (c) the filtering was for philosophical quality, (d) therefore plausibility in this distribution approximates quality. Each step gets its own moment.
Another option: lead with the convergence thesis. Instead of presenting Floridi first and then responding, what if the section stated the thesis early — "we argue that statistical plausibility in a corpus filtered for philosophical quality converges with that quality" — and then showed how this answers Floridi? This would give the reader the destination before the journey, which can be clearer for a complex argument.
Actually, wait. The session file says the synthesis plan (March 16) puts "the virtue-filtered corpus thesis in Section 2 (where it responds to Floridi AND delivers the positive thesis early)." So this is already the idea — Section 2 should both present Floridi and deliver the convergence thesis. But the current execution doesn't stage this properly.
I think the issue is partly that the convergence thesis needs a proper setup paragraph. Something like: "The response turns on a feature of the probability distributions that Floridi describes. Statistical plausibility is not a fixed property of text — it is relative to training data. And the training data, in the case of philosophy, is not a random sample of text."
That would make the turn legible. The reader sees: oh, you're going to argue that the training data matters. Then the next paragraphs fill in: what the training data is (a corpus filtered for quality), what the filtering selected for (the very properties we've been discussing), and what this means (plausibility ≈ quality).
## The grammar analogy
The grammar analogy (children learn grammatical competence from exposure without learning rules) is actually quite strong. The problem is just its introduction ("An analogy may clarify the relationship" — terrible) and possibly its placement. It might work better as the culmination of the convergence argument rather than appearing mid-section. After building the case that statistical plausibility in a quality-filtered corpus approximates quality, the analogy seals it: just as children produce grammatical speech by absorbing patterns left by grammatical norms, LLMs produce philosophical text by absorbing patterns left by philosophical norms. The disanalogy (children understand, LLMs don't) is then honestly faced.
## The self-evidencing point
Currently in Section 1, introduced as yet another listicle item. But this is actually a distinctive and interesting point: philosophical arguments are self-evidencing in Lipton's sense — the argument both explains (why the conclusion holds) and provides the evidence for the explanation's adequacy (by being a good argument). This connects to the text-internal thesis directly: if philosophy is self-evidencing, then text-internal evaluation isn't a compromise but a consequence of the kind of object a philosophical contribution is.
This might work better as the capstone of Section 1's argument rather than appearing mid-list. After establishing that philosophical evaluation concerns text-internal properties (loveliness), the self-evidencing point explains WHY this is so: because in philosophy, the text is both the explanation and the evidence.
## Where does Pigliucci fit?
The session file puts Pigliucci in Section 3 (responding to Zahavy on embodied experience). But Nick's comment in Section 1 suggests Pigliucci should influence how the opening formulation works. I think there are two distinct uses:
1. In Section 1: Pigliucci's point that philosophy "attempts to clarify things, or to analyze in order to bring about understanding, not really to discover new facts, but rather to evoke rational conclusions" supports the text-internal thesis — philosophy evokes rather than discovers, and evocation produces texts. The "nothing prior that it reports" formulation could be refined using Pigliucci's vocabulary: philosophy doesn't report prior facts but evokes constrained conclusions, and the evocation is the text.
2. In Section 3: Pigliucci's point that philosophical "axioms" are empirical provides resources for answering Zahavy — the empirical inputs are propositionally encoded and therefore textually available.
These are different uses and both are legitimate. The Section 1 use would refine the opening; the Section 3 use would answer Zahavy. They're compatible.
## A note on the Introduction's problems
Nick flagged several issues in the introduction too:
- "Even so, there are arguments that LLMs cannot produce work of this kind" — Nick: "no, there are not; again, no writer has written about this." The issue is that Floridi and Zahavy didn't write about philosophy specifically, so saying there are "arguments that LLMs cannot produce [philosophical] work" misrepresents them.
- The Nietzsche/Sorgner attribution error
- "text-focused conception" — Nick doesn't like naming it, LLM-ish
- Ordering: practitioner-focused paragraph might go before text-focused
- The roadmap paragraph is acknowledged as outdated
OK, I think I have enough to give Nick a thorough response. Let me organize this into a coherent analysis.
Actually, let me also think about one more thing: what are the SOURCE-CHECK issues here? The draft-audit skill is supposed to run both voice-fix and source-check. Let me think about what sources might be mischaracterized...
Source-check issues I can identify:
1. Nietzsche/Sorgner — already flagged by Nick. Sorgner didn't write about LLMs.
2. Floridi/Zahavy — the introduction says they argue LLMs "cannot produce work of this kind" (philosophical work). But neither paper addresses philosophy. Floridi is about reasoning generally; Zahavy is about theoretical innovation. The paper should say the arguments "bear on" philosophy if sound, not that they directly address it. (The later ¶ in the intro does say "Neither paper concerns philosophy directly" but the opening sentence of that ¶ contradicts this.)
3. Bengson attribution ambiguity — flagged by Nick.
4. Dellsén example — flagged by Nick (unclear whether it's Dellsén's thought experiment or the paper's).
I don't have the actual Floridi, Zahavy, or Dellsén extractions open to fully verify all quotations, but the block quotes look right based on what I know of these texts. The page numbers are given. I should flag that I haven't verified every quotation against the extraction.
Let me also think about whether the Lipton quotations are accurate. The block quote on p. 59 about likeliest/loveliest — I'd want to check this against the extraction. The block quote on p. 24 about self-evidencing — same. The p. 108 squash analogy — same. I can't fully verify these without reading the Lipton extraction, but the page numbers are specific and the prose style matches Lipton.
Right, I think I have a full picture now. Let me write up the analysis.
## The problems
There are three layers of trouble here, and they compound each other: sentence-level voice failures, paragraph-level clarity failures, and section-level structural failures. Let me work through each.
---
### 1. Voice — two distinct LLM habits
The %%not how i write%% markers cluster around two recurring patterns.
The first is empty signpost sentences — short declarative announcements that tell the reader where the argument is going rather than carrying philosophical content:
- "is not like this" / "The arguments are the contribution." / "Philosophical evaluation has this character." / "The phrase marks an absence." / "This is more than a passing acknowledgement." / "An analogy may clarify the relationship." / "The same applies to philosophy."
These are all the same kind of sentence: metacommentary that a student essay uses to orient itself. Your prose doesn't need these because the ideas carry their own momentum. When you write "Quine's *Two Dogmas* is not like this," the reader gets nothing — you haven't yet said what it IS. A sentence that actually does work would already be inside the contrast: "Quine made a case against the analytic/synthetic distinction — there's no arrangement of facts his paper reports, no prior reality that a different philosopher might have found instead."
The second pattern is the triplicate example — the compulsive LLM habit of listing exactly three instances at the end of every generalisation. You flag this with escalating fury throughout both sections, and rightly. Examples: "state the objection, make a concessive move, identify a flaw" / "handled objections, engaged with existing work, and contributed something non-trivial" / "elegance, coherence, illumination of subject matter" / "argument from analogy, from consequences, from expert opinion" / "move and countermove that constitute the discipline's evaluative practice."
These triplets function as hand-waving. Each is a place where the text should be doing genuine work — explaining what something means, giving one well-developed example, or making the point without recourse to illustration. Instead, three items are rattled off and the text moves on, having gestured at a category without explaining it. The result sounds incantatory rather than argumentative.
Both habits are fixable once identified. But fixing them sentence by sentence won't solve the deeper problems.
---
### 2. The listicle problem — Section 1 after the Lipton paragraph
Your diagnosis is exactly right: "the second half of this section has just turned into fucking listicle paragraphs where you give me a disconnected fact in each paragraph."
Let me trace how the section degrades. It opens with Watson/Crick vs Quine — good, establishes the thesis. Then Lipton's likeliest/loveliest distinction. Then... a sequence of paragraphs each introducing a different theorist:
- Semmelweis (dropped in without setup — "no reader is going to understand why this example is being talked about")
- Williamson on overfitting (purpose unclear — "seems entirely disconnected with what comes before and after it")
- Deep Blue (Gaut's point — "a reader is gonna be entirely lost as to why you're talking about this")
- Dellsén on progress (another author, another paragraph)
- Blind review ("The same applies to philosophy" — flat transition)
- Bengson + Walton (two more sources crammed into one paragraph, neither properly explained)
- Corpus filtering paragraph (tries to synthesise but reader is lost by now)
- Self-evidencing explanation (yet another Lipton concept, "I really would have no idea why you're telling them about this")
The section reads like a literature review — "here's what Person A says, here's what Person B says" — rather than an argument. Each paragraph answers a question the reader wasn't asking, because the question was never posed. The reader can't follow because they don't know what problem each source is solving.
Your comment about the Lipton paragraph itself is also correct: "you're just restating exactly what was in the block quotes." After a block quote, the text needs to BUILD — explain what's at stake, draw out the implication, show why this matters for the argument at hand. Paraphrasing the quote back to the reader wastes their time. And then your separate comment — "you haven't explained what lovely means properly, so now you're having to give extra examples a sentence too late" — is the same issue from a different angle. The Lipton distinction is important enough to deserve proper treatment. It needs room.
---
### 3. Section 2's staging problem
Section 2 has better content but poor staging. The argumentative turn — from "Floridi is right about the mechanism" to "but what the mechanism operates over matters" — is too abrupt. You call it "a very abrupt and disorientating change switch turnaround," and the problem is that the reader needs to see the move coming.
The turn currently happens in a single sentence: "Floridi et al. themselves provide the materials for a response." This is jarring because the previous paragraph just conceded Floridi's point. The reader is oriented toward agreement — "OK, LLMs don't really reason" — and then gets whiplash when the text suddenly says "but actually, Floridi defeats himself."
After the turn, the convergence thesis — "statistical plausibility, within this corpus, converges with philosophical quality" — is compressed into one dense paragraph when it should be the most carefully staged argument in the paper. This is the thesis that does the work. It needs to be built step by step: (a) statistical plausibility is relative to training data, (b) the philosophical training data is a filtered corpus, (c) what it was filtered for is philosophical quality, (d) therefore plausibility in this distribution approximates quality. Each step deserves its own moment. Currently they're jammed together and the reader has to reconstruct the reasoning.
Your comment: "all of this content is good, but it just doesn't seem to me as though it's been properly explained to the reader... the structures here and the ordering of information and the clarity is a fucking disaster." That's the diagnosis. The content is there; the presentation isn't.
---
### 4. The Pigliucci tension
You flag this directly in Section 1: the claim "there is nothing prior that it reports" is too strong given Pigliucci's framework. Having read the chapter, I think you're right to worry.
Pigliucci's position: philosophy is "empirically informed evoking." Its starting points — what he calls its "axioms" — are empirical data about the world. He writes:
> "the basic parameters that philosophers use as their inputs, the starting points of their philosophizing, their equivalent of axioms in mathematics and assumptions in logic (or rules in chess) are *empirical* data about the world."
And:
> "philosophy attempts to clarify things, or to analyze in order to bring about understanding, not really to discover new facts, but rather to evoke rational conclusions arising from certain ways of looking at a given problem or set of facts."
So for Pigliucci, philosophy IS connected to the world — it takes empirical starting points and explores the conceptual landscapes they generate. The philosophical contribution isn't a report of those starting points, but it's constrained by them. The evoked truths have rigid properties.
The current formulation — "there is nothing prior that it reports" — suggests philosophy is entirely untethered from prior reality. But that's not your thesis. Your thesis is that philosophical EVALUATION concerns text-internal properties, not that philosophy has no connection to the world. The "nothing prior" claim is stronger than you need and it hands Zahavy ammunition: if even the paper's own Section 1 says philosophy has nothing prior, then maybe it really is just text all the way down, and Zahavy can say the "text all the way down" view is exactly what makes LLM philosophy hollow.
I can see a few ways to handle this:
Option (i): Refine the Watson/Crick claim using Pigliucci's vocabulary. Something like: "Philosophy doesn't report prior facts — it evokes constrained conclusions from empirically given starting points. But the philosophical contribution lies in the evocation, not the starting points, and the evocation is the text." This preserves the text-internal thesis without denying that philosophy engages with the world.
Option (ii): Acknowledge the connection to the world early (Section 1) and then show it doesn't threaten the text-internal thesis. The empirical constraints enter philosophy as propositions — they're stated in texts, debated in texts, refined in texts. The world-connection is propositionally mediated, and propositions are textual objects. This sets up the later response to Zahavy (Section 3) where Pigliucci's point about propositional starting points does direct work.
Option (iii): Drop the "nothing prior" claim entirely. The Watson/Crick contrast works without it. The point is that Quine's contribution lies in his arguments — not that there's literally nothing the arguments engage with. You can say: "Where Watson and Crick reported an arrangement that existed before they described it, Quine's contribution consists in the arguments he makes. A different philosopher reaching the same conclusions by different arguments would have made a different contribution." That's enough for the text-internal thesis without the metaphysical claim about nothing being prior.
I'm drawn to option (iii) — it's the most economical and avoids a claim that creates problems later. But options (i) and (ii) have the advantage of engaging with Pigliucci's framework early, which might strengthen the paper by showing you've thought about the "connected to the world" concern before the reader raises it.
---
### 5. Source-check issues
A few attribution problems beyond the ones you've already flagged:
In the Introduction: "Even so, there are arguments that LLMs cannot produce work of this kind." Neither Floridi nor Zahavy wrote about philosophy specifically. Floridi is about LLM reasoning generally; Zahavy is about theoretical innovation in science. The next sentence ("Neither paper concerns philosophy directly") actually corrects this, but the opening sentence contradicts it. The reader gets "there are arguments against" followed immediately by "actually, neither paper is about philosophy." This needs to be consistent — I'd say the opening sentence should say something like "there are arguments that LLMs cannot reason in ways that matter for philosophy" or "arguments that, if sound, bear on the question."
The Nietzsche/Sorgner sentence in the Introduction: you've flagged this already. The sentence implies Sorgner was writing about LLMs. He wasn't. The inference about LLMs needs to be clearly marked as yours (or yours and Enrico's): "On Nietzsche's account (as Sorgner reads it), philosophers are creators of values, expressing drives and psychophysiology. On such a view, the question would not arise — an LLM lacks the requisite conditions entirely."
The Dellsén example in Section 1 (the dying scientist): if this thought experiment comes from Dellsén's paper, it needs to be attributed. "Dellsén and colleagues illustrate this with a thought experiment: suppose a scientist..." If it's your illustration of their thesis, it should say so.
The Bengson ambiguity in Section 1: "A paper that scores well on these criteria handles objections in recognisable ways..." — this could be Bengson's claim or yours. Needs explicit marking either way.
I should flag: I haven't re-opened the Floridi, Zahavy, or Lipton extractions to verify every block quote against the source text. The page numbers are specific and the quotations look right in style, but I can't certify them without checking. If you want me to verify the quotations against extractions, I can do that as a separate step.
---
### 6. Structural options — Section 1
I see several possible approaches, and these are genuinely different — not variations on one idea.
Approach A — Reorganise around questions, not sources. Instead of "Lipton says X, Williamson says Y, Bengson says Z," organise around: "What makes philosophy good? [drawing on multiple sources] How is this visible in texts? [drawing on multiple sources] Why does the producer not matter? [drawing on multiple sources]." This weaves sources into the argument rather than giving each their own paragraph. The advantage: solves the listicle problem directly. The risk: it's a full restructure.
Approach B — Thin the sources. The section may be trying to deploy too many theorists in the main text. What if Lipton provided the framework (likeliest/loveliest), Williamson provided the specific articulation of philosophical virtues, and everyone else went to footnotes? The argument becomes: (1) philosophy lives in its arguments; (2) the relevant evaluative standard is loveliness (Lipton); (3) loveliness in philosophy consists in elegance, unity, non-ad-hocness (Williamson); (4) these are text-internal properties; (5) therefore the producer doesn't matter. Five steps, maybe seven paragraphs. The [^criteria] footnote already does some of this consolidation. The advantage: cleaner, faster, the argument is tractable. The risk: loses the richness of showing multiple theorists converging on the same picture.
Approach C — The "one good example" spine. Make Semmelweis the section's anchor rather than one item in a list. Introduce the likeliest/loveliest distinction, illustrate it properly with Semmelweis (with real setup: "Consider one of Lipton's own examples..."), then immediately show what this looks like in philosophy: when we read Quine, we assess whether his case against analyticity is illuminating, whether it handles rival views, whether it unifies — these are loveliness properties. Then the text-internal thesis falls out: if what we evaluate is loveliness, and loveliness is assessable from the text, then evaluation is text-internal. The advantage: the reader has something concrete to hold onto throughout. The risk: the Semmelweis example is from Lipton's book, so we're essentially teaching the reader through Lipton's pedagogy, which might feel derivative.
Approach D — Lead with Williamson rather than Lipton. Instead of going through the likeliest/loveliest apparatus (which requires setup), start from Williamson's direct claim that philosophical theory choice is governed by simplicity, elegance, explanatory power, non-ad-hocness. These are familiar theoretical virtues — a philosopher reading the paper already knows what they are. Then bring Lipton in to explain WHY these virtues matter (they're loveliness properties, not likeliness properties) and what this implies (evaluation is about intrinsic qualities, not about confirmation). The advantage: starts with what the reader already knows, builds outward. The risk: Lipton becomes supplementary rather than structuring, and the paper draws heavily on Lipton later.
Approach E — Combine B and C. Thin the sources AND use Semmelweis as the spine. The section would: (1) Watson/Crick vs Quine (refined with Pigliucci nuance); (2) what makes arguments good — the likeliest/loveliest distinction (Lipton); (3) Semmelweis as the illustration, properly set up; (4) application to philosophy — loveliness evaluation uses Williamson's virtues; (5) these are text-internal, therefore the producer doesn't matter (brief: Gaut, blind review). Everything else to footnotes. Self-evidencing explanation either becomes the section's capstone (explaining WHY evaluation is text-internal) or moves to a footnote.
I don't think there's one right answer among these. It depends on what you want the section to feel like: Approach B is lean and fast, Approach A is rich and integrative, Approach E is concrete and pedagogy-driven. My instinct is that E or B would serve the paper better than the current structure, but I want to hear what you think rather than presuming.
---
### 7. Structural options — Section 2
Section 2's problems are more about staging than structure. The content is in roughly the right order; it just needs better staging of the argumentative turn and more room for the convergence thesis.
The turn needs preparation. Currently: "We grant this characterisation at the level of mechanism" → one sentence later → "Floridi et al. themselves provide the materials for a response." The reader needs to see the pivot point: you're not disagreeing about what LLMs do, you're redirecting attention to what they do it OVER. Something like: "The question, then, is what follows from this concession. If the mechanism is stochastic, does the output inherit that character? That depends on what the stochastic process operates over."
Then the convergence thesis needs to be BUILT, not stated. I'd suggest dedicating three or four paragraphs to it:
- Paragraph 1: Statistical plausibility is relative to training data. What an LLM treats as plausible depends entirely on what it was trained on. (This is a simple point that needs to be made explicit.)
- Paragraph 2: The philosophical training data is not a random sample. It's been filtered — by peer review, by citation, by teaching, by anthologising — for the very properties that constitute philosophical quality.
- Paragraph 3: Therefore, what counts as "plausible" in this corpus is what scores well on philosophical virtues. The convergence isn't accidental; it's a consequence of the filtering.
- Paragraph 4: The grammar analogy — children absorb grammatical norms from grammatical speech; LLMs absorb philosophical norms from philosophical text. Disanalogy honestly faced.
After that, the "just statistics" objection (levels of description / squash analogy) and the borrowed calibration argument follow naturally, because the reader now understands the thesis being defended.
---
### 8. The "et al." formatting rule
For papers with more than two authors: always "Floridi et al." — not "Floridi, Nobre, and Taddeo." This applies on first and subsequent occurrences. The current draft sometimes does this right and sometimes doesn't. This should be a blanket rule I apply whenever writing.
---
### What I haven't covered
I haven't gone through the material after line 18 of Section 2 (where you stopped reading) in the same detail, since your comments don't extend there. The borrowed calibration argument, Model A/B, Floridi's hedge, levels of description, latent ≠ expressed, empirical questions — these all have content worth preserving, but I'd expect similar voice and staging issues. Once we sort out the structural approach for Sections 1 and 2, I can look at those later paragraphs with the same diagnostic eye.
I also haven't addressed the Introduction in depth beyond the specific comments you made. The structural question (which plan to follow for Sections 2-4, per the session file) affects the roadmap paragraph, so that's probably premature to address until you've decided on the overall structure.
Where do you want to start? I'm inclined to think Section 1 is the bigger problem — once that's right, Section 2's issues become more tractable because the reader arrives with the right framework. But you might want to start with the Pigliucci question, since it affects the opening claim and ripples through both sections.
*La tensione tra ciò che il testo dice e ciò che il testo presuppone è esattamente dove si gioca la partita filosofica.*