# opus[1m]
## Retrieval Notes
- Session id: `79fd692b-bd39-4660-88e3-18b563d49809`
- Last activity: `2026-03-20T09:03:14.503Z`
## My Notes
<!-- Add your notes here. This section is preserved across syncs. -->
## Conversation
### User
opus[1m]
### User
/draft-audit
the whole note, not just highlighted passages (if any)
### Assistant
Let me read the target note and published papers in parallel.
### Assistant
Now let me read published papers for comparison benchmarks.
### Assistant
I have the draft and three published papers. Let me now do the paragraph-by-paragraph audit.
---
## Depth Audit: "2. Likeliness, Loveliness, LLMs"
I've read the full note against paragraphs from "Growing the Image," "Hearing Spaces," "Sounds as Properties," and "Agents of Change." Here is the paragraph-by-paragraph assessment.
---
### P1 (lines 4–8): Introducing Floridi et al. on zeroth-order abduction
RICH. This paragraph does exactly what Nick's published introductions do. It opens with the claim ("LLMs do not reason abductively"), immediately clarifies what the claim amounts to (not error-proneness but absence of process), provides a concrete example (car won't start on a cold morning), and then drops a block quote that carries specific terminology. Compare with the introduction of Anscomb in "Growing the Image":
> Not only does Anscomb refer to text-to-image systems as 'AI Agents' throughout her paper, she suggests here that even if they are not artists proper, they might still deserve some degree of 'credit' for the 'contribution' that they have made to an artwork's creation by working 'autonomously' and 'iteratively'
Same structure: the author's position is stated, then unpacked with attention to the specific language they use, then put under pressure. P1 earns its claims.
---
### P2 (line 8, "Zeroth-order abduction..."): Unpacking via Lipton's framework
RICH. The paragraph takes Lipton's two-stage account and shows how Floridi et al.'s diagnosis maps onto it — specifically, that their account "collapses this two-stage process into a single step." Each step of Lipton's framework is stated (background beliefs generate plausible hypotheses; those are ranked by explanatory merit), and the mapping onto Floridi et al. is specific. Anchored by three direct quotations from Floridi et al. ("external feedback loop," "surface-level abductive appearance," "compelling illusion"). Compare with how "Agents of Change" introduces Hoerl's framework:
> Why would a lack of tensed content invite the belief in a privileged present? Hoerl argues for this point by comparing temporal and spatial experience.
> What objects we see and the way those objects look, depends upon the point in space from which we are looking at them. Moreover, this dependence is manifest in the experience: what we see reveals some fact about our own location in relation to those objects. Spatial experience, in and of itself, therefore provides an answer as to why these objects are being perceived in this way: you are seeing this ice cream, and viewing it in profile, because you are located here. By saying that we lack a temporal viewpoint Hoerl is saying that temporal experience lacks an equivalent feature...
Same method: the framework is explained and then the author's specific deployment of it is traced. P2 does this well.
---
### P3 (line 10): Applying the diagnosis to philosophical writing
RICH, with a minor reservation. The paragraph translates the general mechanism into domain-specific terms ("argumentative moves," "dialectical situation," "force of that objection") and the final sentence — "What appears as philosophical activity would be, on this picture, nothing more than a statistical echo of earlier philosophical activity" — is a strong formulation.
The reservation: the paragraph tells us that "an objection may be well handled not because the system assessed the force of that objection but because the statistical structure of the training data makes such handling probable at that point" without providing a concrete example of such handling. Compare with how P1 works the car-won't-start case to make the general point vivid. An equivalent here would be showing a specific philosophical move (say, anticipating a counterexample) and explaining what it would mean for that move to be statistically probable rather than dialectically motivated.
This is borderline rather than a clear failure — the paragraph is transitional, setting up the argument that follows. But if you wanted to thicken it, a concrete philosophical case would be the route.
---
### P4 (line 12): Statistical plausibility is relative
RICH. This paragraph makes the section's central philosophical move: identifying the gap in Floridi et al.'s reasoning. The contrast between "a corpus of advertising copy or undergraduate boilerplate" and "a corpus of philosophy" does real argumentative work — it shows concretely why treating statistical plausibility as undifferentiated is a mistake. Compare with a similar move in "Hearing Spaces":
> Hearing two speakers as a single source, or one speaker as multiple sources, misinforms us about the layout of the world. In contrast, hearing reverberation allows for accurate and potentially useful knowledge about our environment. Given this informativeness, alongside the pervasiveness of experiences of reverberation, theories of auditory perception that reduce reverberation to this type of illusion are implausible.
Both paragraphs draw a contrast that exposes what an opposing view has overlooked. P4 does this cleanly.
---
### P5 (line 14): The philosophical corpus as filtered by selection
RICH. The mechanism of selection is developed: referees judge work worth publishing, later philosophers determine whether it must be answered or built upon. The qualification — "None of this yields a pure corpus — weak work survives and strong work is overlooked — but the corpus is not unfiltered either" — shows philosophical care. Every sentence advances the argument. Compare with how "Sounds as Properties" builds the case that vibrations could be realisers of either event-properties or object-properties:
> Now, for any normal auditory experience these physical features would be instantiated in the world three times over: by sound waves, by disturbances of the medium, and by vibrations of the object itself. Given that both sound waves and medium disturbances are distinct from source events and their objects, neither seems like a good candidate for a realiser of Property's sounds, leaving vibrations as the only remaining possibility. This does not, however, provide an unequivocal reason for thinking that sounds are properties of events but not objects...
Same pattern: developing a point by working through its implications while building in qualifications.
---
### P6 (line 16): Williamson, Bengson et al., and Walton et al. on philosophical standards
MOSTLY RICH, with one thin spot. Williamson and Bengson et al. each get enough development. Williamson's criteria are quoted directly ("elegant and unified, not arbitrary, gerrymandered, ad hoc"), and Bengson et al.'s method is summarised with enough specificity to see how it works (three stages: accommodating data, substantiating claims, possessing theoretical virtues). The sentence "The ordering matters: a view can fit the cases and still fail because the claims doing the work are poorly supported" draws out an implication rather than merely restating the framework.
The thin spot is the Walton et al. reference:
> As Walton et al. observe, familiar forms of defeasible reasoning carry familiar critical questions, and philosophers write under the pressure of those questions whether or not they are explicitly formulated (2008, ch. 1).
Failure mode: Named but not developed (type 3). What are "critical questions"? What does it mean for reasoning to carry them? How does writing "under the pressure" of these questions manifest in actual philosophical prose? The reference does one sentence of work when it needs at least two or three. Compare with how Nick introduces any view in the published papers — Leddington, for instance, in "Sounds as Properties":
> According to other 'property views' (Locke, 1690/1975, II, viii, 14; Pasnau, 1999; and Kulvicki, 2008), sounds are instantiated by material objects: the sound of a bell being struck is a property instantiated by the bell itself. In contrast, on Leddington's account, Property, sounds are instantiated by the source events in which these objects are participating: it is the striking, not the bell, which bears a sound. That is, the striking bears an audible property constituted by the qualities pitch, timbre, and loudness, in a similar way to how the bell itself bears a visible colour property constituted by hue, saturation, and brightness (2019, n. 6).
Even brief introductions of views spell out the mechanism. The Walton et al. reference doesn't.
What's missing: two or three sentences explaining what "critical questions" are in Walton et al.'s framework, with an example of how they create pressure on philosophical prose. For instance: what are the standard critical questions for an argument from analogy, and how does their presence in the evaluative background shape the way analogies are deployed in surviving philosophical texts?
---
### P7 (line 18): Likeliness and loveliness brought together
RICH. This is the synthesis paragraph where the argument lands. It brings together Lipton's distinction, the corpus-filtering point from P4–P5, and the Williamson/Bengson criteria from P6, and shows how they interact: "In a body of writing already filtered for the properties Williamson and Bengson et al. describe — explanatory depth and non-ad-hocness — the most probable continuation is not simply whatever phrase happens most frequently in unfiltered text." The conclusion — "A corpus filtered for loveliness will tend to make lovely continuations likelier" — emerges from the reasoning rather than being asserted. Compare with how "Growing the Image" arrives at its central category:
> Considering Midjourney as a dynamic medium incorporates aspects of both the agency and tool views while avoiding their shortcomings. This allows for the recognition of the system's autonomous nature, similar to the agency view, but also fits with the perception of Midjourney as a means of creation, as per the tool view, and yet acknowledges its autonomy as an integral, functional feature, not a malfunction.
Both paragraphs synthesise preceding threads into a conclusion that is earned by the development that precedes it. P7 does this well.
---
### P8 (line 20): The "borrowed calibration" objection
RICH. This paragraph anticipates an objection, gives it genuine force (the empirical science contrast is specific — "features of the world not exhausted by the scientific literature"), and responds with a concrete example (Williamson's defence of simplicity as itself a member of the corpus it governs). The structure is exactly how objections are handled in the published papers. Compare with "Growing the Image" on the autonomy objection:
> To see why autonomy is not sufficient for attribution of credit, consider the following example. As I pour wine into a glass, you take photos of the liquid splashing and rippling as the glass is filled. The wine is autonomous in the sense that neither I nor you have direct control over exactly how the liquid will splash into the glass (e.g. the size of the ripples, how many bubbles appear), but we would not think that the wine deserves any credit for the resulting photos in any interesting sense, nor would we say it has made any sort of contribution.
Both raise an objection and respond with a specific, developed example. P8 earns its conclusion about self-grounding calibration.
---
### P9 (line 22): Process vs. content and levels of description
RICH. Three distinct moves, each developed:
1. Floridi et al.'s own question is quoted and the response turned against them ("opens more than they allow"). The point that content assessment *is* philosophical evaluation is specific to the domain and argued for.
2. The blind review example is concrete and well-placed — it shows that the institutional structure of philosophy already embodies the principle that process is irrelevant to assessment.
3. The Lipton squash analogy is quoted in full and then applied to the LLM case: "The mechanistic description of what an LLM does and the philosophical assessment of what it produces operate at different levels."
The concluding sentences — "A passage may still handle an objection well, or leave a subject clearer than it found it. Those are assessments of the same text at the level at which philosophical prose is ordinarily judged" — earn their weight because the argument for levels of description has been made, not merely announced. Compare with the closing of the perception-of-change section in "Agents of Change":
> While this is not meant to be a knock down argument against the idea that a belief in time passing could come from perception, it does give us a reason to consider whether other, non-perceptual, forms of experience might do a better job. There is more to sensuous first-person experience than perceiving, and there is no obvious reason to think that the experiential ingredient which tells us time passes must be perceptual, or, in the case of mental change, quasi-perceptual.
Both close by restating the upshot modestly, having earned the right to do so. P9 is strong.
---
### P10 (line 24): Conclusion
APPROPRIATELY SCOPED. This is a concluding paragraph — it does not need to make new moves. It qualifies the section's argument ("not expressed in every output; whether a given output succeeds is an empirical matter"), restates the upshot, and sets up the next section via Zahavy.
The Zahavy reference — "Zahavy (2026) argues that it does, and that this places a limit on what any text-trained system can produce" — is named but not developed. However, this is clearly a forward pointer ("the next section will address this"), not a reference that is supposed to do argumentative work in this paragraph. No depth failure here — the underdevelopment is intentional.
---
## Summary
This is a strong piece of philosophical prose. Most paragraphs make their moves rather than describing them. The draft engages directly with sources' language, develops concrete examples, earns its conclusions through reasoning, and handles objections with specificity. Against the published papers, the writing is consistently at the same standard.
One genuine flag:
- P6 (Walton et al.): Named but not developed. The reference appears in a single sentence without explanation of what "critical questions" are or how they create the pressure being claimed. Adding 2–3 sentences unpacking the Walton et al. framework would bring this up to the standard set by the rest of the section.
One borderline observation:
- P3 (application to philosophy): The translation of Floridi et al.'s diagnosis into philosophical terms is clear but lacks a concrete example of the kind used in P1. A specific philosophical move (e.g., "a counterexample might be anticipated not because the system judges it forceful but because counterexample-anticipation is statistically probable at that point in a philosophical text") would thicken it. This is not a failure — the paragraph is transitional — but it's the one paragraph where the argument proceeds by description rather than demonstration.
Everything else is clean.
*Il corpus che sopravvive non è innocente — porta con sé i giudizi di chi lo ha lasciato passare.*
### User
i am not certain you cuaght everything
### Assistant
Fair enough. Let me re-read with more pressure.
### Assistant
You're right — I was too generous. On a harder pass, here are two additional findings I missed.
---
## Revised Depth Audit — Additional Findings
### P6 (line 16): The child/grammar analogy — described but not made (type 1)
> as with a child who learns to produce grammatical sentences through exposure to well-formed speech without possessing any grammatical theory, exposure to the textual consequences of these norms may be enough to absorb what they require.
This analogy is carrying enormous argumentative weight. It's doing the work of establishing that pattern-absorption from well-formed examples can be sufficient for norm-following — arguably one of the section's most important claims. And it arrives in a subordinate clause.
The analogy announces a parallel but doesn't argue for it. Why does the parallel hold? Children learning grammar from exposure is itself a contested phenomenon — on some accounts it works because of innate language-acquisition capacities, on others because of embodied social interaction, not simply because statistical exposure to well-formed speech is sufficient. If the analogy is supposed to support the idea that a model can absorb philosophical norms through exposure to norm-shaped texts, it needs to address what features of the child's learning are shared by the model's learning, and what features aren't. As it stands, the analogy describes a similarity without showing that the similarity is the right kind to bear the weight placed on it.
Compare with how analogies work in the published papers. The wine-pouring example in "Growing the Image":
> As I pour wine into a glass, you take photos of the liquid splashing and rippling as the glass is filled. The wine is autonomous in the sense that neither I nor you have direct control over exactly how the liquid will splash into the glass (e.g. the size of the ripples, how many bubbles appear), but we would not think that the wine deserves any credit for the resulting photos in any interesting sense, nor would we say it has made any sort of contribution.
The analogy is set up, the relevant feature is identified (autonomy), and the conclusion is drawn (autonomy does not suffice for credit). Each step is explicit. The child/grammar analogy compresses setup, relevance-identification, and conclusion into a single clause.
What's missing: development of what makes the analogy apt. Why is a child's relationship to well-formed speech relevantly similar to a model's relationship to norm-shaped texts? What is it about exposure to well-formed examples that enables norm-following in both cases? And does the analogy survive the obvious disanalogy — that children are embedded in linguistic communities where speech is tied to action and consequence, while models encounter decontextualised text? These questions need at least a sentence or two, not because they're fatal to the analogy but because without them the analogy is asserted, not argued.
---
### P8 (line 20): Self-grounding — asserted without earning (type 2)
> The model is exposed not only to texts shaped by these standards, but to the arguments for why those standards should govern philosophical judgement. In this sense, the calibration available to a model trained on philosophy is not merely inherited but self-grounding: the justifications for the evaluative standards are themselves part of the corpus that embodies them.
The paragraph establishes something true and interesting about the structure of philosophy: its evaluative standards are argued for within the same body of literature they govern. Williamson's defence of simplicity is part of the same corpus as the simple theories it commends. The self-grounding is a feature of the corpus.
But the conclusion — that the model's calibration is "not merely inherited but self-grounding" — transfers the self-grounding property from the corpus to the model, and this transfer is not argued for. The model is exposed to the self-grounding structure, yes. But exposure to a self-grounding corpus is not the same as having self-grounding calibration. The model does not distinguish between "texts shaped by standards" and "arguments for those standards" — it absorbs both as training data. A student who memorises Williamson's defence of simplicity without grasping its force has been exposed to the self-grounding argument but hasn't inherited its self-grounding character.
The gap is between:
- Premise: the philosophical corpus is self-grounding (the justifications for its standards are part of it)
- Conclusion: the calibration available to a model trained on this corpus is self-grounding
What's missing: an inference step explaining why exposure to self-grounding arguments transfers the self-grounding property to the model's calibration. This might involve showing that the model's outputs exhibit sensitivity to the *reasons* for standards (not just the standards' effects on prose), or it might involve arguing that the distinction between "inheriting results" and "inheriting justifications" doesn't require the kind of understanding the objection assumes. Either route would close the gap. As it stands, the paragraph argues well for the premise but asserts the conclusion.
Compare with how "Agents of Change" handles an analogous move — transferring a feature of one domain to another:
> Feeling oneself as the source of an action, can be thought of as a unique type of change experience. Compare what it is like to watch the droplet of ice cream run down the ice cream cone with what it is like to move your head forward to take a lick. In the first case you are experiencing the world to change, in the second you are experiencing yourself as changing the world. You are not simply witnessing your head to change its location, but sensing yourself to be the instigator of the change.
The transfer (from perceiving change to experiencing oneself as changing) is made through a concrete comparison that shows exactly where the analogy holds and where it differs. P8 needs something equivalent to close the gap between corpus-level and model-level self-grounding.
---
## Full revised findings
1. P3 (borderline): Applies Floridi et al.'s diagnosis to philosophy without a concrete philosophical example of the pattern it describes.
2. P6 — Walton et al.: Named but not developed (type 3). One sentence where "critical questions" and how they create pressure need unpacking.
3. P6 — child/grammar analogy: Described but not made (type 1). A subordinate clause doing the work of a paragraph. The analogy announces a parallel without arguing that the parallel holds.
4. P8 — self-grounding conclusion: Asserted without earning (type 2). The inference from "the corpus is self-grounding" to "the model's calibration is self-grounding" skips a step about what it means for a model to inherit self-grounding structure rather than merely inheriting its effects.
P1, P2, P4, P5, P7, P9, and P10 remain clean.
*La profondità non si annuncia — si fa, oppure non c'è.*