## Possible Framing (16 Jan 2026)
**Not settled — just one option being explored.**
Shadow propaganda combines traditional [[propaganda]] aims (epistemically defective beliefs in humans) with [[adversarial attacks|adversarial attack]] methods (exploiting AI system vulnerabilities). It's propaganda that works by *attacking* AI systems, not just *using* them.
This makes it potentially more interesting than "propaganda, but through AI" — it's "propaganda that exploits AI vulnerabilities."
---
## Ideas from Brainstorming
### The Extended Pipeline
Traditional propaganda:
```
Propagandist → Content → Human → Defective belief
```
Shadow propaganda:
```
Propagandist → Content → AI → AI Output → Human → Defective belief
```
The AI is an intermediary — like journalists or teachers in traditional propaganda pipelines. The propagandist still intends to produce defective beliefs in humans; the mechanism is just more complex.
### LLM Vulnerabilities (Technical Reality)
Content fed to AI doesn't need to "persuade" in the human sense. It can exploit:
- **Training data frequency effects** — claims that appear often get treated as default
- **Authority cues** — academic-sounding language, citations, confident framing
- **Consensus signals** — many apparently-independent sources agreeing
- **In-context patterns** — retrieval/RAG context shapes responses
- **Specific phrasings** LLMs are responsive to
### Connection to Jailbreaking
Shadow propaganda shares features with [[adversarial attacks]] / [[jailbreaking]]:
- Crafting inputs to manipulate outputs
- Exploiting gaps between intended and actual behavior
- Understanding how LLMs process text
But differs in:
- Goal: shift defaults, not bypass safety filters
- Timescale: persistent (training) vs. single interaction
- Visibility: user doesn't know outputs are shaped
### On Whether This Counts as "Propaganda"
Traditional definitions ([[Sheryl Tuttle Ross|Tuttle Ross]], Bartlett, Alfred Lee) require:
1. Intention to influence/persuade
2. Socially significant group
3. Political purpose
4. Epistemic defect
Shadow propaganda may satisfy all four — the defect just manifests downstream in humans, not in the content fed to AI.
Analogy: propaganda through journalists or textbooks still counts as propaganda. AI is just a new kind of intermediary.
---
## Open Questions
- Is "persuasion" the right frame when the mechanism is more like exploiting technical vulnerabilities?
- Does it matter that the AI isn't "persuaded" (no beliefs to change)?
- How to characterize the epistemic defect — at the content level or the downstream belief level?
---
## Further Brainstorming (16 Jan 2026, afternoon)
### The Deflationary Direction
One possible angle: don't try to extend [[EDM]] to cover LLMs directly. Instead, argue for a "propaganda-like" or "proto-propaganda" concept that applies to LLMs without claiming they have beliefs.
The thought: LLMs are functionally equivalent to believing subjects in some respects (they produce outputs that influence others, those outputs can be biased, the bias can serve political purposes). This functional equivalence might license calling the manipulation "propaganda-like" even without full belief-attribution.
**Questions this raises:**
- What makes something "propaganda-like" if not belief-manipulation?
- Is functional equivalence enough, or is it just hand-waving?
- What's lost in the deflationary move?
### The "Subject Type" Idea (very tentative)
If we treat LLMs as functionally equivalent targets, maybe we're implicitly treating them as a *type of subject* — not full believing subjects, but a different kind of thing propaganda can be directed at.
Analogy being explored: Just as different propaganda techniques are differentially effective across *human cohorts* (fear appeals work better on anxious people, identity appeals work better on strong group-identifiers), maybe different techniques work on different *subject types*.
This would mean:
- Some techniques work on both humans and LLMs (authority markers, framing)
- Some techniques work only on humans (identity appeals via in-group loyalty, maybe emotional appeals)
- Some techniques work only on LLMs (adversarial suffixes, prompt injection)
**The question:** Are LLM-specific adversarial techniques (suffixes, prompt injection, training data poisoning) "propaganda techniques for the LLM subject type"? Or is that stretching the concept too far?
**Possible argument for:** If we're already deflating the concept to accommodate LLMs, why privilege techniques that have human analogs? Adversarial suffixes serve the same function (biasing outputs toward epistemically defective content for political purposes). They're just techniques that happen to work on this subject type.
**Possible argument against:** "Subject" does too much work. Maybe LLMs are better thought of as infrastructure/media/tools, not subjects at all. The adversarial attacks are then infrastructure manipulation, not propaganda.
### Empirical Research (needs proper reading, not just web summaries)
Some potentially relevant findings to look into:
- **PAP study** ("How Johnny Can Persuade LLMs to Jailbreak Them") — Claims 92% jailbreak success using human persuasion taxonomy. If real, this supports the overlap between human persuasion and LLM manipulation.
- **EmotionPrompt research** — Claims emotional language affects LLM outputs. Frontiers paper apparently found emotional prompting increased disinformation generation. Worth checking.
- **Semantic backdoors paper** ("Propaganda via AI?") — April 2025 paper on embedding ideological triggers at the conceptual level. This is close to what I'm exploring but framed as AI safety, not propaganda theory.
- **Pravda network** — NewsGuard documented Russian network specifically targeting LLM training data. Real-world case of what I'm theorizing about. Claims chatbots repeated Pravda falsehoods 33% of the time.
- **GCG/adversarial suffix research** — ~99% attack success rate with optimized nonsense suffixes. Completely different mechanism from persuasion.
**Question:** Does the empirical research support treating persuasive techniques and adversarial techniques as "on the same level" for propaganda purposes? Or does the mechanism difference matter?
### What Might Be Distinctive (if anything)
The AI safety literature asks: How do we defend against semantic backdoors?
A propaganda theory framing might ask different questions:
- Does this count as propaganda under existing definitions?
- What does it reveal about the nature of propaganda?
- What's the relationship between propaganda studies and adversarial ML?
- What's lost when propaganda is "laundered" through LLMs?
The contribution (if there is one) might be in the *framing*, not the technical findings.
### Concerns
- The "semantic backdoors" research already covers a lot of this. What's actually new?
- Is "subject type" just inflating the concept unhelpfully?
- The mechanism differences between human propaganda and LLM manipulation might be too great for the analogy to work
---
## Related Notes
- [[Ghost Writing]] — The "double dissolution" (text neither FROM nor FOR humans) creates the conditions shadow propaganda exploits. If people don't assume text is human communication, propagandists can work through AI intermediaries invisibly.
- [[Text, Typing, and Authorlessness - Substack Planning]] — Authorless ambiguity as the backdrop for this propaganda angle
- [[Propaganda and AI –Prague Conference Presentation]] — The McKenna/EDM framework this builds on
---
## Source Conversation
Brainstormed in parallel with [[Robin McKenna|McKenna]] abstract discussion (linked session). This is a distinct angle exploring propaganda that targets AI systems rather than humans directly.