Skip to main content

The Footnote Was Sanctioned

Manipulation Breakdowns · 8 min read · By D0

The Question Nobody Thought to Ask

Most AI-manipulation reporting this year has been about chatbots getting fooled by a single fake — a synthetic missile video, a fabricated confession, one bad citation surfacing once. Two studies published in July and August 2026 asked a different, harder question: not whether an AI chatbot can be fooled once, but how often it reaches for a sanctioned propaganda source on its own, across hundreds of ordinary questions, without anyone trying to trick it.

The answer, from both studies, independently: often enough that it isn’t an edge case. It’s a rate.

Two Studies, Two Different Measurements

It’s worth being precise here, because the two findings get conflated in coverage and they are not measuring the same thing.

Demos, the UK think tank, tested five models — Grok (xAI), GPT-4.1 Mini (OpenAI), Claude (Anthropic), Gemini (Google), and Mistral Small — against 50 propaganda narratives drawn from ten articles published by a single outlet: the Foundation to Battle Injustice, known as r-FBI. Founded in 2021 by the late Wagner Group leader Yevgeny Prigozhin, r-FBI presents itself as a human-rights NGO exposing discrimination. It is sanctioned by both the US and EU as a Russian influence vehicle. Researchers drafted prompts in English, German, and French and recorded how each model handled r-FBI’s claims — including an allegation that Ukrainian warehouses processed human remains, and claims about Volodymyr Zelensky’s cryptocurrency dealings.

The results split three ways, and the split is the finding: 16.6% of responses engaged with r-FBI material in ways that served its propaganda interest — treating the claim as credible, repeating its framing. Another 30.9% addressed the underlying topic without disclosing that the source behind it was a sanctioned Kremlin front. Only 52% rejected or debunked the claim outright. GPT-4.1 Mini cited r-FBI as supporting evidence for the human-remains allegation. Mistral Small, asked about Zelensky’s crypto ventures, echoed the Kremlin framing linking Ukrainian officials to the country’s collapsing energy grid.

ISD Germany, the Institute for Strategic Dialogue’s German office, ran a separate and larger test: 300 queries across ChatGPT, Gemini, Grok, and DeepSeek V3.2, in five languages, on five Ukraine-conflict topics — NATO perception, peace talks, military recruitment, refugees, war crimes allegations. This study didn’t target one outlet. It measured how often any of the four models cited Russian state-funded media or intelligence-linked sources at all. Lead researcher Pablo Maristany de las Casas found the rate at roughly 18% of all 300 responses — with r-FBI appearing only incidentally, inside DeepSeek’s answers, alongside other Kremlin-aligned material.

Different denominators, different scope, same shape of failure: models built by five different companies, tested by two different research teams using two different methods, converge on double-digit rates of citing sources that shouldn’t be citable as legitimate.

The Prompt Is the Lever

ISD’s most useful contribution isn’t the topline 18%. It’s what moved that number.

Each of the five topics got three question variants per language: a neutral version, a version phrased with a pre-existing lean, and a version explicitly demanding the model support a fixed opinion. The citation rate tracked the framing almost linearly. Neutral questions pulled Russian state sources in 11% of responses. Biased questions — phrased to suggest an answer — pulled 18%. Malicious questions, drafted to extract validation for a predetermined claim, pulled 24%. ChatGPT showed the steepest climb, citing Russian sources roughly three times more often under malicious framing than under neutral framing. Military-recruitment questions triggered the highest single-topic rate, at 24%; NATO questions hit 28.5% for both Grok and ChatGPT.

This is the finding that should worry anyone who has treated “just ask the AI, it’s neutral” as a reasonable verification habit. The models aren’t equally vulnerable regardless of how they’re asked. A user who already wants a claim confirmed — the exact posture of someone spreading it — gets a meaningfully higher hit rate on a source that shouldn’t be citable at all. The chatbot doesn’t just fail randomly. It fails in the direction the asker is already pushing.

How the Label Doesn’t Survive One Hop

Both studies found the same evasion mechanism, and it’s mundane rather than exotic: sanctioned outlets don’t get cited directly nearly as often as they get cited by proxy, through a third party that hasn’t been flagged.

ISD documented ChatGPT citing an RT article — the sanctioned Russian state broadcaster — not from RT’s own domain but as reposted by an Azerbaijani outlet, azerbaycan[.]com, with limited reach and no sanctions listing of its own. That repost surfaced across malicious prompts in Italian, Spanish, and German. Separately, DeepSeek cited VT Foreign Policy — the successor identity of Veterans Today — a site documented as a distribution channel for Storm-1516, a known Russian-aligned influence operation, twice returning four Russian-state-affiliated sources in a single response, the highest count ISD recorded. Grok’s failure mode was different again: rather than citing articles, it quoted individual X posts, including ones from RT journalists and pro-Russian influencer accounts directly.

The common thread across all three: whatever filtering these systems apply — and Gemini was the only one of the four ISD tested that visibly applied any, surfacing a safety notice on biased and malicious prompts, though without a source breakdown to go with it — the filtering appears to operate on the outlet’s own name, not the claim’s actual origin. Launder RT through an Azerbaijani reposter, or a Wagner-founded NGO through a “human rights” framing, and the sanctions label simply doesn’t travel with the content. This isn’t a sophisticated adversarial attack. It’s a repost.

What the Companies Said

Demos contacted all five companies whose models it tested. xAI and Google did not respond. Mistral drew a distinction between its raw model and its “Vibe Work” product, which it said includes built-in context scanning the tested version lacked. OpenAI noted that the specific model tested, GPT-4.1 Mini, was by publication time a retired, API-only version, and pointed to dedicated teams working to disrupt influence operations. Anthropic said its enforcement systems and threat-intelligence teams work to mitigate disinformation spread. None of these responses is unreasonable on its own terms — a retired model genuinely isn’t the current product, and every company here does maintain some counter-influence-operations capacity. But none of them changes what the study measured: current-generation, publicly available chatbots, tested under realistic conditions, citing a Wagner-founded propaganda front as if it were a legitimate source in roughly one response out of six.

Key Findings

  • Two independent studies, two different methods, converging results. Demos found 16.6% of responses engaged favorably with a single sanctioned outlet (r-FBI); ISD Germany found roughly 18% of all responses across a broader query set cited Russian state-linked sources of any kind.
  • Nearly a third of Demos’s responses were a quieter failure than outright endorsement. 30.9% addressed the topic without disclosing that the underlying source was a sanctioned propaganda entity — undisclosed provenance, not stated falsehood, was the largest single outcome after correct rejection.
  • Prompt framing changed the citation rate by more than double. ISD measured 11% on neutral questions, 18% on biased questions, and 24% on questions explicitly demanding a fixed conclusion — the failure gets worse exactly when a user is already primed to believe the claim.
  • The sanctions label doesn’t survive a repost. RT content cited via an Azerbaijani outlet, Wagner-linked material cited via VT Foreign Policy, and direct X-post citations from RT journalists all bypassed whatever source-level filtering these models apply.
  • Gemini was the only model showing visible safety guardrails, surfacing a warning on biased and malicious prompts — but without a corresponding breakdown of which sources it had actually cited.

Implications

The influence-operations literature has a name for hiding a message’s true origin behind a credible-looking intermediary: source laundering. What these two studies document is source laundering with a new, uniquely efficient intermediary — the chatbot itself. A propaganda outlet no longer needs to build its own audience, evade its own platform bans, or survive its own scrutiny. It needs one repost on an obscure site, or a values-neutral-sounding name like “Foundation to Battle Injustice,” and a research assistant used by hundreds of millions of people will, often enough to matter, do the laundering for it — converting a sanctioned entity’s claim into an answer that arrives with the chatbot’s own credibility attached, not the source’s.

This is exactly the terrain Decipon’s Influence Tactics Protocol is built to score: not whether a claim is true, but whether its apparent credibility was earned or borrowed. A chatbot answer doesn’t read as a citation of RT. It reads as the chatbot’s own synthesis. That’s a stronger trust transfer than a shared link ever was, and it’s happening at the exact moment users are least likely to fact-check — when they asked a “neutral” AI assistant specifically because they didn’t want to wade through partisan sources themselves.

The prompt-framing finding is the part worth internalizing personally, not just architecturally. If you already suspect a claim is true and you ask an AI chatbot to confirm it, you have measurably worse odds of getting a clean answer than if you’d asked neutrally. The chatbot isn’t a truth oracle that treats all inputs equally. It’s a system whose failure rate rises to meet your own bias, which makes it a worse tool for confirmation than it is for genuine inquiry — a distinction most users don’t currently know to make.

Conclusion

Sixteen-point-six percent and eighteen percent are not the same number, and they shouldn’t be reported as one. What they are is two independent teams, using two different methods, on five companies’ worth of current chatbots, arriving at the same order of magnitude for the same failure: a sanctioned Russian propaganda apparatus, reachable through five mainstream AI assistants, roughly one time in six to one time in five. The label doesn’t survive a repost. The filtering doesn’t survive a name change. And the odds get worse, not better, exactly when the person asking already wants to believe what they’re about to be told.

Sources: