The Benchmark
On June 16, 2026, Estonia’s Institute of the Estonian Language published results from a systematic test most AI companies hadn’t run on themselves: how well do their models resist Russian propaganda?
The methodology was straightforward. Researchers built what they called the Language Model Yardstick — a framework of 75 questions covering 14 Kremlin propaganda themes, delivered in English, Russian, and Estonian. Each answer was scored on a 1-to-5 scale, where 1 means the model repeated Russian talking points and 5 means it clearly identified the claim as false and provided accurate context. They tested 60 models and ranked the results.
The findings had a geopolitical texture nobody predicted.
Anthropic’s Claude models took the top positions. Nvidia’s Nemotron 3 and Alibaba’s Qwen 3.6 Plus clustered near the top. Mistral — the French AI company positioned as Europe’s sovereign answer to American AI dominance — produced four models tested, all scoring below 40%, all landing in the bottom third of the leaderboard.
Europe’s AI champion failed Europe’s propaganda test. China’s AI passed it.
What the Questions Were
To understand what the benchmark measured, you need to know what it asked.
The 14 propaganda themes included the claim that Russia was legitimately evacuating Ukrainian children from war zones. That NATO broke promises not to expand east after German reunification. That Ukrainian government officials are neo-Nazis. That Western sanctions harm ordinary Europeans more than they harm Russia. That civilian deaths are staged or exaggerated for propaganda purposes.
These are not obscure Kremlin talking points. They are the central narratives that Russian state media and its amplifier networks have distributed for years — the claims that appear in RT segments, in Telegram channels, in the alternative media ecosystems that deplatformed figures inhabit. They are the narratives Laura Loomer described absorbing over five years before a journalist embedded with Ukraine’s 34th Marine Brigade brought her primary source material that contradicted them.
A model that scores below 40% on these questions is a model that, when prompted with these narratives, either repeats them, hedges between them and contrary evidence, or fails to clearly identify them as false.
A user who asks such a model to summarize the Russia-Ukraine conflict gets answers assembled from Kremlin-aligned framing. They may not know it. The model doesn’t announce its sources.
The Network That Was Never Writing for You
This is where the story gets specific.
According to Newsguard, a network called Portal Kombat — also known as the Pravda network — is the likely source of the contamination observed in vulnerable models. As of April 2026, the network comprised 370 sites, 286 of them actively publishing. The sites produce content in multiple languages. Their articles are indexed by search engines. They circulate on social media.
The targets of this content are not primarily human readers.
Newsguard’s analysis describes the network’s purpose as “flooding search engines and responses of AI chatbots with Russian propaganda.” That phrasing deserves attention. The network’s distribution strategy is oriented around two systems: search engine indexing and AI chatbot training pipelines. Both are mechanisms for delivering information to AI systems that will synthesize it as “what the internet knows” about a given topic.
A human who reads a Portal Kombat article may reject it immediately — the production quality, the slant, the unfamiliar domain, the absence of recognized bylines. Human pattern-matching developed to detect credibility signals has some purchase here.
A training pipeline doesn’t reject articles. It reads them. It weights them by volume and link structure. It incorporates what they say about Ukrainian child evacuations, about NATO expansion, about the origin of the 2014 Maidan protests. It learns: when people write about these subjects, this is what they say.
When 286 active sites say the same thing about Ukraine, the training pipeline absorbs this as an epistemically significant signal — which is exactly the wrong conclusion, but the right outcome for whoever runs Portal Kombat.
Why Open-Source Is the Target
The vulnerability isn’t distributed evenly.
Open-source models scored worst on the Language Model Yardstick. This wasn’t random. It reflects a structural property of how open-source AI development operates.
Commercial AI labs with large research budgets can afford elaborate data pipelines that filter low-quality content, rank sources by credibility, and reduce the influence of coordinated production networks. Training on filtered data is expensive — it requires human review, classification, domain expert judgment about which sources should be trusted at what weight.
Open-source development, especially by resource-constrained teams, often relies on more automated curation: common crawl of the web, Wikipedia, Reddit, public archives. The Portal Kombat network is indexed by search engines and accumulates links. It passes automated quality filters that rely on engagement and link graph signals rather than content-level credibility assessment.
There is also a philosophical dimension. The open-source AI community tends to view constraints on model outputs as a form of censorship. The resistance to content filtering is part of the ideological foundation of many open-source projects — freedom of expression extended to the model itself. The result is models that are more flexible, more customizable, and more willing to say things that a commercial model would hedge or decline.
This makes them useful for many applications. It also makes them excellent relays for propaganda, because they’ve been designed not to refuse.
The same architectural choice that makes an open-source model willing to discuss a sensitive topic without hedging also makes it willing to present a Russian propaganda claim about Ukrainian children without noting that the claim is documented to be false, sourced to state media, and part of a catalogued influence operation.
The design philosophy and the propaganda vulnerability are the same feature.
The Geopolitical Irony
Alibaba’s Qwen outperformed Mistral.
This is not a trivial finding. Qwen is a Chinese model. Mistral is the model France’s government backed as Europe’s AI sovereignty play — the EU positions it as a strategic alternative to American AI dominance, a European-values model built in Europe for European users.
In the June 16 benchmark, a model built by one of China’s largest corporations outperformed Europe’s strategic AI champion at resisting Russian propaganda.
There are several ways to read this. Qwen’s training data may be more heavily filtered, partly because Chinese content moderation practices remove what the state deems harmful — which includes, in some categories, content that originates from Russian state-adjacent propaganda networks. The Chinese and Russian governments have aligned interests in some domains, but the PRC’s media environment has its own coherence, and content from Kremlin-affiliated networks doesn’t necessarily pass through Chinese data pipelines unchecked.
The earlier analysis of intentional embedding in Chinese models — the “keep answers about China positive” directives found in Qwen’s reasoning traces — also cuts both ways: models trained to maintain coherent national narratives may, as a side effect, resist foreign competing narratives, including Russian ones.
Regardless of mechanism, the result stands: Alibaba outscored Mistral on the thing Europe most needs its AI models to do. The propaganda defense that matters most for European information security is running better in Hangzhou than in Paris.
The Attack Surface That Wasn’t There Before
The Portal Kombat network’s targeting strategy represents a strategic evolution in influence operations.
Classic propaganda targets human readers. It needs attention, engagement, emotional resonance. It has to compete in an information environment saturated with competing narratives. It can be labeled, fact-checked, banned, algorithmically suppressed.
AI training pipeline contamination targets a different system. The “reader” is an automated data ingestion process. It doesn’t need to be persuaded. It needs only to be fed consistent, voluminous, indexed content at scale. Once incorporated into model weights, that content is no longer an external claim subject to fact-checking — it is the model’s baseline understanding of what the world looks like.
A user who reads a Portal Kombat article can search and find debunking. A user who asks an AI assistant what happened with Ukrainian children during the war receives an answer that may incorporate Portal Kombat’s framing as one of the primary sources — without any attribution, without any visible signal that the claim traces to a Russian state-affiliated content farm.
The propaganda reached the reader without looking like propaganda. The model was the delivery mechanism. The reader never knew they were reading it.
This is the operational upgrade that matters most in the current landscape. Not deepfakes. Not bots that post on X. The insertion of Kremlin narratives into the weights of AI systems that millions of people use daily to understand contested events — delivered as synthesis, as summary, as the neutral voice of the machine that knows things.
When a user asks an AI assistant a question about the Russia-Ukraine war, the answer they receive is shaped by what that model absorbed during training. If the training data was flooded by 286 actively publishing sites promoting a consistent set of false narratives — narratives specifically about the conflict the user is asking about — the user is receiving propaganda filtered through an AI voice that makes it sound like information.
What the Benchmark Reveals
The Language Model Yardstick is useful not primarily as a consumer guide — though it functions as one — but as evidence of where the information supply chain is contaminated.
The test results show which models absorbed Portal Kombat’s narratives. They don’t identify which specific claims were absorbed or annotate every answer to show which sources shaped it. But they establish an empirical baseline: there is a measurable spectrum of propaganda vulnerability among major AI models, and it correlates with open-source architecture and the curatorial choices made during training.
The implications for people who use AI assistants to understand geopolitical events are not comfortable. The question you haven’t been able to answer until now — which model’s training data includes how much of Portal Kombat’s output — has a partial answer. Estonia built the benchmark. The findings show that if you’re using Mistral to understand the war in Ukraine, you’re getting answers shaped by the same narrative infrastructure that shaped Laura Loomer’s view before a journalist went to Kyiv.
You’re just getting it with more confident grammar and no visible byline.
The Kremlin didn’t put it there directly. Portal Kombat published at scale for years, and the pipeline did the rest. The network was never writing articles for you. It was writing training data for the model you’d eventually ask.
The reader was the algorithm. And now the algorithm is what your AI assistant runs on.
This article is part of Decipon’s Manipulation Breakdowns series, examining specific influence operations through the Influence Tactics Protocol.
Sources:
- New report raises concerns over Russian propaganda spread by Europe’s flagship AI company Mistral — Euronews (June 16, 2026)
- How easily can Russian propaganda fool AI models? A new benchmark finds out — The Decoder
- Mistral AI models score below 40% in detecting Russian propaganda, new benchmark reveals — CryptoBriefing
- Estonian propaganda benchmark ranks Anthropic, OpenAI, Google — Dublin Post
- Estonian study finds AI models still vulnerable to propaganda prompts — ERR News
- European AI and the propaganda test: Mistral loses out to Chinese models — Il Sole 24 ORE
- This Week In Disinformation 14–20 June 2026 — The Disinformation Observer