Skip to main content

The Video Was Real

Manipulation Breakdowns · 7 min read · By D0

Introduction

In late July 2026, as Gen Z-led protests over paper leaks and exam fraud spread across Delhi, a video began circulating that showed Defence Minister Rajnath Singh threatening student protesters with “the full force of the state.” Days later, a second clip surfaced: Prime Minister Modi, in what looked like his own address, making a similar threat. A third showed Commerce Minister Piyush Goyal, caught by reporters outside Parliament, issuing what sounded like a direct warning to the demonstrators.

All three were fabrications. But they weren’t the kind of fabrication most people have been trained to look for. The footage was real — Modi’s actual address about cracking down on paper leaks, Goyal’s actual appearance before the press, Singh’s actual public remarks. What changed was the audio. Someone stripped the original soundtrack and replaced it with fabricated speech, timed to the movement of real mouths, reassembled into statements none of these men made. Delhi Police eventually traced the campaign to more than 400 Pakistan-linked accounts and took them offline. By then the clips had already done their circulating.

A few days apart, a separate and cruder operation ran a different play: an AI-generated video of Finance Minister Nirmala Sitharaman, entirely synthetic this time, appearing to endorse an investment scheme promising 70,000 rupees a day in returns on a 22,000 rupee stake. Government fact-checkers debunked it quickly. It’s worth holding both cases side by side, because they represent two different manipulation techniques wearing the same trust costume, and the first one is harder to catch than most people assume.

Two Different Forgeries, One Shared Target

Start with what the Sitharaman video is: a full synthesis. Nothing about it is real. The face, the voice, the setting, the words — all generated. This is the deepfake most audiences have been warned about, and it’s the one fact-checkers are best equipped to catch, because forensic tools built to detect synthetic media are trained precisely for this signature. Frame-level artifacts, unnatural blinking, audio that doesn’t quite match the acoustic signature of a real recording space — the tells exist, even if they’re getting fainter every model generation.

The Rajnath Singh and Modi clips are a different animal entirely, and they’re the more instructive case. Technically, this is closer to what researchers call a cheapfake — a manipulation that doesn’t require generative synthesis at all, just selective editing of real material. The video track is authentic. The visual forensics that catch synthetic faces find nothing wrong, because there is nothing synthetically wrong. Only the audio was replaced, and it was replaced carefully enough to sync against real lip movement, which modern AI voice-cloning and lip-sync tools now do without much friction.

This matters because most public debunking instinct has been trained on the wrong question. “Is this video AI-generated?” is the question forensic tools answer well. “Does this video show something that actually happened?” is the question that actually protects you, and it’s a harder one — because a video can pass every synthetic-media detector and still be a complete fabrication of what was said.

The Influence Tactics at Work

Three tactics did the load-bearing work here, and they’re worth naming precisely because they explain why real footage, altered, is more dangerous than synthetic footage from scratch.

  • Borrowed authenticity. A cheapfake inherits all the credibility markers of the original recording — the real setting, the real cadence of a real public appearance, the background noise of an actual press scrum outside Parliament. None of that has to be manufactured. It’s simply repurposed. Audiences who’ve been taught to distrust obviously synthetic content have no trained reflex against this, because nothing about the visual layer triggers suspicion.
  • Authority displacement. The manipulation doesn’t ask you to trust a stranger. It asks you to trust officials you already have an opinion about — a defence minister, a prime minister, a commerce minister — and it weaponizes whatever trust or distrust you already carry toward them. A supporter sees the clip and feels betrayed by a leader turning on young people. An opponent sees it and feels confirmed in exactly the suspicion they already held. The forgery doesn’t need to persuade anyone of something new. It needs to activate a belief that’s already loaded.
  • Timing against a live grievance. None of these clips appeared in a vacuum. They landed in the middle of an active, emotionally charged protest movement, when audiences were primed to look for evidence that officials were dismissive or hostile toward the protesters. A fabricated threat delivered during a real confrontation reads as confirmation, not as a claim that needs checking. Urgency suppresses verification, and a live protest supplies urgency for free.

Why the Scam Video Is the Easier Case

The Sitharaman investment-scam deepfake follows a more familiar and, frankly, more tractable pattern: manufactured authority-adjacent credibility, applied to fraud. “Finance Minister” is doing the same trust-transfer work here that “nurse” did for the AI-generated persona that spent thirteen months scamming a following before Instagram caught it — a real job title functioning as borrowed credibility for a claim that has nothing to do with the job itself. A finance minister endorsing a specific consumer product is, on its face, implausible; government officials don’t do celebrity endorsements for investment schemes, and once you know that pattern, the claim collapses on its own logic, independent of any forensic analysis of the footage.

The cheapfake threat videos don’t collapse on their own logic. A defence minister making a public statement about protesters is entirely plausible as an event. The fabrication lives entirely inside the specific words attributed to him, and there’s no structural implausibility to lean on — you have to actually check what he said, against a primary source, to catch it. That’s a much higher bar than “would a minister really do this,” and it’s the reason this category of manipulation is going to outlast the crude synthetic version. Full synthesis gets easier to catch as detection tools mature. Selective audio replacement over real footage doesn’t have the same forensic signature to catch, because there’s less to forensically examine — the fabrication is a splice, not a generation.

What Actually Verifies This

The corrective isn’t a better deepfake detector, though those help for the fully synthetic cases. It’s a different verification reflex, one built around a single question: does an official transcript or independent recording of this specific statement exist?

  • Trace the primary source. A real public statement by a sitting minister will have been captured by wire services, official government channels, or multiple independent journalists in the room. If a viral clip attributes specific words to someone and no other source captured those words, that absence is the tell — not any artifact in the video itself.
  • Check whether the clip’s content matches the clip’s stated context. The Modi cheapfake reused footage from an address about paper-leak enforcement and grafted on fabricated audio about threatening protesters. The mismatch between the known occasion for that speech and the claimed content of the fabricated version is detectable by anyone willing to look up what the actual address was about.
  • Treat plausibility as a low bar, not a high one. “A minister could plausibly say this” is true of almost any fabricated quote a manipulator would bother making. Plausibility filters out only the laziest forgeries. It does nothing against a forgery built specifically to sound like something the target might say — which is every forgery worth making.

Conclusion

The instinct to ask “is this AI” is not wrong, but it’s increasingly incomplete. Forgers have already adapted past it, keeping the parts of a video that pass inspection — the real face, the real room, the real event — and fabricating only the part that does the persuading: the words. Four hundred accounts and two governments’ worth of fact-checkers were needed to run down one week of this in Delhi. The tools that will actually scale against it aren’t better detectors for synthetic pixels. They’re a public habit of asking where a quote came from before believing what it did to you.