The most reliable way to check whether a video was generated by AI in September 2026 is to stop examining the video.
That sounds like giving up. It is closer to the opposite. Every signal that lives inside the file, the visible watermark, the embedded provenance metadata, the artifacts a detector looks for, has turned out to be either removable, routinely stripped in ordinary distribution, or quietly dependent on the first two. What survives is the part nobody thinks of as verification: who posted it, what it claims, and whether the claim holds up anywhere else.
Here is what changed, and which of the four things people believe about spotting AI video still hold.
What is actually generating this footage now
The lineup turned over almost completely in the last eighteen months, which matters because most advice about spotting AI video was written against models that no longer exist.
| Model | Maker | Status, September 2026 | Provenance marking |
|---|---|---|---|
| Veo 3.1 | Stable since October 15, 2025; 4K and native synchronized audio | SynthID watermark in both tracks | |
| Sora 2 | OpenAI | App shut down April 26, 2026; API ends September 24, 2026 | Visible moving watermark plus C2PA metadata |
| Seedance 2.5 | ByteDance | Released July 2026; 30-second native audio-video, up to 50 reference inputs | No public commitment we could verify |
| Gen-4 | Runway | Released 2025; character and scene consistency | No public commitment we could verify |
| LTX-2 | Lightricks | Released October 2025; open source, built-in audio | Open weights, so no marking can be enforced |
Two entries in that last column say we could not verify a commitment. That means exactly what it says: we did not find a published statement, not that none exists. The open-source row is different and more consequential. When the weights are public, any marking step is something the operator can simply remove from the pipeline, and no policy fixes that.
Note the Sora row, because it undercuts a lot of 2025-era advice. OpenAI announced the discontinuation in March 2026, the consumer app closed in April, and the API shuts down on September 24, 2026. The rotating Sora watermark that everyone learned to look for is about to stop appearing on new material entirely.
Claim one: if it is AI, it will have a watermark on it
Sora 2 shipped with a visible, moving watermark specifically to make its output identifiable. Within days of the launch, third-party tools that removed it were widespread. By early October 2025 the removers were common enough to be unremarkable, which is roughly the whole lifespan of that defense.
Google took the harder route. SynthID is imperceptible rather than visible, embedded into the pixels and the audio, and it survives more handling than a logo in the corner. Google's own SynthID documentation is clear about the boundary, though: it marks content from Google's models. Veo, Imagen, Lyria, Gemini. A clip from any other generator carries nothing for it to find.
So the absence of a watermark tells you close to nothing. It is consistent with a real video, with a video from a model that never marked its output, with an open-weights model whose operator skipped the step, and with a marked video that someone ran through a remover. Four very different situations, one identical appearance.
Claim two: Content Credentials will tell me where it came from
C2PA Content Credentials are cryptographically signed provenance records, and the adoption list is genuinely impressive. OpenAI, Google, Adobe, Microsoft and Stability on the generation side. Sony, Leica, Canon and Nikon in camera firmware. Google Search, YouTube, Meta, TikTok, LinkedIn, Pinterest and X on the distribution side.
The gap is between adopting the standard and preserving the data. Credentials are metadata, and ordinary web infrastructure re-compresses and rewrites files constantly. A credential that does not survive a platform's upload pipeline is not there when a reader goes looking, which is why Adobe built cloud-side records under the name Durable Content Credentials as a fallback for exactly this failure.
The standard's own limitation is sharper than the technical one, and it is the part most explanations skip:
Absence of credentials proves nothing about authenticity. Only their presence verifies documented provenance. Summary of the C2PA model, per the Content Credentials documentation
Presence is not a clean win either. A valid signature attests to origin and integrity, not truth, so a staged or misleading video can carry a perfectly valid credential. And the signing chain has been broken in practice twice that is publicly documented: a Nikon Z6III firmware flaw in August 2025 that allowed authentic and inauthentic content to be combined under a valid signature, and an Android exploit disclosed in August 2026 where root access could instruct the system to sign arbitrary data. Google classified the second one as "won't fix (infeasible)."
Claim three: I will just run it through an AI video detector
This is the claim with the most uncomfortable evidence behind it. A benchmark study called RobustSora, first posted in December 2025 and revised in May 2026, built a de-watermarked test set specifically to find out what video detectors are keying on. The finding is that a meaningful share of their performance rides on the watermark rather than on generation artifacts.
Watermark manipulation induces accuracy changes of −9.4 to +1.6 pp (mean 6.6 pp; p<0.01 for 7/10 models on each task). RobustSora: De-Watermarked Benchmark for Robust AI-Generated Video Detection, arXiv, revised May 2026
Read that alongside claim one and the circularity is hard to miss. Remove the watermark, and the detector built to catch unwatermarked fakes gets measurably worse at its job. Add a watermark to genuine footage, and some models move the other way. The tool most people reach for when the watermark is gone is partly the same tool as the watermark.
Detection research is also chasing a moving target. Each new generation of models narrows the perceptual gap, and detectors trained on last year's outputs degrade against this year's. This is not a video-specific problem, and it is the same reason different detectors disagree about the same text, and why false positives cluster around particular kinds of input. A probabilistic classifier gives you a signal, never a verdict.
We should be direct about our own position here. Detecting AI is a text detector. Our detection stack analyses written language, and we do not offer a video detector, so nothing in this article is a pitch for one. Anyone selling certainty about a video file is selling something the current research does not support.
Claim four: you can spot it if you look closely enough
In February 2026 a clip of Tom Cruise and Brad Pitt fighting on a rooftop went viral. It was made with ByteDance's Seedance 2.0, and the response was not a wave of viewers pointing out the tells. It was the Motion Picture Association issuing a statement.
In a single day, the Chinese AI service Seedance 2.0 has engaged in unauthorized use of U.S. copyrighted works on a massive scale. Motion Picture Association, quoted by Variety, February 2026
Disney alleged the model had been trained on a pirated library of its characters. ByteDance pledged fixes within days, and by August 2026 the MPA and ByteDance had reached an agreement on guardrails. What did not happen at any point was mass recognition that the footage was synthetic. The dispute was about training data and likeness rights, because the realism was not in question.
The failure runs in the other direction too, and this is the newer problem. Once an audience knows convincing fakes are cheap, real footage starts getting accused. A rule of thumb that produces confident judgments in both directions is not a rule of thumb, it is a coin flip with extra steps. Whether a clip is generated is also a separate question from who holds rights to what is in it, which is its own growing mess around how platforms police copyright at scale.
The claim that survived: check the audio, then check the story
One piece of received wisdom did hold up, and it is the least glamorous one.
Start with the audio track, because it is now a separate and better-instrumented surface. The current generation of models produces synchronized speech rather than silent clips or generic sound effects, and Veo 3.1's native audio is the reason so much 2026 output arrives with someone talking. Google's Gemini app will check an upload of up to 100 MB and 90 seconds and report the two tracks independently, which it announced in December 2025. The output looks like this:
SynthID detected within the audio between 10-20 secs. No SynthID detected in the visuals. Example response from the Gemini app's SynthID check, Google
A positive result there is real information. A negative one only rules out Google's models, so treat it as one reading rather than an answer.
Then verify the claim instead of the pixels. Almost every AI video that causes harm is doing so because of something it asserts: a person said a thing, an event happened, a product works a certain way. That assertion can be checked against sources that exist outside the file, and unlike the file, it cannot be re-encoded away. Our fact checker works on the text of a claim for exactly this reason, and the reporting around the Seedance clip is a good illustration: what settled that story was journalism about training data and studio agreements, not frame analysis.
Provenance, when it is present, is still the strongest single signal available. It is just not present often enough to build a workflow on. So the order that actually works in September 2026 runs backwards from the intuition: identify the claim, check whether it is corroborated anywhere, run the audio through a SynthID check if it might be a Google model, look for Content Credentials and take their presence seriously and their absence as nothing at all, and treat any detector score as a probability that got worse the moment someone touched the watermark.
If that feels like less certainty than you wanted, it is an accurate reflection of the tooling. The honest version of video verification in 2026 is a set of weak signals combined carefully, and the people most likely to be fooled are the ones who found one strong signal and stopped.
