Text-to-video models predict pixels; they don't simulate the world. Watch for shadows that fall the wrong way, water and hair that move against their own momentum, and cars or crowds that drift instead of pushing off the ground. A single reflection that doesn't track the object it belongs to is often the tell.
Guide Β· Media Literacy
How to Tell If a Video Is AI-Generated
Text-to-video models like Sora, Veo, and Kling can now produce clips that look like ordinary phone footage. Here's a practical way to test one β the visual tells worth checking, the provenance signals most people ignore, and the detectors that help without being trusted blindly.
First, a distinction that saves a lot of confusion. A fully synthetic video is generated from a text prompt β nobody filmed it; the whole scene was invented by a model. A deepfake takes real footage and swaps or alters a face. They fail in different ways, so it helps to know which one you're looking at. This guide is mostly about the first kind; for the second, see our deepfake detection guide.
The honest starting point, per the University of Queensland's Sora-era verification guide: the gap between real and generated has narrowed to the point where no single trick is reliable on its own. What still works is stacking checks β source, then eyes, then provenance, then tools β so that a fake has to survive all of them.
1. Start with the source, not the pixels
Before you squint at a single frame, ask where the clip came from. Who posted it first? Does that account have a history, or was it created last week? Is the same footage being reported by any outlet that does its own verification, or does every copy trace back to one anonymous repost?
Most misleading video isn't exposed by forensics β it's exposed by context. A clip with no credible origin, spreading only through accounts that never explain where they got it, deserves suspicion long before you analyze a pixel. This is the cheapest check and the one that catches the most.
2. Slow it down and watch the seams
Download the clip (or use your player's frame-step keys) and step through it slowly. Generators are good at the impression of a scene and bad at holding it together over time and at the edges. These are the places they tend to break:
The old giveaways haven't fully gone away. Count fingers across a few frames β synthetic hands still gain or lose one during fast motion. Teeth can merge or shift count between smiles. Pay special attention to boundaries: where a hand grips a cup, where hair meets a collar, where fingers cross. Edges are where the model has to commit, and where it slips.
Signs, license plates, book spines, phone screens, brand logos β generative models render these as plausible-looking gibberish that dissolves if you pause on it. If a street sign is readable in one frame and scrambled in the next, that instability is a strong signal the whole scene was generated.
Sora 2 and similar tools generate sound alongside picture, and it often arrives suspiciously polished: no room reverb, no background hum, no incidental noise like footsteps or fabric. Real phone footage is messy. A studio-clean soundtrack under handheld-looking video is a mismatch worth questioning.
Synthetic faces still blink at odd intervals and glance in directions the action doesn't call for. On a genuine smile, the skin around the eyes creases (the Duchenne marker); AI faces often smile with the mouth alone. Skin can also read waxy or too symmetrical in close-up. None of these is proof on its own β but two or three together should raise your guard.
3. Look for provenance signals
A growing share of AI output carries a signed record of how it was made. C2PA Content Credentials embed a tamper-evident manifest β which tool produced the file, when, and what edits followed. You can inspect it by dropping the file into the official Content Credentials Verify tool. Separately, Google's SynthID stamps an invisible watermark into video, audio, and images from many major models; as of 2026, OpenAI, Google, and others have aligned on pairing SynthID watermarking with C2PA metadata.
When a credential is present, it's strong evidence of origin. The catch, as the Global Investigative Journalism Network stresses, is the reverse: a screenshot, a re-encode, or a platform upload usually strips this data out. So use provenance to confirm, never to clear β a clip with no credentials is simply unlabeled, not verified.
4. Run it through a detector β and read the score carefully
Automated detectors give you a second opinion. Free and widely used ones include Hive's AI-generated content detector and Deepware Scanner for deepfakes. Journalists lean on the InVID-WeVerify plugin, which extracts keyframes and bundles several verification tools in one place. Some research systems, like Intel's FakeCatcher, look for the faint pulse of blood flow in skin (rPPG) that generated faces don't reproduce.
Just don't hand any of them the final say. Detectors trail the generators that beat them β a tool trained before the latest model release can be fooled by clips it's never seen β and they throw false positives on ordinary footage. Run more than one, and treat a lopsided, agreeing result as a lead you still confirm with your own eyes and the source check.
5. Reverse-search a few keyframes
Not every "fake" video is generated β plenty are real clips ripped from their original context and recaptioned. Pull a few clear frames (InVID-WeVerify extracts them for you, or just pause and screenshot) and run them through reverse image search. If the footage turns up in an older post about a different event, you've caught a miscaptioning, not a generation. Our reverse image search guide walks through which engine to reach for and how to find the earliest copy.
Where this gets hard β and what to do
The clip is short, low-res, and re-compressed
Social platforms crush video, and compression hides exactly the artifacts you're hunting for. A 6-second, 480p repost gives both your eyes and detection tools little to work with. Try to find a longer, higher-quality version before judging β and if you can't, weight the source and provenance checks more heavily than the pixel inspection.
A detector says '92% AI' (or '92% real')
Treat any single detector score as one weak vote, not a verdict. These models produce false positives on genuine footage and miss newer generators they weren't trained on. Run two or three different tools, and let a confident agreement β plus your own frame inspection β carry more weight than one number.
There are no content credentials
A missing C2PA manifest or SynthID watermark does not mean a video is real. Metadata is routinely stripped when a file is screenshotted, re-encoded, or uploaded to a platform. Provenance can confirm origin when present; its absence proves nothing either way.
It's a real person's face on a generated body
That's a face-swap deepfake, a different problem from a fully synthetic clip. The face may pass while the scene around it fails, or vice versa. When a specific real person is involved, lean on our deepfake guide and, above all, on finding the original source rather than judging the face alone.
The habit that beats any single tool
No detector, watermark, or visual tell is reliable alone, and each new model release chips away at whichever one you trusted most. What holds up is the stack: establish the source, watch the footage closely, check for provenance, corroborate with a couple of tools, and reverse-search the frames. A genuine clip clears all five easily. A generated one has to fool every layer β and it rarely does. When you can't resolve it, the safe move is the same as always: don't share it as fact.
Already have a claim, not just a clip?
If a video is attached to a specific claim that's going around, someone may have checked it already. FAXTR searches 100+ fact-checking organizations across 11 languages in one box β free, no login β so you can see whether that claim already has a published verdict.
Go to the verifier β