The unsettling thing about a chatbot's mistakes is how good they look. There's no broken grammar, no obvious tell β a fabricated statistic arrives in the same calm, well-formatted sentence as a correct one. That's not a bug someone forgot to fix. As OpenAI researchers argued in a 2025 paper, hallucinations fall out of how these models are trained and graded: they're rewarded for producing a plausible answer rather than for admitting they don't know, so when they're uncertain they guess β fluently.
You don't need to understand the math to protect yourself from it. What you need is a habit: a short sequence you run on any answer that actually matters. The steps below go from fastest to most involved, and most bad answers fall apart by step three.
The 7-step check
Treat the answer as a claim, not a source
A chatbot doesn't retrieve a fact and hand it to you β it predicts the next words that sound right. That means a confident paragraph and a wrong one look identical from your side of the screen. So the first shift is mental: whatever the model says is a claim to be checked, the same as a stranger's tweet, not a citation you can lean on. Crucially, checking it does not mean asking the AI "are you sure?" β a model will often double down or invent a fresh justification for the same error. Verification has to come from outside the conversation.
Isolate the parts that can actually be wrong
Most answers are a mix of safe general knowledge and a few load-bearing specifics. Pull out the pieces that would change your decision if they were false: a number, a date, a name, a direct quote, a legal or dosage detail, a "studies show" statistic. General explanation ("compound interest grows faster over time") rarely needs checking; the specific figure attached to it ("at 7% your money doubles every 6 years") does. Fact-check the specifics, not the vibe.
Open every source it cites β don't just read the footnote
When a model gives you links, titles, or a study name, click through to the original before you trust any of it. This is where AI fails most visibly: a 2025 Tow Center study at Columbia Journalism Review tested eight AI search tools across 1,600 queries and found they cited sources incorrectly more than 60% of the time, sometimes inventing URLs outright. For a paper or article, search the exact title in Google Scholar or the publisher's site. No exact-title match is a red flag. And read what the source actually says β a real paper cited for a claim it never made is subtler and more dangerous than a fully invented one.
Know where models are weakest β and press there
Hallucinations aren't random; they cluster. Recent events after the model's training cutoff, niche or local facts, exact quotations, precise statistics, case law, and anything involving a specific named person are the high-risk zones. If your answer lives in one of those, raise your guard. A useful move is to ask the model directly for its training cutoff and whether it searched the live web for this answer β if it's working from memory on a question about last month, treat the specifics as unconfirmed by default.
Cross-check against something independent
Take the isolated claim from step 2 and run it through a channel the model doesn't control: a plain web search, the primary document itself, or a second, different AI model. Two independent models agreeing raises your confidence a little; a model agreeing with a primary source you opened yourself raises it a lot. For a claim that already circulates online, a fact-check search is often the fastest route β FAXTR checks 100+ fact-checking organizations in one query, so you can see whether a viral statistic the AI repeated already has a verdict.
Make the model expose its own uncertainty
Before you accept an answer, ask the model to argue against it: "What's the strongest case that this is wrong?" or "Which parts of this are you least confident about, and why?" Because models are tuned to be helpful, they'll often surface caveats and weak spots on request that they glossed over in the first pass. This won't catch everything β a model can be confidently wrong about its own confidence β but it reliably flags the claims most worth taking to an outside source.
For high-stakes answers, end with a human or a primary source
Medical, legal, financial, and safety decisions are exactly where a plausible-but-wrong answer does real damage, and exactly where models sound most authoritative. Use the chatbot to get oriented and to draft the questions worth asking β then confirm with the actual regulation, the official documentation, the manufacturer, or a qualified professional before you act. The goal isn't to never use AI; it's to never let a generated sentence be the last link in the chain when the stakes are real.
Red flags that a claim is invented
You won't run the full routine on every answer. These are the tells worth training yourself to notice β any one of them means the specific claim deserves a proper check before you use it.
- βA source that vanishes when you search its exact title, or a link that 404s or redirects to a homepage.
- βSuspiciously round numbers and specific-sounding statistics with no traceable origin.
- βReal authors paired with a paper or journal they never actually published in.
- βUnwavering confidence on a question about recent events the model may not have seen.
- βThe model "correcting" itself into a new, equally confident answer each time you push back.
- βQuotes attributed to a named person that you can't find anywhere outside the chat.
Why "ask the AI to double-check" doesn't count
It's tempting to reply "is that accurate?" and take the reassurance. But the model that produced the error is the last thing that can reliably catch it β it has no separate store of ground truth to consult, so it'll often restate the mistake with new conviction or, just as unhelpfully, cave and "correct" a fact that was right. A verification step only counts when the evidence comes from somewhere the model doesn't control: the primary document, an independent search, a second tool, or a person who knows the field.
What this guide is not
None of this says AI answers are worthless β they're often a fast, useful way to get oriented, draft, or summarize. The point is calibration: knowing which parts of an answer have earned your trust and which are still just plausible text. Run the check where it matters, skip it where it doesn't, and never let a generated sentence be the last link in the chain when something real depends on it.
Check a claim in seconds
Repeating a statistic an AI gave you? FAXTR searches 100+ fact-checking organizations in one query β free, no login required.