AI chatbots write fluent, confident answers, and that is exactly the problem. A wrong answer reads just like a right one. This guide gives you a repeatable 7-step routine and shows how it would have caught four well-documented AI mistakes, from a fined law firm to a government report that led to a partial refund.
Why do AI chatbots get facts wrong?
Chatbots generate text that looks like a plausible answer. They do not “look up” facts the way a database does, so they can produce details that sound right but are invented. This tendency to fabricate information is commonly called hallucination. Some chatbots can search the web, which helps, but even then they can misread or misquote a source. The safe habit is to treat every AI answer as a draft to be checked, not a finished fact.
What do real AI mistakes teach us?
These four cases are well documented, and each one maps to a step in the routine below.
Case 1: The lawyers who cited cases that did not exist (2023)
In Mata v. Avianca, a New York attorney used ChatGPT for legal research. The brief cited court decisions, such as Varghese v. China Southern Airlines, that opposing counsel could not find anywhere. When the lawyer asked the chatbot whether one of the cases was real, it said yes and claimed the cases could be found in reputable legal databases. The judge later fined the lawyers and their firm $5,000 and found they had continued to stand by the fake opinions after the court questioned them.
What would have caught it: Step 4 (confirm the citation exists in a real legal database) and Step 6 (never use the chatbot to verify itself).
Case 2: The airline chatbot that contradicted the airline’s own policy (2024)
In Moffatt v. Air Canada, a customer asked the airline’s website chatbot about bereavement fares after a family death. The chatbot said he could claim the reduced fare retroactively within 90 days, but the airline’s own linked policy page said retroactive claims were not allowed. A British Columbia tribunal rejected the airline’s argument that the chatbot was responsible for its own statements, noted that the airline is responsible for all information on its website, and ordered it to pay the customer about CA$812 in damages, interest and fees.
What would have caught it: Step 3 (compare with the primary source, here the official policy page the chatbot itself linked to). It also shows why businesses must check what their own chatbots say.
Case 3: The “first picture of an exoplanet” claim (2023)
In a promotional demo for Google’s Bard chatbot, the AI said the James Webb Space Telescope took the very first pictures of a planet outside our solar system. NASA’s records show that the European Southern Observatory’s Very Large Telescope captured the first such image in 2004. CNN reported that Alphabet’s shares fell about 7.7% that day, wiping roughly $100 billion off its market value.
What would have caught it: Step 1 and Step 5. “First,” “only” and “largest” claims are easy to isolate, and a quick check against a second source such as NASA exposes them.
Case 4: The consulting report with fake references (2025)
Deloitte Australia agreed to partially refund the Australian government for a AU$440,000 report that contained errors, including a fabricated quote attributed to a federal court judgment and references to academic papers that do not exist. A university researcher, Chris Rudge, identified the problems and said he found up to 20 errors. A revised version of the report disclosed that a generative AI language system had been used in preparing it.
What would have caught it: Step 4 (check every quote and reference against the original) and Step 7 (record what was verified before delivery). A professional reviewer checking even a sample of the footnotes would have found the problem.
Author’s note: [PLACEHOLDER: add 2–3 sentences from your own experience, for example a time a chatbot gave you a wrong date, link or statistic, what gave it away, and how long the check took. First-hand detail is the strongest trust signal on this page.]
What is the 7-step verification routine?
Use this every time a chatbot answer will affect a decision, a document or something you plan to share.
Step |
Action |
Time |
|---|---|---|
1 |
Isolate the claims |
1–2 min |
2 |
Rank them by risk |
1 min |
3 |
Find the primary source |
2–5 min |
4 |
Verify every citation, quote and link |
2–5 min |
5 |
Cross-check with an independent source |
2–5 min |
6 |
Stress-test the chatbot |
2 min |
7 |
Decide and record |
1 min |
Step 1: How do you isolate the claims?
Read the answer and underline anything that can be proven true or false: names, dates, numbers, quotes, “first/only/most” statements, legal or medical rules, prices, and cited sources. Ignore opinions and general advice for now. A ten-sentence answer usually contains three to six checkable claims.
Step 2: How do you rank claims by risk?
Decide how much a mistake would cost. A wrong restaurant opening hour is minor. A wrong drug interaction, tax rule or contract clause is not. Use the time budget below to decide how deep to go.
Risk level |
Examples |
Minimum check |
|---|---|---|
Low |
Trivia, brainstorming, casual recommendations |
One quick search (about 1 minute) |
Medium |
Work documents, blog posts, school or business research, product comparisons |
Primary source plus one more source (5–10 minutes) |
High |
Medical, legal, financial, safety, anything published under your name or sent to a client |
Full routine plus a qualified professional (15+ minutes) |
Step 3: What counts as a primary source?
A primary source is where the information originates, not someone summarising it. Examples: the official policy or pricing page, the law or court decision itself, the original study, a government or regulator website, the company’s own filing. Search for the claim and open the original document. In the Air Canada case, the correct answer was on the airline’s own policy page.
Step 4: How do you verify citations, quotes and links?
- Click every link. Does it open? Does the page really say what the chatbot claims?
- Search the exact title of any book, paper or article in a search engine, library catalogue or scholarly database. If nothing appears, treat it as fake.
- Check authors, journal and year. Fabricated references often combine a real author with an invented title or a real journal with a wrong volume.
- Check quotes against the original text. Paste a distinctive phrase into a search engine in quotation marks.
- For legal cases, look up the case name and citation in an official court site or legal database and read the actual decision.
A citation that “looks right” in format proves nothing. The chatbot in the Avianca case produced case names, citations and summaries in perfect legal style.
Step 5: How do you cross-check with an independent source?
Find at least one other reliable source that does not simply copy the first. Good choices: official bodies, major news organisations, peer-reviewed research, recognised fact-checking organisations, and subject-matter experts. Be wary of pages that repeat the same wording. They may themselves be AI-generated or copied from one another. If two reliable sources disagree, say so and report the range instead of picking one.
Step 6: How do you stress-test the chatbot?
- Ask it to show its work: “Which exact source supports each claim? Say if you are unsure.”
- Re-ask in a new chat with different wording. If the answer changes, treat the unstable parts as suspect.
- Ask for the opposite view or for reasons the answer might be wrong.
- Do not ask “Are you sure?” or “Is this real?” and treat a yes as proof. In the Avianca case, the chatbot confirmed its own fake cases when asked. Its confidence is not evidence.
Step 7: How do you decide and record the result?
Mark each claim as Verified, Unverified or Wrong. Remove or rewrite anything you could not verify, add the date you checked, and for high-risk content get a qualified professional to sign off. Keep a short log so you or a colleague can trace what was checked. Copy this template:
Claim |
Risk (L/M/H) |
Source checked (link) |
Second source |
Status |
Date |
|---|---|---|---|---|---|
[e.g., “Policy allows refunds within 90 days”] |
H |
[official policy URL] |
[customer service confirmation] |
Verified / Unverified / Wrong |
[date] |
What are the most common red flags in AI answers?
Red flag |
Why it matters |
What to do |
|---|---|---|
Very specific numbers, dates or quotes with no source |
Precision can be invented |
Find the original document |
“First,” “only,” “largest” or “never” claims |
Superlatives are easy to get wrong (see the Bard case) |
Check against an authoritative record |
Citations that cannot be found |
Possible fabricated references (see Cases 1 and 4) |
Search the exact title; drop it if missing |
Links that lead to errors or unrelated pages |
The link may be guessed rather than real |
Search for the source yourself |
Policy, price, legal or medical rules stated as certain |
These change by place and date (see Air Canada) |
Confirm with the official page or a professional |
Recent events or “latest” information |
The model may be out of date unless it searched the web |
Check a current, dated source |
Answers that change when you re-ask |
Signals low reliability |
Verify manually before using |
How do you fact-check an AI answer in 60 seconds?
When you are short on time and the stakes are low, use this shortcut: underline the one claim that matters most, open the official or original source for it, and confirm it matches. If you cannot confirm it in a minute, do not use it. Save the full 7 steps for anything that will be published, sent to a client or used to make a decision.
Your 7-step checklist
- ☐ I listed the checkable claims (names, numbers, dates, quotes, sources)
- ☐ I ranked them low, medium or high risk
- ☐ I found the primary source for the important claims
- ☐ I opened every link and confirmed every citation and quote exists
- ☐ I cross-checked with at least one independent reliable source
- ☐ I re-asked in a new chat and did not rely on “are you sure?”
- ☐ I recorded the result and removed anything unverified
Frequently asked questions
Why do AI chatbots make up facts?
Chatbots generate text that sounds plausible based on patterns, not by checking a verified fact database. When they lack reliable information they can still produce confident-sounding but false details, often called hallucinations.
Can I trust an AI chatbot if it gives sources?
Not automatically. Sources can be wrong, misquoted, outdated or entirely invented. Open each source and confirm that it exists and actually says what the chatbot claims.
Is it enough to ask the chatbot whether its answer is correct?
No. A chatbot can confirm its own mistakes with equal confidence. Verification must come from an independent source such as an official website, a court record or a published study.
How can I tell if a citation is fake?
Search for the exact title, author and publication in a search engine or library database. If you cannot find the work, or the author, journal or year do not match, treat the citation as unreliable and do not use it.
Are chatbots with web search more accurate?
They can be more current, but they can still misread, misquote or choose weak sources. Check the pages they cite, just as you would for any other answer.
How long should fact-checking take?
About a minute for low-stakes questions, 5 to 10 minutes for work or publishing, and longer, plus expert review, for medical, legal or financial decisions.
Sources and further reading
- LawNext: Court imposes sanctions on lawyers who filed bogus cases after relying on ChatGPT (Mata v. Avianca)
- NBC News: ChatGPT cited “bogus” cases for a New York federal court filing
- Mondaq: Airline ordered to compensate a BC man because its chatbot provided inaccurate information (Moffatt v. Air Canada, 2024 BCCRT 149)
- NPR: Google shares drop $100 billion after its new AI chatbot makes a mistake
- Associated Press: Deloitte to partially refund Australian government for report with apparent AI-generated errors








