Why AI Chatbots Make Things Up, and How to Catch It
What NIST and OpenAI say causes AI "hallucinations," why confident answers can still be wrong, and practical checks before you use AI output at work.
Anyone who uses an AI chatbot at work for long enough will eventually get an answer that sounds completely sure of itself and is simply wrong: a statistic that does not exist, a policy that was never written, a book with the wrong author. This is usually called a "hallucination." It is not a rare glitch. Both the US government's AI risk guidance and the company behind ChatGPT describe it as a built-in tendency of how these systems work. Understanding why it happens makes it much easier to catch.
What the experts call it
The National Institute of Standards and Technology (NIST) uses the term confabulation in its Generative AI Profile, published in July 2024. NIST defines it as the production of confidently stated but erroneous or false content, which is colloquially known as hallucinations or fabrications. NIST's definition also includes outputs that drift away from the prompt or the material you provided, and outputs that contradict something the system said earlier in the same conversation.
OpenAI, in a September 2025 explainer on its own research, describes hallucinations as plausible but false statements generated by language models, and states that ChatGPT also hallucinates. It gives an example: when a widely used chatbot was asked for the title of a specific researcher's PhD dissertation, it confidently produced three different answers, and none of them was correct.
Why it happens: prediction, not lookup
A chatbot does not look facts up in a database by default. NIST explains that generative models produce outputs that approximate the statistical patterns of their training data; for example, large language models predict the next word or piece of a word in a sentence. That process often produces accurate text, but NIST notes it can also produce text that is factually wrong or internally inconsistent. NIST calls confabulations a natural result of the way these models are designed.
OpenAI's explainer adds a useful distinction. Models rarely make spelling mistakes or leave parentheses unmatched, because those follow consistent patterns in the training text. But arbitrary, rarely mentioned facts, such as a particular person's birthday, cannot be predicted from patterns, so that is where false answers tend to appear.
NIST also flags where the risk is highest: open-ended prompts that ask for long answers, and subjects that require specialized expertise.
Why it sounds so confident
OpenAI argues that part of the problem is how models are graded. Most evaluations score models on accuracy alone, which works like a multiple-choice test with no penalty for guessing. If a model does not know someone's birthday and guesses a date, it has a 1 in 365 chance of being right. If it admits it does not know, it is guaranteed zero points. Over thousands of questions, the guessing model scores better, so models learn to guess.
OpenAI shows the trade-off with results from one of its factual question tests:
| Model | Declined to answer | Correct | Wrong |
|---|---|---|---|
| gpt-5-thinking-mini | 52% | 22% | 26% |
| OpenAI o4-mini | 1% | 24% | 75% |
The older model was correct slightly more often, but it almost never declined, so it was wrong almost three times as often (75 percent versus 26 percent). The lesson for users is that a model that always has an answer is not necessarily a more reliable one.
NIST warns about a second effect. AI outputs can include confabulated logic or citations that appear to justify the answer, and a model may lay out reasoning steps even when the final answer is wrong. That polished explanation can make people trust it more than they should. NIST also describes automation bias, the tendency to defer too much to automated systems, which makes confabulation risks worse.
Where this bites at work
The riskiest uses are the ones where a wrong detail travels:
- Facts, figures and dates in a report or slide deck.
- Quotes, citations, links and references, which may not exist at all.
- Policy, legal, tax or medical statements presented as rules.
- Summaries of a long document, where a detail can be invented or a point reversed.
- Anything about specific people, small companies or niche topics, where training data is thin.
NIST specifically notes that confabulation risks deserve close monitoring when AI is used in consequential decisions.
A practical checking routine
None of these steps is complicated, and together they catch most problems.
- Treat every specific fact as unverified. Numbers, names, dates, quotes and titles should be checked against a source you trust before they leave your hands.
- Ask for sources, then open them. A citation is only useful once you have confirmed it exists and says what the AI claims. Fabricated references are a known failure.
- Give it the material. When the answer should come from a document, provide the document and ask it to answer only from that text. Then spot-check the answer against the original, since NIST notes outputs can still diverge from the input.
- Invite uncertainty. Tell the tool it is fine to say it does not know, and ask it to flag anything it is unsure about. OpenAI's own guidance favors expressing uncertainty over guessing.
- Ask twice. If a fact matters, ask in a fresh conversation or phrase the question differently. Different answers to the same question are a red flag, as in OpenAI's dissertation example.
- Do not trust the reasoning more than the answer. A neat step-by-step explanation is not proof.
- Use AI for drafting and structure, not as the final authority. Let it organize, rephrase and suggest; keep verification with a person.
Key takeaways
- NIST calls AI hallucinations "confabulation": confidently stated but false content, including drift from your input and self-contradiction.
- They happen because models predict likely text rather than look up facts; rare, specific facts are the weakest spot.
- OpenAI says accuracy-only grading rewards guessing, which is one reason wrong answers sound confident.
- Plausible reasoning and citations can make a wrong answer more convincing, so verify sources directly.
- Check every specific fact, provide source material, and allow the tool to admit when it does not know.