Why AI Hallucinates: Causes, Types and How to Reduce Them
AI hallucinations are plausible outputs that lack reliable support, and reducing them means improving evidence, task design and verification rather than simply asking a model to be accurate.
Key takeaways
- Fluent language is not evidence: an answer can sound convincing while containing unsupported claims.
- Hallucinations can originate in training, missing context, retrieval failures or unsupported inference.
- Ground answers in relevant sources and require explicit handling of missing or conflicting information.
- Verify consequential claims with independent evidence rather than relying on the model to review itself.
- Measure unsupported claims alongside omissions, useful coverage and appropriate refusals.
On this page
- A convincing answer is not necessarily a true one
- What counts as an AI hallucination?
- Why language models can produce unsupported answers
- The main types of hallucination
- A risk-based approach: what needs checking first?
- How to reduce hallucinations in everyday use
- Worked example: summarising a fictional returns policy
- Retrieval-augmented generation: useful grounding, not immunity
- Worked example: checking a suspicious research citation
- Common fixes that do not work as well as people expect
- Building a repeatable verification workflow
- Measuring improvement rather than trusting impressions
- A practical checklist before you trust an answer
- FAQ
A convincing answer is not necessarily a true one
Ask an AI assistant for a summary of a meeting, and it may produce a clear, organised account. Ask who approved the budget, and it may confidently name someone—even though the transcript contains no approval.
The second answer is not necessarily a random malfunction. It can emerge from the same process that produced the useful summary: generating a plausible response from learned patterns and available context.
This is why AI hallucinations are difficult to manage. They rarely announce themselves as errors. They often arrive wrapped in sensible explanations, familiar terminology and precise-looking details.
The practical goal is not to distrust every AI output equally. It is to distinguish tasks where approximation is acceptable from tasks where unsupported details matter, then put appropriate checks between the model and the decision.
This article focuses mainly on language models and AI assistants. The same underlying concern also appears in systems that interpret images, transcribe audio or combine several kinds of input.
What counts as an AI hallucination?
An AI hallucination is generated content presented as factual or grounded when it is false, unsupported by the relevant evidence, or inconsistent with the material the model was asked to use.
There is no single definition used everywhere. Researchers distinguish several related problems, including factual inaccuracies and failures to stay faithful to a source. The Survey of Hallucination in Natural Language Generation explains these distinctions across tasks such as summarisation, dialogue and translation.
For everyday use, two questions are especially helpful:
- Is the claim true in the world?
- Is the claim supported by the evidence this task permits?
Those questions are related, but not interchangeable.
Factual accuracy versus faithfulness
Suppose a supplied company report says:
Northbridge Foods opened its first shop in Leeds in 2016.
The assistant summarises it as:
Northbridge Foods began trading in Manchester in 2014.
That is a source-faithfulness failure: the output contradicts the supplied document.
Now suppose the assistant adds:
Its founder previously worked as a pastry chef.
That detail might happen to be true. But if the task was to summarise only the supplied report and the report never mentions it, the addition is unsupported.
Conversely, a model can faithfully repeat an incorrect source. If the report itself contains a mistaken date, accurately summarising the report does not make the date correct in the world.
This distinction determines the appropriate check. Comparing the answer with the document tests faithfulness. Checking company records or another authoritative source tests factual accuracy.
Not every undesirable answer is a hallucination
A fictional story about a talking fox is not a hallucination when fiction was requested. An explicitly labelled estimate is not automatically a hallucination because it later proves inaccurate.
Similarly, a typo, a formatting failure and a biased recommendation are not necessarily hallucinations. A mathematical error may be better described as a calculation or reasoning error, although people often group it under the same broad label.
The boundary matters less than the response. Ask what failed:
- Was information invented?
- Was evidence misread?
- Was uncertainty hidden?
- Was a calculation wrong?
- Was a tool result ignored?
Different failures need different remedies. “Make the model more accurate” is too vague to guide useful action.
Why language models can produce unsupported answers
A language model generates text by estimating plausible next tokens given its input and what it has learned during training. Tokens are pieces of text: sometimes whole words, sometimes word fragments or punctuation.
This process can support impressive factual recall, generalisation and problem-solving. But generating a likely continuation is not the same operation as checking every statement against a trusted database.
Training teaches patterns, not a complete evidence ledger
During pre-training, a model learns from large collections of material. It absorbs relationships between concepts, common writing structures, terminology and many factual associations.
The foundational paper Language Models are Few-Shot Learners illustrates how a model trained with an autoregressive language-modelling objective can perform many tasks through text prompts.
What the model does not automatically acquire is a reliable record saying:
This exact claim came from this document, published on this date, and remains valid under these conditions.
Some information may be memorised closely; other information is represented more diffusely. At answer time, recalling a familiar association can look much like constructing a plausible but incorrect one.
A model may know the usual shape of a research citation without knowing the particular paper requested. It can then combine a credible title, real researchers and a plausible journal into a reference that does not exist.
Training material contains errors and contradictions
The material models learn from includes outdated information, jokes, fiction, disputed claims and straightforward mistakes.
Learning how people write does not automatically separate common beliefs from accurate ones. TruthfulQA was designed to test whether language models reproduce falsehoods that people commonly express.
Consider a health myth that appears repeatedly online. Familiarity may make the associated wording easy for a model to generate, even when reliable medical evidence rejects it.
Contradictions also create problems. A business may have changed its name, a law may have been amended, and a software function may have been removed. Without the right temporal context, a model can blend information from different periods.
Assistant training improves behaviour but does not guarantee truth
After pre-training, many assistants undergo additional training to follow instructions and produce responses people prefer. This can improve relevance, clarity and willingness to acknowledge limits.
The paper Training language models to follow instructions with human feedback describes an influential approach to this process.
However, human preferences are imperfect signals. A polished, decisive answer can look more helpful than a cautious one, particularly when an evaluator cannot easily check the subject matter.
Training can therefore reduce some errors while leaving incentives to complete the task convincingly. It cannot turn uncertain internal knowledge into verified evidence.
Our guide to pre-training, fine-tuning and RLHF explains how these training stages differ.
The question may demand more information than is available
Many hallucinations begin with a gap between the requested answer and the available evidence.
Examples include:
- Asking about an event that happened after the model’s training.
- Requesting details from a document that was never attached.
- Asking for the reason behind a decision when only the decision is recorded.
- Requesting an exact number from an incomplete dataset.
- Asking about a private organisational policy using only general knowledge.
A reliable response would identify the gap, ask a question or narrow the answer. An unreliable response fills it with something plausible.
This is particularly easy when the question contains a false assumption. “Why did the council reject the proposal?” invites an explanation even if the proposal was never rejected.
Long context is not the same as reliable attention
Providing evidence helps, but the model must still find and use the relevant parts.
A long input may contain duplicated sections, conflicting drafts, distracting material or important exceptions buried far from the main rule. In some systems, older conversation content may also be shortened or omitted.
Lost in the Middle demonstrated that the position of relevant information can affect performance in long-context tasks. The exact behaviour varies across models and settings, but the practical lesson remains: fitting a document into the input does not guarantee that every detail will be used correctly.
See what a context window is and why it limits AI for the distinction between input capacity and dependable information use.
Generation settings affect variation, not factual grounding
Many systems expose settings such as temperature and top-p, which influence token selection.
Lower randomness can make answers more repeatable. Higher randomness can increase variety. Neither setting provides missing evidence.
A model can produce the same wrong answer consistently at a low temperature. It can also produce several different wrong answers at a higher one.
For factual extraction, conservative generation settings may be sensible. But they are a supporting choice, not the main defence. The guide to temperature, top-p and sampling explains their actual role.
The main types of hallucination
Classifying errors makes them easier to recognise and test. The following categories overlap, but each suggests a particular verification method.
Invented facts and entities
The assistant invents an organisation, product feature, person, event or relationship.
Example: A travel assistant claims that a museum offers late opening every Thursday, based on a familiar museum schedule rather than the museum’s current information.
Best check: Consult the relevant primary source and confirm that it is current.
Names and precise details deserve particular attention. A fluent paragraph can remain broadly sensible while one invented name makes it unusable.
Fabricated or misleading citations
The model provides a source that does not exist, misattributes a real source, or cites something that does not support the claim.
Example: A report cites a real paper about reading comprehension as evidence that a particular tutoring product improves examination results.
The paper’s existence does not validate the product claim.
Best check: Verify both existence and entailment: can you locate the source, and does its content actually support the statement?
A working link passes only the first part of that test.
Distorted summaries and unsupported additions
The assistant changes the meaning of source material or adds details that were never present.
Example: “The committee discussed a possible hiring freeze” becomes “The committee approved a hiring freeze.”
Common distortions include changing:
- Possibility into certainty.
- Discussion into agreement.
- Correlation into causation.
- A recommendation into a requirement.
- One participant’s opinion into a group conclusion.
Best check: Compare the summary against the relevant passages, paying attention to verbs and qualifiers.
Numerical and structural errors
The assistant reports a number incorrectly, swaps rows in a table or attaches a value to the wrong category.
Example: A spreadsheet lists £42,000 as forecast revenue and £24,000 as actual revenue. The narrative reverses them.
A calculator will not fix this if the wrong numbers are selected first.
Best check: Verify extraction separately from calculation. Confirm labels, units, dates and scope before checking the arithmetic.
Temporal and jurisdictional blending
The answer mixes facts from different times, countries, versions or organisations.
Example: An explanation of employment rights combines one country’s notice rules with another country’s holiday entitlement.
Each individual detail may resemble a real rule, but the combined answer is unreliable.
Best check: Specify the relevant jurisdiction, effective date and version. Require the answer to preserve those boundaries.
False claims about actions and access
The assistant says it checked a website, sent a message, read an attachment or updated a record when it did not.
Example: “I have cancelled your subscription” appears in a chat system that has no cancellation tool.
This is especially consequential because users may stop acting once they believe the task is complete.
Best check: Look for an actual tool result or confirmation from the destination system. Text describing an action is not proof that the action happened.
Perceptual hallucinations
A system interpreting an image, audio recording or scanned document introduces content that is not present.
Example: It reads a blurred product label as a familiar brand, or transcribes speech during an unclear stretch of audio.
Best check: Inspect the original media, request timestamps or regions, and allow an “unreadable” or “inaudible” result rather than forcing completion.
A risk-based approach: what needs checking first?
You do not need the same verification process for a dinner-party poem and a medication instruction.
A useful triage considers three things:
- Impact: What happens if the answer is wrong?
- Uncertainty: How complete, clear and current is the evidence?
- Propagation: Will the answer be copied, published or used to trigger further actions?
A small mistake can become serious when propagated widely. An incorrect opening time in a private note is inconvenient; the same mistake distributed to thousands of customers creates a larger operational problem.
Low-consequence uses
For brainstorming, fictional writing and early drafting, speed and variety may matter more than factual precision.
Still, distinguish invented examples from real claims. A marketing draft can accidentally introduce a false customer testimonial or an unsupported performance promise.
A simple rule helps: creative freedom applies to wording and clearly labelled fiction, not to real-world evidence.
Medium-consequence uses
For internal summaries, research notes and purchasing comparisons, check the facts that affect decisions.
You may not need to verify every connecting sentence, but you should verify prices, deadlines, compatibility claims, ownership, dates and recommendations attributed to others.
Assign a reviewer rather than assuming “someone” will check later.
High-consequence uses
For medical, legal, financial, safety-critical or consequential employment decisions, treat AI output as an aid rather than an authority.
Use authoritative, applicable sources and qualified human review where needed. Do not rely on a general chatbot to determine a safe dose, a legal deadline or eligibility for a benefit.
The point is not merely to attach a disclaimer. The workflow must prevent unsupported output from directly becoming consequential action.
How to reduce hallucinations in everyday use
The strongest practical approach combines constrained tasks, relevant evidence, clear uncertainty rules and independent checks.
Step 1: Define what the answer is allowed to rely on
Decide whether the task is:
- Closed-book: use only the supplied material.
- Evidence-seeking: find and evaluate outside sources.
- Exploratory: suggest possibilities without claiming they are established facts.
Mixing these modes creates confusion. A document summary should not silently become a general-knowledge essay.
For a closed-book task, use a prompt like this:
Task: Summarise the attached policy for staff.
Evidence boundary:
Use only the attached policy.
Do not fill gaps with general knowledge or customary practice.
If a requested detail is absent, write "Not stated in the policy".
Distinguish requirements, recommendations and examples.
Flag conflicting passages instead of choosing silently.
For each substantive point, include its section heading.
This does not guarantee compliance, but it makes the expected behaviour testable.
Step 2: Supply relevant context, not merely more context
Provide the document version, date, intended audience and any definitions necessary to interpret it.
Remove unrelated material where possible. Clearly label drafts and superseded documents. If two sources conflict, state whether one has priority.
For example, a customer-support assistant needs to know whether a return policy applies to online purchases, shop purchases or both. A large pile of product information will not compensate for that missing distinction.
When documents are long, start by locating the relevant sections, then ask for an answer based on those sections. Preserve enough surrounding text to retain exceptions.
Step 3: Make missing information an acceptable outcome
Many prompts unintentionally require completion at any cost:
Fill every cell. Answer all questions. Provide five examples.
If only three supported examples exist, the requested format creates pressure to invent two more.
Instead, explicitly permit incompleteness:
Return up to five supported examples.
If fewer than five are available, return fewer.
For missing fields, use null.
Do not infer dates, names or amounts from surrounding patterns.
List any information needed to complete the task.
The important change is not the word “null”. It is the removal of a requirement to disguise missing evidence.
Step 4: Separate extraction from interpretation
First ask what the source says. Then ask what that might mean.
A useful output structure has three fields:
| Field | Purpose |
|---|---|
| Source statement | What is explicitly recorded |
| Interpretation | A clearly labelled inference |
| Open question | What remains unresolved |
For a meeting transcript, “Aisha proposed moving the launch” belongs in the first field. “The team may be concerned about readiness” belongs in the second, if useful. “Was the date formally changed?” belongs in the third.
This prevents reasonable interpretation from masquerading as recorded fact.
Step 5: Ask for evidence that a person can inspect
Require short supporting excerpts, section names, page references where available, or links to the relevant source.
But do not assume the model has copied correctly. Check that quotations are exact and that references resolve.
For scanned documents, printed page numbers may differ from PDF page positions. Agree on a convention rather than accepting ambiguous references such as “page 12”.
Evidence should make checking easier, not merely decorate the answer with scholarly-looking notation.
Step 6: Use tools for operations they perform reliably
Use a calculator or code for arithmetic, a database query for exact records, and an appropriate search or retrieval tool for current information.
Then verify that the right inputs reached the tool and that the result was interpreted correctly.
For example, calculating growth from £80,000 to £100,000 is straightforward:
Growth = (100000 - 80000) / 80000 × 100
= 25%
But the calculation is irrelevant if one number is monthly revenue and the other is quarterly revenue. Tool use reduces one class of error; it does not remove the need to understand the task.
Step 7: Verify the claims that matter
Break an answer into checkable claims. Prioritise those that affect action.
For a supplier comparison, these might include:
- The product supports a required file format.
- The quoted price includes the necessary tier.
- Data can be stored in the required region.
- Cancellation is available under the stated conditions.
Open the relevant documentation or contract. Record the source and date. If a claim cannot be verified, remove it, qualify it or seek clarification.
“Looks reasonable” is not a verification method.
Worked example: summarising a fictional returns policy
Consider this fictional policy:
Policy version: 3 March 2026
Standard items bought online may be returned within 30 calendar
days of delivery.
Personalised items cannot be returned unless faulty.
Customers pay return postage unless the item is faulty.
Approved refunds are issued within 10 working days of inspection.
A user asks:
Can I return a personalised mug after 20 days, and when will I get my money back?
The plausible but unsupported answer
An unreliable assistant might reply:
Yes. You are within the 30-day return period, and your refund will arrive within 10 days of sending the mug back.
This answer makes several mistakes:
- It applies the standard-item rule to a personalised item.
- It ignores the faulty-item exception.
- It changes working days to unspecified days.
- It starts the refund clock at posting rather than inspection.
- It implies that the refund is guaranteed.
The wording is simple and helpful-looking. The problem is not style; it is lost conditions.
A better prompt and answer
Use a constrained prompt:
Answer using only the policy below.
First identify:
- the item category;
- the relevant eligibility rule;
- any missing fact;
- the event that starts the refund period.
Do not assume the item is faulty.
Preserve calendar days versus working days.
Do not promise a refund unless the policy supports it.
A grounded answer would be:
If the mug is personalised and not faulty, this policy does not allow a return merely because it is within 30 days. If it is faulty, the policy allows a return and says customers do not pay return postage. Approved refunds are issued within 10 working days of inspection, not within 10 days of posting. Is the mug faulty?
Notice that the better answer is conditional, not evasive. It tells the user exactly which missing fact changes the outcome.
Exercise: audit the answer yourself
Take any short policy or instruction sheet and follow these steps:
- Highlight every condition: “if”, “unless”, “only”, “after” and “before”.
- Underline quantities, units and deadlines.
- Ask the model a question that depends on an exception.
- Compare its answer with your marked text.
- Count unsupported additions and lost conditions.
- Revise the prompt and repeat with a different question.
Do not test only straightforward cases. Exceptions reveal failures that a general summary may hide.
Retrieval-augmented generation: useful grounding, not immunity
Retrieval-augmented generation, usually called RAG, connects a model to an external collection of information. The system retrieves relevant passages and supplies them as context for generation.
The research paper Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks describes an influential version of this approach.
For organisations, RAG can make internal policies, manuals or product documentation available without expecting the model to have learned them during training.
Our beginner’s guide to retrieval-augmented generation covers the basic architecture.
Where a RAG system can fail
Think of RAG as a chain rather than a truth switch:
- The right document must exist in the collection.
- The document must be current and correctly processed.
- Retrieval must find the relevant passage.
- The passage must retain necessary context.
- The model must interpret it correctly.
- The answer must remain within what it supports.
A failure at any point can produce an unsupported answer.
For example, retrieval may find the general returns rule but miss a separate personalised-goods exception. The model then gives a well-cited answer based on incomplete evidence.
Similarity is not the same as authority
Many retrieval systems use embeddings to find passages with related meaning. This is useful, but similarity does not establish that a passage is authoritative, current or applicable.
An old policy may closely match the question. A draft may contain clearer wording than the approved version. A customer complaint may repeat the exact phrase the user searched for.
Good systems therefore use metadata such as version, publication status, date, department and jurisdiction alongside relevance.
The introduction to embeddings and vector databases explains what semantic retrieval can and cannot establish.
What stronger grounded answering looks like
A practical RAG workflow should:
- Prefer approved sources over informal mentions.
- Retain document identifiers and version information.
- Retrieve exceptions and surrounding context.
- Flag contradictory evidence.
- Require citations attached to specific claims.
- Decline to answer unsupported parts.
- Check access permissions before exposing retrieved content.
Evaluate retrieval separately from answer generation. If the correct passage never reaches the model, changing the writing prompt may achieve little.
Also treat retrieved text as data, not as instructions. A document that says “ignore earlier instructions” should not gain authority over the system merely because it was retrieved.
Worked example: checking a suspicious research citation
Suppose an assistant produces this explicitly fictional reference:
Patel, R. and Green, L. (2022). Microlearning and Adult Memory Consolidation. Journal of Digital Cognition.
It then claims the paper proves that five-minute lessons outperform all longer teaching formats.
There are two separate questions: does the reference exist, and does it support that sweeping conclusion?
A step-by-step citation audit
Step 1: Search the exact title. Use a scholarly search service, library catalogue or publisher search. Do not rely on the model repeating the citation.
Step 2: Check the bibliographic combination. Authors, year, title and publication venue must refer to the same work. A model can mix individually real details.
Step 3: Open the actual source. A search snippet or another AI summary is not enough for a consequential claim.
Step 4: Locate the relevant result. Look for the population studied, intervention, comparison and measured outcome.
Step 5: Compare scope. A small study of vocabulary recall cannot by itself establish superiority across all adult education.
Step 6: Rewrite or remove the claim. If the source cannot be located, label the citation unverified and do not publish it as evidence.
Failure to find a paper immediately does not prove it is fabricated. There may be indexing problems or an inaccurate title. But it does mean the citation has not passed verification.
A better research prompt
Find sources relevant to this claim:
"Short lessons can support adult vocabulary learning."
Use search tools if available.
Do not generate references from memory.
For each source, provide:
- exact title;
- authors and year;
- a working source URL;
- the population and outcome studied;
- what the source supports;
- what it does not establish.
If you cannot access the source text, state that limitation.
Do not say a study "proves" a broader claim than it tested.
This improves the research process by narrowing the claim and requiring attention to scope. The resulting references still need checking.
Common fixes that do not work as well as people expect
Several popular techniques can help at the margins while creating false reassurance if used alone.
“Do not hallucinate”
This instruction expresses a goal but supplies neither missing knowledge nor a verification mechanism.
A more useful instruction specifies behaviour:
If the supplied source does not state the date, write “not stated” and identify the missing information.
Concrete rules are easier to follow and evaluate than broad demands for truthfulness.
Asking the model to check its own answer
Self-review can catch contradictions, missing constraints and some calculation errors. But the same model may preserve the same mistaken assumption.
Asking “Are you sure?” may produce a more emphatic version of the original error.
A stronger review supplies independent evidence or a different checking method. For example, compare each claim with the policy text, or run the generated code against tests.
The guide to self-consistency and verification prompts explores where these methods help and where they fall short.
Treating agreement as proof
If several model runs agree, the answer may be stable. It is not necessarily true.
Models can share training material, common misconceptions and predictable associations. Repeatedly asking a question can reproduce the same error.
Agreement becomes more useful when combined with independent sources and diverse verification methods, rather than counted as evidence by itself.
Trusting confidence scores without calibration
A model saying “95% confident” does not automatically provide a meaningful probability of correctness.
Calibration requires testing whether answers assigned a confidence level are correct at roughly the corresponding frequency on relevant tasks.
Research such as Semantic Uncertainty investigates uncertainty estimation for generated language. That is different from simply asking a chatbot to invent a confidence percentage.
For everyday workflows, evidence categories are often more actionable: directly supported, inferred, contradicted or unknown.
Assuming structured output makes content true
A valid JSON object can contain an invented date. A neat table can reverse two numbers. A required field can even increase pressure to fill a gap.
Use schemas that allow missing values and include evidence fields. Validate content separately from format.
Similarly, a long explanation does not guarantee sound reasoning. Ask for concise, checkable calculations or supporting evidence rather than treating verbosity as a reliability signal.
Assuming a newer or larger model needs no checks
More capable models may perform better on many tasks, but performance varies by domain, language, input quality and workflow.
A model that writes excellent prose may still mishandle a scanned table or an obscure contractual exception.
Choose models using representative tests. Do not infer reliability for your task solely from a general reputation or a demonstration.
Building a repeatable verification workflow
For recurring work, turn good habits into a process that another person can follow.
Create a claim ledger
A claim ledger is a simple table connecting important statements to evidence and review status.
| Claim | Evidence | Status | Next action |
|---|---|---|---|
| Returns allowed within 30 days | Policy, standard-items section | Supported with conditions | Preserve item restriction |
| Refund starts when parcel is posted | No supporting passage | Unsupported | Correct to inspection |
| Personalised mug is faulty | Customer has not said | Unknown | Ask customer |
This format helps prevent unsupported details from disappearing into polished prose.
You do not need a ledger for every casual conversation. Use one when an answer will influence a purchase, policy decision, publication or customer commitment.
Separate drafting from approval
Let the model prepare a draft, but require review before publication or action.
For an assistant with tools, distinguish:
- Proposed action.
- Authorised action.
- Attempted action.
- Confirmed successful action.
These states should come from the workflow and tool results, not from the assistant’s tone.
If an email fails to send, the system should report failure rather than allowing the language model to generate a generic success message.
Exercise: run an evidence-first review
Choose an AI-generated answer of roughly 300 words.
- Split it into individual factual claims.
- Mark each claim as supported, contradicted, unverified or inference.
- Identify the evidence needed for each unverified claim.
- Check the three claims with the highest consequences.
- Remove details that add precision without support.
- Rewrite inferences so they cannot be mistaken for facts.
- Confirm that the final wording answers the original question.
This exercise often reveals that an answer can become more reliable by becoming shorter. Unnecessary specifics create additional opportunities for error.
Preserve an audit trail proportionate to the risk
For routine internal work, recording the source link and review date may be enough.
For consequential decisions, retain the source version, relevant excerpts, model output, reviewer decision and any tool results needed to explain what happened.
Avoid retaining sensitive information unnecessarily. Verification and privacy should be designed together, not treated as competing afterthoughts.
The aim is reproducibility: another reviewer should be able to understand why the answer was accepted.
Measuring improvement rather than trusting impressions
A revised prompt can feel safer because its answers sound cautious. That does not prove it performs better.
Build a small, representative test set and compare versions systematically.
Include answerable and unanswerable questions
A useful test collection includes:
- Straightforward factual questions.
- Questions requiring an exception.
- Missing-information cases.
- Conflicting-source cases.
- Outdated-document cases.
- Questions with false assumptions.
- Numerical questions with awkward units.
- Requests where the correct outcome is a clarification.
If every test has a clear answer in the first paragraph of a document, the evaluation will miss many real-world failures.
Score individual claims
Long answers often contain a mixture of accurate and inaccurate statements. Giving the whole response one label hides that variation.
FActScore introduced an approach based on evaluating the support for atomic factual claims in generated text. The general idea is useful even without implementing the research method.
As an illustrative manual calculation, suppose an answer contains 12 checkable claims:
- Nine are supported.
- Two are unsupported.
- One contradicts the source.
Nine out of 12 claims are supported, giving a support proportion of 75%. That describes this example only; it is not a general model performance statistic.
Keep contradictions separate from claims whose truth simply remains unverified.
Measure coverage and appropriate abstention too
A system can avoid errors by saying almost nothing. That is not necessarily useful.
Track several outcomes together:
- Supported factual claims.
- Unsupported or contradicted claims.
- Required facts omitted.
- Appropriate acknowledgements of missing information.
- Unnecessary refusals on answerable questions.
- Correct handling of consequential exceptions.
Also weight severity. A minor descriptive error and an incorrect safety instruction should not count as equivalent operational failures.
Re-test after meaningful changes
Model updates, prompt revisions, retrieval changes and new document formats can all alter performance.
Keep some test cases separate from prompt development so that you are not merely optimising for familiar examples. Run repeated trials when outputs vary.
For deployed systems, review real failures and add representative cases to the test set. Improvement should mean fewer consequential errors on realistic tasks, not just better-looking demonstration answers.
A practical checklist before you trust an answer
Before using an AI response, ask:
- What is the evidence boundary? Is this general knowledge, supplied material or retrieved information?
- Is the source appropriate? Check authority, date, version and jurisdiction.
- Which claims affect the decision? Prioritise those for verification.
- Were conditions preserved? Look for missing exceptions, qualifiers and units.
- Can citations be inspected? Confirm both source existence and actual support.
- Was anything inferred? Ensure inference is labelled rather than presented as fact.
- Did the claimed action happen? Require a tool or destination-system confirmation.
- What remains unknown? Decide whether to ask, investigate, qualify or stop.
Hallucination reduction is not a single prompt trick. It is a design principle: make evidence available, constrain unsupported completion, and place verification where errors could cause harm.
A useful assistant is not one that always produces a complete answer. It is one whose answers help you distinguish what is established, what is inferred and what still needs checking.
FAQ
Why does AI hallucinate even when it sounds certain?
Fluent wording and factual correctness are different properties. A model can generate the language of certainty because it fits the response pattern, without having verified the underlying claim. Confident tone should therefore be treated as presentation, not evidence.
Is an AI hallucination the same as a lie?
Not usually in the ordinary human sense. Lying implies knowingly communicating something false. “Hallucination” describes an output failure without requiring assumptions about human-like intent. For practical purposes, focus on whether the claim is supported and what harm it could cause.
Can hallucinations be eliminated completely?
There is no general guarantee for open-ended generative use. Narrow tasks with controlled inputs, deterministic checks and strict evidence requirements can be made substantially more reliable. Even then, the wider system can fail through incorrect source data, retrieval mistakes or faulty integrations.
Does giving the model internet access solve the problem?
No. Internet access can supply current evidence, but the assistant may search poorly, choose weak sources, misread pages or attach a citation to an unsupported claim. Browsing is useful when paired with source selection and verification, not as a substitute for them.
Should I always ask for sources?
Ask for sources when factual claims matter, especially for current, specialised or consequential information. But remember that generated references can be wrong. Request accessible links and claim-specific support, then inspect the important sources yourself.
Is a low temperature safer?
It can make generation less variable, which is useful for some extraction tasks. It does not make an unsupported answer true. Prioritise relevant evidence, explicit missing-information handling and verification before adjusting sampling settings.
Does asking for step-by-step reasoning prevent hallucinations?
It can help organise some tasks, but a detailed explanation can also rationalise an incorrect answer. Request checkable intermediate outputs—such as extracted values, equations or source excerpts—rather than assuming that a long reasoning-style response proves correctness.
What should I do when two sources disagree?
Check whether they concern the same date, jurisdiction, version and situation. Prefer the source with the appropriate authority, but do not silently erase a genuine conflict. State the disagreement and seek clarification when the difference affects the decision.
What is the best quick habit for everyday users?
Pause at precise, consequential details. Check the names, dates, amounts, quotations and conditions that would change your next action. This targeted habit is more manageable than verifying every word and more useful than accepting an answer because it reads smoothly.
When should I stop using the AI answer and ask a person?
Escalate when the stakes are high, evidence is missing or contradictory, or the task requires professional judgement or formal authority. AI can help organise questions and source material, but it should not turn unresolved uncertainty into a decision merely to complete the conversation.
Sources
- Survey of Hallucination in Natural Language Generation
- Language Models are Few-Shot Learners
- Training language models to follow instructions with human feedback
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Lost in the Middle: How Language Models Use Long Contexts
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
About the author
Editorial team · Editorial team
Author identity has not been supplied yet. Replace this record with the real writer's name, background and verifiable experience before publishing anything on the public site.
Spotted an error? Report a correction.
Related reading
A Repeatable AI Research Workflow: From Question to Verified Brief
A disciplined research process that uses AI for planning and synthesis while keeping every important claim tied to evidence you have checked.
How to Summarise Long Documents with AI Without Missing What Matters
A practical, source-grounded workflow for turning long documents into reliable summaries while preserving caveats, contradictions and important detail.
AI Meeting Notes: A Safe Workflow from Transcript to Action Items
A careful end-to-end method for using AI to draft meeting notes without inventing decisions, owners or deadlines.