Hallucinations
Do hallucinations make AI useless?
Errors matter, but usefulness depends on task, verification, and failure cost.
"AI hallucinates, so it is useless."
What this page actually tests
Hallucination rates make generative AI broadly unsuitable for factual work unless users add verification, grounding, or domain review.
Wording note: Useless claims zero value across every task. The realistic concern is that unreliable factual generation blocks or raises the cost of important uses unless the workflow adds verification and constraints.
Hallucinations limit usefulness; they do not erase it.
Misleading. Hallucinations make bare chatbot output unsuitable for many factual and high-stakes uses, and verification can be expensive. They do not make the technology useless across constrained, low-risk, creative, or reviewable tasks.
Why people repeat it
The concern is common because models can present false facts and citations fluently, leaving users to detect errors that may be difficult or expensive to notice.
What the sources support
Fact: The GPT-4 technical report describes GPT-4 as a Transformer model pretrained to predict the next token and notes factual reliability limitations despite improved performance.
Baseline: The comparison is not "perfect truth machine" versus "trash"; it is model output with verification versus workflows where errors are unacceptable.
Evidence conclusion: The evidence proves blind trust is a bad idea. It does not prove every assisted drafting, coding, classification, or summarization use is worthless.
Source: GPT-4 Technical Report
Fact: Nature research on semantic entropy reported hallucination-detection performance around AUROC 0.790, beating several baselines, while also noting limits for some error types.
Baseline: A measurable detection method is different from pretending hallucinations are either solved or impossible to manage.
Evidence conclusion: The evidence supports testing and mitigation in specific workflows, not the claim that the entire technology category has zero utility.
Source: Detecting hallucinations in large language models using semantic entropy
Fact: NIST lists confabulation and information integrity as generative AI risks to govern alongside privacy, security, bias, and misuse.
Baseline: Risk-management frameworks do not treat every risk as a ban; they match controls to use case and harm level.
Evidence conclusion: The conclusive lesson is boring but useful: high-stakes use needs controls, and low-stakes use still needs verification proportional to the task.
Source: Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile
Fact: A field study of more than 5,000 customer-support agents found generative-AI assistance increased resolved chats per hour by about 14% on average, with larger gains for less-experienced workers.
Baseline: The assistant operated inside a bounded support workflow with historical examples and measurable outcomes; it was not an unchecked general-purpose authority.
Evidence conclusion: Measured productivity in one real workflow is enough to reject 'useless,' while saying nothing about whether the same system should answer medical or legal questions alone.
Source: Generative AI at Work
Fact: A preregistered experiment with 758 consultants found AI improved speed and quality on tasks within its capability frontier but made users more likely to give wrong answers on a task outside that frontier.
Baseline: Usefulness changed with task fit and human judgment rather than with the mere presence of a chatbot.
Evidence conclusion: The evidence supports bounded use and evaluation, not either extreme of universal trust or universal uselessness.
Source: Navigating the Jagged Technological Frontier
Source balance
Checked both sides before calling it.
Supports the claim
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile - NIST treats confabulation and information integrity as real generative AI risks.
- Detecting hallucinations in large language models using semantic entropy - Research documents hallucination behavior and methods to detect uncertainty.
Challenges or narrows it
- Generative AI at Work - A bounded customer-support deployment produced a measured 14% average productivity gain.
- Navigating the Jagged Technological Frontier - AI improved performance on in-frontier tasks while the same experiment documented failure outside that frontier.
- GPT-4 Technical Report - The model shows useful capabilities alongside documented limitations.
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile - The risk-management framing implies mitigation and bounded use, not total uselessness.
Baseline context
- GPT-4 Technical Report - Provides capability and limitation context.
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile - Provides risk-management categories for high-stakes use.
Assessment: The claim is misleading. Hallucinations materially restrict factual and high-stakes use, but the evidence also supports useful bounded workflows; confirming the limitation is not the same as confirming the uselessness conclusion.
Where critics may still have a point
- For legal, medical, financial, or safety-critical advice, a confident wrong answer can be unacceptable without expert review.
- Retrieval, citations, and system prompts do not automatically solve hallucinations.
- Low-stakes usefulness does not justify high-stakes deployment without measurement.
Hallucinations limit usefulness; they do not erase it.
Current models can generate confident falsehoods, so factual workflows need grounding, abstention, testing, or human review. Those controls narrow where AI is useful and add cost. Evidence from bounded workflows still shows real value, so the conclusion that hallucination makes AI useless does not follow from the documented failure mode.
Why this verdict: The sources confirm a major reliability constraint, but evidence of useful bounded workflows directly contradicts the slogan's conclusion that the entire technology is useless.
Sources
-
GPT-4 Technical Report
Used for: Model capability, training-objective, and limitation framing.
-
Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile
Used for: Generative AI risk categories and mitigation framing.
-
Detecting hallucinations in large language models using semantic entropy
Used for: Hallucination detection research and reliability measurement context.
-
Generative AI at Work
Used for: Measured productivity evidence from a bounded customer-support deployment with more than 5,000 agents.
-
Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality
Used for: Cross-checking useful in-scope performance against the risk of worse answers outside the model's capability frontier.