The Confident MachinePre-print

Part II: Anatomy of a Hallucination · Page 5 of 14

Defining Hallucination

What counts as a hallucination

Not every wrong answer is a hallucination: a stale fact or requested fiction doesn't count. The definition is narrower, fluent output presented as fact in a system with no mechanism to check what it says against what's real.

Researchers split that failure into three categories.[16] When a source text exists, an intrinsic hallucination contradicts it directly; an extrinsic hallucination adds a specific, checkable detail the source never confirms or denies. With no source at all, a wrong answer to an open question is a factual error. The exercise below adds a fourth bucket alongside those three, for answers that aren't hallucinations at all.

Six real, documented incidents, four categories. Classify each one:

Classify each statement: a category is marked right or wrong as soon as you pick one.

0/6 answered

In a 2023 launch demo, Google's Bard said the James Webb Space Telescope took the first-ever image of a planet outside our solar system; ground-based telescopes had already done that in 2004.[19]

Ask an AI coding assistant for help, and it sometimes recommends installing a specific add-on that doesn’t exist. Attackers exploit this: a researcher registered one repeatedly-invented add-on name and watched it get installed over 30,000 times.[20]

Air Canada's support chatbot told a customer he could apply for a bereavement fare after booking. The airline's actual policy didn't allow that.[21]

Summarizing a BBC push notification, Apple Intelligence reported that a murder suspect had shot himself. The article said no such thing.[22]

Asked who the reigning champion is, a model trained months ago names last year’s winner: correct when it learned it, out of date now.

Asked to write a short story, the model invents a character who never existed.

Six real, documented incidents, each sourced. Every statement gets immediate feedback and its category's practical implication.

Real-world consequences

These failures carry different stakes. An unchecked factual error enabled a supply-chain attack: attackers pre-registered a hallucinated software package name and had it installed over 30,000 times.[20] A claim that contradicts an available source carries legal exposure, as when an airline was held liable for its chatbot's invented refund policy.[21] An unverifiable addition spreads specific misinformation attached to real reporting, as when a phone assistant appended a fabricated detail to a genuine news alert.[22] A stale or fictional answer carries none of that risk. Neither one is a fabrication: the failure is scope or timing, not invention.