Detecting Hallucinations in Real Conversations
This presentation examines AuthenHallu, the first hallucination-detection benchmark built entirely from authentic human-LLM interactions. Drawing from one million real conversations, the researchers annotated 800 query-response pairs to reveal how often and where language models produce incorrect or inconsistent outputs in everyday use. The work challenges existing benchmarks that rely on artificially induced errors and demonstrates that current models struggle to reliably detect hallucinations in natural dialogue, particularly when responses contradict user input or previous context.Script
In medicine and law, a single incorrect answer from a language model can have serious consequences. Yet most hallucination benchmarks test models on deliberately induced errors or simulated conversations, not the messy reality of how people actually use these systems.
The researchers built AuthenHallu from the LMSYS Chat dataset containing one million authentic human-Large Language Model conversations. After filtering for English dialogues, removing toxic content, and clustering by topic, they selected 400 representative conversations. Three trained annotators then manually labeled all 800 query-response pairs for hallucinations.
The results were striking. Hallucinations appeared in 251 of the 800 responses, a rate of 31 percent. Fact-conflicting hallucinations were most common with 157 instances, but the breakdown by topic revealed something more specific: math problems and date calculations had hallucination rates of 60 percent, while medical and health queries reached 36 percent.
When the researchers tested six instruction-tuned models on hallucination detection, even the best performer, Qwen 3 32B, achieved only 64 percent F1 score. Mistral 7B managed just 38 percent. Combining models through ensemble voting did not improve on the single best model, suggesting the models make correlated mistakes.
Categorizing hallucinations proved even harder. Models performed reasonably well on fact-conflicting errors, with Gemma 3 27B reaching 79 percent F1. But context-conflicting hallucinations, where a response contradicts something said earlier in the same conversation, stumped every system. All ensemble results stayed below 10 percent F1 for this category.
AuthenHallu reveals a critical gap between artificial benchmarks and real-world model behavior. When language models face genuine user questions about dates, calculations, or medical facts, they hallucinate far more often than controlled tests suggest, and current detection methods remain unreliable. Explore the full findings and create your own research videos at EmergentMind.com.