Interpreting the causes of hallucinations in large language models
Develop interpretability methods and causal diagnostics that explain why large language models hallucinate, including identifying query- and model-specific mechanisms that lead to hallucinations in retrieval-augmented generation systems used for legal research.
References
Interpreting why an LLM hallucinates is an open problem.
Although xAI has made progress in revealing how LLMs produce outputs, we still lack clear insight into why specific hallucinations emerge.
For instance, although the Phi-4 model, which has approximately twice as many parameters as Llama 3.1 8B Instruct and Mistral 7B Instruct, exhibits notably different behavior, it remains unclear whether these differences are attributable to model size or other architectural factors.
Still, it remains an open question how internal model dynamics differ when hallucination is caused by genuine knowledge gaps rather than misaligned context sharing.
We hypothesize that confident hallucination may generalize beyond medicine; any domain where models have strong parametric priors and receive outcome-only rewards could exhibit this pattern.