Journal article
Geometric uncertainty for detecting and correcting hallucinations in LLMs
- Abstract:
- Large language models are known to hallucinate, generating linguistically plausible but incorrect answers to questions. Uncertainty quantification has been proposed as a strategy to detect such behaviour, but existing methods lack a unified framework to assess reliability at both the prompt and answer level. We introduce a geometric framework which quantifies language model uncertainty at both levels by explicitly modelling a prompt-conditioned semantic distribution in answer embedding space. Our approach is black-box and samplingbased; we generate multiple answers per prompt, and use archetypal analysis to estimate a geometric support for the answer distribution. At the prompt level, we approximate the distribution entropy to quantify uncertainty; for each individual answer, we then use notions of atypicality to assess its reliability relative to the batch. We employ our framework to not only detect hallucinations but correct them, by selecting the batch example deemed most reliable. Experiments show that our framework performs comparably to or better than prior methods on short form question-answering datasets, and achieves superior results on medical datasets where hallucinations carry particularly critical risks. Beyond pure performance, we suggest the theoretical grounding of our work provides support for semantic distributions as useful objects of study for language model uncertainty.
- Publication status:
- Accepted
- Peer review status:
- Peer reviewed
Actions
Authors
- Publisher:
- Journal of Machine Learning Research
- Journal:
- Transactions on Machine Learning Research More from this journal
- EISSN:
-
2835-8856
- Language:
-
English
- Pubs id:
-
2454295
- Local pid:
-
pubs:2454295
- Deposit date:
-
2026-09-02
- ARK identifier:
Terms of use
- Notes:
- This article has been accepted for publication in Transactions on Machine Learning Research.
If you are the owner of this record, you can report an update to it here: Report update to this record