papersSEP 12 04:00 UTC
arXiv Paper Probes Knowledge Attribution to Distinguish Hallucination Types in LLMs
A new arXiv preprint proposes probing methods to trace where large language models source their knowledge, aiming to separate two kinds of hallucination. The authors distinguish faithfulness violations, where a model mishandles context it was given, from factuality violations, where its answers stem from incorrect stored knowledge. The work targets better attribution of model outputs to internal knowledge versus supplied context.