Visual TL;DR. LLM Hallucinations leads to Limited Detection. Limited Detection needs new approach MoE Architectures. MoE Architectures generate Internal MoE Signals. Internal MoE Signals exploited by InnerExpert Framework. MoE Architectures basis for InnerExpert Framework. InnerExpert Framework enables Per-Token Detection. Per-Token Detection results in Superior Performance.
- LLM Hallucinations: generating factually incorrect content, often at answer or sentence level
- Limited Detection: current methods fail to pinpoint exact source of falsehoods, hindering precise interventions
- MoE Architectures: utilize routing mechanism to activate sparse subsets of 'experts' within each layer
- Internal MoE Signals: router entropy, expert disagreement, and usage patterns previously overlooked for detection
- InnerExpert Framework: first method designed to exploit MoE internal signals for hallucination detection
- Per-Token Detection: leverages MoE internal signals for granular, per-token hallucination detection
- Superior Performance: achieving state-of-the-art results with high efficiency in hallucination detection
Visual TL;DR
