Visual TL;DR. Longer LLM Contexts leads to Challenging 'More Better'. Challenging 'More Better' by Uzunoglu et al. Study. Longer LLM Contexts causes Information Paradox. Information Paradox results in Increased Context Reliance. Longer LLM Contexts shows Diminishing Returns. Diminishing Returns seen in Pretraining Scenarios. Increased Context Reliance contributes to Performance Degradation. Diminishing Returns leads to Performance Degradation.
- Longer LLM Contexts: prevailing wisdom assumes more data always equates to better performance
- Information Paradox: abundant relevant information reduces incentive to encode parametrically
- Increased Context Reliance: models become detrimentally reliant on immediate context during inference
- Diminishing Returns: performance gains are not linear with increasing context window size
- Performance Degradation: hinders parametric knowledge, leading to overall performance degradation
- Challenging 'More Better': new research challenges the assumption that more context is always better
- Uzunoglu et al. Study: researchers Uzunoglu, van Durme, and Khashabi propose this paradox
- Pretraining Scenarios: tasks like NLU and closed-book QA show improvement only up to intermediate length
Visual TL;DR
