Visual TL;DR. Parth Asawa (UCB PhD) presents at Rethink Evaluation. Rethink Evaluation critiques Current AI Benchmarks. Current AI Benchmarks leads to Fails Continual Learning. Fails Continual Learning requires Need New Benchmark. Need New Benchmark defines True Learning Measure. Parth Asawa (UCB PhD) at AI Engineer Fair.
- Parth Asawa (UCB PhD): UC Berkeley PhD student challenges current AI evaluation methods for LLMs
- Current AI Benchmarks: evaluate LLMs on isolated tasks, ignoring past experiences and learning
- Fails Continual Learning: this method cannot measure how models learn and adapt over time
- Need New Benchmark: calls for a new approach to truly assess models' learning ability
- True Learning Measure: performance should improve as a function of prior experience, not restart
- Rethink Evaluation: critical look at how large language models are currently evaluated
- AI Engineer Fair: presented his critical findings at the AI Engineer World's Fair
Visual TL;DR
