UC Berkeley PhD Student Challenges AI Evaluation Methods
Parth Asawa, a PhD student at UC Berkeley, argues that current AI evaluation methods are insufficient for measuring continual learning and calls for a new benchmark approach.

Visual TL;DR
UC Berkeley PhD student challenges current AI evaluation methods for LLMs
From the articleAt the AI Engineer World's Fair, Parth Asawa, a PhD student at UC Berkeley, presented a critical look at how artificial intelligence, particularly large language models (LLMs), are evaluated.
critical look at how large language models are currently evaluated
From the article 3 mentionsAsawa highlighted how the typical evaluation of LLMs involves presenting them with a series of independent tasks.
From the articleAt the AI Engineer World's Fair, Parth Asawa, a PhD student at UC Berkeley, presented a critical look at how artificial intelligence, particularly large language models (LLMs), are evaluated.
evaluate LLMs on isolated tasks, ignoring past experiences and learning
From the article 5 mentionsAsawa proposed three key criteria for designing effective continual learning benchmarks:
this method cannot measure how models learn and adapt over time
From the article 9+ mentionsAsawa argued that the current approach, which often treats each task in isolation, fails to measure a crucial aspect of intelligence: the ability to learn and adapt over time, a concept known as continual learning.
calls for a new approach to truly assess models' learning ability
From the article 4 mentionsShared Structure: To enable improvement from prior experience, tasks need to have underlying shared latent structures that models can exploit.
From the article 3 mentionsHe contrasted this with what a true measure of learning ability should look like: performance that improves as a function of prior experience, rather than a scattered, inconsistent performance.
Contents(4)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.