1 articles with this tag
Lyft's Nick Ung and Ashe discuss building effective AI agent evaluations, emphasizing realistic user simulation, actionable metrics, and statistical rigor.