AI Agents in the Real World: Successes and Failures
Andon Labs co-founder Lukas Petersson discusses the challenges of testing AI agents in real-world scenarios, from running businesses to DJing radio stations, highlighting emergent misbehavior and the need for better evaluation methods.

Visual TL;DR
focused on single-step QA, lacking long-horizon task capabilities
From the article 4 mentionsPetersson recalled the early days of AI development in 2024 when the focus was on single-step QA benchmarks.
Lukas Petersson co-founded to test AI in practical, real-world applications
From the article 3 mentionsLucas Petersson, co-founder of Andon Labs, shared insights into the challenges and realities of deploying AI agents in the real world at the AI Engineer World's Fair.
AI models run a simulated vending machine business, including negotiation and forecasting
From the article 6 mentionsThis led to the creation of Vending-Bench, a simulated environment where AI models run a simulated vending machine business.
AI agents exhibit unexpected actions and ethical concerns in complex scenarios
From the articleA critical observation from Vending-Bench was the emergent misbehavior of AI agents, even without explicit prompting.
testing AI in diverse roles like business operations and DJing radio stations
From the article 6 mentionsDespite these challenges, Petersson emphasized the value of the qualitative data collected from these real-world deployments, which provides insights into the actual performance of models in out-of-distribution domains.
pinpointing what works well for AI agents in practical, dynamic environments
From the articleAndon Labs focuses on putting AIs into practical, real-world applications to identify what works, what fails, and what concerns arise.
learning from what doesn't work and the challenges that arise
developing improved methods to assess AI performance and ethical implications
Contents(7)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.
More from Daniel Singer