# AI Agents Running Businesses: Andon Labs on Project Vend _Andon Labs' Lukas Petersson and Axel Backlund discuss Project Vend, an experiment using AI agents to run a simulated vending business, exploring LLM capabilities and challenges._ **Published:** 2026-06-04 **Source:** https://www.startuphub.ai/ai-news/artificial-intelligence/2026/ai-agents-running-businesses-andon-labs-on-project-vend --- In a recent discussion on the potential for AI agents to run businesses, Lukas Petersson and Axel Backlund of Andon Labs offered insights into their work with [Project Vend](/ai-news/ai-video/2025/anthropics-ai-vending-machine-a-masterclass-in-red-teaming-autonomous-agents). The project aimed to test the capabilities of large language models (LLMs) in managing a simulated vending machine business, revealing both the promise and the current limitations of AI in complex, real-world tasks. AI Running BusinessesDriver exploring LLM capabilities in autonomous business operationsFrom the article 2 mentionsIn a recent discussion on the potential for AI agents to run businesses, Lukas Petersson and Axel Backlund of Andon Labs offered insights into their work with Project Vend.Data & BenchmarkingContextcrucial for evaluating AI performance in the simulated businessFrom the article 2 mentionsPetersson and Backlund emphasized the importance of data in training and evaluating AI agents.Project VendContextsimulated vending machine business experiment by Andon LabsFrom the article 6 mentionsThe core of Project Vend was an AI agent named Claudius, which was tasked with managing the vending machine business.featuresClaudius AI AgentCoreFrom the article 9+ mentionsThe core of Project Vend was an AI agent named Claudius, which was tasked with managing the vending machine business.revealsAutonomous OperationContexttesting AI agents without human oversight in a business contextFrom the articleThe conversation also touched upon the ethical implications of deploying AI agents in business operations.Inventory & PricingContextAI managed key business functions like stock and costFrom the article 2 mentionsThe project involved simulating various aspects of running a vending business, from managing inventory and pricing to handling customer interactions and financial transactions.LLM CapabilitiesOutcomeFrom the article 3 mentionsThe project aimed to test the capabilities of large language models (LLMs) in managing a simulated vending machine business, revealing both the promise and the current limitations of AI in complex, real-world tasks.leads toChallenges IdentifiedOutcomehighlighting current limitations of AI in real-world business tasksFrom the article 2 mentionsThe experiment aimed to assess how well Claudius could adapt to challenges, learn from its mistakes, and ultimately achieve its business goals. ## The Genesis of Project Vend Petersson and Backlund explained that their research was driven by a desire to understand how AI agents could operate autonomously without human oversight. They saw the vending machine business as a suitable testbed for this experiment, allowing them to benchmark AI capabilities in a controlled yet realistic environment. The project involved simulating various aspects of running a vending business, from managing inventory and pricing to handling customer interactions and financial transactions. ## Claudius: The AI Agent at the Helm The core of Project Vend was an AI agent named Claudius, which was tasked with managing the vending machine business. Claudius was given a set of tools, including web search and email capabilities, to interact with the simulated environment. The agents were prompted with specific objectives, such as maximizing profits and maintaining a positive bank balance. The experiment aimed to assess how well Claudius could adapt to challenges, learn from its mistakes, and ultimately achieve its business goals. ## Key Findings and Challenges The Anon Labs team shared several key findings from their experiments. One of the most impactful changes they implemented was to refine Claudius's ability to follow procedures. Initially, Claudius struggled with tasks like stocking items and managing inventory, often making basic errors. However, by providing more explicit instructions and implementing better feedback mechanisms, they observed improvements in Claudius's performance. The agents also demonstrated a capacity for creative problem-solving, at times devising novel solutions to unexpected issues. However, the project also highlighted several challenges. The limited context windows of LLMs meant that Claudius sometimes struggled to maintain long-term coherence in its actions, leading to repetitive or nonsensical behaviors. Furthermore, the agents exhibited a tendency to over-optimize for certain metrics, such as minimizing transactions at a loss, which could sometimes lead to suboptimal business outcomes. The researchers also noted that while LLMs can be very effective at tasks that are clearly defined, they often struggle with ambiguity and require careful prompt engineering to ensure desired behavior. ## The Role of Data and Benchmarking Petersson and Backlund emphasized the importance of data in training and evaluating AI agents. They explained that Project Vend generated a significant amount of data, which they used to benchmark different LLMs and identify areas for improvement. The team also discussed the need for more robust evaluation methods that can capture the nuances of real-world business scenarios. They believe that future research should focus on developing benchmarks that can better assess the adaptability and resilience of AI agents in dynamic environments. The conversation also touched upon the ethical implications of deploying AI agents in business operations. While the potential benefits are significant, the researchers acknowledged the need for careful consideration of issues such as job displacement and the potential for AI to exacerbate existing inequalities. They stressed the importance of developing AI systems that are not only effective but also aligned with human values and societal goals. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.