Personalized AI Agents Now Have a Benchmark

A new iOSWorld benchmark reveals AI agents' struggles with personalized, multi-app tasks, highlighting the need for richer context and advanced reasoning capabilities.

Illustration of interconnected iOS apps and data streams representing the iOSWorld benchmark.
The iOSWorld benchmark provides a realistic simulation environment for evaluating AI agents' ability to interact with personalized user data across multiple iOS applications.
Visual TL;DR
AI Agents StruggleDriver
current AI agents fail at personalized, multi-app tasks
From the article 2 mentionsThe quest for truly intelligent personal AI agents hinges on their ability to move beyond stateless instruction following to deeply understand and reason over a user's unique identity, history, and preferences.
Impersonal SandboxesDriver
From the articleCurrent benchmarks, however, fall short by operating in impersonal sandboxes, failing to reflect the rich, interconnected data residing on a user's device.
Need for ContextContext
AI needs to understand user identity, history, preferences
From the articleIt provides the community with a vital tool to rigorously assess and advance the capabilities of AI agents in realistic, personalized contexts.
Introducing iOSWorldCore
From the article 3 mentionsTo address this critical gap, researchers have introduced iOSWorld, the first interactive, native iOS simulator benchmark.
Realistic User DataCore
simulates digital life with 26 interconnected apps
From the articleCurrent benchmarks, however, fall short by operating in impersonal sandboxes, failing to reflect the rich, interconnected data residing on a user's device.
Improved AI AgentsEffect
enables better reasoning over personalized user context
From the article 2 mentionsThe quest for truly intelligent personal AI agents hinges on their ability to move beyond stateless instruction following to deeply understand and reason over a user's unique identity, history, and preferences.
New Benchmark TasksContext
133 tasks across single, multi-app, and memory tiers
From the article 5 mentionsThe release of iOSWorld as an open-source benchmark, complete with apps, seeded data, tasks, rubrics, and evaluation code, marks a pivotal moment for AI research.

The quest for truly intelligent personal AI agents hinges on their ability to move beyond stateless instruction following to deeply understand and reason over a user's unique identity, history, and preferences. Current benchmarks, however, fall short by operating in impersonal sandboxes, failing to reflect the rich, interconnected data residing on a user's device.

Bridging the Personalization Chasm with iOSWorld

To address this critical gap, researchers have introduced iOSWorld, the first interactive, native iOS simulator benchmark. This novel environment is built around a persistent user identity and encompasses 26 newly developed iOS apps. These apps feature interconnected data streams, including transactions, messages, travel records, social connections, and financial activity, creating a realistic simulation of a user's digital life. iOSWorld is structured with 133 tasks across three difficulty tiers: single-app tasks (27), multi-app tasks spanning 2 to 8 apps (60), and memory and personalization tasks requiring inference from personal data (46).

Performance Realities and the Power of Context

Evaluations on the iOSWorld benchmark using frontier and open-source models highlight significant challenges. The best-performing configuration achieved only 52% overall accuracy, dropping to a stark 37% on multi-app tasks. Notably, privileged vision+XML access provided substantial gains of up to 26 percentage points for frontier models, underscoring the importance of richer input modalities. Smaller models, however, did not demonstrate similar benefits from this enhanced accessibility-tree input, suggesting architectural or training limitations.

The release of iOSWorld as an open-source benchmark, complete with apps, seeded data, tasks, rubrics, and evaluation code, marks a pivotal moment for AI research. It provides the community with a vital tool to rigorously assess and advance the capabilities of AI agents in realistic, personalized contexts.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.