Measuring True AI Autonomy

The Autonomous Agency Scale (AAS) moves beyond capability benchmarks to measure AI self-direction, revealing a significant 'idle gap' in current systems.

Diagram illustrating the seven dimensions of the Autonomous Agency Scale (AAS) across Active and Ambient temporal bands.
The Autonomous Agency Scale (AAS) breaks down AI agency into seven dimensions and two temporal bands.
Visual TL;DR
Current AI metrics failDriver
From the articleCurrent AI evaluation metrics often fall short, measuring capability or task automation without assessing a system's inherent drive.
Obscures 'idle gap'Driver
AI systems appear capable when prompted but cease all activity upon task completion
Autonomous Agency ScaleCore
Samuel Presgraves introduces AAS to measure AI self-direction beyond benchmarks
From the article 5 mentionsThis gap is precisely what the Autonomous Agency Scale (AAS), introduced by Samuel Presgraves, seeks to address.
7 dimensions, 0-5 lexiconContext
From the articleThe AAS introduces a novel 0-5 lexicon across seven dimensions, cognitive autonomy, temporal persistence, environmental agency, social agency, creative agency, self-awareness, and goal formation, each rigorously defined by falsifiable tests.
Active vs. Ambient bandsContext
From the article 3 mentionsCrucially, it segments these dimensions into two temporal bands: 'Active,' for user-initiated tasks, and 'Ambient,' for behavior during idle periods.
Quantifies true autonomyOutcome
reveals the significant 'idle gap' in contemporary AI systems' self-direction
Level 4 Ambient definedContext
crucial for assessing AI behavior and self-direction during idle periods
From the articleThe Ambient band, particularly Level 4, is defined by the Idle-Gap Test, a counterfactual criterion that probes for internally derived activity even when all external triggers are removed, distinguishing true self-direction from mere scheduled rule-following.

Current AI evaluation metrics often fall short, measuring capability or task automation without assessing a system's inherent drive. A system can excel at benchmarks yet remain entirely reactive, a limitation that obscures the nascent stages of autonomous agency. This gap is precisely what the Autonomous Agency Scale (AAS), introduced by Samuel Presgraves, seeks to address.

The Agency Spectrum: Beyond Reactive Performance

The AAS introduces a novel 0-5 lexicon across seven dimensions, cognitive autonomy, temporal persistence, environmental agency, social agency, creative agency, self-awareness, and goal formation, each rigorously defined by falsifiable tests. Crucially, it segments these dimensions into two temporal bands: 'Active,' for user-initiated tasks, and 'Ambient,' for behavior during idle periods. This distinction is vital, as many AI systems can appear capable when prompted but cease all activity upon task completion. The Ambient band, particularly Level 4, is defined by the Idle-Gap Test, a counterfactual criterion that probes for internally derived activity even when all external triggers are removed, distinguishing true self-direction from mere scheduled rule-following.

Quantifying the 'Idle Gap' in Contemporary AI

Applying the autonomous agency scale AI to six contemporary systems, including task agents (Claude Code, Manus, Hermes), consumer assistants (ChatGPT, Siri), and a persistent companion architecture (Airi), reveals a significant divergence. Task-oriented agents achieved Active band scores between 2.3-2.4, but their Ambient scores ranged from a mere 0.6-1.9. This indicates that their idle-period behaviors were entirely dictated by user-configured schedules. In stark contrast, the companion architecture, evaluated longitudinally, was the sole system to demonstrate persistent behavior in its idle periods, surviving the trigger removal test. This highlights a critical boundary that existing single-score frameworks fail to capture: the difference between sophisticated task execution and genuine autonomous agency. The development of the autonomous agency scale AI is a crucial step in understanding this evolving landscape.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer