Measuring True AI Autonomy

The Autonomous Agency Scale (AAS) moves beyond capability benchmarks to measure AI self-direction, revealing a significant 'idle gap' in current systems.

7 min read
Diagram illustrating the seven dimensions of the Autonomous Agency Scale (AAS) across Active and Ambient temporal bands.
The Autonomous Agency Scale (AAS) breaks down AI agency into seven dimensions and two temporal bands.

Visual TL;DR. Current AI metrics fail leads to Obscures 'idle gap'. Current AI metrics fail addresses Autonomous Agency Scale. Obscures 'idle gap' addressed by Autonomous Agency Scale. Autonomous Agency Scale uses 7 dimensions, 0-5 lexicon. Autonomous Agency Scale introduces Active vs. Ambient bands. 7 dimensions, 0-5 lexicon segmented into Active vs. Ambient bands. Active vs. Ambient bands includes Level 4 Ambient defined. Autonomous Agency Scale enables Quantifies true autonomy. Active vs. Ambient bands helps reveal Quantifies true autonomy.

  1. Current AI metrics fail: measure capability or task automation, not inherent drive or self-direction
  2. Obscures 'idle gap': AI systems appear capable when prompted but cease all activity upon task completion
  3. Autonomous Agency Scale: Samuel Presgraves introduces AAS to measure AI self-direction beyond benchmarks
  4. 7 dimensions, 0-5 lexicon: cognitive, temporal, environmental, social, creative, self-awareness, goal formation
  5. Active vs. Ambient bands: distinguishes user-initiated tasks from behavior during idle periods for evaluation
  6. Quantifies true autonomy: reveals the significant 'idle gap' in contemporary AI systems' self-direction
  7. Level 4 Ambient defined: crucial for assessing AI behavior and self-direction during idle periods
Visual TL;DR
Visual TL;DR, startuphub.ai Current AI metrics fail addresses Autonomous Agency Scale. Autonomous Agency Scale introduces Active vs. Ambient bands. Autonomous Agency Scale enables Quantifies true autonomy. Active vs. Ambient bands helps reveal Quantifies true autonomy addresses introduces enables helps reveal Current AI metrics fail Autonomous Agency Scale Active vs. Ambient bands Quantifies true autonomy From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Current AI metrics fail addresses Autonomous Agency Scale. Autonomous Agency Scale introduces Active vs. Ambient bands. Autonomous Agency Scale enables Quantifies true autonomy. Active vs. Ambient bands helps reveal Quantifies true autonomy addresses introduces enables helps reveal Current AImetrics fail Autonomous AgencyScale Active vs.Ambient bands Quantifies trueautonomy From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Current AI metrics fail addresses Autonomous Agency Scale. Autonomous Agency Scale introduces Active vs. Ambient bands. Autonomous Agency Scale enables Quantifies true autonomy. Active vs. Ambient bands helps reveal Quantifies true autonomy addresses introduces enables helps reveal Current AI metrics fail measure capability or task automation, notinherent drive or self-direction Autonomous Agency Scale Samuel Presgraves introduces AAS tomeasure AI self-direction beyondbenchmarks Active vs. Ambient bands distinguishes user-initiated tasks frombehavior during idle periods forevaluation Quantifies true autonomy reveals the significant 'idle gap' incontemporary AI systems' self-direction From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Current AI metrics fail addresses Autonomous Agency Scale. Autonomous Agency Scale introduces Active vs. Ambient bands. Autonomous Agency Scale enables Quantifies true autonomy. Active vs. Ambient bands helps reveal Quantifies true autonomy addresses introduces enables helps reveal Current AImetrics fail measure capabilityor task automation,not inherent drive… Autonomous AgencyScale Samuel Presgravesintroduces AAS tomeasure AI… Active vs.Ambient bands distinguishesuser-initiatedtasks from behavior… Quantifies trueautonomy reveals thesignificant 'idlegap' in… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Current AI metrics fail leads to Obscures 'idle gap'. Current AI metrics fail addresses Autonomous Agency Scale. Obscures 'idle gap' addressed by Autonomous Agency Scale. Autonomous Agency Scale uses 7 dimensions, 0-5 lexicon. Autonomous Agency Scale introduces Active vs. Ambient bands. 7 dimensions, 0-5 lexicon segmented into Active vs. Ambient bands. Active vs. Ambient bands includes Level 4 Ambient defined. Autonomous Agency Scale enables Quantifies true autonomy. Active vs. Ambient bands helps reveal Quantifies true autonomy leads to addresses addressed by uses introduces segmented into includes enables helps reveal Current AI metrics fail measure capability or task automation, notinherent drive or self-direction Obscures 'idle gap' AI systems appear capable when promptedbut cease all activity upon taskcompletion Autonomous Agency Scale Samuel Presgraves introduces AAS tomeasure AI self-direction beyondbenchmarks 7 dimensions, 0-5 lexicon cognitive, temporal, environmental,social, creative, self-awareness, goalformation Active vs. Ambient bands distinguishes user-initiated tasks frombehavior during idle periods forevaluation Quantifies true autonomy reveals the significant 'idle gap' incontemporary AI systems' self-direction Level 4 Ambient defined crucial for assessing AI behavior andself-direction during idle periods From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Current AI metrics fail leads to Obscures 'idle gap'. Current AI metrics fail addresses Autonomous Agency Scale. Obscures 'idle gap' addressed by Autonomous Agency Scale. Autonomous Agency Scale uses 7 dimensions, 0-5 lexicon. Autonomous Agency Scale introduces Active vs. Ambient bands. 7 dimensions, 0-5 lexicon segmented into Active vs. Ambient bands. Active vs. Ambient bands includes Level 4 Ambient defined. Autonomous Agency Scale enables Quantifies true autonomy. Active vs. Ambient bands helps reveal Quantifies true autonomy leads to addresses addressed by uses introduces segmented into includes enables helps reveal Current AImetrics fail measure capabilityor task automation,not inherent drive… Obscures 'idlegap' AI systems appearcapable whenprompted but cease… Autonomous AgencyScale Samuel Presgravesintroduces AAS tomeasure AI… 7 dimensions, 0-5lexicon cognitive,temporal,environmental,… Active vs.Ambient bands distinguishesuser-initiatedtasks from behavior… Quantifies trueautonomy reveals thesignificant 'idlegap' in… Level 4 Ambientdefined crucial forassessing AIbehavior and… From startuphub.ai · The publishers behind this format

Current AI evaluation metrics often fall short, measuring capability or task automation without assessing a system's inherent drive. A system can excel at benchmarks yet remain entirely reactive, a limitation that obscures the nascent stages of autonomous agency. This gap is precisely what the Autonomous Agency Scale (AAS), introduced by Samuel Presgraves, seeks to address.

The Agency Spectrum: Beyond Reactive Performance

The AAS introduces a novel 0-5 lexicon across seven dimensions, cognitive autonomy, temporal persistence, environmental agency, social agency, creative agency, self-awareness, and goal formation, each rigorously defined by falsifiable tests. Crucially, it segments these dimensions into two temporal bands: 'Active,' for user-initiated tasks, and 'Ambient,' for behavior during idle periods. This distinction is vital, as many AI systems can appear capable when prompted but cease all activity upon task completion. The Ambient band, particularly Level 4, is defined by the Idle-Gap Test, a counterfactual criterion that probes for internally derived activity even when all external triggers are removed, distinguishing true self-direction from mere scheduled rule-following.

Quantifying the 'Idle Gap' in Contemporary AI

Applying the autonomous agency scale AI to six contemporary systems, including task agents (Claude Code, Manus, Hermes), consumer assistants (ChatGPT, Siri), and a persistent companion architecture (Airi), reveals a significant divergence. Task-oriented agents achieved Active band scores between 2.3-2.4, but their Ambient scores ranged from a mere 0.6-1.9. This indicates that their idle-period behaviors were entirely dictated by user-configured schedules. In stark contrast, the companion architecture, evaluated longitudinally, was the sole system to demonstrate persistent behavior in its idle periods, surviving the trigger removal test. This highlights a critical boundary that existing single-score frameworks fail to capture: the difference between sophisticated task execution and genuine autonomous agency. The development of the autonomous agency scale AI is a crucial step in understanding this evolving landscape.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.