Autonomy at Levels: Structured Decision-Making
- Autonomy at Levels is a framework that defines decision-making authority as a graded continuum rather than a binary attribute.
- It categorizes 21 decision domains into three areas—Technical Core, Technical Supporting, and Work Organization & Management—to balance self-organization with coordinated control.
- The model applies across domains like agile software, AI, and cybersecurity, offering actionable insights for diagnosing and optimizing distributed governance.
Autonomy at levels is the treatment of autonomy as a graded allocation of decision authority rather than a binary property. In large-scale agile software development, the term denotes five nested levels at which decisions can be made—individual, team, leader/expert role, organization, and project—together with 21 autonomy categories grouped into three areas, so that autonomy can be analyzed as a structured distribution of authority across interdependent teams rather than as a single attribute (Lassenius et al., 4 Mar 2025). Across adjacent literatures, comparable level-based schemes classify autonomy by decision-making locus, self-direction, execution authority, planning sophistication, and human oversight, with staged progressions appearing in cybersecurity AI, radiology, spacecraft, AI agents, AGI, process execution management, and autonomous scientific facilities (Mayoral-Vilches, 30 Jun 2025, Ghuwalewala et al., 2021, Baker et al., 2 Mar 2025, Feng et al., 14 Jun 2025, Morris et al., 2023, Aalst, 2022, Houx, 11 Jan 2026).
1. Rationale for level-based autonomy
In single-team agile development, more autonomy is generally considered better. In large-scale agile development, however, several teams collaborate on the same software with technical and social dependencies, so the central problem is not maximal autonomy but the balance between autonomy and organizational control. The preliminary taxonomy in "Towards a Taxonomy for Autonomy in Large-Scale Agile Software Development" identifies five levels and 21 categories of autonomy grouped into three areas, with the stated purpose of helping practitioners identify the “sweet spot” between self-organization and top-down control, specify which decisions teams may take and which must be escalated, compare and tailor scaling frameworks such as SAFe and LeSS, and diagnose autonomy bottlenecks in programs and transformation efforts (Lassenius et al., 4 Mar 2025).
The motivating failure modes are explicit. Unbounded autonomy in large-scale agile can lead to coordination breakdowns, reduced quality, and misalignment with organizational goals; too little autonomy undermines motivation, creativity, and rapid problem solving. The framework therefore does not prescribe a universal optimum. Instead, it decomposes autonomy by decision category and by decision locus. This suggests that autonomy is best treated as a governance problem over coupled decision rights, not as a monolithic “on/off” switch.
A closely related theoretical move appears in work on AGI and AI agents. "Levels of AGI for Operationalizing Progress on the Path to AGI" treats autonomy as an axis orthogonal to performance and generality, characterized by decision-making locus, self-direction, and human oversight; "Levels of Autonomy for AI Agents" similarly frames autonomy as a deliberate design decision, separate from capability and operational environment, and anchors each level in the role the user plays when interacting with the agent (Morris et al., 2023, Feng et al., 14 Jun 2025). In both cases, the underlying assumption is the same: autonomy is not exhausted by competence.
2. The five decision levels in large-scale agile development
The agile taxonomy distinguishes five nested levels at which decisions can be made. No numerical metrics or thresholds are prescribed; the levels are distinguished by the role or unit that has authority and by the scope of the decision (Lassenius et al., 4 Mar 2025).
| Level | Who decides | Typical scope |
|---|---|---|
| Individual | An individual team member | One’s own work |
| Team | The agile team as a whole | Internal organization of team work |
| Leader/Expert Role | Leadership or specialist roles | Cross-team or expertise-heavy decisions |
| Organization | Centralized functions above teams | Standards, budgets, governance |
| Project | Temporary project or program wrapper | Project trade-offs, scope, milestones, budgets |
At the individual level, decisions concern one’s own work, such as which task to pick next, a personal work-from-home day, or a pairing partner. The authority boundary is explicit: an individual cannot override team commitments or cross-team coordination rules. At the team level, the agile team decides how it organizes its own work, including task assignment, sprint length, “Definition of Done,” and internal scheduling. Decisions affecting other teams, such as architecture changes or release cadence, typically require escalation (Lassenius et al., 4 Mar 2025).
The leader/expert role level covers decisions made by formal or informal leadership or specialist roles such as team lead, architect, UX expert, security officer, or Release Train Engineer in SAFe. These decisions require domain expertise or cross-team influence but do not derive from corporate strategy. The organization level covers line management, Communities of Practice, HR, enterprise architecture boards, or TPU (Test Process Unit), and includes standards, budgets, governance, reporting mechanisms, enterprise toolsets, and company-wide security requirements. The project level is cross-cutting: project steering committees, project managers, or project offices decide project-level scope changes, milestone gates, trade-offs between teams, and project budgets; this level may override team or individual preferences for the sake of delivery (Lassenius et al., 4 Mar 2025).
A notable feature of the taxonomy is that these levels are nested but not strictly hierarchical in the simple sense of “higher is stronger.” Project authority is temporary and cross-cutting; expert authority is domain-specific; team authority is operational and local. This suggests that the problem of autonomy in large-scale agile is not merely decentralization, but structured coexistence among multiple decision centers.
3. Three areas and twenty-one categories of autonomy
The taxonomy groups 21 categories into three areas: Technical Core, Technical Supporting, and Work Organization & Management. The categories are not equivalent in risk or coupling, and their “typical level mappings” vary by organization (Lassenius et al., 4 Mar 2025).
In Technical Core, the seven categories are Requirements, Architecture, Design, Construction, Testing, Operations, and Maintenance. Requirements concern what features or stories enter the backlog and are typically controlled by the team via the Product Owner, though road-map prioritization may be organizational. Architecture concerns high-level system structure and technologies and is typically controlled by expert roles and organization. Design, Construction, and much of Testing are usually team-level. Operations and Maintenance may sit with the team, a DevOps team, or an Operations department. When teams control most of the Technical Core, the framework states that they have high technical autonomy; loss of autonomy here forces teams into “follow me” mode (Lassenius et al., 4 Mar 2025).
In Technical Supporting, the categories are Software Configuration Management, Quality, and Security. These are typically more centralized: version-control strategy and branching model often sit at organization level; quality definitions and peer-review guidelines may be set by a Quality Assurance group; security policies and vulnerability management are typically controlled by an InfoSec function. The stated rationale is cross-team consistency and risk management, although too much centralization delays teams (Lassenius et al., 4 Mar 2025).
In Work Organization & Management, the eleven categories are Process, Models & Methods, Infrastructure & Tools, Budgeting, Resource Allocation, Scheduling, Measurement & Monitoring, Work Location, Work Hours, Team Composition, and Collaborators. These categories define the conditions under which teams self-organize. Process may be set by the organization or tailored by the team; models and methods may be organizational or team-level; infrastructure and tools are often centralized; budgeting and resource allocation typically sit with organization or project; scheduling is split between sprint-level team control and broader release horizons controlled by organization or project; measurement and monitoring may be local or enterprise-wide; work location, work hours, team composition, and collaborators vary across individual, team, project, and organizational constraints (Lassenius et al., 4 Mar 2025).
The principal methodological claim is therefore decompositional: autonomy should be analyzed category by category, with each category mapped to the lowest possible level that still maintains cross-team alignment and risk control. The paper explicitly advises against treating autonomy as a monolithic property.
4. Illustrative allocations in scaling frameworks and organizations
The illustrative application in the agile paper compares four cases: LeSS, SAFe, Organization A, and Organization B (Lassenius et al., 4 Mar 2025).
Under LeSS, requirements are owned by the teams’ Product Owners; architecture is handled by teams guided by a common architect; and most other decisions, including design, construction, testing, and scheduling, rest at team level, while process, budgeting, and resource allocation are centrally controlled. Under SAFe, the product backlog and program backlog are split; architecture decisions are taken by “Architects,” identified as an expert role; and release train cadence, process, tools, budgeting, and monitoring are set at ART level (Lassenius et al., 4 Mar 2025).
Organization A, described as a medium-sized programme, moved testing and operational maintenance into a shared QA and Ops team, corresponding to expert roles, while construction, design, and requirements remained with squads. Teams could choose their own working-hours pattern and collaboratively decide team composition, but tooling and process were centrally enforced. Organization B, described as an experimental transformation, pushed almost all decision categories down to teams or individuals—including process, metrics, work location, and team composition—but ran into coordination and predictability problems. The explicit lesson drawn is that unbounded autonomy across all 21 categories is not sustainable at scale (Lassenius et al., 4 Mar 2025).
These cases clarify that the taxonomy is intended as both a diagnostic tool and a design aid. It does not rank frameworks by a single autonomy score. Instead, it exposes how different scaling approaches distribute authority across distinct categories, and how those allocations interact with inter-team dependencies. A plausible implication is that autonomy in large-scale agile is best understood as a patterned portfolio of delegated rights rather than as a single maturity trajectory.
5. Cross-domain taxonomies of autonomy levels
Comparable autonomy taxonomies recur across domains, but their criteria differ according to operational risk, task structure, and the acceptable role of human oversight. In cybersecurity AI, "Cybersecurity AI: The Dangerous Gap Between Automation and Autonomy" defines six levels from Level 0 to Level 5 using a capability vector for autonomous planning, scanning, exploiting, and mitigating. Level 2 corresponds to LLM-assisted planning suggestions with continuous human monitoring; Level 3 to semi-automated execution in known scenarios; Level 4 to coverage of the full adversarial lifecycle in well-defined operational domains; and Level 5 to an aspirational “AI hacker” requiring zero human intervention for any task in any environment. The paper states that current “autonomous” pentesters operate at Level 3–4 and that true Level 5 autonomy remains aspirational (Mayoral-Vilches, 30 Jun 2025).
Radiology presents an analogous but clinically grounded scale. "Levels of Autonomous Radiology" defines six levels from no automation to full automation. Level 2 is partial automation in which AI proposes but does not decide; Level 3 is conditional automation in which AI can autonomously interpret only under a narrow set of predefined conditions with the radiologist “on call”; Level 4 handles the vast majority of routine cases without any human in the loop; and Level 5 is a fully autonomous end-to-end system from image acquisition to finalized report, with continual learning, self-monitoring, and self-updating (Ghuwalewala et al., 2021). The recurring mid-level pattern is conditional autonomy inside a bounded operational domain with fallback to a human.
A similar progression appears in process and laboratory settings. "Six Levels of Autonomous Process Execution Management (APEM)" moves from manual orchestration through detection, recommendation, limited automated response, and fully autonomous behavior in known contexts to end-to-end autonomy. "Benchmarking Autonomy in Scientific Experiments" defines the BASE Scale from Level 0 to Level 5 and identifies the Inference Barrier at Level 3, where decision-making shifts from scalar feedback to semantic digital twins under a latency budget, thereby extending the decision manifold from spatial exploration to temporal gating (Aalst, 2022, Houx, 11 Jan 2026). In both, the decisive transition is from reactive or scripted behavior to semantically informed intervention.
AI-centered taxonomies make the role of the human operator even more explicit. "Levels of Autonomy for AI Agents" defines five levels characterized by user roles—operator, collaborator, consultant, approver, and observer—and links each role to corresponding control mechanisms such as invocation gating, dynamic delegation, consultation triggers, approval conditions, and emergency off-switches (Feng et al., 14 Jun 2025). "Levels of AGI for Operationalizing Progress on the Path to AGI" defines autonomy as the degree of control and initiative ceded by humans to an AI system, emphasizing decision-making locus, self-direction, and human oversight, and ranging from “No AI” through “AI as a Tool,” “Consultant,” “Collaborator,” and “Expert” to “AI as an Agent” (Morris et al., 2023).
Space systems support both top-down and bottom-up conceptions. "Spacecraft Autonomy Levels" proposes SAL 0–5, a six-level scale from “Basic Spacecraft Controllability” to “Autonomous,” with decision authority, execution authority, planning sophistication, and ground-intervention frequency as distinguishing criteria (Baker et al., 2 Mar 2025). "Autonomy at Levels for Spacecraft" instead proposes a bottom-up engineering paradigm in which autonomy elements are embedded at every level of spacecraft decomposition—systems, subsystems, assemblies, components—and every control loop is wrapped by an outer autonomy loop (Baker et al., 19 Aug 2025). A different theoretical perspective appears in "Constitutive Components for Human-Like Autonomous Artificial Intelligence," which defines Core Functions, the Integrative Evaluation Function, and the Self Modification Function, and on that basis proposes reactive, weak autonomous, and strong autonomous levels (Yamada, 15 Jun 2025).
Taken together, these frameworks indicate that “autonomy at levels” is not a single taxonomy but a family resemblance across taxonomies. Common elements include graded delegation, operational domains, explicit fallback conditions, and the coupling of capability claims to oversight regimes.
6. Measurement, dynamic allocation, and evaluative frameworks
Not all autonomy taxonomies are equally operationalized. In the large-scale agile taxonomy, the framework remains conceptual and does not prescribe concrete formulas or numerical thresholds for autonomy (Lassenius et al., 4 Mar 2025). By contrast, some adjacent work proposes explicit measures.
"A Measure for Level of Autonomy Based on Observable System Behavior" defines an observed level of autonomy as an edit distance between a sequence of human-equivalent actions and a sequence of observed system actions , namely , where is the Damerau–Levenshtein edit distance. The paper maps the resulting observational score to discrete autonomy levels and proposes the measure as a runtime, observation-based predictor of autonomy “in the wild” (Pittman, 2024).
"Learning to Optimize Autonomy in Competence-Aware Systems" provides a more decision-theoretic formalization. It defines an autonomy model , where is the set of autonomy levels, restricts which levels are allowed in each state–action pair, and gives the autonomy cost of transitioning between levels. The competence-aware system augments states and actions so that a policy selects both what to do and how autonomously to do it, minimizes total expected cost including base execution cost and human-assistance cost, and under stated assumptions converges to a competence map
This is a direct formalization of level selection as an online optimization problem (Basich et al., 2020).
Dynamic allocation has also been studied empirically. In "An Analysis of Human-Robot Information Streams to Inform Dynamic Autonomy Allocation," the autonomy system is partitioned into discrete levels 0 corresponding to teleoperation, autonomous stopping, and blended autonomy, with a classifier 1 predicting whether to shift up, not shift, or shift down. The reported result is that interaction features between human and robot are the most informative for predicting when to shift autonomy levels; among classical learners, interaction-only features reach 85.4% balanced accuracy and all streams 91.9% (Miller et al., 2021). This suggests that level transitions are not merely predefined thresholds but can themselves be learned.
In medical AI, evaluation is explicitly level-conditioned. "Evaluating Medical LLMs by Levels of Autonomy" aligns L0–L3 with different permitted actions and therefore different evidence requirements: informational tools are evaluated with factual accuracy, readability, and calibration; information transformation and aggregation require extraction and provenance metrics; decision support requires safety, selective prediction, and subgroup analysis; supervised agents require tool-use correctness, human acceptance rate, and auditability (Ye et al., 20 Oct 2025). The general principle is that autonomy claims become credible only when their metrics match the actions permitted at that level.
7. Governance, misconceptions, and implications
A recurrent misconception is the collapse of automation into autonomy. The cybersecurity taxonomy argues that the industry often combines “automated” and “autonomous” AI, creating dangerous misconceptions about system capabilities. Its central warning is that organizations deploying mischaracterized “autonomous” tools risk reducing oversight precisely when it is most needed (Mayoral-Vilches, 30 Jun 2025). The same concern appears, in domain-specific form, in radiology, process execution management, medical LLM evaluation, and scientific facilities: middle levels are not equivalent to full autonomy, and bounded domains matter (Ghuwalewala et al., 2021, Aalst, 2022, Ye et al., 20 Oct 2025, Houx, 11 Jan 2026).
As autonomy increases, the associated risk profile changes. In the AGI framework, lower autonomy levels foreground risks such as de-skilling, misinformation, over-trust, manipulation, and rapid societal change, while higher levels raise concerns about mass displacement, misalignment, concentration of power, and existential threats. The framework explicitly states that capabilities enable but do not mandate particular autonomy paradigms, and that designers can choose lower autonomy than the maximum capabilities would allow (Morris et al., 2023). This makes autonomy not only a capability descriptor but a deployment choice.
Several works therefore tie level taxonomies directly to governance. "Levels of Autonomy for AI Agents" proposes AI autonomy certificates, digitally signed documents issued by a third-party body declaring that an agent may operate at autonomy level 2, after the developer submits an “autonomy case” and the governing body verifies the claim through testing (Feng et al., 14 Jun 2025). "Evaluating Medical LLMs by Levels of Autonomy" similarly links each level to a blueprint for metric choice, evidence assembly, and claim reporting, and couples higher levels to stronger oversight, provenance, and audit requirements (Ye et al., 20 Oct 2025). In spacecraft, the proposed autonomy levels are explicitly presented as useful in education and communication with lawmakers, government officials, and the general public (Baker et al., 2 Mar 2025).
A plausible implication is that autonomy-at-levels frameworks function simultaneously as engineering abstractions, HCI design tools, validation schemas, and public accountability devices. Their enduring value lies less in any single numbering system than in the insistence that autonomy claims be decomposed into who decides, who acts, under what constraints, with what fallback, and under whose oversight.