What does this framework help you decide?
Choose the simplest approach that can do the job effectively. Answer a few questions about the work, the value of independent decisions, the consequences of mistakes and the controls you have in place.
You'll get a suggested approach, a level of autonomy, human-oversight guidance and an explanation of the trade-offs. It's a decision aid, not permission to deploy.
Describe the work
Move the sliders to reflect the actual task, not the technology you want to deploy. All questions use a 0–4 scale.
1 · Is agency actually needed?
Higher values increase the potential value of independent planning and action.
2 · Is an AI model needed?
Variable or unstructured inputs can justify AI even when an agent is unnecessary.
3 · How much authority is safe to delegate?
These questions constrain recommended autonomy; they do not create a need for an agent.
4 · What additional value does autonomy deliver?
Assess incremental benefit over the best feasible workflow, not the value of the task overall.
5 · How mature are the governing controls?
Score controls that are implemented and evidenced, rather than planned capabilities.
6 · Human oversight preference
Choose an operating pattern to test against the assessment. This preference never overrides required approvals or foundational controls.
7 · Non-negotiable controls
Shared vocabulary
Precise terms help avoid treating uncertainty as though it were the same thing as autonomy.
Scoring and decision rules (fully transparent)
Agency need is the average of four task questions; AI need averages two interpretation questions. Delegation concern averages impact, inverted reversibility and scope. Incremental value averages two value questions; evidenced control maturity averages three maturity questions. The indicative permissible-autonomy ceiling is min(4, 4 − delegation concern + (control maturity − 2) × 0.75), floored at 0; mandatory approval caps the ceiling at 2, and missing foundational controls caps it at 1. This is a discussion aid, not a quantitative risk model or approval. Low agency need (<1.6) favours workflow or AI-assisted workflow. Higher agency need can justify an agent only if incremental autonomy value is at least 1.8 and permissible autonomy is at least 1.6; otherwise favour a workflow or model-assisted workflow until controls improve. High autonomy requires agency need ≥2.8, incremental value ≥2.8, and a permissible-autonomy ceiling ≥2.8. Where need exceeds the ceiling, the output explicitly flags an autonomy–assurance gap. Thresholds are illustrative, not validated risk scoring or deployment authorization.
Human oversight: a separate design dimension
The four architectural options do not prescribe a single human-oversight pattern. Separate who plans from who authorises actions, monitors execution and owns the outcome.
| Architecture | Illustrative oversight pattern |
|---|---|
| Workflow automation | Often unattended execution plus exception handling. May include a mandatory human approval step. |
| AI-assisted workflow | Human review of uncertain AI outputs or high-impact downstream actions; routine, low-risk inference may be automatic. |
| Constrained agent | Agent plans within limits; humans authorise specified actions, exceptions, or changes of scope. Monitoring and stop mechanisms support oversight between gates. |
| Higher-autonomy agent | Bounded unattended execution, live monitoring, escalation triggers, interruptibility and post-action audit; approval remains mandatory wherever policy requires it. |
Critical distinction: human review is meaningful only if reviewers have enough information, time, competence and authority to challenge, change, or halt the action. Adding a checkbox or approval button is not, by itself, a risk control.
Behind the model: evidence, misuse and sources
An evidence-informed architecture conversation, not a scientifically validated agent-selection test. Expand any section to inspect the reasoning and follow the original references.
The science behind it Research lineage
1. Levels of automation are a design choice
Parasuraman, Sheridan and Wickens (2000) distinguish information acquisition, information analysis, decision/action selection and action implementation. They argue that each function can receive a different level of automation and should be evaluated for human performance consequences. This supports asking what an AI system may decide or execute, rather than applying a single “agentic / not agentic” label to the whole system. Original research ↗
2. Agency and uncertainty are separate dimensions
Deterministic behaviour concerns repeatability; probabilistic models express uncertainty; stochastic processes involve randomness. None, by itself, tells us who can select actions or approve execution. The distinction in this canvas is a conceptual architecture model, not a taxonomy claimed to be prescribed by NIST or ISO.
3. Governance changes permissible delegation
The NIST AI Risk Management Framework organises risk work into Govern, Map, Measure and Manage, with activities throughout the system lifecycle. Its Generative AI Profile extends that work to generative-AI-specific risks. ISO/IEC 42001 specifies requirements for an organisational AI management system. These sources support the principle that deployment authority should depend on context, risk assessment, evidence and continuing control operation—not model capability alone. NIST AI RMF ↗ · NIST GenAI Profile ↗ · ISO/IEC 42001 ↗
4. More actions introduce agent-specific failure modes
OWASP's Top 10 for Agentic Applications (2026) documents threats associated with agents that plan and act through tools. NIST has also studied agent hijacking through indirect prompt injection. These are reasons to evaluate permissions, tool boundaries, monitoring, approvals and recovery separately from the quality of an agent's prose or predictions. OWASP ↗ · NIST agent hijacking ↗
5. Human approval is not the same as effective human oversight
Human-factors research documents automation bias: people may over-rely on automated advice and fail to notice errors. A nominal human-in-the-loop checkbox cannot substitute for meaningful review time, information, authority and the ability to interrupt or reverse an action. Automation-bias review ↗ · Verification-complexity review ↗
What this instrument is — and is not Methodology
This is a transparent, qualitative decision heuristic. Its sliders operationalise four architectural distinctions: predictability, uncertainty/AI need, agency need and risk-adjusted autonomy. It compares required autonomy with an illustrative permissible-autonomy ceiling, and asks whether incremental value makes the additional delegation worth discussing.
Evidence boundary: the underlying principles are supported by research and standards; this canvas's particular questions, weights, thresholds, formula and four recommendation labels are original design choices. They have not been calibrated, independently validated, certified, or demonstrated to predict incidents or financial outcomes. A “3 / 4” is an ordinal discussion input, not a 75% probability, a risk rating or a control-effectiveness measurement.
Changing an answer can change the suggested architecture. Treat that as sensitivity analysis rather than precision. For consequential cases, compare multiple candidate architectures, collect evidence, document assumptions and require the organisation's normal architecture, risk and change approvals.
See “Scoring and decision rules (fully transparent)” above for the actual formula and cut-offs used by this version.
How this can be misused Failure modes
- Score laundering: presenting a heuristic recommendation as “science says we must deploy an agent.” Countermeasure: record alternatives, assumptions, dissent and the human decision owner.
- Slider gaming: overstating task complexity or control maturity to get a preferred result. Countermeasure: attach evidence, use a cross-functional review and preserve the original assessment.
- Equating controls on paper with operating controls: a policy or checkbox does not establish tested containment, identity, recovery or monitoring. Countermeasure: require demonstrations, evaluations and named operational owners.
- Decorative human approval: the reviewer cannot realistically verify the output, challenge it or stop execution. Countermeasure: design actionable review, appropriate information, time and override capability.
- Conflating inference with agency: a probabilistic classifier is not necessarily an agent, and constraining an agent's tools does not make its model deterministic. Countermeasure: map inference, planning, decision rights and action authority separately.
- False comfort from deterministic workflows: predefined steps can still implement flawed rules, act on erroneous AI output or have excessive permissions. Countermeasure: assess actual failure impact for every option.
- Hiding residual risk: a high control-maturity slider must not cancel a disallowed action, legal duty or unacceptable consequence. Countermeasure: use independent hard gates and organisational risk acceptance; never allow a high average score to override them.
- Missing value and opportunity cost: choosing the least autonomous option by default can prevent the system from solving the problem. Countermeasure: compare measurable outcomes of workflow, AI-assisted workflow and bounded agent trials.
External sources Follow the evidence
- Parasuraman, R., Sheridan, T. B. & Wickens, C. D. (2000). A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics A, 30(3), 286–297. DOI ↗ — framework for levels and functions of automation.
- NIST (2023). Artificial Intelligence Risk Management Framework 1.0. Official publication ↗ — govern, map, measure and manage.
- NIST (2024). Generative AI Profile, NIST AI 600-1. Official publication ↗ — generative-AI risk considerations.
- ISO/IEC (2023). ISO/IEC 42001: Artificial intelligence management system. Official standard page ↗ — organisational management-system requirements (full standard may require purchase).
- OWASP (2026). Top 10 for Agentic Applications. Official project ↗ — agent-specific security risks and guidance.
- NIST (2025). Strengthening AI Agent Hijacking Evaluations. Official research blog ↗ — indirect prompt injection and agent action risk.
- Goddard, K., Roudsari, A. & Wyatt, J. C. (2012). Automation bias: a systematic review of frequency, effect mediators, and mitigators. Open access ↗ — human reliance on automated advice.
- Lyell, D. & Coiera, E. (2017). Automation bias and verification complexity: a systematic review. Open access ↗ — challenges of verifying automated outputs.
- Decision-Method Selector — the companion helper for choosing how a decision should be made; shares this site's transparent “science / misuse / sources” disclosure.