THE SHORT ANSWER

Require a person when an action is irreversible, financial, externally representative, permission-changing, sensitive, high-impact or outside tested conditions. Let low-risk reversible steps proceed only within clear scopes and limits. An approval should show the proposed action, essential evidence, uncertainty and consequence.

Place approval before consequence

Example approval decisions
ActionDefault controlReason
Read a permitted referenceProceed and logLow-impact and reversible
Draft an external emailProceed to previewNo communication has occurred
Send or publishConfirm content, audience and identityRepresents a person or organisation
Spend or transfer moneyExplicit authorised approvalFinancial consequence
Change access or delete dataExplicit approval plus verificationSecurity or irreversible impact
Make a high-impact decisionHuman decision with evidenceAccountability cannot be delegated to a score

Evidence & context: Model Context Protocol · NIST

Ask when the system leaves its tested boundary

Escalation can be triggered by missing information, conflicting sources, an unfamiliar input, failed validation, repeated tool errors, exhausted budget or a policy exception. A confidence score can route attention, but it should not be the only control for a consequential action.

Design a clear unavailable path. If the authorised person does not respond, the system should wait or fail safely rather than silently widening its own permission.

Make approval meaningful

  • Show the exact proposed action and target.
  • Summarise the evidence and material uncertainty.
  • State what will change and whether it can be reversed.
  • Offer approve, revise and reject paths.
  • Bind approval to this action, not a broad future category.
  • Record the decision without exposing unnecessary sensitive content.

Frequent low-value prompts create approval fatigue. Use deterministic rules for routine safe operations and reserve people for decisions that require authority or judgment.

Keep responsibility visible

A person in the loop is not a decorative final click. The reviewer needs enough time, context and authority to disagree. Teams should define who owns policy, tool access, incidents and the final business outcome.

Evidence & context: Anthropic

Sources & further reading

  1. Model Context Protocol tools

    Model Context Protocol. The current official tools specification checked 13 September 2026. Draft details can change; OpenSkool relies on the durable separation between tool discovery, model selection and host-controlled execution.

  2. Generative Artificial Intelligence Profile (NIST AI 600-1)

    NIST. Risk-management guidance, including confabulation. It does not establish a universal error rate.

  3. Building effective agents

    Anthropic. A provider's engineering taxonomy of agents and workflows, not a universal industry definition. We use the conceptual distinction, not its changing product recommendations.

Examples and exercises are illustrative unless attributed to a source. No independent expert review is claimed.

A correction, a counterexample or an experience worth sharing?

Join the conversation ↗