futureofagents.org
RESEARCH NOTE / VOL. 01 · 09 OCT 20267 REFERENCES / 7 SECTIONS

AGENT SECURITY / HUMAN CONTROL

Delegation is not
permission.

Tool-using AI agents cross boundaries between generated decisions and real systems. A rigorous account begins by separating what a model asks from what its credentials and policies allow.

BY THE AGENT SYSTEMS RESEARCH DESKEditorial field note / Reviewed 2026-10-09RESEARCHED / CITED / OPEN TO REVISION

A tool call is a request across a trust boundary

A model can propose a tool invocation, but the ability to name a tool does not grant authority to perform a protected action. MCP describes discoverable tools with schemas, including calls into external systems. The specification explicitly discusses the value of human oversight; it also warns that tool annotations should not be assumed trustworthy when supplied by untrusted servers.[1]

A schema describes inputs; authorization decides whether they may be used

Input validation can catch malformed parameters, but it does not decide whether a specific user is entitled to read a repository, change a firewall, or export customer records. The MCP authorization specification describes an authorization model; a real deployment still needs a correctly implemented identity and permission boundary.[2][1]

Excessive agency increases the impact of ordinary mistakes

OWASP describes excessive agency as a risk arising when an LLM-based system has too much functionality, permission, or autonomy for its intended purpose. Hallucinated intentions, ambiguous instructions and manipulated retrieved content can all expose the same overpowered tool surface. The relevant remedy is to restrict capabilities—not just urge the model to behave responsibly.[3]

Human approval needs to be transaction-specific

An approval screen is meaningful when the operator sees which tool will run, the target resource, the intended change and its likely effects before the operation executes. The OpenAI Agents SDK documents a pause-and-resume approval flow in which protected tool calls create interruptions; this is an example of an implementable control rather than evidence that any specific deployment is secure.[4]

Logs and tests must outlive the model run

A fluent final answer cannot establish that a database write succeeded, that no unapproved files changed, or that rollback works. Independent state checks, separate audit storage, testable error paths and adversarial runs are needed. NIST’s AI Risk Management Framework provides a broader voluntary risk-management structure, but site-specific assurance still requires system-level evidence.[5][6]

Multi-agent handoffs do not remove responsibility

Delegating from one agent to another creates an additional boundary. A handoff can change who generates recommendations, but it should not implicitly grant more credentials or bypass the original human approval policy. A report on a multi-agent workflow should enumerate each agent’s tools, permitted operations, shared state and escalation path.[7][3]

A reproducible evaluation must include denied and failed calls

We propose evaluating agent systems on matched legitimate and adversarial tasks, recording approved actions, denied writes, false completion claims, intervention counts and recovery behavior. This is a proposed testing framework, not a dataset we have already run. Until real outcomes are published, no quantitative claim of agent safety or effectiveness is warranted.[5][1]

EDITORIAL / VERSION RECORD

Revision record

v1.0 — 09 October 2026: Initial methodological field note published, based on the cited standards, security guidance and SDK documentation. No original benchmark, penetration test or production telemetry is claimed. Accepted consequential corrections will be described here with source and date.

RESEARCH / PROVENANCE

Works Cited

7 REFERENCES
  1. 01
    Model Context Protocol maintainers — MCP Specification (2026-07-28) — Tools ↗2026-07-28 · Technical specification, not an independent security audit.
  2. 02
    Model Context Protocol maintainers — MCP Specification (2026-07-28) — Authorization ↗2026-07-28 · Normative authentication/authorization protocol details; deployments require correct configuration.
  3. 03
    OWASP GenAI Security Project — LLM06:2025 Excessive Agency ↗2025 · Threat taxonomy and mitigation guidance rather than measurement of any particular agent.
  4. 04
    OpenAI — Agents SDK — Human in the Loop (TypeScript) ↗accessed 2026-10-09 · Vendor SDK documentation illustrates approval patterns; implementations need evaluation.
  5. 05
    National Institute of Standards and Technology — AI Risk Management Framework ↗2023–2026 · Voluntary general risk framework, not proof of agent reliability.
  6. 06
    OpenAI — Agents SDK — Running Agents (Python) ↗accessed 2026-10-09 · Vendor implementation guide for runs, checkpoints and orchestration.
  7. 07
    OpenAI — Agents SDK — Agents (Python) ↗accessed 2026-10-09 · Vendor reference for agent configuration and handoffs.

Document dates and versions matter. These sources support the described security principles; vendor documentation should not be mistaken for independent verification of a product.

CORRECTIONS / EVIDENCE

Can you disprove or refine a claim?

We review reproducible tests, security research, contradictory specifications and implementation evidence. The research record should improve when better information arrives.

Prepare an evidence submission ↗