
Agent assurance
AI agent risk assessment
How to control permissions, tool use, human approval, spending, identity, prompt injection, and shutdown before an agent reaches production.
By Kesho Partners
14 minute read
Assess an AI agent by the authority it can exercise, the environments and data it can reach, the reversibility of its actions, and the quality of controls around failure. Define a machine-enforced authority matrix for every tool action; use least-privilege identity, explicit approval for consequential actions, transaction and rate limits, untrusted-input isolation, complete action logs, incident triggers, and a tested shutdown path. A capable model is not a sufficient control.
Agent risk is authority multiplied by uncertainty
An AI assistant can generate a poor answer. An agent connected to email, code, payment, identity, customer, or infrastructure tools can turn a poor decision into an external action. The risk boundary therefore includes the model, orchestration, memory, identity, tools, data, approval interfaces, execution environment, and people supervising it.
NIST describes AI agent systems as capable of planning and taking autonomous actions that affect real systems or environments. Its 2026 security work focuses on threats created when model outputs are combined with software functionality, including adversarial data, insecure models, misaligned actions, and the need to constrain and monitor access.
Map the full agent system
A risk assessment that lists only the foundation model misses the components that grant authority. Draw the execution path from user request to plan, retrieval, memory, tool selection, credential use, external action, observation, and final response. Mark trust boundaries and every place untrusted content can influence the plan.
| Element | Record | Risk question |
|---|---|---|
| Objective | Allowed tasks and explicit prohibited outcomes | Can success be measured without encouraging harmful shortcuts? |
| Identity | Agent, user, service and delegated identities | Whose authority is exercised and how is it authenticated? |
| Tools | Functions, parameters, environments and credentials | What can the agent read, create, modify, send, execute or delete? |
| Data | Sources, sensitivity, provenance and retention | Can untrusted or sensitive data enter prompts, memory or outputs? |
| Control flow | Planning, retries, delegation and stopping conditions | Can the agent loop, escalate, or bypass intended gates? |
| Human role | Approver competence, context and response time | Can a person make a meaningful decision before the action occurs? |
| Observability | Traces, tool calls, decisions, alerts and retention | Can operation and failure be reconstructed? |
Rate consequence and control failure separately
Do not reduce agent risk to a single generic score. Assess the maximum plausible consequence of each action, the likelihood that the agent attempts it incorrectly, the probability that preventive controls fail, and the ability to detect and reverse the result. The highest-impact actions deserve specific scenarios and release criteria.
- Reach
- Number of records, users, systems, environments, recipients, or transactions affected.
- Sensitivity
- Personal, confidential, regulated, security, payment, or authentication data involved.
- Irreversibility
- Whether an action can be cancelled, rolled back, recalled, compensated, or recovered.
- Velocity
- How many actions can occur before detection, approval, or automatic containment.
- Dependence
- Whether important services continue safely when the model, vendor, tool, or network fails.
- Adversarial exposure
- Whether emails, files, websites, tickets, users, or retrieved content can manipulate the agent.
Design human approval as a security control
Human approval fails when the approver sees an abstract summary, lacks time or expertise, or approves a plan that the agent can later change. The approval screen should display the exact consequential action, target, key inputs, source evidence, uncertainty, limits, and a clear reject path.
Bind approval to an immutable action payload or validated transaction. If the recipient, amount, command, attachment, or parameters change, require approval again. Record the approver, decision, timestamp, information shown, and executed action.
Approval is not a universal remedy. High-volume review can become rubber-stamping, and some actions remain unsuitable even with a person in the loop. Reduce authority and automate deterministic checks before adding approval workload.
Assume untrusted data can contain hostile instructions
Indirect prompt injection occurs when an agent ingests malicious instructions hidden in content such as email, files, web pages, or retrieved documents. NIST calls this agent hijacking and identifies the lack of separation between trusted instructions and untrusted data as a central problem.
NIST experiments added scenarios involving remote code execution, data exfiltration, and automated phishing. Those experimental results are not universal failure rates. The durable lesson is that security evaluation must be task-specific, adaptive, and repeated, while production controls must constrain the impact of a successful hijack.
- Minimize authority
- Do not expose tools or data unnecessary for the approved task.
- Separate channels
- Label and isolate trusted policy, user intent, retrieved data and tool output where the architecture allows.
- Validate actions
- Apply deterministic policy, schema, target, data-loss and transaction checks outside the model.
- Sandbox execution
- Run code, browsers and files in isolated environments with network and resource restrictions.
- Test adaptively
- Use held-out, task-specific attacks and revise them as models and defenses change.
Limit blast radius with hard controls
Limits should be enforced outside the agent's reasoning loop. Define maximum transactions, recipients, records, tool calls, runtime, tokens, spending, retries, concurrency, and data volume per session and period. Use lower initial limits and expand only when evaluation and production evidence justify it.
| Control | Purpose | Trigger response |
|---|---|---|
| Transaction cap | Bound financial or contractual exposure | Reject and escalate |
| Rate and recipient limit | Prevent rapid misuse or mass communication | Pause session and alert |
| Data-volume limit | Reduce bulk access and exfiltration | Block query or output |
| Tool-call and runtime budget | Stop loops and unbounded consumption | Terminate execution |
| Environment boundary | Prevent test actions reaching production | Deny by identity and network policy |
| Circuit breaker | Contain anomalous or harmful behavior | Revoke credentials and disable tool path |
Make release conditional on residual authority
Release approval should state the tasks the agent may perform, the environments and users included, the tools and limits enabled, known failure modes, monitoring thresholds, incident ownership, and the evidence supporting residual risk acceptance. A later increase in authority is a new release decision.
- Block release
- Unknown system boundary, shared privileged credentials, mutable approvals, untested shutdown, missing action logs, or uncontrolled destructive actions.
- Restrict release
- Limited user cohort, read-only or draft mode, low transaction caps, mandatory approval, and heightened monitoring.
- Expand authority
- Only after stable task, security, intervention, incident, cost and recovery evidence across a defined observation period.
Test shutdown and recovery before launch
The incident plan should cover harmful actions, attempted privilege escalation, prompt injection, data exposure, runaway cost, tool misuse, model or supplier degradation, and loss of traceability. Define who can disable the agent, revoke credentials, block tools, preserve evidence, notify affected owners, reverse actions, and approve restoration.
Run a live exercise in the target environment. A shutdown control that exists only in architecture diagrams has not been proven. Measure time to detect, contain, revoke, reconstruct, recover, and communicate.
Related service
AI agent architecture and assurance
Kesho helps teams design authority boundaries, production controls, evaluation suites, release gates, and operating evidence for AI agents.
Explore the service