KESHO PARTNERS
Engineering team reviewing an AI agent workflow and control boundaries

Agent assurance

AI agent risk assessment

How to control permissions, tool use, human approval, spending, identity, prompt injection, and shutdown before an agent reaches production.

By Kesho Partners

14 minute read

Assess an AI agent by the authority it can exercise, the environments and data it can reach, the reversibility of its actions, and the quality of controls around failure. Define a machine-enforced authority matrix for every tool action; use least-privilege identity, explicit approval for consequential actions, transaction and rate limits, untrusted-input isolation, complete action logs, incident triggers, and a tested shutdown path. A capable model is not a sufficient control.

Agent risk is authority multiplied by uncertainty

An AI assistant can generate a poor answer. An agent connected to email, code, payment, identity, customer, or infrastructure tools can turn a poor decision into an external action. The risk boundary therefore includes the model, orchestration, memory, identity, tools, data, approval interfaces, execution environment, and people supervising it.

NIST describes AI agent systems as capable of planning and taking autonomous actions that affect real systems or environments. Its 2026 security work focuses on threats created when model outputs are combined with software functionality, including adversarial data, insecure models, misaligned actions, and the need to constrain and monitor access.

Sources: [1], [2]

Map the full agent system

A risk assessment that lists only the foundation model misses the components that grant authority. Draw the execution path from user request to plan, retrieval, memory, tool selection, credential use, external action, observation, and final response. Mark trust boundaries and every place untrusted content can influence the plan.

Minimum agent system record
ElementRecordRisk question
ObjectiveAllowed tasks and explicit prohibited outcomesCan success be measured without encouraging harmful shortcuts?
IdentityAgent, user, service and delegated identitiesWhose authority is exercised and how is it authenticated?
ToolsFunctions, parameters, environments and credentialsWhat can the agent read, create, modify, send, execute or delete?
DataSources, sensitivity, provenance and retentionCan untrusted or sensitive data enter prompts, memory or outputs?
Control flowPlanning, retries, delegation and stopping conditionsCan the agent loop, escalate, or bypass intended gates?
Human roleApprover competence, context and response timeCan a person make a meaningful decision before the action occurs?
ObservabilityTraces, tool calls, decisions, alerts and retentionCan operation and failure be reconstructed?

Build an authority matrix before connecting tools

The authority matrix is the core control artifact. Define permissions by action and consequence, not broad tool name. Reading one approved customer record is different from exporting all customer records; drafting a refund is different from executing it.

Enforce the matrix in the tool gateway, identity layer, policy engine, and transaction system. A system prompt that tells the model not to perform an action is behavioral guidance, not an authorization boundary.

Example agent authority matrix
Action classDefault authorityRequired controlExample
Read boundedAutomaticPurpose-scoped query, field filtering, audit logRead one assigned support case
DraftAutomaticNo external side effect, labelled draft, retained traceDraft a customer response
Reversible writeConditionalValidation, narrow scope, undo path, alert thresholdUpdate an internal case status
External communicationHuman approvalPreview exact recipient and content, prevent post-approval mutationSend email to a customer
Financial or contractualHuman approval plus limitPayee validation, amount cap, segregation of dutiesIssue a refund
Privileged or destructiveProhibited or exceptionalIsolated environment and separate break-glass processRun production shell command or delete records

Rate consequence and control failure separately

Do not reduce agent risk to a single generic score. Assess the maximum plausible consequence of each action, the likelihood that the agent attempts it incorrectly, the probability that preventive controls fail, and the ability to detect and reverse the result. The highest-impact actions deserve specific scenarios and release criteria.

Reach
Number of records, users, systems, environments, recipients, or transactions affected.
Sensitivity
Personal, confidential, regulated, security, payment, or authentication data involved.
Irreversibility
Whether an action can be cancelled, rolled back, recalled, compensated, or recovered.
Velocity
How many actions can occur before detection, approval, or automatic containment.
Dependence
Whether important services continue safely when the model, vendor, tool, or network fails.
Adversarial exposure
Whether emails, files, websites, tickets, users, or retrieved content can manipulate the agent.

Give the agent its own identity and least privilege

Do not let an agent silently inherit the full permissions of a developer, administrator, or end user. Use a distinct workload identity, short-lived credentials, explicit delegation, and policy checks that consider the user, agent, tool, action, resource, environment, and transaction context.

NIST's AI Agent Standards Initiative includes work on software and AI-agent identity and authorization. The field is still developing, so organizations should rely on established identity and access-management principles while tracking emerging agent-specific standards and protocols.

Authenticate
Identify the user, agent instance, workload, tool and service at each boundary.
Authorize
Evaluate every consequential tool call against explicit policy and current context.
Delegate
Record whose authority is being used, for which purpose, resource, duration and limits.
Revoke
Expire or remove credentials immediately when a session, role, incident or approval ends.
Separate
Use distinct development, test and production identities and deny cross-environment access.

Sources: [1], [3]

Design human approval as a security control

Human approval fails when the approver sees an abstract summary, lacks time or expertise, or approves a plan that the agent can later change. The approval screen should display the exact consequential action, target, key inputs, source evidence, uncertainty, limits, and a clear reject path.

Bind approval to an immutable action payload or validated transaction. If the recipient, amount, command, attachment, or parameters change, require approval again. Record the approver, decision, timestamp, information shown, and executed action.

Approval is not a universal remedy. High-volume review can become rubber-stamping, and some actions remain unsuitable even with a person in the loop. Reduce authority and automate deterministic checks before adding approval workload.

Assume untrusted data can contain hostile instructions

Indirect prompt injection occurs when an agent ingests malicious instructions hidden in content such as email, files, web pages, or retrieved documents. NIST calls this agent hijacking and identifies the lack of separation between trusted instructions and untrusted data as a central problem.

NIST experiments added scenarios involving remote code execution, data exfiltration, and automated phishing. Those experimental results are not universal failure rates. The durable lesson is that security evaluation must be task-specific, adaptive, and repeated, while production controls must constrain the impact of a successful hijack.

Minimize authority
Do not expose tools or data unnecessary for the approved task.
Separate channels
Label and isolate trusted policy, user intent, retrieved data and tool output where the architecture allows.
Validate actions
Apply deterministic policy, schema, target, data-loss and transaction checks outside the model.
Sandbox execution
Run code, browsers and files in isolated environments with network and resource restrictions.
Test adaptively
Use held-out, task-specific attacks and revise them as models and defenses change.

Sources: [2], [4], [5]

Limit blast radius with hard controls

Limits should be enforced outside the agent's reasoning loop. Define maximum transactions, recipients, records, tool calls, runtime, tokens, spending, retries, concurrency, and data volume per session and period. Use lower initial limits and expand only when evaluation and production evidence justify it.

Containment controls
ControlPurposeTrigger response
Transaction capBound financial or contractual exposureReject and escalate
Rate and recipient limitPrevent rapid misuse or mass communicationPause session and alert
Data-volume limitReduce bulk access and exfiltrationBlock query or output
Tool-call and runtime budgetStop loops and unbounded consumptionTerminate execution
Environment boundaryPrevent test actions reaching productionDeny by identity and network policy
Circuit breakerContain anomalous or harmful behaviorRevoke credentials and disable tool path

Sources: [2], [5]

Make release conditional on residual authority

Release approval should state the tasks the agent may perform, the environments and users included, the tools and limits enabled, known failure modes, monitoring thresholds, incident ownership, and the evidence supporting residual risk acceptance. A later increase in authority is a new release decision.

Block release
Unknown system boundary, shared privileged credentials, mutable approvals, untested shutdown, missing action logs, or uncontrolled destructive actions.
Restrict release
Limited user cohort, read-only or draft mode, low transaction caps, mandatory approval, and heightened monitoring.
Expand authority
Only after stable task, security, intervention, incident, cost and recovery evidence across a defined observation period.

Test shutdown and recovery before launch

The incident plan should cover harmful actions, attempted privilege escalation, prompt injection, data exposure, runaway cost, tool misuse, model or supplier degradation, and loss of traceability. Define who can disable the agent, revoke credentials, block tools, preserve evidence, notify affected owners, reverse actions, and approve restoration.

Run a live exercise in the target environment. A shutdown control that exists only in architecture diagrams has not been proven. Measure time to detect, contain, revoke, reconstruct, recover, and communicate.

Related service

AI agent architecture and assurance

Kesho helps teams design authority boundaries, production controls, evaluation suites, release gates, and operating evidence for AI agents.

Explore the service