
AI evaluation
How to evaluate AI agents in production
A measurement framework for agent outcomes, trajectories, tool use, security, reliability, latency, cost, human intervention, and production drift.
By Kesho Partners
15 minute read
Evaluate AI agents at four levels: whether the business task was completed correctly, whether the execution trace and tool calls were valid, whether safety and security controls held under adversarial and repeated attempts, and whether production performance remains reliable and economical. Use representative task suites, deterministic checks where possible, expert review for ambiguous outcomes, release thresholds by risk tier, and continuous monitoring tied to rollback and authority reduction.
A good final answer can hide a bad execution path
Agents make plans, retrieve data, call tools, observe results, retry, and change external systems. Evaluating only the final text misses wrong recipients, unnecessary data access, invalid tool arguments, excessive retries, policy violations, and near-miss security failures.
Evaluation must cover the complete task trajectory and the resulting business state. The target is not a universally intelligent agent; it is an agent that performs a defined set of tasks within explicit quality, authority, security, reliability, latency, and cost boundaries.
Write the evaluation contract first
For each task family, define the starting state, permitted information, allowed tools, expected outcome, prohibited actions, completion criteria, time and cost budget, approval requirements, and recovery behavior. This creates an executable contract between product, engineering, risk, security, and operations.
| Field | Example definition |
|---|---|
| Task | Resolve an eligible billing adjustment for one customer |
| Starting state | Authenticated user, assigned case, current account record |
| Permitted tools | Read account, calculate adjustment, draft response |
| Approval gate | Human approval before financial write or external send |
| Success | Correct adjustment proposed with supporting evidence |
| Prohibited | Access another account, expose sensitive data, execute unapproved refund |
| Budget | Maximum calls, latency, tokens and monetary cost |
| Failure behavior | Stop safely, preserve trace and hand off with a reason |
Use a multidimensional scorecard
A single average score can allow a catastrophic failure to disappear inside good routine performance. Report dimensions separately and apply hard failure conditions to consequential actions. Segment results by task, risk tier, model, tool, language, user cohort, data condition, and attack type.
| Dimension | Primary measure | Release question |
|---|---|---|
| Task outcome | Correct completions / eligible attempts | Did the intended business state result? |
| Trajectory | Valid steps / total material steps | Was the plan efficient, grounded and policy-compliant? |
| Tool use | Correct tool and arguments / tool calls | Were calls authorized, necessary and valid? |
| Safety | Severe violations and policy-failure rate | Did prohibited outcomes remain blocked? |
| Security | Attack success by task and repeated-attempt scenario | Can adversarial input redirect consequential action? |
| Reliability | Completion, timeout, retry and recovery rates | Does performance hold under realistic failure? |
| Human operation | Intervention, rejection, correction and escalation rates | Is oversight effective and sustainable? |
| Efficiency | Latency and cost per successful task | Is successful operation fast and economical enough? |
Build a representative task suite
Start from production workflows and incident scenarios, not generic benchmark prompts. Include common tasks, high-consequence edge cases, ambiguous requests, incomplete data, conflicting records, unavailable tools, changed permissions, malformed outputs, and realistic hand-offs.
Maintain a stable regression set for comparability and a rotating challenge set to reduce overfitting. Keep held-out cases for release decisions. Sample production traces only with appropriate privacy, security, and data-governance controls.
- Routine
- High-volume, well-specified cases that establish baseline quality and cost.
- Boundary
- Requests just inside and outside policy, authority, eligibility or data limits.
- Failure
- Tool timeouts, partial writes, stale data, dependency errors and unavailable humans.
- Consequential
- Low-frequency actions with material customer, financial, security or operational impact.
- Adversarial
- Direct and indirect prompt injection, exfiltration, privilege escalation and resource exhaustion.
Evaluate the trajectory and every material tool call
Use deterministic evaluation for schemas, identity, permissions, arguments, recipients, amounts, state transitions, policy rules, latency, and cost. Use expert review or calibrated model-based graders for semantic quality where deterministic ground truth is unavailable. Periodically compare automated graders with human judgments and investigate disagreement.
| Check | Pass condition |
|---|---|
| Tool selection | Selected tool is permitted and necessary for the task |
| Arguments | Parameters are complete, valid, scoped and supported by evidence |
| Data access | Only authorized records and fields are accessed |
| Grounding | Material claims and actions follow trusted observations |
| Approval | Consequential payload matches the immutable approved payload |
| State | External result is verified rather than inferred from an attempted call |
| Stopping | Agent completes, escalates or stops without unnecessary loops |
Measure security by attack task and consequence
NIST's agent-hijacking work shows why an aggregate attack-success rate is insufficient. Different malicious tasks have different success rates and consequences. A rare data-exfiltration success can matter more than many harmless failures.
NIST also found in a specific AgentDojo experiment that repeated attempts materially changed measured attack success. Those numbers describe that experiment, not a universal property of agents. The general method is still important: where attackers can retry cheaply, evaluate the probability of at least one success across realistic repeated attempts.
- Direct injection
- Malicious user instructions that conflict with policy or attempt to reveal secrets.
- Indirect injection
- Hostile instructions in email, files, websites, retrieved records or tool output.
- Exfiltration
- Attempts to send sensitive data to unauthorized targets or encode it in outputs.
- Privilege escalation
- Attempts to invoke unavailable tools, broaden scope or misuse delegated identity.
- Destructive action
- Attempts to execute code, delete data, alter access or trigger irreversible transactions.
- Repeated attack
- Multiple varied attempts across sessions, tasks, contexts and model stochasticity.
Test the system around the model
Agent reliability includes component communication, throughput, timeouts, retries, idempotency, partial failure, state recovery, versioning, and observability. AWS's Generative AI Lens frames generative-AI workloads across operational excellence, security, reliability, performance efficiency, cost optimization, and sustainability rather than model quality alone.
Inject failures into tools and dependencies. Verify that retries do not duplicate external actions, stale observations do not become decisions, and failed workflows preserve enough state for safe recovery or human hand-off.
Sources: [3]
Measure latency and cost per successful task
Average token cost per run hides failed, abandoned, retried, and human-corrected work. Measure total model, tool, infrastructure, review, and remediation cost divided by correct completed tasks. Segment by task family and percentile because expensive tails often identify loops or poor routing.
| Metric | Definition |
|---|---|
| End-to-end latency | Time from accepted request to verified completion or safe hand-off |
| Tool latency share | Dependency time / total task time |
| Cost per attempt | Total automated execution cost / attempts |
| Cost per success | Automated plus human operating cost / correct completed tasks |
| Retry amplification | Total executions / unique accepted requests |
| Wasted execution | Cost of failed, duplicate, timed-out or policy-blocked runs |
Measure whether human oversight works
A low intervention rate can mean high quality or inattentive oversight. Pair intervention with approval rejection, correction, escalation, review time, agreement, missed-error, and post-action reversal rates. Review whether operators receive enough context and whether workload encourages rubber-stamping.
- Intervention
- Share of tasks where a person changes, stops or redirects execution.
- Rejection
- Share of proposed consequential actions that approvers decline.
- Correction
- Share of outputs or state changes repaired before completion.
- Escape
- Errors not caught by the designed oversight step.
- Review burden
- Median and tail review time, queue depth, abandonment and alert volume.
Set thresholds by task risk
Release thresholds should be approved before the final evaluation. Higher-consequence tasks need harder failure conditions, larger samples, stronger independent review, and narrower initial authority. Do not offset a severe safety or security failure with a high average task score.
| Decision | Condition |
|---|---|
| No release | Any prohibited consequential action, authorization bypass, uncontained data exposure, or untested recovery |
| Restricted release | Quality passes but uncertainty remains; use limited cohort, authority, volume, duration and heightened review |
| Release | All hard gates pass and each quality, security, reliability, latency, cost and oversight threshold is met |
| Rollback | Production hard gate, drift threshold, incident trigger or control failure is breached |
Turn evaluation into production monitoring
Use the same task definitions and core metrics before and after release. Monitor by model, prompt and orchestration version, tool, task, cohort, environment, and authority level. Record denominator changes so apparent improvements are not caused by routing harder cases away.
Trigger investigation or automatic containment on severe policy events, unusual tool sequences, increased retries, cost spikes, latency degradation, rising rejection or correction, missing traces, changed supplier behavior, or distribution shift. Production examples should feed future regression and adversarial suites after governance review.
A dashboard is not the operating loop. Every material signal needs an owner, threshold, response, evidence-retention rule, and authority to pause or roll back the agent.
Related service
AI agent evaluation and production engineering
Kesho helps teams build task suites, trace evaluators, adversarial tests, release thresholds, observability, and production improvement loops for AI agents.
Explore the service