KESHO PARTNERS
Production engineering team monitoring AI agent traces and service metrics

AI evaluation

How to evaluate AI agents in production

A measurement framework for agent outcomes, trajectories, tool use, security, reliability, latency, cost, human intervention, and production drift.

By Kesho Partners

15 minute read

Evaluate AI agents at four levels: whether the business task was completed correctly, whether the execution trace and tool calls were valid, whether safety and security controls held under adversarial and repeated attempts, and whether production performance remains reliable and economical. Use representative task suites, deterministic checks where possible, expert review for ambiguous outcomes, release thresholds by risk tier, and continuous monitoring tied to rollback and authority reduction.

A good final answer can hide a bad execution path

Agents make plans, retrieve data, call tools, observe results, retry, and change external systems. Evaluating only the final text misses wrong recipients, unnecessary data access, invalid tool arguments, excessive retries, policy violations, and near-miss security failures.

Evaluation must cover the complete task trajectory and the resulting business state. The target is not a universally intelligent agent; it is an agent that performs a defined set of tasks within explicit quality, authority, security, reliability, latency, and cost boundaries.

Write the evaluation contract first

For each task family, define the starting state, permitted information, allowed tools, expected outcome, prohibited actions, completion criteria, time and cost budget, approval requirements, and recovery behavior. This creates an executable contract between product, engineering, risk, security, and operations.

Task evaluation contract
FieldExample definition
TaskResolve an eligible billing adjustment for one customer
Starting stateAuthenticated user, assigned case, current account record
Permitted toolsRead account, calculate adjustment, draft response
Approval gateHuman approval before financial write or external send
SuccessCorrect adjustment proposed with supporting evidence
ProhibitedAccess another account, expose sensitive data, execute unapproved refund
BudgetMaximum calls, latency, tokens and monetary cost
Failure behaviorStop safely, preserve trace and hand off with a reason

Use a multidimensional scorecard

A single average score can allow a catastrophic failure to disappear inside good routine performance. Report dimensions separately and apply hard failure conditions to consequential actions. Segment results by task, risk tier, model, tool, language, user cohort, data condition, and attack type.

Production agent evaluation scorecard
DimensionPrimary measureRelease question
Task outcomeCorrect completions / eligible attemptsDid the intended business state result?
TrajectoryValid steps / total material stepsWas the plan efficient, grounded and policy-compliant?
Tool useCorrect tool and arguments / tool callsWere calls authorized, necessary and valid?
SafetySevere violations and policy-failure rateDid prohibited outcomes remain blocked?
SecurityAttack success by task and repeated-attempt scenarioCan adversarial input redirect consequential action?
ReliabilityCompletion, timeout, retry and recovery ratesDoes performance hold under realistic failure?
Human operationIntervention, rejection, correction and escalation ratesIs oversight effective and sustainable?
EfficiencyLatency and cost per successful taskIs successful operation fast and economical enough?

Build a representative task suite

Start from production workflows and incident scenarios, not generic benchmark prompts. Include common tasks, high-consequence edge cases, ambiguous requests, incomplete data, conflicting records, unavailable tools, changed permissions, malformed outputs, and realistic hand-offs.

Maintain a stable regression set for comparability and a rotating challenge set to reduce overfitting. Keep held-out cases for release decisions. Sample production traces only with appropriate privacy, security, and data-governance controls.

Routine
High-volume, well-specified cases that establish baseline quality and cost.
Boundary
Requests just inside and outside policy, authority, eligibility or data limits.
Failure
Tool timeouts, partial writes, stale data, dependency errors and unavailable humans.
Consequential
Low-frequency actions with material customer, financial, security or operational impact.
Adversarial
Direct and indirect prompt injection, exfiltration, privilege escalation and resource exhaustion.

Evaluate the trajectory and every material tool call

Use deterministic evaluation for schemas, identity, permissions, arguments, recipients, amounts, state transitions, policy rules, latency, and cost. Use expert review or calibrated model-based graders for semantic quality where deterministic ground truth is unavailable. Periodically compare automated graders with human judgments and investigate disagreement.

Trace-level checks
CheckPass condition
Tool selectionSelected tool is permitted and necessary for the task
ArgumentsParameters are complete, valid, scoped and supported by evidence
Data accessOnly authorized records and fields are accessed
GroundingMaterial claims and actions follow trusted observations
ApprovalConsequential payload matches the immutable approved payload
StateExternal result is verified rather than inferred from an attempted call
StoppingAgent completes, escalates or stops without unnecessary loops

Measure security by attack task and consequence

NIST's agent-hijacking work shows why an aggregate attack-success rate is insufficient. Different malicious tasks have different success rates and consequences. A rare data-exfiltration success can matter more than many harmless failures.

NIST also found in a specific AgentDojo experiment that repeated attempts materially changed measured attack success. Those numbers describe that experiment, not a universal property of agents. The general method is still important: where attackers can retry cheaply, evaluate the probability of at least one success across realistic repeated attempts.

Direct injection
Malicious user instructions that conflict with policy or attempt to reveal secrets.
Indirect injection
Hostile instructions in email, files, websites, retrieved records or tool output.
Exfiltration
Attempts to send sensitive data to unauthorized targets or encode it in outputs.
Privilege escalation
Attempts to invoke unavailable tools, broaden scope or misuse delegated identity.
Destructive action
Attempts to execute code, delete data, alter access or trigger irreversible transactions.
Repeated attack
Multiple varied attempts across sessions, tasks, contexts and model stochasticity.

Sources: [1], [2]

Test the system around the model

Agent reliability includes component communication, throughput, timeouts, retries, idempotency, partial failure, state recovery, versioning, and observability. AWS's Generative AI Lens frames generative-AI workloads across operational excellence, security, reliability, performance efficiency, cost optimization, and sustainability rather than model quality alone.

Inject failures into tools and dependencies. Verify that retries do not duplicate external actions, stale observations do not become decisions, and failed workflows preserve enough state for safe recovery or human hand-off.

Sources: [3]

Measure latency and cost per successful task

Average token cost per run hides failed, abandoned, retried, and human-corrected work. Measure total model, tool, infrastructure, review, and remediation cost divided by correct completed tasks. Segment by task family and percentile because expensive tails often identify loops or poor routing.

Efficiency measures
MetricDefinition
End-to-end latencyTime from accepted request to verified completion or safe hand-off
Tool latency shareDependency time / total task time
Cost per attemptTotal automated execution cost / attempts
Cost per successAutomated plus human operating cost / correct completed tasks
Retry amplificationTotal executions / unique accepted requests
Wasted executionCost of failed, duplicate, timed-out or policy-blocked runs

Measure whether human oversight works

A low intervention rate can mean high quality or inattentive oversight. Pair intervention with approval rejection, correction, escalation, review time, agreement, missed-error, and post-action reversal rates. Review whether operators receive enough context and whether workload encourages rubber-stamping.

Intervention
Share of tasks where a person changes, stops or redirects execution.
Rejection
Share of proposed consequential actions that approvers decline.
Correction
Share of outputs or state changes repaired before completion.
Escape
Errors not caught by the designed oversight step.
Review burden
Median and tail review time, queue depth, abandonment and alert volume.

Set thresholds by task risk

Release thresholds should be approved before the final evaluation. Higher-consequence tasks need harder failure conditions, larger samples, stronger independent review, and narrower initial authority. Do not offset a severe safety or security failure with a high average task score.

Example release logic
DecisionCondition
No releaseAny prohibited consequential action, authorization bypass, uncontained data exposure, or untested recovery
Restricted releaseQuality passes but uncertainty remains; use limited cohort, authority, volume, duration and heightened review
ReleaseAll hard gates pass and each quality, security, reliability, latency, cost and oversight threshold is met
RollbackProduction hard gate, drift threshold, incident trigger or control failure is breached

Turn evaluation into production monitoring

Use the same task definitions and core metrics before and after release. Monitor by model, prompt and orchestration version, tool, task, cohort, environment, and authority level. Record denominator changes so apparent improvements are not caused by routing harder cases away.

Trigger investigation or automatic containment on severe policy events, unusual tool sequences, increased retries, cost spikes, latency degradation, rising rejection or correction, missing traces, changed supplier behavior, or distribution shift. Production examples should feed future regression and adversarial suites after governance review.

A dashboard is not the operating loop. Every material signal needs an owner, threshold, response, evidence-retention rule, and authority to pause or roll back the agent.

Related service

AI agent evaluation and production engineering

Kesho helps teams build task suites, trace evaluators, adversarial tests, release thresholds, observability, and production improvement loops for AI agents.

Explore the service