Engineers preparing an AI product for production
KESHO PARTNERSAll insights

AI product development

The AI production release gate: nine decisions before launch

A practical release standard for product and engineering teams moving an AI workflow from a convincing prototype into live operations.

By Kesho Partners

10 minute read

An AI feature is ready for production only when the team has accepted evidence across nine areas: intended use, task quality, failure modes, security, permissions, latency, unit cost, human intervention, and production monitoring. Passing an average quality score is insufficient. The release decision must also define operating limits, fallback behavior, accountable owners, and the conditions that trigger rollback or reassessment.

A prototype answers a different question from production

A prototype asks whether a model can perform a useful task under guided conditions. Production asks whether the complete system can perform that task for real users, with variable inputs, connected data, permissions, failures, costs, and support obligations.

The release gate should evaluate the full workflow rather than the model in isolation. Retrieval, tools, prompts, data pipelines, interfaces, human review, and downstream systems can each improve or undermine the result. AWS notes that changing data requires continuous monitoring to detect and mitigate accuracy and performance issues; release evidence is therefore the start of control, not its endpoint.

Sources: [1], [2]

The nine-part production release gate

For each gate, record the threshold, observed result, evidence link, owner, and decision. A gate can pass, pass with a time-bound condition, or fail. Security, data rights, unsafe agency, and the absence of a viable fallback should not be averaged against strong quality elsewhere.

Kesho AI production release gate
GateRequired decisionMinimum evidence
1. Intended useThe workflow and prohibited uses are explicitSystem card, users, boundaries, assumptions
2. Task qualityPerformance meets workflow-specific thresholdsRepresentative evaluation, baseline, error analysis
3. Failure modesKnown failures have a controlled outcomeAdversarial tests, edge cases, fallback behavior
4. SecurityInput, output, data, and dependencies are protectedThreat model, abuse tests, output validation
5. PermissionsThe system has only the authority it needsTool scopes, approvals, irreversible-action controls
6. LatencyResponse time fits the user and operational needPercentile latency by workflow and dependency
7. Unit costCost per successful outcome supports the business caseModel, tool, retry, review, and support cost
8. Human interventionPeople can understand, override, and escalateReview procedure, authority, interface test
9. MonitoringThe team can detect deterioration and respondSignals, thresholds, alerts, owner, rollback test

Sources: [1], [2], [3]

Evaluate the task and its consequences

Build the evaluation set from the intended workflow: ordinary traffic, difficult cases, malformed input, missing context, adversarial attempts, different user groups, and relevant languages. Keep a simpler baseline so the team can show that AI materially improves the outcome.

Select metrics that reflect consequences. Exactness may matter for extraction; groundedness and citation quality for research; precision and recall for detection; successful completion and correction rate for an agent. Report important slices and failure categories instead of hiding them inside one aggregate score.

Quality threshold
The minimum acceptable result on representative data, including critical slices.
Failure budget
Which failures are tolerable, how often, and with what fallback or review.
Regression rule
What blocks release when a model, prompt, retrieval source, or tool changes.
Business measure
The downstream result: completed work, reduced rework, decision time, loss avoided, or revenue.

Sources: [1], [2]

Test how the system behaves when inputs are hostile

OWASP's 2025 list identifies prompt injection, sensitive-information disclosure, supply-chain weaknesses, improper output handling, excessive agency, misinformation, and unbounded consumption among the key risks for LLM applications. The relevant tests depend on the product architecture, data, users, and tools.

Treat model output as untrusted input before it reaches code, databases, browsers, messages, or external systems. For agents, constrain permissions by task, require approval for consequential actions, cap iterations and spend, and make tool activity visible to operators.

Sources: [2], [3]

Prove fallback, monitoring, and ownership before launch

Run the fallback rather than documenting it. Confirm that users can continue when a provider fails, latency rises, context is unavailable, or quality falls below threshold. A manual route that cannot handle expected volume is not a viable fallback.

Monitoring should connect technical signals with workflow outcomes: quality samples, unresolved requests, user corrections, overrides, tool failures, latency, cost, security events, and downstream harm. Every alert needs an owner and an action. Define when to restrict functionality, switch models, increase review, or roll back.

Sources: [1], [2]

Record the release decision

The release record should identify the tested version of the application, model, prompts, evaluation data, retrieval sources, and tools. It should state which thresholds passed, which exceptions remain, who accepted them, and when approval expires.

Re-run the relevant gates after a material change to intended use, model, data, tools, permissions, user population, or risk. This keeps the gate proportional: teams do not repeat every review for every deployment, but they do reassess when the evidence may no longer describe the live system.

Related service

AI product development from prototype to production

Kesho designs, builds, evaluates, and operates AI products with the controls, integrations, and human workflows required for production use.

Explore the service