
AI product development
The AI production release gate: nine decisions before launch
A practical release standard for product and engineering teams moving an AI workflow from a convincing prototype into live operations.
By Kesho Partners
10 minute read
An AI feature is ready for production only when the team has accepted evidence across nine areas: intended use, task quality, failure modes, security, permissions, latency, unit cost, human intervention, and production monitoring. Passing an average quality score is insufficient. The release decision must also define operating limits, fallback behavior, accountable owners, and the conditions that trigger rollback or reassessment.
A prototype answers a different question from production
A prototype asks whether a model can perform a useful task under guided conditions. Production asks whether the complete system can perform that task for real users, with variable inputs, connected data, permissions, failures, costs, and support obligations.
The release gate should evaluate the full workflow rather than the model in isolation. Retrieval, tools, prompts, data pipelines, interfaces, human review, and downstream systems can each improve or undermine the result. AWS notes that changing data requires continuous monitoring to detect and mitigate accuracy and performance issues; release evidence is therefore the start of control, not its endpoint.
The nine-part production release gate
For each gate, record the threshold, observed result, evidence link, owner, and decision. A gate can pass, pass with a time-bound condition, or fail. Security, data rights, unsafe agency, and the absence of a viable fallback should not be averaged against strong quality elsewhere.
| Gate | Required decision | Minimum evidence |
|---|---|---|
| 1. Intended use | The workflow and prohibited uses are explicit | System card, users, boundaries, assumptions |
| 2. Task quality | Performance meets workflow-specific thresholds | Representative evaluation, baseline, error analysis |
| 3. Failure modes | Known failures have a controlled outcome | Adversarial tests, edge cases, fallback behavior |
| 4. Security | Input, output, data, and dependencies are protected | Threat model, abuse tests, output validation |
| 5. Permissions | The system has only the authority it needs | Tool scopes, approvals, irreversible-action controls |
| 6. Latency | Response time fits the user and operational need | Percentile latency by workflow and dependency |
| 7. Unit cost | Cost per successful outcome supports the business case | Model, tool, retry, review, and support cost |
| 8. Human intervention | People can understand, override, and escalate | Review procedure, authority, interface test |
| 9. Monitoring | The team can detect deterioration and respond | Signals, thresholds, alerts, owner, rollback test |
Evaluate the task and its consequences
Build the evaluation set from the intended workflow: ordinary traffic, difficult cases, malformed input, missing context, adversarial attempts, different user groups, and relevant languages. Keep a simpler baseline so the team can show that AI materially improves the outcome.
Select metrics that reflect consequences. Exactness may matter for extraction; groundedness and citation quality for research; precision and recall for detection; successful completion and correction rate for an agent. Report important slices and failure categories instead of hiding them inside one aggregate score.
- Quality threshold
- The minimum acceptable result on representative data, including critical slices.
- Failure budget
- Which failures are tolerable, how often, and with what fallback or review.
- Regression rule
- What blocks release when a model, prompt, retrieval source, or tool changes.
- Business measure
- The downstream result: completed work, reduced rework, decision time, loss avoided, or revenue.
Test how the system behaves when inputs are hostile
OWASP's 2025 list identifies prompt injection, sensitive-information disclosure, supply-chain weaknesses, improper output handling, excessive agency, misinformation, and unbounded consumption among the key risks for LLM applications. The relevant tests depend on the product architecture, data, users, and tools.
Treat model output as untrusted input before it reaches code, databases, browsers, messages, or external systems. For agents, constrain permissions by task, require approval for consequential actions, cap iterations and spend, and make tool activity visible to operators.
Prove fallback, monitoring, and ownership before launch
Run the fallback rather than documenting it. Confirm that users can continue when a provider fails, latency rises, context is unavailable, or quality falls below threshold. A manual route that cannot handle expected volume is not a viable fallback.
Monitoring should connect technical signals with workflow outcomes: quality samples, unresolved requests, user corrections, overrides, tool failures, latency, cost, security events, and downstream harm. Every alert needs an owner and an action. Define when to restrict functionality, switch models, increase review, or roll back.
Record the release decision
The release record should identify the tested version of the application, model, prompts, evaluation data, retrieval sources, and tools. It should state which thresholds passed, which exceptions remain, who accepted them, and when approval expires.
Re-run the relevant gates after a material change to intended use, model, data, tools, permissions, user population, or risk. This keeps the gate proportional: teams do not repeat every review for every deployment, but they do reassess when the evidence may no longer describe the live system.
Related service
AI product development from prototype to production
Kesho designs, builds, evaluates, and operates AI products with the controls, integrations, and human workflows required for production use.
Explore the service