
Technical due diligence
AI technical due diligence: separating a demo from a defensible product
A seven-part scorecard for investors evaluating whether an AI company can deliver its claims reliably, securely, and at viable economics.
By Kesho Partners
12 minute read
AI technical due diligence should test seven connected dimensions: customer value, data rights and quality, model evidence, system architecture, security and governance, unit economics, and team execution. A strong demo proves that a workflow can work once. A defensible product proves that it can work repeatedly, within stated limits, under production conditions, at a cost the business can sustain.
Begin with the investment thesis, not the technology stack
The diligence scope should follow the claims that support the investment. If the thesis depends on proprietary data, test ownership, collection rights, coverage, quality, and whether the data materially changes performance. If it depends on workflow lock-in, inspect integrations, user behavior, switching costs, and how deeply the product sits inside daily operations.
This prevents a common failure: completing a technically detailed review that never answers whether the technology supports the commercial case. Every finding should connect to growth, margin, customer concentration, regulatory exposure, delivery capacity, or the cost and time required to reach the next milestone.
Write the three to five technical claims that must be true for the investment thesis to hold. Use those claims to set the depth of review.
The seven-dimension AI diligence scorecard
Score each dimension from 1 to 4: unproven, fragile, credible, or defensible. Record the evidence supporting the score and your confidence in that evidence. Do not average away a critical failure: data rights, security, or unacceptable model behavior can remain a deal-level issue even when the overall score looks healthy.
| Dimension | Core question | Evidence to inspect | Material warning |
|---|---|---|---|
| Customer value | Does AI improve a measurable workflow outcome? | Usage cohorts, retention, task completion, user interviews | The product is impressive but optional |
| Data | Can the company lawfully access and maintain the data advantage? | Rights, provenance, quality tests, coverage, refresh process | Critical data is borrowed, brittle, or poorly governed |
| Model evidence | Does performance hold on representative cases? | Evaluation sets, baselines, errors, regressions, human review | Claims rely on selected demos or vendor benchmarks |
| Architecture | Can the whole system meet reliability and scale requirements? | System map, dependencies, observability, failure handling | Core behavior depends on hidden manual work |
| Security & governance | Are foreseeable risks controlled and owned? | Threat models, access, incidents, approvals, oversight | No accountable owner or release evidence |
| Economics | Does cost improve or degrade as usage grows? | Inference cost, gross margin, support burden, vendor terms | Revenue scales more slowly than compute and review cost |
| Team execution | Can this team operate and improve the product? | Roadmap history, ownership, hiring gaps, key-person risk | Critical knowledge sits with one person or supplier |
Test the product's evidence, not the base model's reputation
A capable foundation model does not prove that the application is reliable. Product performance depends on prompts, retrieval, tools, permissions, data, interface design, human review, and failure handling. Ask for evaluations built from the company's actual tasks and users.
The evidence should show a baseline, target thresholds, representative test cases, known failure modes, regression history, and who can approve exceptions. For agentic systems, inspect tool permissions and whether the agent can take irreversible action. OWASP identifies prompt injection, sensitive-information disclosure, supply-chain exposure, improper output handling, excessive agency, misinformation, and unbounded consumption among the major application risks.
- Representative
- Test ordinary cases, difficult cases, adversarial inputs, missing context, and ambiguous requests drawn from the intended workflow.
- Comparative
- Compare against the current human or software process, a simpler non-AI baseline, and alternative model choices.
- Repeatable
- Keep the dataset, scoring method, model version, prompt version, and result so the claim can be reproduced after change.
- Decision-linked
- Translate quality into business consequence: rework, failed transactions, escalation volume, customer harm, or lost time.
Model the economics at the workflow level
Token cost is only one component. Include retrieval, external tools, retries, observability, storage, human review, customer support, and the engineering work required when providers or models change. Then compare total cost with the value of the completed workflow.
Run sensitivity cases for higher usage, longer context, degraded first-pass quality, vendor price changes, and heavier human review. The relevant question is not whether today's gross margin is attractive at low usage. It is whether the operating model becomes stronger as the product succeeds.
| Variable | Base case | Stress case | Decision |
|---|---|---|---|
| Model and tool cost | Observed cost per completed task | 2× usage and longer context | Is margin still acceptable? |
| Human review | Current escalation rate | Quality falls or task complexity rises | Can operations absorb review? |
| Vendor dependency | Current contracted terms | Price, limits, or access changes | Is switching technically and commercially viable? |
| Support burden | Current incidents and manual fixes | Enterprise-scale usage | Does support scale with revenue? |
Turn findings into an investment decision
The final output should distinguish confirmed facts, management assertions, open questions, and assumptions. Rank findings by materiality to the thesis and urgency, not by technical elegance. A weak test suite can be repaired; unclear rights to foundational data may be existential.
For each material issue, state the business implication, recommended action, accountable owner, dependency, and realistic sequence. The investor should leave knowing what changes the decision now, what belongs in deal terms, and what management should address during the first 100 days.
- Proceed
- Evidence supports the thesis and remaining risks fit normal execution.
- Proceed with conditions
- The thesis remains credible, but specific controls, hires, rights, or remediation should be completed.
- Reprice or restructure
- Material cost, dependency, or delivery assumptions differ from the deal model.
- Do not proceed
- A foundational claim is unsupported or the downside cannot be controlled within the investment case.
Related service
Technical due diligence for investors and acquirers
Kesho provides independent assessments of product, architecture, data, AI, security, economics, and engineering capability, tied directly to the investment thesis.
Explore the service