Technology specialists reviewing technical evidence
KESHO PARTNERSAll insights

Technical due diligence

AI technical due diligence: separating a demo from a defensible product

A seven-part scorecard for investors evaluating whether an AI company can deliver its claims reliably, securely, and at viable economics.

By Kesho Partners

12 minute read

AI technical due diligence should test seven connected dimensions: customer value, data rights and quality, model evidence, system architecture, security and governance, unit economics, and team execution. A strong demo proves that a workflow can work once. A defensible product proves that it can work repeatedly, within stated limits, under production conditions, at a cost the business can sustain.

Begin with the investment thesis, not the technology stack

The diligence scope should follow the claims that support the investment. If the thesis depends on proprietary data, test ownership, collection rights, coverage, quality, and whether the data materially changes performance. If it depends on workflow lock-in, inspect integrations, user behavior, switching costs, and how deeply the product sits inside daily operations.

This prevents a common failure: completing a technically detailed review that never answers whether the technology supports the commercial case. Every finding should connect to growth, margin, customer concentration, regulatory exposure, delivery capacity, or the cost and time required to reach the next milestone.

Write the three to five technical claims that must be true for the investment thesis to hold. Use those claims to set the depth of review.

The seven-dimension AI diligence scorecard

Score each dimension from 1 to 4: unproven, fragile, credible, or defensible. Record the evidence supporting the score and your confidence in that evidence. Do not average away a critical failure: data rights, security, or unacceptable model behavior can remain a deal-level issue even when the overall score looks healthy.

Kesho AI technical diligence scorecard
DimensionCore questionEvidence to inspectMaterial warning
Customer valueDoes AI improve a measurable workflow outcome?Usage cohorts, retention, task completion, user interviewsThe product is impressive but optional
DataCan the company lawfully access and maintain the data advantage?Rights, provenance, quality tests, coverage, refresh processCritical data is borrowed, brittle, or poorly governed
Model evidenceDoes performance hold on representative cases?Evaluation sets, baselines, errors, regressions, human reviewClaims rely on selected demos or vendor benchmarks
ArchitectureCan the whole system meet reliability and scale requirements?System map, dependencies, observability, failure handlingCore behavior depends on hidden manual work
Security & governanceAre foreseeable risks controlled and owned?Threat models, access, incidents, approvals, oversightNo accountable owner or release evidence
EconomicsDoes cost improve or degrade as usage grows?Inference cost, gross margin, support burden, vendor termsRevenue scales more slowly than compute and review cost
Team executionCan this team operate and improve the product?Roadmap history, ownership, hiring gaps, key-person riskCritical knowledge sits with one person or supplier

Sources: [1], [2], [3]

Test the product's evidence, not the base model's reputation

A capable foundation model does not prove that the application is reliable. Product performance depends on prompts, retrieval, tools, permissions, data, interface design, human review, and failure handling. Ask for evaluations built from the company's actual tasks and users.

The evidence should show a baseline, target thresholds, representative test cases, known failure modes, regression history, and who can approve exceptions. For agentic systems, inspect tool permissions and whether the agent can take irreversible action. OWASP identifies prompt injection, sensitive-information disclosure, supply-chain exposure, improper output handling, excessive agency, misinformation, and unbounded consumption among the major application risks.

Representative
Test ordinary cases, difficult cases, adversarial inputs, missing context, and ambiguous requests drawn from the intended workflow.
Comparative
Compare against the current human or software process, a simpler non-AI baseline, and alternative model choices.
Repeatable
Keep the dataset, scoring method, model version, prompt version, and result so the claim can be reproduced after change.
Decision-linked
Translate quality into business consequence: rework, failed transactions, escalation volume, customer harm, or lost time.

Sources: [2], [3]

Model the economics at the workflow level

Token cost is only one component. Include retrieval, external tools, retries, observability, storage, human review, customer support, and the engineering work required when providers or models change. Then compare total cost with the value of the completed workflow.

Run sensitivity cases for higher usage, longer context, degraded first-pass quality, vendor price changes, and heavier human review. The relevant question is not whether today's gross margin is attractive at low usage. It is whether the operating model becomes stronger as the product succeeds.

Economics sensitivity test
VariableBase caseStress caseDecision
Model and tool costObserved cost per completed task2× usage and longer contextIs margin still acceptable?
Human reviewCurrent escalation rateQuality falls or task complexity risesCan operations absorb review?
Vendor dependencyCurrent contracted termsPrice, limits, or access changesIs switching technically and commercially viable?
Support burdenCurrent incidents and manual fixesEnterprise-scale usageDoes support scale with revenue?

Turn findings into an investment decision

The final output should distinguish confirmed facts, management assertions, open questions, and assumptions. Rank findings by materiality to the thesis and urgency, not by technical elegance. A weak test suite can be repaired; unclear rights to foundational data may be existential.

For each material issue, state the business implication, recommended action, accountable owner, dependency, and realistic sequence. The investor should leave knowing what changes the decision now, what belongs in deal terms, and what management should address during the first 100 days.

Proceed
Evidence supports the thesis and remaining risks fit normal execution.
Proceed with conditions
The thesis remains credible, but specific controls, hires, rights, or remediation should be completed.
Reprice or restructure
Material cost, dependency, or delivery assumptions differ from the deal model.
Do not proceed
A foundational claim is unsupported or the downside cannot be controlled within the investment case.

Related service

Technical due diligence for investors and acquirers

Kesho provides independent assessments of product, architecture, data, AI, security, economics, and engineering capability, tied directly to the investment thesis.

Explore the service