
Financial services
The third-party AI evidence pack for financial services
Ten evidence areas for procurement, risk, compliance, security, and product teams evaluating an AI vendor or embedded AI capability.
By Kesho Partners
12 minute read
A regulated firm evaluating third-party AI should request evidence across ten areas: intended use, data flows, model and provider dependencies, evaluation, explainability, security, resilience, human oversight, incidents and changes, and exit. The review should be proportional to the use case's materiality and should verify evidence rather than relying on policy statements or a generic security questionnaire.
AI vendors create a different evidence problem
The Bank of England and FCA's 2024 survey found that 75% of responding financial-services firms already used AI and one-third of reported use cases were third-party implementations. Forty-six percent described only a partial understanding of the AI technologies they used, with third-party models a major reason for the gap.
A conventional SaaS review remains necessary, but it may not establish how an AI system was evaluated, where its outputs are unreliable, which upstream models and datasets it depends on, how behavior changes, or when a person can intervene. Those questions matter more as AI enters customer, risk, compliance, and operational decisions.
Sources: [1]
The ten-part vendor evidence request
Use the table as a starting request, then deepen the review according to customer impact, decision authority, data sensitivity, operational criticality, and substitutability. Ask for dated artifacts tied to the proposed product version and deployment, not a broad description of the vendor's general practice.
| Evidence area | What to request | Decision supported |
|---|---|---|
| 1. Intended use | Purpose, users, decisions, limitations, prohibited uses | Is the product appropriate for this workflow? |
| 2. Data flow and retention | Inputs, outputs, locations, subprocessors, training use, deletion | Can data obligations and confidentiality be met? |
| 3. Dependencies | Models, providers, datasets, tools, regions, concentration | Which hidden parties and changes affect the service? |
| 4. Evaluation | Representative tests, thresholds, errors, robustness, regressions | Do performance claims hold for the firm's use? |
| 5. Explainability | Available reasons, evidence, logs, user disclosures | Can users and reviewers understand relevant outcomes? |
| 6. Security | Threat model, access, testing, output controls, incident history | Can foreseeable abuse and disclosure be controlled? |
| 7. Resilience | Availability, fallback, recovery, capacity, provider failure scenarios | Can the service remain within operational tolerances? |
| 8. Human oversight | Review points, authority, override, escalation, appeal | Can people intervene effectively before harm? |
| 9. Incidents and changes | Notification terms, versioning, material-change tests, audit evidence | Will the firm know when its assessment is stale? |
| 10. Exit | Data export, deletion, transition support, replacement feasibility | Can the firm leave without unacceptable disruption? |
Set review depth by materiality
The same vendor can require different assurance for different uses. Summarizing internal policy has a different failure impact from recommending a fraud response, influencing credit, communicating with a vulnerable customer, or operating a critical control.
Classify the proposed use before reviewing the supplier. Consider the scale and reversibility of harm, customer and market impact, data sensitivity, degree of automation, ability to detect errors, operational criticality, and dependence on the service. Document who approved the classification.
- Low materiality
- Limited consequence, easy human verification, no sensitive action, and simple fallback. Use baseline evidence and monitoring.
- Medium materiality
- Meaningful operational or customer effect. Require workflow-specific evaluation, oversight, resilience, and change controls.
- High materiality
- Consequential decisions, critical operations, material exposure, or difficult remedy. Require independent challenge and stronger contractual evidence.
Follow the data and the dependency chain
Four of the five highest current AI risks in the Bank/FCA survey were data-related: privacy and protection, quality, security, and bias or representativeness. Trace customer and firm data through prompts, logs, retrieval stores, support systems, model providers, analytics, and subprocessors. Establish whether any party can use it for training or service improvement.
Map model, cloud, data, and tool dependencies. The survey found concentration among named providers and respondents expected critical third-party dependency to show the largest increase in systemic risk. The firm needs notice of material dependency changes and a tested response when a provider, model, region, or commercial term changes.
Require evidence for the firm's workflow
A vendor benchmark is not automatically representative of the firm's customers, data, language, products, channels, or controls. Agree the evaluation population, critical failure types, thresholds, and monitoring before approval. Where testing requires firm data, establish an appropriately controlled evaluation environment.
Human oversight should identify what reviewers see, what they can change, how much time they have, and when the system must escalate. Measure overrides, disagreements, complaints, and downstream outcomes. A nominal reviewer who routinely accepts an opaque output does not provide meaningful control.
Put continuing evidence into the operating model
Approval should identify which changes require notice, reassessment, or consent: foundation model changes, new subprocessors, different data use, material capability additions, new regions, revised safety controls, or changed service limits. Contractual rights matter only if named owners monitor and exercise them.
Maintain a current system owner, vendor owner, risk classification, evidence set, issue log, review date, and exit plan. Track performance and incidents against agreed thresholds. When evidence is unavailable, record the uncertainty and decide whether compensating controls reduce it sufficiently; do not convert absence of evidence into an assumed pass.
Related service
AI delivery and governance for financial services
Kesho helps banks, insurers, fintechs, and financial infrastructure firms assess AI suppliers, build controlled products, and produce evidence for internal review.
Explore the service