How Do I Assess an AI Vendor's Security?

Assess five things in order: data handling, model provenance, tenancy and isolation, subprocessors and independent assurance. A practitioner Q&A for buyers.

This page answers the questions IT leads, security managers and procurement owners ask when a new AI tool lands on the buying list. It covers what to ask a vendor, which certifications carry weight, how to read a subprocessor list and what a smaller organisation can reasonably skip. Written for teams who need a defensible answer this quarter, not a twelve-month assurance programme.

Reviewed by Jason Holloway.

How do I assess an AI vendor’s security?

Assess five things in order: how the vendor handles your data, where the model came from, how tenancy separates you from other customers, which subprocessors touch your data and what independent assurance exists. Answer those five and you have a defensible position. Everything else is refinement.

We run vendor assessments in that sequence deliberately. Data handling comes first because it decides whether the other four matter. If the vendor cannot tell you where your prompts and outputs are stored, how long they persist and whether they enter training, you have a blocker before the architecture questions.

The second reason the ordering matters is time. Most procurement windows allow a fortnight, not a quarter. Working down the list means that if you run out of time, you have covered the areas most likely to produce a material finding. Our AI Security Gap Analysis applies the same sequence at organisation level rather than per-vendor.

What data-handling questions should I ask an AI vendor?

Four questions, and insist on written answers rather than a sales call summary.

  • Is our input used to train or fine-tune your models, and can that be contractually disabled?
  • What is the retention period for prompts, outputs and logs, and who can read them?
  • In which jurisdictions is our data processed and stored?
  • On termination, what is deleted, when does it happen and how is deletion evidenced?

The answers you want are specific. A blanket claim that they do not train on customer data is weaker than a citation to the clause in the DPA that excludes it, together with a stated log retention window. Vague reassurance from a sales team frequently contradicts the platform documentation, and where the two conflict the documentation is closer to the truth.

A common failure mode is assessing the enterprise tier while staff use the consumer tier. That is a Shadow AI problem rather than a vendor problem, and it needs solving separately.

What is model provenance and why does it matter for vendor risk?

Model provenance means knowing which model is doing the work, who built it and whether it can change without notice. A vendor’s security posture is only as good as the models sitting underneath it.

Many AI products are interfaces over third-party foundation models. That is not a problem in itself, but it changes your risk picture: your data may traverse an organisation you never assessed, and the underlying model can be swapped at the vendor’s discretion. Ask which models they use, whether those are self-hosted or accessed via API, and whether you will be notified of a change.

Provenance also governs behaviour. A model change alters output, and output drives decisions. That is why we treat model behaviour as an assurance question in its own right through AI Behaviour Verification rather than assuming the vendor’s testing covers your use case.

How do I evaluate tenancy and isolation claims?

Ask what specifically is isolated. Multi-tenancy with logical separation is normal and acceptable for most workloads. What you need to establish is whether isolation covers the vector store, the retrieval index, the cache and the logs, not only the primary database.

Retrieval-augmented systems are where this goes wrong. If a vendor indexes your documents to answer questions, that index is a copy of your data in a new location with its own access model. Ask who can query it, whether embeddings are segregated per customer and whether support staff can view retrieved content.

For higher-risk deployments, ask whether a single-tenant option exists and what it costs. The answer frequently reveals more about the architecture than a diagram would.

How do I review an AI vendor’s subprocessor list?

Start by finding the list. It usually sits in an annex to the DPA or on a versioned page the contract references. If neither exists, that is your first finding.

Then separate material subprocessors from peripheral ones. Material means any party that stores, processes or can read customer content, including model providers, hosting and any human review or annotation service. Billing and marketing tooling matters less. Check the jurisdiction of each material entry and the transfer mechanism supporting it, against your own approved list.

Finally, check your rights when the list changes. You want written notice before a new subprocessor is engaged, a stated notice period and a right to object with a termination remedy if the objection is not resolved. A list with no change-notification clause behind it tells you what is true today and nothing about next quarter.

Which AI security certifications are worth trusting?

ISO 27001 and SOC 2 Type II are worth requesting and reading. ISO 42001 is worth asking about, and its absence is not disqualifying yet given how recently it arrived. Self-attested trust pages, badges with no report behind them and AI safety commitments published as marketing copy carry no assurance value.

The distinction is whether an independent party tested something and documented what they found. Request the SOC 2 report, not the logo, then read the scope statement and the exceptions section. A certification scoped to the vendor’s corporate IT while the AI platform sits outside that boundary is common and easy to miss.

Alignment with the NIST AI Risk Management Framework or the OWASP LLM Top 10 signals engineering maturity, though neither is a certification. Where the EU AI Act applies to the product classification, ask how they are preparing.

What does a worked assessment look like in practice?

An illustrative case: a finance team wants an AI meeting-summarisation tool.

Data handling: recordings and transcripts retained ninety days, excluded from training under the enterprise DPA, processed in the EU. Acceptable. Provenance: two third-party foundation models via API, no contractual notice on model change. Finding, raised for contract negotiation. Tenancy: logical separation confirmed for storage, but support staff can access transcripts during ticket investigation. Finding, mitigated by requiring named-approver access. Subprocessors: eleven listed, two in jurisdictions outside the approved list. Finding, escalated. Certifications: ISO 27001 held, scope confirmed to include the platform.

Outcome: approved with three contractual conditions and a review at renewal. Two days of work, documented, defensible. The point is not that the vendor was perfect. The point is that the residual risk was named and owned.

What can a smaller organisation reasonably skip?

Four things are safe to skip.

  • Penetration testing the vendor’s platform yourself, which they will refuse anyway.
  • Building a bespoke hundred-question questionnaire when a fifteen-question version covers the same ground.
  • Demanding single-tenancy for low-risk internal tools.
  • Requiring ISO 42001 as a mandatory condition today.

What you cannot skip: written confirmation on training and retention, a current subprocessor list, one independent assurance report and a named internal owner for the tool. Those four are the minimum defensible position, achievable inside a normal procurement cycle. Proportionality is the discipline, because a tool summarising public marketing copy does not warrant the depth of one processing client financial records.

Where should vendor assessment sit in our wider AI governance?

Vendor assessment is one input to AI governance, not a substitute for it. It tells you whether a specific tool is acceptable. It does not tell you what your organisation is running, who authorised it or how those decisions are reviewed.

Organisations that assess vendors well while lacking an inventory still carry unmanaged risk, because the tools that never reached procurement are the ones nobody assessed. Start with discovery, then tier, then assess. Related reading: what to ask suppliers about their AI features.

Assess the estate, not only the vendor

We apply the same five-part sequence at organisation level, so the tools that never reached procurement are covered too.