Title card: How CISOs evaluate AI vendors. One workflow, Data route, Permissions, Rejection test, Logs + exit.

How CISOs evaluate AI vendors: the questions that produce evidence.

Short answer. CISOs evaluate AI vendors by testing one real workflow instead of scoring a questionnaire. Trace where the data actually goes, who holds the admin keys, what happens when a human rejects a proposal, and whether the raw logs export without a vendor’s help. A ten-day pilot against a scoped account produces more usable evidence than a security review answered by the vendor’s own sales engineer.

Security teams already run this discipline for a new payment processor or a new identity provider. An AI vendor earns the same treatment. The questions below are built to produce a record you can check yourself.

Why do vendor security questionnaires fall short?

A questionnaire captures what a vendor says about itself, in the vendor’s own words, answered by whoever in the company is fastest with a spreadsheet. It rewards a well-drafted policy document over a well-run system. Nothing in a completed SIG (the Shared Assessments SIG questionnaire) or CAIQ (the Cloud Security Alliance CAIQ questionnaire) tells you whether a rejected AI proposal actually stops before it reaches your systems. Nor does it tell you who besides your own staff can log into your account today.

A questionnaire still earns a place early, as a filter to cut a long vendor list down to a short one. It cannot substitute for watching the product behave under your own account. That gap is why an evidence sequence belongs in the process before a contract is signed, the same reasoning we set out in how to prove AI controls to an auditor.

How do CISOs evaluate AI vendors, step by step?

  1. Scope one workflow. Pick a single real task the AI will perform, and leave the vendor’s full feature list for later. A narrow scope is easier to test cleanly and tells you more than a broad tour of the product.
  2. Trace the data route. Find out which model or service actually processes a prompt for that workflow, whether it sits inside a boundary you control, and which subprocessors the vendor names for it. Get the answer in writing, tied to your account, the same question we walk through in self-hosted alternatives to Microsoft Copilot.
  3. Test identity and permissions. Confirm which of your own people or service accounts can act through the AI, and check that scope against what the workflow genuinely needs. Excess permission is a design flaw you can find before you sign.
  4. Run the approval and rejection test. Approve one proposal and compare the executed action to what was shown. Reject a second one and check the target system directly for any change. A vendor demonstration usually shows the approval; ask to see the rejection, the same test we set out in human approval workflows for AI agents.
  5. Request the logs and the export. Pull the raw record for one governed action (an AI action that went through approval) from your own admin console. A report the vendor generated for you does not count as the record. Confirm it holds the identity, the proposal and the decision together.
  6. Confirm the exit. Get the offboarding process in writing: what happens to your data and configuration the day you stop paying, and how you would verify deletion independently of the vendor’s word.

Each step produces proof: a written data-route answer, a permission list, a rejected proposal that verifiably changed nothing, a raw log export, an exit clause. A questionnaire produces none of these on its own.

Which questions produce a claim, and which produce evidence?

Question The weak answer The evidence to ask for instead
Are you SOC 2 compliant? “Yes, we’re SOC 2 Type II certified.” The report’s scope section, confirming your specific deployment and configuration sit inside it, distinct from the vendor’s core platform alone
Do you have human-in-the-loop controls? “Yes, a human reviews every output before it executes.” A rejected proposal from your own pilot account, plus an independent read of the target system confirming nothing executed
Is our data used to train your models? “We don’t train on customer data.” The specific model endpoint and provider handling your prompts for one real run, named in writing
Can we get our logs? “Logs are available on request.” The raw log for one action, pulled from your own admin console right now, without a support ticket
Who can access our account? “Access is tightly controlled.” A support ticket opened against your own test account, showing what vendor staff can see or do without asking you first
What happens if we leave? “You can export your data any time.” The offboarding clause itself, naming a deletion timeline and a way to verify it independently

What does a realistic 10-day evaluation look like?

  1. Day 1 to 2. Choose one real but non-critical workflow, and get a scoped pilot account with a single administrator identity your own staff control.
  2. Day 3. Submit a synthetic task with no real business data in it, and ask which model endpoint processed it. Get the answer in writing.
  3. Day 4. Map the permissions the AI actually holds in the pilot account against what the workflow genuinely needs, and flag anything broader.
  4. Day 5 to 6. Approve one proposal and compare the action that executed against what was shown on screen.
  5. Day 7. Reject a second proposal, then check the target system directly. Do not rely on the vendor’s own status screen for this.
  6. Day 8. Pull the raw log for one of these actions from your own console, and confirm it carries the identity, the proposal and the decision together.
  7. Day 9. Open a support ticket against the pilot account and record what vendor staff can see or reach without asking you first.
  8. Day 10. Get the exit and offboarding process in writing, and score the vendor against the sequence above before the pilot account closes.

Ten days is enough for one workflow, tested properly. It is not enough to evaluate a vendor’s entire platform, and you should not try.

What red flags should you check for?

  • Check whether the SOC 2 or ISO/IEC 42001 scope statement names the specific module you plan to use, or only the vendor’s core platform.
  • Check whether “human in the loop” is backed by a proposal-bound approval record in your pilot account, or only described in a policy document.
  • Check whether the log the vendor hands you is the raw record or a summary report generated for the occasion.
  • Check whether your support ticket reveals a standing access path for vendor staff that the contract does not mention.
  • Check whether the offboarding clause names a specific deletion timeline, or only a general commitment to work with you on exit.
  • Check whether the vendor can name the exact model endpoint and any subprocessor handling your prompts, in writing, for the workflow you tested.

How do you score the result?

Score each of the six steps as evidenced, claimed or failed. Evidenced means you saw the actual record or behaviour yourself: the rejected proposal, the raw log, the support ticket’s result. Claimed means the vendor described it but your pilot did not verify it. Failed means the vendor could not produce it, or the pilot contradicted what was claimed.

A vendor needs an evidenced result on identity and permissions, the approval and rejection test, and the logs before it goes forward for anything touching regulated data. A claimed result on the data route or the exit is a conditional pass. Ask the vendor to close the gap in writing, and retest before renewal.

What do the standard frameworks add to this?

None of them replace running the test yourself, but each backs a piece of it. The NIST AI Risk Management Framework organises this kind of evaluation under its Govern and Measure functions, calling for accountability structures and documented test results. The OWASP Top 10 for LLM Applications names the specific failure modes your pilot is checking for, including LLM06:2025 Excessive Agency. That is the risk that a system holds more function, permission or autonomy than the task needs. The joint Guidelines for Secure AI System Development from the UK’s NCSC and the US’s CISA cover the same ground from the provider’s side. That is useful context for judging how seriously a vendor treats secure design before the pilot even starts.

Questions buyers ask.

How long should a CISO spend evaluating an AI vendor?

Ten working days against one scoped workflow is enough to produce real evidence: a data-route answer, an approval and rejection test, a raw log export and an exit clause in writing. A longer evaluation often means the scope crept back to the vendor’s whole platform instead of one workflow.

Is a SOC 2 or ISO 42001 certificate enough to approve an AI vendor?

No. Both describe a vendor’s management system and control environment. Neither describes your specific deployment. Check that the certificate’s scope statement names the module you plan to use, and still run the approval, rejection and log tests against your own pilot account.

Which single test catches the most overstated vendor controls?

The rejection test. Reject a proposal in the pilot account, then check the target system directly instead of trusting the vendor’s own status screen. A vendor whose demonstration shows only approvals has not shown you the control that actually matters.

Should a CISO test a vendor’s incident response the same way?

Yes, on the same principle. Ask for a specific record from a past drill or a real incident, distinct from a policy describing the intended process. If the vendor cannot produce a timeline, an identity trail and a resolution record from an actual test, treat the answer as unverified.

I wrote a short brief for exactly this conversation at workspace.handvantage.com/for-ciso, built for a CISO who needs to explain an AI platform decision to a board in one sitting. The same evidence sequence sits behind the twelve tests in the Agentic AI Procurement Handbook. Handvantage builds Vantage Workspace, a self-hosted AI workspace built so that AI actions go through human approval and leave a record. Weigh my view accordingly, and run the same pilot on us that you would run on anyone else. This is an evaluation method, not legal advice, and your own procurement and compliance function should confirm which controls your organisation actually requires.

The workspace site applies the same sequence to an AI workspace with agents: the CISO evaluation guide for AI workspace and agent governance.

Josh Olayemi · Founder, Handvantage · September 2026 · About the author

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *