FREE TOOLS › TOOL 03
Ten questions for your next AI-security vendor call
Take this into any evaluation, including ours. Each question has an answer worth hearing and an answer that should slow you down. Tick as you go and the tally at the bottom tells you where you landed.
What component makes the decision to block?
Listen for:
A deterministic engine. Rules, compiled patterns, statistical heuristics. Something that produces the same answer for the same input every time.
Worry if you hear:
A model, a classifier, a judge, an evaluator, or the phrase AI-native used as an answer rather than a description.
Why it matters:
If a probabilistic component holds the enforcement position, published research has demonstrated evasion of up to 100% against comparable production systems. That may be acceptable to you. It should not be a surprise.
What is your measured added latency at p99, and will you put it in a contract?
A specific number, a description of the test conditions, and a willingness to be measured on your hardware during a pilot.
It depends. Negligible. Nobody has ever complained. Or a marketing number with no test conditions attached.
Why it matters. Most major platform products publish no latency figure at all. Published figures range from 19 milliseconds for a small encoder classifier to about 530 milliseconds for a three-rail model-based stack.
Which version of MITRE ATLAS is your coverage claim measured against?
A specific release, such as 2026.07, and a list of which techniques are excluded and why.
A percentage with no version. Or 100% of MITRE ATLAS with no qualifier, which means they have not read it.
Why it matters. ATLAS ships monthly content releases and holds 16 tactics and 178 techniques as of release 2026.07, more than double an earlier version many vendor pages still quote.
Have you re-mapped to the OWASP LLM Top 10 released on 3 August 2026?
Yes, and an accurate description of what changed: Excessive Agency moved from sixth to third, and LLM08 is now Hidden Context Exposure.
Confusion, or a coverage claim that matches the 2025 list exactly.
Why it matters. This is a freshness test more than a coverage test. A vendor whose framework mapping is a year stale has a detection team that is busy with something else.
What happens to my policies when I change model provider?
Nothing. Policy is enforced at a layer above the provider and travels with your stack.
Anything involving remapping, reconfiguration or a services engagement.
Why it matters. You will swap models within eighteen months, for cost, capability, compliance or an outage. Coupling your security posture to a procurement decision is a choice worth making knowingly.
Can I run this against my own production traffic before I pay you?
Yes, non-blocking, in a defined timeframe, at no cost, producing a written report.
A sandbox demo, a reference architecture, a proof of concept that requires a signed order form first.
Why it matters. This is the question that most cleanly separates vendors who can demonstrate something from vendors who can describe something.
Can an AI component in your product author or enforce a live policy?
No. It can propose. A deterministic check and a human approve before anything goes live.
Descriptions of self-learning, adaptive or autonomous policy generation, presented as a benefit.
Why it matters. A system that can rewrite its own enforcement rules based on model output has a path from prompt injection to policy change. Ask them to walk you through why it does not.
Show me the log entry for a blocked request.
A rule identifier, the matched condition, a timestamp, and a tamper-evident chain a human can read and an auditor can verify.
A confidence score, a risk rating, or a dashboard tile with no underlying record of why.
Why it matters. Every regulatory regime currently in force converges on evidence. A confidence score is not evidence, it is a number a model produced about itself.
What can your product not stop?
A specific, unhesitating answer with named categories, such as operating-system and network-layer paths, or attacker-side reconnaissance.
Hesitation, a pivot to feature count, or any version of nothing, we cover everything.
Why it matters. This is the integrity question. A vendor who cannot name their own scope boundary either does not know it or has decided not to tell you, and both are disqualifying for a security control.
Who owns you, and what happens to my deployment if that changes?
A direct answer on ownership plus a willingness to negotiate continuity terms: escrow, assignment provisions, notice on change of control, sunset commitments.
Reassurance without contract language. A five-year independence promise, which no founder can make.
Why it matters. Fifteen firms in this category were acquired between August 2024 and June 2026, including six of the vendors that defined it. Ask everyone, including us.
Your tally
answered to your satisfaction.
No verdict bands here, deliberately: we don't define any, and a vendor can fail question 1 and still be the right purchase for your environment. The value is in asking, hearing the answer, and knowing what you are accepting.
Work through the list and this will fill in.
This checklist is deliberately biased toward the things we think matter, which is worth knowing while you use it. A vendor can fail question 1 and still be the right purchase for your environment. The value is in asking, hearing the answer, and knowing what you are accepting.