Ask for the denominator before you ask for the percentage. A number on its own, 92%, 99.9%, 0.68%, tells you nothing until you know what it is a percentage of, over what time window, and measured by whom. This article gives six questions for any AI security vendor, a pass/fail scorecard to fill in on the call, and agreed criteria for stopping a proof of concept, not only extending one.
Why the percentage alone is meaningless
A claim like "our AI resolves 95% of incidents" is not evidence. It is a headline waiting for a footnote. To turn it into evidence you need four things stated alongside it: the denominator (95% of what population, exactly), the sample (whose data, how much of it, over what period), the method (how was each case scored, and by whom), and the failure definition (what counts as a miss). Strip any one of those out and the number cannot be checked.
This matters more for AI security tools than for most software categories, because the product's entire pitch is usually a number: fewer false positives, faster detection, more incidents handled automatically. If the vendor will not show you the denominator behind that number, you are being sold a marketing claim dressed as a measurement.
What six questions separate a measured claim from a marketing one?
Use this scorecard on every vendor call. Each question has a pass condition. A vendor who cannot answer a question, or answers it with a different question, fails that row.
| # | Question | Pass condition | Fail signal |
|---|---|---|---|
| 1 | What is the denominator? | A named population (e.g. "all cases opened") with an exact or bounded count | "Most", "the majority", or no population stated |
| 2 | What is the sample and time window? | Named customer set (even if anonymised) and a specific rolling window (e.g. 30 days, 12 months) | "Historically" or no time window given |
| 3 | Who scored it, and how? | A defined method: automated count, audited by a named third party, or independently benchmarked | "Our customers tell us" or an uncited internal survey |
| 4 | What counts as a failure? | A written definition of a miss, a false positive, or an escalation, agreed before the test started | The failure definition is decided after the result, or not defined at all |
| 5 | Is this claim independently verified, and for what layer? | A named test body (e.g. an AV testing lab), with the scope of what it tested stated explicitly | An independent badge is implied to cover the whole product when it tested one component |
| 6 | What is the stopping rule? | Written criteria for when a proof of concept fails, agreed before it starts | Only success criteria exist; failure is decided informally at the end |
Copy this table into your evaluation pack. Send it to every shortlisted vendor before the first demo. The answers, or the silences, will do most of your shortlisting for you.
Worked examples: reading SenseOn's own numbers against the scorecard
Here is how SenseOn's own published numbers read against the six questions above, including the parts that do not fully close the loop.
"92.5% of incidents resolved by AI under human governance." Denominator: total cases raised across all customer environments, over a rolling 30-day window. Method: cases resolved without human escalation, divided by cases opened. Read this claim carefully: it is a resolution rate under human governance, not an accuracy figure. It says nothing about how many of the 7.5% that were escalated turned out to be genuine threats versus noise. Row 4 (failure definition) is answered: an escalation is the counted event.
"0.68% true-positive density across 30M+ cases." Of more than 30 million cases investigated across all customer environments over a rolling 12-month window, 0.68% were confirmed true positives. Read the denominator in the same breath as the figure: it is true-positive density, the share of a very large investigated population that turned out to be a real threat, not a false-positive rate. SenseOn does not publish a false-positive rate, and this number cannot be inverted into one, because a true-positive-density figure and a false-positive rate answer different questions from different denominators. If a vendor's own materials ever appear to derive a false-positive rate from a true-positive-density number, that is the error to catch. Ask which denominator the figure you are being shown uses, every time.
"Mean time to detect and respond <20 min." Aggregate across managed-service customer environments, rolling 30-day window. This is an average across environments of varying size and complexity, not a guarantee for any single environment; ask the vendor for the range, not just the mean.
"~40% data reduction at the edge." Pre-processing volume compared with post-processing volume, measured in gigabytes per day. This is an engineering efficiency figure, not a detection-quality figure; the two are easy to conflate in a sales conversation and worth separating.
SE Labs AAA rating and AV-Comparatives A+. Both are independent tests, and both are scoped to endpoint protection testing. Neither tests, and neither claims to test, the quality of an AI agent's investigation or response decisions. If a vendor cites an endpoint-testing badge as evidence for its AI decision layer, that is row 5 failing: the badge is real, the scope claimed for it is not. As of this writing, no independent benchmark of agentic SOC decision quality exists in the market for any vendor to cite, SenseOn included; treat any claim to the contrary as a fail on row 5.
The wider SIEM detection gap, for context. CardinalOps' 2025 State of SIEM Report, based on a sample of more than 13,000 detection rules and more than 2.5 million logs, found that enterprise SIEMs carry detections for 22% of MITRE ATT&CK techniques, leaving 78% uncovered. This is third-party research, not a SenseOn measurement, and it is cited here as an example of a claim with its sample size stated in the same sentence as the finding, exactly what row 1 and row 2 of the scorecard ask every vendor, including the one publishing a report like this, to provide.
When should you stop a proof of concept, not just extend it?
Most POC failures are never declared. They drift into a second extension, then a third, because nobody agreed in advance what a failure looks like. Fix that before the POC starts, not after it stalls, with three written stopping conditions:
- A coverage stopping rule. Name the log sources, environments or attack techniques the POC must exercise. If the vendor cannot demonstrate detection against an agreed, pre-listed technique set within the test window, that is a stop, not a request for more time.
- A noise stopping rule. Set a maximum number of analyst hours per week the tool may cost your team during the test. If it exceeds that budget without a documented plan to tune it down, stop.
- A verification stopping rule. Agree who checks the vendor's own reported numbers against your own environment's data, and how. If the vendor's dashboard and your independent count disagree by more than an agreed margin, and the vendor cannot explain why, stop.
Write all three into the POC agreement before you sign it. A vendor confident in its own numbers will not object to being held to them.

The scorecard, one more time
Denominator. Sample and time window. Method. Failure definition. Independent scope. Stopping rule. Six questions, one page, filled in on every vendor call before you buy. If a claim survives all six with a specific, checkable answer, it has earned the right to be in your business case. If it does not, it is marketing, and marketing is not evidence.
See the full methodology behind every number SenseOn publishes, read the CISO's AI Accountability Playbook for the wider set of questions a board will ask, and compare this method against a broader vendor-evaluation framework in How to Evaluate Security Vendors Without the Hype.