AI pentesting works when its output survives four checks: the same target run twice produces findings that match, your team's manual validation workload actually falls, a human sits on the escalation path for anything ambiguous or destructive, and the findings it surfaces get fixed, not filed. Score a vendor against those four before coverage claims or price enter the conversation.
The four-measure scorecard
| Measure | What it checks | Weak signal | Strong signal |
|---|---|---|---|
| Repeatability across runs | Same target, same scope, run twice: do the findings match? | Finding count or severity shifts materially between runs, with no explanation offered | Findings converge run over run, and any drift is explained (target changed, a patch landed) |
| Validation effort left with your team | Once a finding is flagged, how much manual confirmation does your team still have to do? | Every finding needs your own reproduction before anyone trusts it, so the workload has moved, not shrunk | Findings arrive with evidence, request and response, or reproduction steps, your team can check quickly rather than remake from scratch |
| Human escalation path | What happens when the tool is uncertain, or hits something destructive? | The tool proceeds on ambiguous or high-risk actions with no defined stop point | A named person or role is the escalation point, uncertain and high-risk actions pause for them, and that pause is logged |
| Proof findings reach remediation | Do flagged findings turn into tickets that get closed, or do they sit in a report? | Findings live in a PDF or dashboard with no link to a ticket or an owner | Each finding maps to a tracked ticket, an owner and a close date, and that mapping can be pulled back out later |
A worked example
Run the same external-facing test twice, two weeks apart, nothing else changed on your side. Fifteen findings the first run, eleven the second, four gone with no explanation. That is a repeatability fail: score it weak, and ask the vendor why before you look at anything else. Of the eleven that held, three needed a day of manual reproduction before your team trusted them, which means validation effort is not falling, it is moving from the tool to your analysts. One finding touched a production system directly: check whether a human paused it before it ran, and whether that pause is on record anywhere you can retrieve later. Two months on, pull the ticket tracker: if none of the eleven map to a closed ticket, the tool found things, it did not fix anything. Four measures, one afternoon of checking, and you know more about a vendor than any coverage claim in their deck tells you.
SenseOn does not sell, resell or benchmark AI pentesting. Automated exploitation tooling and the AI that governs SOC detection and response decisions are different categories of product, and this article is not a way in to a pitch for either. The scorecard above is offered as a starting method: run it against any vendor you are evaluating, including ones that have nothing to do with SenseOn.
Related reading: