Autonomous Pentesting Benchmark Report 2026
We ran an autonomous pentest on a public app, then measured it against the field. One public target, third-party reference data, every Strobes figure backed by 31,400 logged telemetry events.
Strobes AI · Benchmark 2026
Autonomous Pentesting
Benchmark
One public target. Third-party reference data. Full Strobes run telemetry.
Results at a Glance
Attribution and sources
Strobes results were produced by a Strobes-run assessment of Fider v0.33.0 and are reported by Strobes, backed by recorded run telemetry.
Aikido and XBOW figures come from Doyensec’s independent assessment of the same public target (Aikido-sponsored) and are used here as external reference data. Doyensec did not assess the Strobes run.
Fider v0.33.0 is publicly available. Other teams can assess the same version. Results may vary based on configuration, scope, tooling, models, prompts, execution controls, and methodology.
The results, backed by recorded run telemetry

One public target. Recorded run telemetry.
Over June 10 to 11, 2026, Strobes AI ran a fully autonomous assessment of Fider v0.33.0, a production-grade open-source feedback platform with real authentication, file uploads, webhooks, OAuth, and an admin console. The security firm Doyensec assessed the same application in an Aikido-sponsored engagement and published results for two commercial AI security platforms. Because the target is public and versioned, anyone can stand up the same instance and check the work.
The platforms were assessed against the same public application target. Testing configuration, execution conditions, and methodology may differ.
189s to verified admin takeover
A confirmed multi-step attack chain, executed end to end with zero human intervention. The combined six-scanner field confirmed zero exploitable findings on the same target.
OTP endpoint, no rate limit
Recon surfaces an unprotected one-time-password endpoint, the entry into the admin account.
Code brute-forced in 189s
The admin session is captured in 189 seconds against the unprotected code.
Session verified and replayed
The captured genuine admin session is confirmed and replayed for continuous access.
Pivot to webhooks
Continuous admin access pivots into the outbound webhook functionality.
SSRF to cloud metadata
Blind SSRF reaches AWS IMDS. The outbound webhook hits the cloud metadata service, exposing instance credentials.
Measured against six scanners and two AI pentesting platforms

70 to 75% lower cost per validated finding
Figures are per validated finding. Scanner pricing reflects public list rates; the Strobes figure is measured AI-credit consumption on the same target. The cost bases differ, so read the comparison as directional even though the gap is large.

$235 per finding
Scanner A: 13 validated findings at $4,000 per scan.

$167 per finding
XBOW: 24 validated findings at $4,000 per scan.

$22 to $27 per finding
Strobes AI: 45 findings on about $1.1k of credits, 70 to 75% lower cost per validated finding.
Why the gap is structural, not tuning
A signature scanner emits fixed candidates. Strobes AI authenticates, holds session state, chains weaknesses into a shared exploit, reasons about business logic, and routes each result to a confirm-or-discard step, so every candidate runs through a separate validation agent that must confirm it before it is recorded. The credit is the mechanism behind the zero false positives.
From quarterly pentest to continuous validation
If autonomous assessment costs what this benchmark shows, your release cadence becomes the only limit. Talk to us about running this against your own target across network, API, cloud, or source-level review in one workspace.
Join 150+ security teams already reducing exposure with Strobes



