
Most agentic pentesting demos run the same 40 minutes. Clean workspace, a target the vendor knows intimately, an exploit chain that lands first try, a report with no false positives, then a slide about how the category is changing.
None of that is dishonest. It’s a rehearsal, and rehearsals go well.
The bill arrives in month three. The agent can’t reach the internal apps holding your highest-risk data. Findings land in Jira as a title and a severity, so your developers open tickets asking what to actually fix. Legal discovers screenshots of production data sat in a third-party region for 90 days. The contract is signed and this year’s pentest budget is already committed.
Ten days of structured testing catches all of it. This is the day-by-day structure, what to verify yourself rather than take on trust, and a weighted scorecard for comparing two vendors on the same target.

A demo can’t show you the three things that decide whether a platform works for you: failure modes, coverage gaps, and how findings behave inside your workflow.
Failure modes get edited out by definition. You watch the happy path, not the agent hitting a WAF, losing its session to a custom SSO redirect, or timing out halfway through a chain. Those are weekly conditions in a real environment.
Coverage gaps are the more expensive blind spot. A demo report shows what the agent found. It rarely shows what it skipped, which endpoints it couldn’t authenticate into, or which routes it abandoned after a 429. Vendors optimize the findings list. The exclusions list tells you how much of your attack surface actually got tested.
Then workflow. Findings look excellent inside the vendor’s interface, where evidence, request history, and remediation guidance sit one click apart. What matters is the ticket in your board, read by a developer who has never opened the platform and has 40 minutes before standup.
Strobes runs demos too. Every question in this post applies to us.
If you haven’t built a shortlist yet, start with the eight technical criteria in the agentic pentesting complete guide and the current platform capability breakdown. This post picks up where the shortlist ends.
Work in this order: objective, success criteria, scope, prerequisites. Each one derives from the one before it. Most POCs go wrong because someone picks a convenient target first and reverse-engineers the criteria to fit it, which produces a test nobody can fail and a result nobody can act on.
Time estimate: 10 working days of testing, inside an elapsed calendar of three to six weeks for most organizations. Those are different numbers and conflating them is how POCs get abandoned halfway.
The ten days are yours to control. The calendar around them is not. A read-only cloud audit role commonly takes one to two weeks to provision. A data processing agreement rarely clears legal in under two. Internal network connectivity waits on a change window. None of that is testing time, and none of it compresses because you asked nicely. Start the provisioning tickets and the legal questionnaire the day you decide to run a POC, then schedule the ten days once the slowest of them has a date.
Write one sentence describing the decision this POC lets you make. Real examples:
Decide whether we can replace two of our four annual application pentests.
Decide whether we can get continuous validation across 40 AWS accounts without adding headcount.
Decide whether we can test every release candidate instead of every quarter.
Decide whether we can retire the in-house tooling we have been maintaining ourselves.
The objective determines everything downstream. A POC built to answer the first question and one built to answer the second share a scorecard and very little else.
Success criteria are what make you satisfied on day 10, and they belong to your requirement rather than to the category. Agree them with whoever owns the budget before any vendor sees anything.
If the objective is replacing manual application pentests, the criteria look like finding parity against your last manual engagement plus coverage of the specific finding classes your testers flagged. If the objective is continuous cloud coverage, finding parity is irrelevant, and the criteria become IAM privilege escalation path discovery, CIS Benchmark drift detection, and time from account onboarding to first validated result. Same platform, same scorecard, completely different definition of a pass.
POC SUCCESS CRITERIA (agreed ____/____/____ , frozen for the duration) Objective .................. ____________________________________ Surface .................... application / API / cloud / internal / code / ASM Engagement model ........... standard / red team / assumed breach / path validation Max time to first validated finding ....... ___ hours Ticket quality threshold .................. ___/3, read cold by a developer Disqualifying condition ................... any week 1 criterion scoring 0
Each criterion needs three things written down: what it is, how you will measure it, and what result counts as a pass. A criterion without a measurement method is a preference. This is the shape that survives a disagreement on day 9:
| # | Success Criterion | Measurement Method | Target |
|---|---|---|---|
| 1 | Finding quality on our own code | Review findings from two repositories against our last manual review | Real vulnerabilities with reproduction steps, and no unexplained variance between two runs |
| 2 | Validation and prioritization | Sample 20 findings and check each one for evidence that a payload executed | Every finding carries execution evidence; false positive rate stated with its methodology |
| 3 | Attack path generation | Ask for the chain from an external asset to a named internal system | A path we can follow and verify ourselves, or a clear statement that none was found |
| 4 | Autonomous network testing | Run against the agreed internal segments with the connector deployed | Assessment completes, exclusions named, findings a developer can act on without translation |
| 5 | Bring-your-own model | Configure our own provider credentials and run a full engagement on them | Engagement completes on our keys, with token usage visible to us |
| 6 | Workflow and reporting | Follow three findings from discovery to a ticket a developer can act on | Tickets score 2 or better on our rubric; reports usable without a human translator |
Yours will not look like this, because criteria come from your objective and that table came from someone else’s. What transfers is the shape: numbered, measurable, with a target agreed before anyone connects.
One more thing about who you are actually comparing against. Most buyers frame this as vendor A versus vendor B, but the decision your finance team is approving is usually displacement: retiring a scanner, dropping two of four annual pentests, or shelving in-house tooling. If that’s true for you, one of your criteria should be parity with the incumbent, measured on the same target. A platform that beats the other vendor and loses to your existing annual pentest has still failed.
Note where row 5 sits. If your organization already holds its own model provider credentials, bring-your-own-model stops being a week 2 governance question and becomes a week 1 criterion, because the entire engagement runs on your keys and your bill.
Then freeze it. Every vendor will produce results that can be argued as a pass, and a frozen criterion is the only thing that settles the argument on day 9.
Scope serves the criteria. Reverse that and you’ll scope whatever was easy to provision, then find out on day 8 that the easy thing couldn’t have proven your criterion either way.
| If your success criterion is | Scope has to include |
|---|---|
| Finding parity with last year’s manual application pentest | The same application, authenticated, with the same role set your manual testers had |
| IAM privilege escalation path discovery | A cloud account with a read-only audit role attached, and written authorization to enumerate assume-role chains |
| Coverage of internal apps that never appear in external scans | A connectivity method provisioned and tested before day 1, not promised for day 3 |
| Detecting vulnerable dependencies that are actually reachable | Read access to the repository and dependency manifests, not just the running application |
| Time from asset discovery to first validated finding | Seed domains, plus a decision on whether newly discovered assets are in scope by default |
Then write scope as a file both vendors receive, so you can diff their behavior against it later. The shape varies by engagement type.
The file changes. The habit of writing it down doesn’t.
Every POC needs this core, regardless of what you’re testing:
☐ Written objective and frozen success criteria
☐ Signed authorization from security, legal, and infrastructure for autonomous exploitation
☐ Two vendors on the same scope, in the same window
☐ A named support channel, the vendor contacts staffing it with their roles, an agreed response time in writing, and a backup contact by severity level
☐ A closure call on the calendar before day 1, with the frozen criteria as its agenda
Both of those last two get agreed at kickoff or not at all. In a compressed window a day lost waiting on a blocked credential costs you a tenth of the POC, so you want a named channel, a stated response time, and somebody to escalate to when the primary contact is on leave. Book the closure call in the same meeting and put the frozen criteria on the agenda, along with one question for the end: what carries into the paid engagement. Scope file, credential configuration, integration field mapping. If none of it survives, you’re paying to redo onboarding you already did. Mutual sign-off against the frozen criteria is what should gate the move to commercial discussion, so treat that review as the decision point rather than a formality. A POC that closes with the vendor presenting a summary deck instead of walking your criteria line by line has closed in the vendor’s favor by default.
One shortcut worth taking: ask each vendor for their pre-kickoff intake questionnaire. A good one tells you which prerequisites you forgot.
That checklist is your half. A POC has two sides, and the half nobody documents is what the vendor owes you and by when. Write both columns before day 1. Yours: credentials, repository or network scope, connector or runner support, nominated users, and timely feedback. Theirs: a provisioned workspace, integration configuration, deployment support, technical enablement, progress tracked against the frozen criteria, and a facilitated final review. A vendor who will not commit to their column in writing is showing you how the paid engagement will run.
Two items on that list need expanding. Authorization for an autonomous agent executing exploit payloads against live systems takes more than security sign-off, so if your organization requires formal governance review before an AI agent touches production, work through the seven authorization checks for agentic pentesting before the POC rather than during it. And run the two vendors in parallel rather than back to back, because sequential POCs take a quarter and your read on the first one fades before the second finishes.
Then add whatever your engagement type actually requires. The first set is scoped by surface.
| Surface | Add to the core |
|---|---|
| Web application | Credentials for two or more roles, a decision on SSO and MFA handling, and a target on your real auth stack, not a stripped staging build |
| API | OpenAPI or Postman collection, token issuance method, and the rate limit ceiling you’ll allow |
| Cloud (AWS, GCP, Azure) | Read-only audit or assume-role, in-scope account IDs, region list, CIS Benchmark version, and written authorization for privilege escalation enumeration |
| Internal network or Active Directory | Connectivity provisioned and tested, subnet ranges, a domain user credential, Kerberoasting authorization, and a change window |
| Source code | Repository read access, target branch, and whether commit history is in scope for secret detection |
| External attack surface | Seed domains, authorized IP ranges, and the default in-or-out rule for assets the agent finds itself |
The second set is scoped by engagement model. These run across several rows above, so they stack on top of the surface prerequisites rather than replacing them.
| Engagement Model | Add to the Core |
|---|---|
| Red team | An objective defining what winning means (domain admin, a named data store, a specific transaction), rules of engagement, whether detection evasion is authorized, whether the SOC is informed, and a white-cell contact reachable out of hours for deconfliction |
| Assumed breach | The foothold specified precisely: which host, which credential, what privilege level, who provisions it. Plus the lateral movement boundary, target end state, and whether persistence is permitted (usually not, in a POC) |
| Attack path validation | An asset inventory with criticality ratings, since “path to a critical asset” is undefined without one. Plus named end-state assets, your control inventory (segmentation, EDR, WAF) so you can tell a blocked path from a missed one, and simultaneous authorization across every surface the path may cross |
Attack path validation carries the heaviest day-0 load in this post. A path running from external asset to web application to internal network to cloud IAM needs all four scopes authorized before day 1, so the provisioning burden is the union of four rows rather than one.
Red team POCs need a different scorecard emphasis. Findings aren’t the criterion. Whether the agent reached the objective is, and so is whether your SOC saw it happen, which puts detection metrics from your own logging into the success criteria. That review runs longer than a findings review, so either extend the window or narrow the objective.
Provisioning is where day 0 becomes day 3. Cloud roles and internal connectivity almost never land same-day, so raise those tickets before you schedule the vendors.
Week 1 answers whether the output is real and whether the coverage is honest. It’s the gate. A platform that fails days 1 to 5 doesn’t earn days 6 to 10.
Run the platform against your target. Then stop reading the findings list as a list and start treating three of the findings as claims you’re going to independently verify.
Reproduce a finding by hand. This is the highest-value hour of the entire POC, and it takes about four minutes per finding. Take the three highest-severity findings, pull the raw request out of each evidence package, and replay it yourself with curl or Postman using the low-privilege session the agent used.
You’re checking one thing: does the response you get back match the response the agent recorded? If a viewer-role session really does return a success on an approval endpoint, the finding is verified and you never have to take the vendor’s word for anything again. If you get a 403 where the agent claimed a 200, stop and ask why before reading another finding. Ask them to sit with you while you do this. How a vendor behaves during independent verification of their own output is informative on its own.
Then run the full engagement a second time against the same target and compare the two finding lists. Identical is the strong result. Differences aren’t automatically disqualifying, because rate limits and target state changes cause legitimate variance. What matters is whether the vendor can account for every difference from their own audit log. A vendor who can’t explain variance in their own output has an observability problem you inherit on signature.
Interrogate the false positive number. Any vendor claims zero. Three questions turn the claim into something you can score:
| Ask | A weak answer sounds like | A strong answer sounds like |
|---|---|---|
| Measured against what target? | “Our internal benchmark suite” | A named application, version pinned |
| Across how many findings? | “Consistently zero” | A confirmed count over a stated total, across a named number of runs |
| Did your team filter before we saw it? | Hesitation, or “our analysts review output” | “No. Here’s the unfiltered run” |
Check the novel-versus-CVE split. Ask for the breakdown. If every finding maps to a public CVE, you’re evaluating a scanner with better reasoning about which signatures to run. Agentic behavior shows up in the findings no signature would produce: chained Low-severity issues that combine into privilege escalation, access control gaps specific to your role model, logic flaws that exist only in your codebase.
Give it a business logic problem. This is the criterion most buyers skip because they don’t know how to construct the test. Three that are quick to build on almost any application:
Cross-tenant read. Create two accounts in different organizations. Note an object identifier from the first, then request it while authenticated as the second. Does the agent find that class of issue on its own, without you pointing at the endpoint?
Out-of-order workflow. Take any multi-step process with a gate in the middle, a payment step, an approval, an eligibility check. Complete step one, then call the step-three endpoint directly. A logic flaw here never appears in a dependency scan.
Residual access after a role change. Grant a user an elevated role, have them start a session, then downgrade the role without ending the session. Check what they can still do. This is a real finding class and no signature describes it.
For a cloud engagement the equivalent isn’t a checkout flow, it’s an assume-role chain that grants more than the policy author intended. Two outcomes are acceptable either way. The agent finds something, or it reports plainly that it couldn’t reason about the path. Silence is the failure state, and silence is common, because a platform that says nothing looks identical to a platform that found nothing.
Day 3 gate. Couldn’t reproduce a finding by hand? Couldn’t get the false positive methodology in writing? Stop. Workflow testing a platform whose output you can’t verify is a waste of week 2.
Short block, sharp question. The findings list is half the report. The exclusions list is the half that tells you what you actually bought.
Ask for the full list of every host, IP, or account the agent sent traffic to during the engagement, then compare it against the scope file you wrote on day 0. Two comparisons matter, and they answer different questions.
Anything in scope that was never touched is a coverage gap, and the report should already have named it. Anything touched that was never in scope is a scope enforcement failure, which is the more serious of the two, because it means scope lives in the agent’s prompt rather than in the platform. Prompts can be reasoned around.
If the vendor can’t produce that list at all, you’ve learned something more important than either comparison.
Then read the coverage section of the report for four things: auth-gated routes it couldn’t access, endpoints it backed off from after rate limiting, JavaScript it couldn’t resolve into routes, and dynamic content it couldn’t parse. All four should appear without you asking. A report with no exclusions section is either testing everything, which is unlikely, or not tracking what it missed.
Last, mark a finding resolved and watch what happens. Automatic retest is table stakes. Persistent memory across runs is the differentiator, because continuous testing without it hands you the same report every month with no sense of whether you’re improving.
Week 2 asks whether the platform survives contact with your organization.
Measure time to first validated finding on your target. Wall clock from engagement start to the first finding you’d actually escalate. Take both timestamps from the audit trail rather than the dashboard, since dashboards tend to show when a finding was surfaced rather than when the agent confirmed it.
This number predicts adoption better than any feature comparison, because it decides whether your team runs the platform on every significant deployment or quietly falls back to quarterly.
Ask for a failure, not a success. Put a WAF rule in front of the target. Rate-limit an endpoint to 5 requests per minute. Point it at the custom auth flow you know is unusual. What you want is loud failure: the reason logged, an alternative approach attempted, the gap surfaced in the report.
A run that completes with no findings and no explanation is worse than a crash, because it looks like good news. Ask the vendor to demonstrate this rather than describe it. The ones who’ve done it before will have a story ready. The ones who haven’t will offer to schedule a follow-up call.
Score ticket quality with an actual developer. Not “does it integrate with Jira,” which they all do. Borrow any developer on your team for fifteen minutes and have them read three tickets cold, with no context from you and no access to the vendor’s interface:
| Score | Ticket Content Criteria |
|---|---|
| 0 | Title and severity only |
| 1 | Description provided, but evidence requires opening the platform. |
| 2 | Reproduction steps and affected endpoint; fix guidance is missing or generic. |
| 3 | Reproduction steps, affected endpoint or code path, specific fix guidance, and severity with reasoning. |
Average under 2 across three tickets means every finding will route through a human translator before it reaches a developer. This is also why the vendor writes into your ticketing system during the POC rather than after it. Make them configure the evidence field mapping in week 1 so week 2 scores what you would actually receive, not a demo ticket. Check the integration depth against the tools your team already lives in.
Test the human handoff with your analyst, not their engineer. Every serious platform pauses on high-impact actions or hands off when authentication needs judgment. Hand that pause to a mid-level analyst on your team. If completing it requires a call with the vendor’s services team, you’ve found a cost line that isn’t in the quote. Human-in-the-loop validation covers why the gate exists before you evaluate how well it works.
Run three targets at once. You’re checking for graceful degradation and honest queueing. Per-target time that triples under load, or queue behavior nobody mentioned in the sales cycle, will shape your testing cadence more than any published benchmark will.
This block is mostly paperwork, which is why buyers skip it until two weeks before signature. Send these questions on day 0 and use days 9 and 10 to test the parts that are testable.
Get written answers, not call assurances, on:
Where prompts, findings, screenshots, and target data physically reside, and under whose control
Retention period for each of those, and the deletion process on contract end
Whether your data trains models, and whether that’s contractual or policy
Which models power the agent today, what notice you get on a swap, and whether bring-your-own-model is supported
Sub-processors with access to engagement data
What happens to your findings if you don’t sign. You’ve just handed a vendor a working list of your exploitable weaknesses, and general retention terms rarely address the walk-away case or whether the list survives in sales notes
Ask for the pricing model in the same batch, and ask it as a scoping question rather than a commercial one. You need to know whether the unit you are being priced on resembles the unit you tested. If pricing is per target and your POC ran against one, you have measured nothing about the cost of forty. If it is per asset, ask how assets are counted when the agent discovers new ones itself. Then ask specifically what sits outside the license: connector or runner deployment, services time when a handoff needs their engineer, charges for retesting after a fix, and, if you are bringing your own model credentials, the token spend that lands on your provider bill rather than theirs. Those four are where the surprises live, and none of them appear on a pricing page.
Then test three things directly.
1. Export the full audit log and confirm it contains every request and payload rather than a findings summary, because a summary can’t answer the only question an auditor asks: what did the agent do inside our environment?
2. Pull the kill switch on a running assessment and time it, watching what state the target is left in and whether partial results survive. A platform that discards everything on cancellation is punishing you for exercising control.
3. Reconfigure approval thresholds mid-engagement. Approval on every action makes the platform unusable, approval on nothing makes it unsafe, and the answer you want is granular per action class with the agent continuing other work while a request sits pending.
Diff the two vendors before you score either of them. Then score all 19 criteria 0 to 3, weight week 1 heaviest, and compare on totals rather than highlights. A platform that wins on integrations and loses on evidence quality has lost.
You are about to collect a false positive rate, a time to first finding, and a novel-versus-CVE split, and then discover you have no idea whether any of them is good. There is no published industry baseline for agentic pentesting, the vendor-published numbers are measured on targets they chose, and a rate that is excellent on a stock application may be poor on yours. Absolute thresholds are not available to you and anyone offering one is guessing.
Judge relatively instead. Three reference points, all measured on the same target:
| Reference | What it tells you |
|---|---|
| Vendor A | One data point, uninterpretable alone |
| Vendor B | Turns every metric into a comparison instead of a number |
| Your last manual engagement on that target | The only baseline that reflects your environment and your risk appetite |
Then read direction rather than magnitude. More findings is not better if the extra ones are unvalidated. A lower false positive rate is not better if it was achieved by suppressing anything uncertain, so always pair it with total findings. Faster time to first finding is better only if the finding is one you would escalate. And a higher novel-versus-CVE ratio is better up to the point where you cannot verify the novel ones, at which point it is just unfalsifiable output.
Write your own thresholds into the criteria table on day 0 based on what your last engagement produced. A number you set before seeing any vendor results is worth more than an industry average you found afterwards.
Two reports of 20 findings and 12 findings tell you almost nothing on their own. What you need is the overlap, because the findings only one platform caught are the entire reason you ran two.
No two vendors share a report schema, so this is manual work on a spreadsheet, and it’s worth doing carefully rather than quickly. Export both finding lists, then match them on affected location plus weakness class plus HTTP method. Three things will trip you up:
Path templating. One vendor writes /api/v2/orders/{id}/approve, the other writes /api/v2/orders/1042/approve. These are the same finding. Normalize identifiers out of paths before you compare anything, or the diff will report divergence that doesn’t exist.
CWE disagreement. An access control issue lands as CWE-284 for one platform and CWE-863 for the other. Group by weakness family rather than exact ID, and read the descriptions when two entries look adjacent.
Chained versus atomic reporting. One vendor reports a three-step privilege escalation as one Critical finding. The other reports three Mediums. Same discovery, incomparable rows. Match on the end state achieved, not the count.
Do the matching by hand for anything High or Critical. Automating it is where a diff starts producing confident wrong answers, and a false “vendor B missed this” is exactly the mistake this exercise exists to avoid. Then read the three sets rather than the two totals:
| Set | What it tells you |
|---|---|
| Found by both | Your baseline. Neither vendor gets credit for these |
| Only vendor A | A's coverage advantage, and B's blind spot |
| Only vendor B | The same in reverse |

If procurement will only approve one vendor, which is common, you still need a second set to diff against. Two workable substitutes. Use your most recent manual pentest report on the same target as the comparison set, which gives you a human baseline rather than a peer one and answers the displacement question directly. Or hold back a scope segment you have already tested by other means, let the agent run against it blind, and compare against what you already know is there. Neither is as strong as two vendors on one target, and both are far better than reading one report on its own.
The rule that follows: one Critical sitting in the “only B” column outranks A’s larger total. A missed Critical is the failure mode you bought the platform to prevent, and a bigger findings count doesn’t compensate for it. Score coverage transparency against these three sets rather than against each vendor’s own exclusions list, because now you have an independent measure of what each one couldn’t see.
| Block | Days | Criteria | Weight | Why |
|---|---|---|---|---|
| Finding quality | 1 to 3 | 5 | 40% | Untrustworthy output makes everything downstream irrelevant |
| Coverage gaps | 4 to 5 | 4 | 25% | Determines how much of your surface was actually tested |
| Operational fit | 6 to 8 | 5 | 25% | Predicts weekly use versus quarterly abandonment |
| Trust and control | 9 to 10 | 5 | 10% | Low weight, hard blocker when it fails |
Define the anchors so two reviewers scoring the same vendor land in the same place:
0 = Absent. Vendor can't demonstrate it. 1 = Claimed. Marketing says yes, POC evidence is thin. 2 = Demonstrated. Works, with caveats you'd accept. 3 = Proven. Verified on your target, documented, repeatable.

Any 0 on a week 1 criterion is disqualifying regardless of total.
When the totals come out close, and they often do, treat that as the scorecard failing to discriminate rather than as a tie you have to break arbitrarily. Go in this order. Compare the week 1 subscores only, since finding quality is the thing you are buying. If those are also level, go back to the three diff sets and ask which vendor’s blind spot you can live with. If it is still level, break the tie on ticket quality and human handoff, because those two are what your team touches every week, and a platform your team avoids using scores zero in practice regardless of what the sheet says. Never break a tie on total findings count.
The 19 criteria stay constant across engagement types. The test cases underneath them don’t. “Business logic” means a payment workflow on an application POC and an IAM escalation path on a cloud POC, and you score the same criterion against whichever one your scope actually contains.
These weights are ours, and you should change them if your objective differs. Raise trust and control if you are regulated. Raise operational fit if adoption is your real risk rather than detection.
Three failures account for most ambiguous outcomes.
Cause: The agent never authenticated. Custom SSO, MFA, or a non-standard session mechanism blocked it, and the run finished against unauthenticated surface only.
Fix: Read the coverage report before the findings list. Confirm the agent reached an authenticated session, supply credentials through the vendor’s vault, re-run. And note the second-order finding here: if the platform completed a run without telling you it failed to authenticate, that’s a scoring event under coverage transparency, not just a configuration problem.
Cause: The integration is pushing a summary rather than the evidence package, usually because nobody configured field mapping during the POC.
Fix: Have the vendor configure it during the POC. If full evidence can’t reach your ticket, remediation routes through a human translator permanently.
Cause: Data residency, retention, or model-training terms surfaced on day 9 instead of day 0.
Fix: Send the data-handling questionnaire on day 0 and run it alongside testing. This is the most common reason a two-week POC becomes a ten-week procurement cycle, and it has nothing to do with the technology.
You finish day 10 able to say whether each frozen success criterion was met, with evidence rather than impressions behind every answer:
A false positive rate measured on your stack, with the methodology written down
A named list of what the agent couldn’t reach, and why
At least one High or Critical finding you replayed by hand and confirmed
Three tickets read cold and scored by a developer on your team
Written answers on residency, retention, model transparency, and audit export
A time-to-first-validated-finding number per vendor
A total score per vendor, week 1 weighted heaviest
The last item matters more than which platform you pick. Someone will eventually ask how you chose. “The demo was impressive” doesn’t survive the follow-up question. “Ten-day controlled evaluation, two vendors, same production-mirror target, 19 criteria scored, here’s the sheet” does.
Running two vendors in parallel used to be impossible. Two manual engagements on one target inside a fortnight doesn’t work on cost or scheduling. When assessments finish in hours, a comparative POC costs attention instead of budget, and that shift in pentest economics is what makes any of this practical.
Ten working days of active testing is enough for most environments, sitting inside an elapsed calendar of three to six weeks. Agentic platforms finish assessments in hours, so agent runtime is never the constraint. Provisioning and legal review are. Budget the ten days for testing and treat everything before them as lead time you start immediately. If a vendor asks for 60 days of testing, ask which specific test needs that long.
Yes, with platform-level scope enforcement, human approval gates on exploit execution, encrypted credential handling with automatic revocation, and a complete audit trail. Production or a production-mirror is the better target, because a sandbox tells you almost nothing about how the agent handles your real auth stack and business logic.
As a warm-up sanity check only, and score nothing on it. Every platform in the category finds documented bugs in documented applications. Your custom auth flows, role hierarchy, and business logic are what separate a shortlist, and none of that exists in a training target.
Only PCI DSS actually mandates penetration testing. Requirement 11.4 in v4.0 calls for it at least annually and after significant changes to the cardholder data environment, with more frequent segmentation validation for service providers. SOC 2, ISO 27001, and HIPAA do not name penetration testing as a requirement. SOC 2 auditors commonly accept a well-scoped test as evidence for the logical access and monitoring criteria, ISO 27001:2022 points at it under Annex A 8.8 on technical vulnerability management, and HIPAA’s risk analysis obligation is often satisfied in part the same way.
So the accurate answer is that an agentic pentest can serve as the evidence in all four cases, and is mandatory in one. Test it during the POC by asking for a compliance-formatted report generated from your own test data, with findings mapped to the control language your auditor uses, rather than a sample from the vendor’s library.
Agentic pentesting: the complete guide: the eight technical criteria for building your shortlist
What to evaluate before you let an AI agent exploit your systems: the governance checks to clear before day 0
Agentic pentesting with Strobes AI: the full methodology running against a live web application
Best AI pentesting tools in 2026: current platform coverage and capability breakdown
Adversarial exposure validation: where continuous validation fits in a CTEM program
Run the POC against a live target in your environment
Scope a two-week agentic pentesting POC and grade it with the scorecard above. Or start with the agentic pentesting solution overview.