Pwnkemon Benchmarks: Methodology & Results
Last updated:
This page documents how we benchmark Pwnkemon, the self-serve agentic AI pentesting platform, against reproducible targets — and publishes the results once the run is complete. We publish the method before the numbers on purpose: a benchmark is only credible if you can reproduce it.
Results
Results are not published yet. The benchmark harness and methodology below are final; the numbers will appear here once the run is complete. We won't fill this table with estimated or illustrative figures.
Targets
Deliberately vulnerable, publicly reproducible environments:
- PortSwigger Web Security Academy labs (web / API vulnerability classes).
- OWASP Juice Shop (modern single-page app with a broad vulnerability set).
- A deliberately vulnerable Active Directory lab (network / AD paths).
Metrics
- Found. Distinct issues the agent surfaced.
- Validated. Findings with a confirmed, reproducible proof.
- False positives. Reported issues that were not real.
- Time (min). Wall-clock time to a finished report.
- Cost (USD). Total metered cost of the run.
Versions
Each published result records the Pwnkemon platform version, the scan tier used, and the target versions/commit hashes, so a run can be tied to an exact software state.
How to reproduce
- Stand up each target at the pinned version listed with the result.
- Verify ownership of the target in Pwnkemon (DNS TXT or HTTP-file challenge).
- Run the stated scan tier against each target.
- Score findings as found / validated / false-positive against the target's known-issue set.
- Record wall-clock time and metered cost from the scan record.
Methodology notes
Validation counts only findings with a reproducible proof, not raw alerts. False positives are scored against each target's documented vulnerability set. This is our own benchmark of our own product, so we publish the method, the targets and the raw per-target numbers to let you check our work rather than take a single headline figure on trust.
Related: AI pentesting tools compared · Benchmarks · Pricing · GitHub Action