Autonomous AI penetration testing benchmarks
Autonomous AI penetration testing benchmarks with real results only. Darkmoon, the open source AI pentest tool, was run against public vulnerable labs across cloud, infrastructure, IoT and web, and every finding here carries a working proof in a published report.
Darkmoon, the open source autonomous AI penetration testing tool, found 318 vulnerabilities across 17 runs on public labs, and proved 96 of them with a real exploit.
S.03By surfacePick an attack surface
Each surface has its own keyword page with the concrete results for its labs, the exploit chains, and links to the full write-ups.
Cloud
AWS, Azure and GCP identity, storage and metadata chains.
Infrastructure & CI/CD
Jenkins, Redis, PostgreSQL, Terraform, Vault, GitLab and the Docker socket.
IoT & firmware
OWASP IoTGoat, static firmware and a live appliance.
Web application
OWASP Juice Shop, black-box, six-campaign escalation.
S.04HighlightsThe six biggest runs
Sorted by each report's own findings total. Every tile opens the long-form write-up; the full leaderboard is just below.
Six weekly campaigns escalating from 14 to 57 findings on a black-box OWASP Juice Shop, local LLM, proof per finding.
Guessed root token, image-layer secrets, and a Docker-socket container escape reading the host /etc/shadow.
An exposed terraform.tfstate and Ansible inventory chained to an AdministratorAccess CI key.
A deleted blob recovered, ROPC minted a token with no MFA, and the directory fully enumerated.
An admin PAT with api and sudo scopes, an unmasked AWS secret in CI/CD variables, and open signup.
PostgreSQL COPY TO PROGRAM RCE, pg_shadow and mysql.user hashes, and live session tokens.
S.05LeaderboardEvery run, every finding
Scroll horizontally on narrow screens. Each row links to its long-form write-up and to the raw report it is drawn from.
| Lab / target | Surface | Findings | Severity | Exploited | Model | Evidence |
|---|---|---|---|---|---|---|
| AWS · huge-logistics S3camp_20260802_03bfc675 | Cloud | 9 | 5C1H2M1L | 5 | claude-opus-4-6 | Write-upReport |
| AWS · pwnedlabs EBS/S3camp_20260802_611cece1 | Cloud | 9 | 2H5M2L | 0 | claude-opus-4-6 | Write-upReport |
| Azure · Entra ID tenantcamp_20260802_7eee391f | Cloud | 28 | 11C10H4M3L | 12 | claude-opus-4-6 | Write-upReport |
| Azure · BloodHound / priv-esccamp_20260802_9d245c0c | Cloud | 19 | 8C6H5M | 11 | claude-opus-4-6 | Write-upReport |
| Azure · Key Vault (extract)camp_20260802_38118fb8 | Cloud | 16 | 2C6H8M | 6 | claude-opus-4-6 | Write-upReport |
| Azure · Key Vault (pivot)camp_20260802_59e4e905 | Cloud | 7 | 3C2H2M | 4 | claude-opus-4-6 | Write-upReport |
| GCP · SSRF to metadatacamp_20260802_656007d3 | Cloud | 4 | 3C1H | 3 | claude-opus-4-6 | Write-upReport |
| GCP · public GCS bucketcamp_20260802_3ced7196 | Cloud | 5 | 3C1H1M | 4 | claude-opus-4-6 | Write-upReport |
| Vault + registry + Docker socketcamp_20260801_2bd90d3f | Infrastructure & CI/CD | 41 | 15C15H9M2L | 8 | claude-opus-4-6 | Write-upReport |
| Terraform + AWS + Ansiblecamp_20260801_b1b96939 | Infrastructure & CI/CD | 34 | 21C5H8M | 16 | claude-opus-4-6 | Write-upReport |
| GitLab CE 19.2.1camp_20260801_a719641d | Infrastructure & CI/CD | 26 | 4C8H10M2L2I | 2 | claude-opus-4-6 | Write-upReport |
| PostgreSQL 16 + MySQL 5.6camp_20260801_c0151524 | Infrastructure & CI/CD | 22 | 6C10H6M | 13 | claude-opus-4-6 | Write-upReport |
| Redis 7.4.10 (unauth)camp_20260801_96be38b9 | Infrastructure & CI/CD | 9 | 3C5H1M | 5 | claude-opus-4-6 | Write-upReport |
| Jenkins 2.541.3 (security off)camp_20260801_b6ad197d | Infrastructure & CI/CD | 3 | 3C | 3 | claude-opus-4-6 | Write-upReport |
| IoTGoat firmware imagecamp_20260802_dd187131 | IoT & firmware | 20 | 4C7H7M2L | 1 | claude-opus-4-6 | Write-upReport |
| IoTGoat live devicecamp_20260802_257f75f1 | IoT & firmware | 9 | 4C3H2M | 3 | claude-opus-4-6 | Write-upReport |
| OWASP Juice Shopcamp_20260426_3d2f | Web application | 57 | 8C24H21M4L | proof per finding | Local (Ollama / llama.cpp) | Write-upReport |
DisclaimerThis is Darkmoon's own benchmark on public training labs.
It is not a third-party certification or an analyst endorsement. The offensive runs are produced by the open source Darkmoon CLI; the web dashboard and the remediation-to-PR loop are paid Pro capabilities. Every number is drawn from a published report and can be reproduced or contested. Severity is shown as C / H / M / L / I. Counts are each report's own total, so a report that lists an aggregated finding twice counts it twice. The raw reports live in the darkmoon-research corpus and the Darkmoon-Benchmarks repository.
S.07FAQQuestions buyers ask
Short, factual answers about the method, the numbers and what is open source versus paid Pro.
What is an autonomous AI penetration testing benchmark?
What is an autonomous AI penetration testing benchmark?
It is a repeatable run of an autonomous AI pentester against a known vulnerable lab, where every finding carries the exact command and its raw output so the result can be reproduced or contested. Darkmoon, the open source autonomous AI penetration testing tool, publishes each run as a full report rather than a headline number.
Are these AI security testing benchmark results real?
Are these AI security testing benchmark results real?
Yes. Every row is transcribed from a published report in the darkmoon-research evidence corpus or the Darkmoon-Benchmarks repository. The severity split and the exploited count are the report's own summary table, not an estimate. Where a report does not record the model, this page says so rather than guessing.
Does Darkmoon run on a local model, and does it keep data on my side?
Does Darkmoon run on a local model, and does it keep data on my side?
The open source CLI can run entirely on a local model (Ollama or llama.cpp), and the Privacy Gateway tokenizes your real IPs, hosts and credentials so the model works on placeholders while the real values stay on your perimeter. The OWASP Juice Shop web benchmark on this page was run on a local model.
How does Darkmoon compare to other AI pentest tools?
How does Darkmoon compare to other AI pentest tools?
Darkmoon is an open source, self-hosted, local-first alternative to Strix, XBOW, PentAGI and PentestGPT, with proof of exploitation per finding across web, cloud, Active Directory, Kubernetes, CI/CD, databases and IoT. See the tool-by-tool comparison for the full feature matrix and honest credit where competitors lead.
What is open source and what is paid Pro?
What is open source and what is paid Pro?
The offensive engine that produced every benchmark on this page is the open source Darkmoon CLI (GPL-3.0): it finds, proves, runs locally and keeps data on your side. The web dashboard and the remediation-to-pull-request loop are paid Pro capabilities.
S.08NextRun the same benchmarks on your own labs
Darkmoon is open source (GPL-3.0), self hosted and local first. Clone it, point it at a lab you own, and read every line. A star helps other teams find it.