S.01Benchmarks

Autonomous AI penetration testing benchmarks

Autonomous AI penetration testing benchmarks with real results only. Darkmoon, the open source AI pentest tool, was run against public vulnerable labs across cloud, infrastructure, IoT and web, and every finding here carries a working proof in a published report.

17
autonomous lab runs published (cloud, infrastructure, IoT, web)
318
findings across every run, from each report's own total
96
findings exploited with a working proof (the Juice Shop run proves each finding individually)
4
attack surfaces benchmarked end to end

Darkmoon, the open source autonomous AI penetration testing tool, found 318 vulnerabilities across 17 runs on public labs, and proved 96 of them with a real exploit.

S.05Leaderboard
Every run, every finding

Scroll horizontally on narrow screens. Each row links to its long-form write-up and to the raw report it is drawn from.

Lab / targetSurfaceFindingsSeverityExploitedModelEvidence
AWS · huge-logistics S3camp_20260802_03bfc675Cloud95C1H2M1L5claude-opus-4-6Write-upReport
AWS · pwnedlabs EBS/S3camp_20260802_611cece1Cloud92H5M2L0claude-opus-4-6Write-upReport
Azure · Entra ID tenantcamp_20260802_7eee391fCloud2811C10H4M3L12claude-opus-4-6Write-upReport
Azure · BloodHound / priv-esccamp_20260802_9d245c0cCloud198C6H5M11claude-opus-4-6Write-upReport
Azure · Key Vault (extract)camp_20260802_38118fb8Cloud162C6H8M6claude-opus-4-6Write-upReport
Azure · Key Vault (pivot)camp_20260802_59e4e905Cloud73C2H2M4claude-opus-4-6Write-upReport
GCP · SSRF to metadatacamp_20260802_656007d3Cloud43C1H3claude-opus-4-6Write-upReport
GCP · public GCS bucketcamp_20260802_3ced7196Cloud53C1H1M4claude-opus-4-6Write-upReport
Vault + registry + Docker socketcamp_20260801_2bd90d3fInfrastructure & CI/CD4115C15H9M2L8claude-opus-4-6Write-upReport
Terraform + AWS + Ansiblecamp_20260801_b1b96939Infrastructure & CI/CD3421C5H8M16claude-opus-4-6Write-upReport
GitLab CE 19.2.1camp_20260801_a719641dInfrastructure & CI/CD264C8H10M2L2I2claude-opus-4-6Write-upReport
PostgreSQL 16 + MySQL 5.6camp_20260801_c0151524Infrastructure & CI/CD226C10H6M13claude-opus-4-6Write-upReport
Redis 7.4.10 (unauth)camp_20260801_96be38b9Infrastructure & CI/CD93C5H1M5claude-opus-4-6Write-upReport
Jenkins 2.541.3 (security off)camp_20260801_b6ad197dInfrastructure & CI/CD33C3claude-opus-4-6Write-upReport
IoTGoat firmware imagecamp_20260802_dd187131IoT & firmware204C7H7M2L1claude-opus-4-6Write-upReport
IoTGoat live devicecamp_20260802_257f75f1IoT & firmware94C3H2M3claude-opus-4-6Write-upReport
OWASP Juice Shopcamp_20260426_3d2fWeb application578C24H21M4Lproof per findingLocal (Ollama / llama.cpp)Write-upReport

Disclaimer
This is Darkmoon's own benchmark on public training labs.

It is not a third-party certification or an analyst endorsement. The offensive runs are produced by the open source Darkmoon CLI; the web dashboard and the remediation-to-PR loop are paid Pro capabilities. Every number is drawn from a published report and can be reproduced or contested. Severity is shown as C / H / M / L / I. Counts are each report's own total, so a report that lists an aggregated finding twice counts it twice. The raw reports live in the darkmoon-research corpus and the Darkmoon-Benchmarks repository.

S.07FAQ
Questions buyers ask

Short, factual answers about the method, the numbers and what is open source versus paid Pro.

What is an autonomous AI penetration testing benchmark?

It is a repeatable run of an autonomous AI pentester against a known vulnerable lab, where every finding carries the exact command and its raw output so the result can be reproduced or contested. Darkmoon, the open source autonomous AI penetration testing tool, publishes each run as a full report rather than a headline number.

Are these AI security testing benchmark results real?

Yes. Every row is transcribed from a published report in the darkmoon-research evidence corpus or the Darkmoon-Benchmarks repository. The severity split and the exploited count are the report's own summary table, not an estimate. Where a report does not record the model, this page says so rather than guessing.

Does Darkmoon run on a local model, and does it keep data on my side?

The open source CLI can run entirely on a local model (Ollama or llama.cpp), and the Privacy Gateway tokenizes your real IPs, hosts and credentials so the model works on placeholders while the real values stay on your perimeter. The OWASP Juice Shop web benchmark on this page was run on a local model.

How does Darkmoon compare to other AI pentest tools?

Darkmoon is an open source, self-hosted, local-first alternative to Strix, XBOW, PentAGI and PentestGPT, with proof of exploitation per finding across web, cloud, Active Directory, Kubernetes, CI/CD, databases and IoT. See the tool-by-tool comparison for the full feature matrix and honest credit where competitors lead.

What is open source and what is paid Pro?

The offensive engine that produced every benchmark on this page is the open source Darkmoon CLI (GPL-3.0): it finds, proves, runs locally and keeps data on your side. The web dashboard and the remediation-to-pull-request loop are paid Pro capabilities.

S.08Next
Run the same benchmarks on your own labs

Darkmoon is open source (GPL-3.0), self hosted and local first. Clone it, point it at a lab you own, and read every line. A star helps other teams find it.