Benchmarks

Autonomous AI penetration testing benchmarks

Autonomous AI penetration testing benchmarks with real results only. Darkmoon, the open source AI pentest tool, was run against public vulnerable labs across cloud, infrastructure, IoT and web, and every finding here carries a working proof in a published report.

17

autonomous lab runs published (cloud, infrastructure, IoT, web)

318

findings across every run, from each report's own total

96

findings exploited with a working proof (the Juice Shop run proves each finding individually)

4

attack surfaces benchmarked end to end

Darkmoon, the open source autonomous AI penetration testing tool, found 318 vulnerabilities across 17 runs on public labs, and proved 96 of them with a real exploit.

Leaderboard

Every run, every finding

Lab / targetSurfaceFindingsSeverityExploitedModelEvidence
AWS · huge-logistics S3camp_20260802_03bfc675Cloud95C1H2M1L5claude-opus-4-6Write-up Report
AWS · pwnedlabs EBS/S3camp_20260802_611cece1Cloud92H5M2L0claude-opus-4-6Write-up Report
Azure · Entra ID tenantcamp_20260802_7eee391fCloud2811C10H4M3L12claude-opus-4-6Write-up Report
Azure · BloodHound / priv-esccamp_20260802_9d245c0cCloud198C6H5M11claude-opus-4-6Write-up Report
Azure · Key Vault (extract)camp_20260802_38118fb8Cloud162C6H8M6claude-opus-4-6Write-up Report
Azure · Key Vault (pivot)camp_20260802_59e4e905Cloud73C2H2M4claude-opus-4-6Write-up Report
GCP · SSRF to metadatacamp_20260802_656007d3Cloud43C1H3claude-opus-4-6Write-up Report
GCP · public GCS bucketcamp_20260802_3ced7196Cloud53C1H1M4claude-opus-4-6Write-up Report
Vault + registry + Docker socketcamp_20260801_2bd90d3fInfrastructure & CI/CD4115C15H9M2L8claude-opus-4-6Write-up Report
Terraform + AWS + Ansiblecamp_20260801_b1b96939Infrastructure & CI/CD3421C5H8M16claude-opus-4-6Write-up Report
GitLab CE 19.2.1camp_20260801_a719641dInfrastructure & CI/CD264C8H10M2L2I2claude-opus-4-6Write-up Report
PostgreSQL 16 + MySQL 5.6camp_20260801_c0151524Infrastructure & CI/CD226C10H6M13claude-opus-4-6Write-up Report
Redis 7.4.10 (unauth)camp_20260801_96be38b9Infrastructure & CI/CD93C5H1M5claude-opus-4-6Write-up Report
Jenkins 2.541.3 (security off)camp_20260801_b6ad197dInfrastructure & CI/CD33C3claude-opus-4-6Write-up Report
IoTGoat firmware imagecamp_20260802_dd187131IoT & firmware204C7H7M2L1claude-opus-4-6Write-up Report
IoTGoat live devicecamp_20260802_257f75f1IoT & firmware94C3H2M3claude-opus-4-6Write-up Report
OWASP Juice Shopcamp_20260426_3d2fWeb application578C24H21M4Lproof per findingLocal (Ollama / llama.cpp)Write-up Report

Severity is shown as C / H / M / L / I. Counts are each report's own total, so a report that lists an aggregated finding twice counts it twice. The raw reports live in the darkmoon-research corpus and the Darkmoon-Benchmarks repository.

FAQ

Questions buyers ask

What is an autonomous AI penetration testing benchmark?

It is a repeatable run of an autonomous AI pentester against a known vulnerable lab, where every finding carries the exact command and its raw output so the result can be reproduced or contested. Darkmoon, the open source autonomous AI penetration testing tool, publishes each run as a full report rather than a headline number.

Are these AI security testing benchmark results real?

Yes. Every row is transcribed from a published report in the darkmoon-research evidence corpus or the Darkmoon-Benchmarks repository. The severity split and the exploited count are the report's own summary table, not an estimate. Where a report does not record the model, this page says so rather than guessing.

Does Darkmoon run on a local model, and does it keep data on my side?

The open source CLI can run entirely on a local model (Ollama or llama.cpp), and the Privacy Gateway tokenizes your real IPs, hosts and credentials so the model works on placeholders while the real values stay on your perimeter. The OWASP Juice Shop web benchmark on this page was run on a local model.

How does Darkmoon compare to other AI pentest tools?

Darkmoon is an open source, self-hosted, local-first alternative to Strix, XBOW, PentAGI and PentestGPT, with proof of exploitation per finding across web, cloud, Active Directory, Kubernetes, CI/CD, databases and IoT. See the tool-by-tool comparison for the full feature matrix and honest credit where competitors lead.

What is open source and what is paid Pro?

The offensive engine that produced every benchmark on this page is the open source Darkmoon CLI (GPL-3.0): it finds, proves, runs locally and keeps data on your side. The web dashboard and the remediation-to-pull-request loop are paid Pro capabilities.

This is Darkmoon's own benchmark on public training labs. It is not a third-party certification or an analyst endorsement. The offensive runs are produced by the open source Darkmoon CLI; the web dashboard and the remediation-to-PR loop are paid Pro capabilities. Every number is drawn from a published report and can be reproduced or contested.

Run the same benchmarks on your own labs

Darkmoon is open source (GPL-3.0), self hosted and local first. Clone it, point it at a lab you own, and read every line. A star helps other teams find it.