S.01Web benchmark

Web app pentest benchmark: OWASP Juice Shop

Web app pentest benchmark on OWASP Juice Shop: an autonomous AI pentest that escalated from 14 to 57 findings across six black-box campaigns on a local model, with proof of exploitation for every finding. Real results, published and reproducible.

57
vulnerabilities in the final campaign
28.5 min
on a local model
6
black-box campaigns
14 → 57
findings, first to last campaign

Darkmoon, the open source autonomous AI penetration testing tool, found 57 vulnerabilities on a black-box OWASP Juice Shop in the final campaign, each with a working proof, in 28.5 minutes on a local model.

S.03Escalation
Six campaigns on the same target

Re-run weekly against the same OWASP Juice Shop, the autonomous agent found more each time as it learned the target, from 14 findings to 57.

MEDIUM
Campaign 1 of 6 · 2026-03-22 · 12 min
14 findings

camp_20260322_e7f8

HIGH
Campaign 2 of 6 · 2026-03-29 · 16.33 min
28 findings

camp_20260329_a3b4

HIGH
Campaign 3 of 6 · 2026-04-05 · 22 min
38 findings

camp_20260405_c9d0

CRITICAL
Campaign 4 of 6 · 2026-04-12 · 25.67 min
44 findings

camp_20260412_e5f6

CRITICAL
Campaign 5 of 6 · 2026-04-19 · 32 min
49 findings

camp_20260419_a1b2

CRITICAL
Campaign 6 of 6 · 2026-04-26 · 28.5 min
57 findings

camp_20260426_3d2f

S.04Result
The headline run

The row links to the long-form OWASP Juice Shop write-up and to the raw run file in the benchmarks repository.

Lab / targetFindingsSeverityExploitedModelEvidence
OWASP Juice Shopcamp_20260426_3d2f578C24H21M4Lproof per findingLocal (Ollama / llama.cpp)Write-upReport

Disclaimer
Darkmoon's own benchmark on the public OWASP Juice Shop lab.

The offensive run is produced by the open source Darkmoon CLI; the web dashboard and the remediation-to-PR loop are paid Pro. The run file and the per-campaign escalation are published in the Darkmoon-Benchmarks repository. Full run file: juice-shop-2026-04-26.md. The offensive run above is the open source CLI. On the same OWASP Juice Shop, the paid Pro remediation engine turned 57 findings into 57 pull requests and demonstrated 42 end to end, with every excluded case disclosed.

S.07FAQ
Web benchmark questions

What is the OWASP Juice Shop autonomous pentest benchmark?

It is a web app pentest benchmark where Darkmoon, run autonomously and black-box against a default OWASP Juice Shop image, escalated from 14 findings to 57 across six weekly campaigns on the same target, with proof of exploitation per finding. The final run took 28.5 minutes on a local model.

How many vulnerabilities did the Juice Shop run find?

The final campaign found 57 findings on OWASP Juice Shop, split 8 critical, 24 high, 21 medium and 4 low, each proven individually. Earlier campaigns on the same target found 14, 28, 38, 44 and 49, so the run improves as it learns the target.

Was the web benchmark run on a local model?

Yes. The OWASP Juice Shop run used a local model (Ollama or llama.cpp) through the open source Darkmoon CLI, black-box. The web dashboard and the remediation-to-PR loop are paid Pro capabilities.

Can Darkmoon also fix the findings it proves?

On the same OWASP Juice Shop target, the paid Pro remediation engine turned 57 findings into 57 pull requests and demonstrated 42 of them end to end (fix, compile, live exploit-retest, human-reviewed PR), with the excluded cases disclosed in full. Nothing is auto-merged.

S.08Next
Run the Juice Shop benchmark yourself

Open source, self hosted and local first. Spin up OWASP Juice Shop, point Darkmoon at it, and read every proof. A star helps other teams find it.