S.01Remediation benchmark · Pro

Automated vulnerability remediation, proven end to end

Most autonomous pentesters stop at the finding. Darkmoon's Pro remediation engine turns each one into an AI remediation pull request, and we published the honest score: 42 of 57 findings fixed and retested end-to-end on OWASP Juice Shop.

57
findings, each mapped to its own addressable PR (#1–#57)
42
demonstrated end-to-end with a clean live exploit-retest
27
distinct fixes across the 57 PRs (deduplication disclosed)
0
pull requests merged, every one awaits human review

S.03The loop
Finding, fix, retest, review

The same four steps run on every finding. A fix only counts toward the 42 when the original exploit is re-run against a live instance and confirmed closed.

01

Finding

A vulnerability the autonomous pentest already proved with a working exploit is mapped back to the exact source file(s) that govern it.

02

Fix

The remediation agent generates a minimal, root-cause patch against the real Juice Shop code, not a suppression or a config toggle.

03

Compile + retest

The branch is type-checked with tsc, then the original exploit is re-run against a live instance. The finding must be confirmed closed.

04

Human-reviewed PR

A pull request is opened for a human to review. It is AI-generated code, disclosed as such, and never auto-merged.

S.04Honest limits
The 15 we did not count

42 is the number that survives a clean runtime retest. The other 15 are disclosed here rather than folded into a bigger headline.

9 of 15
Why it is excluded from the 42
Partial draft PRs

Close a narrower exploit at runtime but are partial or architectural relative to the full finding.

1 of 15
Why it is excluded from the 42
PR #52, ineffective fix

Did not close the exploit on retest. The fix is ineffective and is flagged honestly.

1 of 15
Why it is excluded from the 42
PR #3, not runtime-retested

Verbose errors: changes behaviour only under NODE_ENV=production, so it was not runtime-retested.

4 of 15
Why it is excluded from the 42
Frontend / asset PRs

#44, #46, #18/#39, frontend or static-asset changes not exercised by the API-level retest harness.

15 of 57
Why it is excluded from the 42
Total excluded

57 findings − 42 demonstrated = 15 disclosed above.

Method
How each PR was checked

Five checks, in order, on every one of the 57 pull requests. Everything needed to repeat them is published.

  1. # five checks, in order, on every PR
  2. 01Maps to a real, exploit-proven finding in the pentest dataset.
  3. 02The diff targets the real Juice Shop file(s) governing the finding, no suppression.
  4. 03Type-checks with tsc (24 distinct backend branches, all exit 0).
  5. 04The exploit is reproduced on main then re-run against the fix branch on a live v19.2.1 instance.
  6. 05Scanned for secrets; PR bodies disclose the fix is AI-generated and awaits human review.

Disclaimer
This is Darkmoon's own reproducible benchmark on a public lab.

It is not a third-party certification or an analyst endorsement. Every number above is drawn from the published dossier and can be reproduced or contested.

S.07Next
Test the loop on your own code

Darkmoon is open source (GPL-3.0) and self hosted. The offensive engine is free; the remediation engine is a Pro capability.