Latest
Integrations & workflow
Darkmoon inside the tools your team already runs: CI/CD, IDEs, automation and SecOps.
- Learn more
DarkMoon in Splunk: autonomous pentest results as a SOC data source
|Stream campaigns, findings, remediation PRs and retest verdicts into Splunk over HEC as four CIM-aligned sourcetypes, six SOC dashboards including MITRE ATT&CK coverage, and a Send to DarkMoon alert action that turns a correlation search into a non-destructive offensive-validation loop. Redaction-safe by design.
- Learn more
DarkMoon in Grafana: a live security-posture dashboard from the Pro API
|Two ways to see DarkMoon posture in Grafana: an Infinity-datasource dashboard you import in minutes, or a signed Go-backend app plugin that keeps the token server-side. Overview to campaign to finding to evidence metadata to remediation PR, with safe fields only.
- Learn more
DarkMoon in n8n: an autonomous pentest node for your SOAR workflows
|A community n8n node and trigger for DarkMoon: launch campaigns, read findings and evidence metadata, get retest verdicts and posture timeseries, register signed webhooks, and open a remediation PR through an opaque credential reference that never carries a secret and never auto-merges.
- Learn more
DarkMoon in Jenkins: a pipeline step that gates the build on real findings
|A darkmoonScan() pipeline step that wraps the DarkMoon CLI on the agent, fails or marks the build unstable from the severity summary, maps findings to SARIF 2.1 for Warnings Next Generation, and keeps the token in the Jenkins Credentials store, never on the command line.
- Learn more
DarkMoon in GitLab: a CI/CD Catalog component with native MR reports
|Include one component and every pipeline runs an autonomous pentest, emits a Code Quality report for the merge request widget, optionally a SAST report, and fails the job on findings per your policy. The full report stays opt-in.
- Learn more
DarkMoon in GitHub Actions: autonomous pentest in the pull request
|A Marketplace action that autodetects edition, launches or attaches to a campaign, streams to completion, writes a severity table to the job summary, uploads SARIF to Code scanning, comments the result on the PR, and fails from findings, never from an exit code. Tokens masked, evidence stripped.
- Learn more
DarkMoon in JetBrains: campaigns, findings and reports in the IDE
|A native IntelliJ-platform tool window for DarkMoon: campaigns, a sortable and filterable vulnerabilities table with a detail pane, and redaction-safe reports. The JWT lives in PasswordSafe, Pro-only features degrade cleanly on OSS, and the plugin never opens a PR on its own.
- Learn more
DarkMoon everywhere: pentesting in your IDE, CI/CD, automation and SecOps
|One autonomous pentest engine, wired into where security teams already work: a GitHub Action and GitLab component in CI, VS Code and JetBrains in the IDE, an n8n node for automation, Splunk and Grafana for SecOps, and a foundation SDK. What runs on the open-source CLI and what needs the paid Pro REST API.
- Learn more
Autonomous penetration testing in your CI/CD pipeline
|Wire autonomous penetration testing into CI/CD: run Darkmoon on an ephemeral environment, gate the build on exploited findings, and attach a proof report.
Platform & architecture
How the engine is built: the gateway, the agents, the graph, the reports and the remediation loop.
- Learn more
The orbital infrastructure graph: attack paths you can walk, in the browser and in VR
|Why we rebuilt the infrastructure map as an orbital, guided, WebXR-ready graph: exposure rings, per-asset findings, an animated MITRE attack path, and a headset view that needs no Meta or Apple SDK. Our reasoning, the attempts that failed, and where we think security visualisation is going.
- Learn more
Inside the Privacy Gateway: the model never sees your real infrastructure
|How Darkmoon tokenizes every sensitive value locally before it reaches the model, the 819-line command gateway, the per-session vault, the fail-closed prompt socket that tokenizes even the launch prompt, and the honest limits of deterministic placeholders.
- Learn more
From finding to fix to exploit-retest: DarkMoon's closed remediation loop
|Darkmoon's Pro remediation loop turns a confirmed or exploited finding into a fix that only counts when the original exploit is re-run in an ephemeral sandbox and no longer fires, then opens a pull request for human review across 8 SCM providers. Never auto-merged.
- Learn more
How DarkMoon investigates a target: signal, specialist agents, and a governed 142-tool boundary
|From a target string to qualified findings: fingerprinting technology signals, dispatching 50 specialist agents by target class, the adversarial EXPLOITED/CONFIRMED/UNCONFIRMED rubric, and the single build-enforced allow-list of 142 security tools that bounds what every agent can run.
- Learn more
SAST that ships the fix: sandbox-validated remediation pull requests
|A deeper look at Darkmoon's Pro remediation pipeline: the three-phase SAST audit that locates the root cause with codemap, the injected-LLM patch plus an overreach judge, the ephemeral 127.0.0.1 Docker sandbox with exploit and regression gates, the four honest dispositions, the encrypted credential store with an opaque CREDENTIAL_REF, and the dashboard PR column. Human-reviewed, never auto-merged.
- Learn more
Scheduled and recurring autonomous pentests: how the Darkmoon scheduler works
|A background task polls every 60 seconds for due campaigns and launches the pentest orchestrator, with none, daily, weekly or monthly recurrence and one-shot schedules that disable themselves after firing. The exact model, from the API routes and the recurrence logic as implemented.
- Learn more
The Darkmoon dashboard: watching findings land as the agent works
|Real-time monitoring as built: the agent pushes each finding and infrastructure node to disk the moment it discovers it, the dashboard polls at 5 seconds while a run is active and 15 seconds when idle, sub-agent activity is streamed out of opencode's own database, and orphaned runs are reconciled on restart.
- Learn more
Why Darkmoon builds its reports server-side, deterministically, from proof
|The report body is never trusted to the model. It is assembled server-side from the findings the agent pushed: a management summary, an executive summary, a findings table, per-finding detail with MITRE ATT&CK and ISO 27001 mappings, and a prioritized remediation roadmap. Placeholders are rehydrated at the single write point, and the run can target standard, HackerOne or Bugcrowd output.
- Learn more
From finding to fix: autonomous remediation with human-reviewed pull requests
|Darkmoon's Pro remediation agent turns each confirmed finding into a minimal fix: it reproduces the exploit against a fresh build, patches the root cause, re-runs the exploit and its variants plus the repo's tests in an ephemeral Docker sandbox, then opens a pull request for human review. It never merges on its own.
- Learn more
Multi-agent AI penetration testing without giving the model a shell
|How multi-agent AI reshapes autonomous penetration testing: a controlled MCP tool layer, credential-gated dispatch of specialist agents, and a bounded executor.
- Learn more
Why proof of exploitation beats AI vulnerability scores
|AI vulnerability scanners hand you scores and probabilities. Here is why proof of exploitation, with the exact payload and raw output, is what a buyer should demand.
Industries & regulation
What a penetration test covers and proves for law firms, hospitals, SaaS teams, MSPs and regulated entities under NIS2, DORA, SOC 2 and HIPAA.
- Learn more
How much does a penetration test cost in 2026, and what €799 buys
|The variables that move a penetration testing price (scope, surface types, depth, report, retest, legal overhead, human versus autonomous), what Darkmoon's €799 flat-rate engagement includes, and the four cases where a bespoke human engagement is still the right buy.
- Learn more
SOC 2 penetration testing requirements for SaaS startups
|SOC 2 does not name a mandatory penetration test, yet nearly every auditor asks for one. What the Trust Services Criteria actually require, what the report must contain to count as evidence, how to scope a multi-tenant SaaS test, and how to make it repeatable in CI.
- Learn more
Law firm cybersecurity checklist: the attack paths that actually get exploited
|A checklist organised by attack path rather than by product: identity and Microsoft 365, client portals, the document management system, remote access, suppliers and the funds workflow. With the professional secrecy basis (art. 66-5, RIN art. 2, ABA 1.6(c)) and what a penetration test must prove on each path.
- Learn more
Hospital penetration testing: scope, HIPAA evaluation and NIS2 evidence
|What a penetration test of a healthcare estate covers (portals, HIS, PACS and DICOM reachability, VPN, Active Directory, legacy imaging workstations), what it excludes (medical devices in clinical use), and how it serves the HIPAA evaluation standard, NIS2 Article 21(2)(f) and the French HDS context.
- Learn more
NIS2 for MSPs and MSSPs: Annex I, the size rule and Implementing Regulation 2024/2690
|Managed service providers and managed security service providers are named in NIS2 Annex I. The verbatim size rule, what Implementing Regulation 2024/2690 adds, how supply-chain obligations flow to and from your clients, what to prove, and the French transposition status as of October 2026.
- Learn more
NIS2 vs ISO 27001 vs DORA: what overlaps, what does not
|A directive, a certifiable management-system standard and a sector regulation are three different objects. Lex specialis between NIS2 and DORA, what each text says about testing (only DORA names a penetration test), what overlaps, what does not, and where one validated test report fits in all three files.
- Learn more
Darkmoon and NIS2: turning autonomous offensive validation into evidence
|NIS2 asks essential and important entities to run risk management, test the effectiveness of their measures, and keep evidence. Here is how autonomous offensive validation helps demonstrate those obligations, honestly, using the ISO 27001, NIST SP 800-115 and MITRE ATT&CK mappings the reports actually carry. It helps demonstrate; it does not make you compliant.
- Learn more
Continuous pentesting for NIS2 and DORA: why one audit a year is not enough
|NIS2 and DORA push regulated organisations toward regular testing of their security measures, not a single annual snapshot. Here is why point-in-time pentests age badly against a moving estate, and how autonomous, repeatable testing complements the mandatory human red team without replacing it.
Alternatives & landscape
Honest comparisons with the other AI pentesting tools, and why local-first matters.
- Learn more
PentestGPT alternative: from an LLM advisor to an autonomous, local pentester
|PentestGPT is an excellent LLM assistant that guides a human operator through a test. If you want the model to actually run the engagement, through a controlled tool layer, on a local model, with Active Directory and Kubernetes coverage and proof per finding, here is the honest comparison.
- Learn more
Open source XBOW alternative: self hosted, local, auditable
|XBOW made autonomous AI pentesting famous, but it is closed, cloud hosted and enterprise priced. Here is the open source, self hosted, local-LLM path, with a reproducible OWASP Juice Shop benchmark.
- Learn more
The local, self hosted AI pentester: autonomous testing without a cloud model
|Most autonomous AI pentesters send your code, IPs and traffic to a hosted model. Here is how to run autonomous penetration tests entirely locally, with deterministic tokenization so nothing sensitive ever leaves.
- Learn more
PentAGI alternative: local first, with Active Directory and Kubernetes
|PentAGI is a strong self hostable multi-agent pentester built around cloud LLMs. If you need a local-LLM default, a privacy gateway, and Active Directory plus Kubernetes coverage, here is the comparison.
- Learn more
HexStrike AI alternative: from MCP tool runner to a full autonomous platform
|HexStrike AI is a powerful MCP tool runner. Darkmoon ships the orchestration methodology, runs on a local LLM, keeps data off the model, and covers AD and Kubernetes. An honest comparison.
- Learn more
Strix alternative: the local-first open source AI pentester
|Strix is a strong open source AI pentester, but it runs against a cloud model. If you need autonomous AI penetration testing that keeps your code and traffic on your own infrastructure and also covers Active Directory and Kubernetes, here is the honest comparison.
- Learn more
Shannon alternative: autonomous AI pentesting that stays local
|Shannon posts an impressive white-box benchmark, but it reads your source and calls a cloud model. Here is a black-box, local-LLM alternative with a reproducible OWASP Juice Shop benchmark, plus Active Directory and Kubernetes coverage.
- Learn more
How to run an AI pentest without sending your data to the LLM
|Autonomous AI pentesting normally ships your real IPs, hostnames and credentials to a hosted model. Here is the deterministic tokenization and local rehydration design that avoids it.
- Learn more
The open source AI pentest tools worth knowing in 2026
|An honest field guide to the open source autonomous AI penetration testing tools in 2026: what each one is good at, licensing, scope, and how to pick.
- Learn more
NodeZero alternative: the self hosted open source path
|If you evaluated NodeZero or Pentera but need self hosted, auditable and open source autonomous security testing, here are the real options and the trade offs.
Techniques by domain
Method pieces: web, API, GraphQL, CMS, Active Directory, cloud, IaC, brokers, firmware and LLM endpoints.
- Learn more
LLM endpoint penetration testing: 8 OWASP LLM Top 10 findings, proven end to end
|Darkmoon's new llm agent detected an OpenAI-compatible endpoint and found 8 distinct issues mapped to the OWASP LLM Top 10: system-prompt leakage with hardcoded credentials, prompt injection and jailbreak, insecure output handling, unbounded consumption, SSRF and unauthenticated access, each with the exact request and raw response.
- Learn more
How the Darkmoon LLM agent pentests an AI inference endpoint, with garak
|Auto-detection of OpenAI-compatible, Ollama, vLLM and TGI endpoints, dispatch like the GraphQL and Kubernetes agents, then fingerprint, capability profiling, an optional bounded garak pass and adaptive OWASP-LLM attacks with a per-run canary and explicit detectors, bounded throughout to avoid denial of service.
- Learn more
Pentesting an LLM without leaking your own infrastructure to the LLM
|Why an offensive LLM-endpoint agent ships with privacy-gateway hardening: pre-model prompt tokenization from terminal and UI, server-side report rehydration and a per-session override, measured on a real run where the model saw the target 0 times across 3.5 MB of traffic.
- Learn more
Kubernetes penetration testing with AI: from a mounted token to root on the node
|How an autonomous agent walks a Kubernetes cluster: RBAC abuse, ServiceAccount token theft, kubelet API, secrets, container escape, each proven when executed.
- Learn more
Autonomous API penetration testing: JWT abuse, BOLA and real attack chains
|Autonomous API penetration testing that maps the surface, abuses JWTs, finds BOLA, reuses tokens and chains findings into real impact, with proof for each step.
- Learn more
Active Directory penetration testing with autonomous AI: BloodHound, Kerberoasting, DCSync and ADCS
|How the Darkmoon active-directory agent runs AD penetration testing end to end: BloodHound attack paths, AS-REP and Kerberoasting, DCSync, and ADCS ESC1 to ESC16.
- Learn more
GraphQL penetration testing beyond introspection: BOLA, batching and mutations, proven
|GraphQL penetration testing past introspection: BOLA and IDOR through node IDs, query batching and alias brute force, depth and complexity abuse, and mutation testing.
- Learn more
Beyond WPScan: autonomous AI penetration testing for WordPress and five more CMS platforms
|Autonomous AI penetration testing for WordPress, Drupal, Joomla, Magento, PrestaShop and Moodle. Six dedicated agents that prove exploitation instead of matching versions.
- Learn more
Terraform and Ansible security: from an IaC misconfiguration to a proven attack path
|Static IaC scanners flag Terraform and Ansible misconfigurations. Darkmoon's terraform and ansible agents prove which ones chain into an obtained credential.
- Learn more
Pentesting Kafka, RabbitMQ and MQTT with AI agents: the broker as the shortest path to your data
|Autonomous security testing for Kafka, RabbitMQ and MQTT: management APIs, anonymous access, weak topic authorization and secrets, with proof of impact.
- Learn more
IoT firmware penetration testing with AI: the method
|The full methodology an autonomous agent follows on embedded targets, from unpacking an image to rooting a live appliance. Start here, then read the two case studies: the static image analysis and the live device run.
- Learn more
Autonomous cloud penetration testing: SSRF, metadata and service-account chains
|What it takes for an AI agent to turn one SSRF or one leaked key into cloud initial access on AWS, Azure and GCP, metadata tokens, gopher smuggling, Key Vault and storage exfiltration.
Case studies & benchmarks
Real runs against real targets, with the exact findings, what was proven and what was refused.
- Learn more
Benchmark: autonomous specialist agents against nine vulnerable labs, with proof and honest limits
|A reproducible end-to-end validation run of the Darkmoon agent roster against local vulnerable labs, on the Pro stack with claude-opus-4-6. Redis, PostgreSQL and MySQL, Vault, a container registry, the Docker socket, Terraform, AWS on LocalStack, Ansible and Jenkins, with real finding counts, what was exploited, what we deliberately demoted, and what could only be command-validated.
- Learn more
Autonomous web app pentest against OWASP Juice Shop: 7 findings, 4 exploited with proof
|An authorized autonomous scan of a locally hosted OWASP Juice Shop: UNION SQL injection that dumped the user table, a login authentication bypass that minted an admin JWT, null-byte path traversal, a stored-XSS sanitizer bypass, JWT hash exposure and verbose errors, each graded EXPLOITED or CONFIRMED by demonstrated impact.
- Learn more
Jenkins with security disabled: unauthenticated script console RCE, proven end to end
|A Jenkins 2.541.3 controller running with SecurityRealm None. An autonomous agent went from anonymous HTTP to code execution on the controller and exfiltrated the master encryption key. Three findings, all critical, all exploited.
- Learn more
Redis with no password: what the agent got, and what it refused to claim
|An unauthenticated Redis 7.4.10 instance assessed autonomously: 9 findings, 5 exploited. The classic RDB-write RCE was demoted to mitigated because Redis 7.x blocked it, and we published that rather than the bigger number.
- Learn more
PostgreSQL COPY TO PROGRAM RCE and a MySQL 5.6 audit in one autonomous run
|22 findings against PostgreSQL 16.14 and MySQL 5.6.51: command execution as the postgres user, arbitrary file read and write, password hash extraction from pg_shadow and mysql.user, and live session tokens.
- Learn more
An exposed terraform.tfstate, and the secrets an autonomous agent pulled out of it
|A Terraform state file served over plain HTTP with no authentication, an Ansible inventory in the open, and a LocalStack AWS account with an AdministratorAccess CI user. 16 unique findings, 10 of them critical.
- Learn more
Vault, a container registry and the Docker socket: one run, 41 findings
|Three specialist agents dispatched in parallel against HashiCorp Vault, an OCI registry and an exposed Docker Engine socket. Root token guessed, image layers unpacked for credentials, and a container escape that read the host /etc/shadow.
- Learn more
Anonymous S3 listing to IT admin: a credential chain walked without a human
|A publicly listable S3 bucket led to a PowerShell script with hardcoded IAM keys, which led to a credential export with six more sets, which led to an it-admin key. Five critical findings, all exploited, plus a card data exposure.
- Learn more
A public EBS snapshot and a public bucket: an AWS assessment with zero exploitation
|9 findings against a live AWS account with a deliberately limited IAM user. Nothing was exploited and the report says so: the value here is what a constrained identity can still map, and how much of it is confirmable.
- Learn more
Entra ID: a deleted blob, ROPC without MFA, and a tenant takeover path we did not walk
|Anonymous blob versioning recovered a deleted archive with hardcoded credentials, ROPC minted a token with no MFA challenge, and Graph enumeration exposed a service principal holding Privileged Authentication Administrator.
- Learn more
Helpdesk to Global Admin in Azure: the chain an autonomous agent reconstructed
|A password in an Entra ID custom security attribute, a second one in VM userData, and Global Admin credentials in a storage blob. The agent chained four identities and captured the lab flag.
- Learn more
Reading Azure Key Vault secrets with a stolen password, then pivoting with them
|Two campaigns against the same Key Vault: the first extracted three plaintext contractor passwords through over-permissive RBAC, the second reused one to authenticate as that contractor and reach customer card data.
- Learn more
A GCS bucket name in an HTML comment, 500 customer records out
|The bucket name was commented out in the careers page. Object fuzzing found a password protected backup archive, and the password was cracked from a wordlist built out of the same website.
- Learn more
SSRF to gopher to GCP metadata: stealing a service account token end to end
|A profile image fetcher took a user supplied URL. Direct metadata access was blocked by the required header, so the agent smuggled the header through a gopher:// payload, stole the token and emptied the bucket.
- Learn more
Auditing a GitLab instance through its own API
|13 unique findings against GitLab CE 19.2.1: an admin PAT with api and sudo scopes, an unmasked AWS secret in CI/CD variables, an exposed runner registration token, open signup and no branch protection.
- Learn more
Live device case study: an unauthenticated root backdoor, found twice by two independent runs
|Two separate campaigns against the same live OWASP IoTGoat appliance. Both reached root through the backdoor on port 5515, both pulled /etc/shadow, and one cracked the Mirai default credential over SSH.
- Learn more
Static firmware analysis, case study: 20 findings without ever powering the device
|No device, no network, just the IoTGoat firmware image unpacked. The agent found the backdoor binary, the telnet daemon, the command injection endpoint and a Mirai default credential before anything was ever powered on.
- Learn more
What an autonomous pentest looks like when the site is actually hardened
|16 findings, zero critical, zero high, nothing exploited. The most useful report we published this week is the boring one, because it shows what the agent does when there is no way in.
- Learn more
We aimed an autonomous AI attacker at OpenNHP. On the protected plane it found nothing.
|A joint, reproducible test with the OpenNHP project: Darkmoon, a fully autonomous AI pentester, run against the OpenNHP public demo. On the exposed demo surface it behaved like any capable attacker and produced 51 findings. On the NHP-protected hosts it could not even discover a target, so the attack chain never began.