Players

Offensive AI security companies

Autonomous pentest AI splits into four tiers competing on different bases, and the money is lopsided toward offense.

evidence: medium12 minupd 2026-08-29offensive-securitypentestfundingcompetitive-map

There is no single market leader in autonomous offensive security — there are four tiers that barely compete head-to-head, because they are selling different things to different buyers under different risk models. Deterministic-core incumbents sell auditability. LLM-agent pure-plays sell a growth story. Bug-bounty platforms sell trust in their existing marketplace. Frontier labs build the best systems in existence and refuse to sell them at all. A founder building defensive tooling needs to understand all four, because they define what "AI can do offensively" actually means in practice — and they collectively raised far more capital than the defensive side has, which should worry you.

#The market at a glance

Key numbers
  • XBOW: $120M Series C at >$1B valuation, March 18 2026, plus a $35M strategic extension including Accenture, Samsung and NVIDIA (XBOW; Bloomberg; GeekWire)
  • Horizon3.ai: $250M Series E at >$2B valuation, August 3 2026, over 7,000 customers, 120% YoY ARR growth (Horizon3.ai; TechCrunch)
  • Pentera: $60M Series D, March 2025, over 1,000 CISOs as customers (Pentera; TechCrunch) — treat the ~$1B valuation and ~$100M ARR figures reported elsewhere as [unverified], not on Pentera's own site
  • RunSybil: $40M, March 18 2026, led by Khosla Ventures with Anthropic's own Anthology Fund participating (RunSybil; Fortune)

#Master comparison table

Company HQ Founded Funding (disclosed) Valuation Architecture Model dependency Proof point Weakness
[[company-xbow|XBOW]] Seattle (nominal — see XBOW) Jan 2024 ~$155M in 2026 rounds (XBOW) >$1B LLM agent, exploit-chain construction, up to 48-step chains (appsecsanta) Undisclosed base model(s) HackerOne US leaderboard #1, Apr–Jun 2025 (XBOW) — heavily qualified, see XBOW Human-filtered submissions; no independent benchmark found
Horizon3.ai (NodeZero) San Francisco ~2019 $100M Series D (2025) + $250M Series E (2026) ≈ $350M+ >$2B Deterministic exploit engine; LLM used only for advisory/summarization via Bedrock Claude, LLaMA, Mistral — named explicitly, guidance-only (Horizon3.ai) NSA CAPT program participant; 7,000+ customers (TechCrunch) "Autonomy" claim rests on a narrower, more conservative technical base than XBOW's
Pentera Petah Tikva, Israel pre-LLM (as Pcysys) $60M Series D, Mar 2025 (fintech.global) [unverified] Deterministic attack engine; "Pentera Peer" is an AI copilot layer only (Pentera) Undisclosed Over 1,000 CISOs as customers (Pentera) Explicitly rejects "probabilistic methods" — a smaller claim than "autonomous hacker"
RunSybil US 2025 $40M, Mar 2026 Not disclosed Continuous autonomous pentesting, chains vulns and auth boundaries Undisclosed Named customers: Cursor, Turbopuffer, Notion, Baseten (Fortune) No disclosed ARR or independent benchmark; credibility rests on founder pedigree
Terra Security Israel 2025 $8M seed + $30M Series A ≈ $38M (SecurityWeek) Not disclosed Claims fine-tuned agent "swarms," human-in-the-loop Claimed fine-tuned, not disclosed at model-card level "Multiple Fortune 500" customers, unnamed Fine-tuning claim unverifiable; no benchmark
MindFort US 2025 (YC X25) $3M+ seed (MindFort) Not disclosed Claims a purpose-built model ("MF-1") plus "HillClimb" recursive-learning infra MF-1 confirmed to exist on mindfort.ai; "self-post-trained" claim not found independently (verification) Ships fixes as pull requests via CI/CD integration Smallest funding in the well-known set; earliest stage
Synack (Sara) US GA May 2026 N/A (division of Synack) N/A Hybrid: AI breadth + human Red Team depth Undisclosed Chained 3-vuln exploit "without human guidance" in early deployment (GlobeNewswire) Explicitly not claiming full autonomy
Bugcrowd (Savant Pathseeker) US Launched Jul 28 2026 N/A (product line) N/A "Frontier AI models" plus proprietary practitioner-built skill layer Undisclosed frontier model(s), explicitly not trained on researcher data Positions AI for baseline coverage, humans for complex findings (Help Net Security) Marketplace incumbent defending share, not a frontier-capability play
HackerOne (Hai) US Product line N/A N/A "Agentic AI" layered on existing marketplace Undisclosed CEO on record: does not train on researcher submissions (The Register) Faced researcher backlash over data-provenance trust
AISLE US/EU Emerged from stealth Oct 2025 Not disclosed Not disclosed "Cyber Reasoning System" — discovers and auto-patches via a Verifier Agent Claims parity with Anthropic Mythos using "much cheaper models" (AISLE) 3-to-3 FreeBSD zero-day parity with Mythos Preview; first AI-native CVE Numbering Authority (CSA Labs) Funding and team claims beyond the three named founders are largely unverified

#Tier 1: deterministic-core incumbents

Pentera and Horizon3.ai are the biggest companies in this space by funding, valuation, and customer count, and both make a deliberately smaller technical claim than the pure-plays. Horizon3.ai states outright: "NodeZero never uses GenAI to create or execute exploits. Every action is deterministic, pre-validated, and tested internally" — LLMs (Claude, LLaMA, Mistral, served via Bedrock) are used only for advisory reasoning and summarization (Horizon3.ai). Pentera's CEO frames it the same way: a "deterministic attack engine" wrapped in an AI copilot, explicitly rejecting "probabilistic methods that can, in some cases, introduce major risk" (Pentera).

This is not a weaker product — it is a different bet. Keeping exploit generation out of the LLM call sidesteps model-provider usage-policy risk entirely, and it is the most auditable answer available to a buyer who needs to explain to a regulator exactly what fired against their network. It is also why Horizon3.ai has NSA CAPT program involvement and both companies sell comfortably into government-adjacent and regulated enterprise — see Who buys, and what they pay on how a 5,000-person regulated buyer runs a security review before signing anything in this category.

#Tier 2: LLM-agent-first pure-plays

XBOW, RunSybil, Terra Security, MindFort, and smaller entrants like Vulnetic bet the other direction: give an LLM (or a fine-tuned/purpose-built model) an open-ended reasoning loop and let it construct exploit chains itself. XBOW is clearly ahead on funding and public proof points — see XBOW for the full dossier — but "ahead on narrative" and "ahead on independently verified capability" are not the same claim, and no head-to-head third-party benchmark comparing XBOW against RunSybil, Terra, or MindFort was found in this research pass. MindFort is the one company in this tier making a specific, checkable technical claim (a purpose-built model called MF-1, confirmed to exist on its own site, though the "self-post-trained" framing is unverified beyond MindFort's own description — verification log).

Notably, Anthropic's own venture arm — the Anthology Fund — is an investor in RunSybil (Fortune). See Where the gaps actually are and Dual-use risk and what it costs you for why a frontier lab funding an offensive pure-play is a more interesting signal than it first appears.

#Tier 3: bug-bounty platforms bolting on AI

HackerOne (Hai), Bugcrowd (Savant Pathseeker), Synack (Sara), and Cobalt (Autonomous Pentest) are not trying to out-build XBOW. They are defending an existing human-researcher marketplace by adding an AI tier, while explicitly promising not to train on researcher submissions — a direct response to researcher backlash that surfaced publicly in February 2026 when HackerOne's CEO had to go on record: "HackerOne does not train generative AI models, internally or through third-party providers, on researcher submissions or customer confidential data" (The Register). Intigriti and Bugcrowd have turned data-provenance into a competitive wedge against HackerOne on exactly this point.

This tier wins on distribution and existing enterprise trust relationships, not frontier capability. None of them discloses proprietary model training with anything like MindFort's specificity.

#Tier 4: frontier labs that don't sell it

Google's Big Sleep and Anthropic's Claude Mythos Preview are not commercial products, and that is the single most consequential fact in this entire market. Mythos Preview developed 181 working exploits for Firefox JS-engine bugs versus 2 for the prior Opus 4.6 model — a roughly 90x jump — and found a 27-year-old OpenBSD TCP SACK bug, a 16-year-old FFmpeg bug, and a 17-year-old FreeBSD RCE (CVE-2026-4747), among others (Anthropic). Access is restricted to a coordinated-disclosure partner program for critical-infrastructure organizations and open-source maintainers, with professional validators confirming findings before disclosure and a 90+45 day disclosure timeline. (Note: research notes referencing a "50 → 200 partner organizations" figure or a program named "Project Glasswing" for this were not independently found and are excluded here per verification.)

AISLE is the one independent company plausibly operating at this tier commercially, claiming 3-to-3 FreeBSD zero-day parity with Mythos Preview "on much cheaper models" (AISLE) — worth tracking closely, though its funding and most team claims beyond the three named founders (Ondrej Vlcek, Jaya Baloo, Stanislav Fort) are unverified.

So what

This tier "wins" on raw capability by a wide margin but explicitly refuses to compete commercially. The most capable offensive-AI systems that exist are not for sale — every commercial vendor in Tiers 1–3 is implicitly benchmarked against a ceiling it cannot buy access to. See What the frontier labs do themselves.

#What an engagement actually costs

None of the three highest-profile autonomous pentest vendors — XBOW, Horizon3.ai, Pentera — publish list pricing; all three are enterprise-sales, quote-only (buyers-and-pricing notes; Horizon3.ai). XBOW's own language is a direct positioning shot at the traditional model: "usage-based pricing that scales with your coverage, not a fixed annual engagement," and it lists on AWS, GCP, Oracle, and Azure marketplaces. That is a deliberate attempt to change the unit of sale from "an engagement" to "coverage," not just to undercut the old unit's price.

For comparison, the market it is trying to disrupt: traditional manual web-app pentests run $5,000–$50,000 per engagement, network pentests $150–$1,000 per device, and the overall penetration-testing services market is sized at roughly $2.72B in 2026 growing at a 15.3% CAGR to $5.54B by 2031 (Mordor Intelligence). Third-party comparison sites (not the vendors themselves) cite XBOW around $4,000 per test against budget competitors like Hacktron at $350 — treat that as directional, competitor-sourced SEO pricing, not confirmed list price. See Who buys, and what they pay for the full pricing tables across the AI security stack, including how buyer size and regulatory status change the sales cycle from days (100-person company) to 3–9 months (5,000-person regulated buyer).

#Consolidation, so far

The clearest completed M&A event adjacent to this category is Check Point's acquisition of Lakera (AI red-teaming for LLM apps, not classic pentesting), confirmed via Lakera's site now reading "©1994-2026 Check Point Software Technologies Ltd." — the ~$300M deal value reported in press is [unverified]. Otherwise the offensive pure-play category — XBOW, Horizon3.ai, Pentera, RunSybil, Terra — remains unconsolidated venture-funded competition as of August 2026; no evidence surfaced of any of them acquiring each other. CrowdStrike and Palo Alto Networks, the best-capitalized pure-play cybersecurity vendors, are conspicuously not building or buying into autonomous offense — their 2025–2026 AI M&A (Pangea, SGNL, CyberArk, Chronosphere, Protect AI) is entirely defensive/AI-app-security. See Who funds this and at what price for the sector-wide funding picture.

#What this means for the defensive side

Verdict

The offensive tier of this market has raised roughly $2B+ in disclosed funding across a handful of companies (Horizon3.ai alone at $2B valuation, XBOW past $1B) while defensive AI-native code security — the category a founder building defensive tooling competes in — is a fragmented field of smaller raises against a market Gartner sizes at under $3B total for "securing AI" in 2026 (Gartner). Offense is easier to demo, easier to fund, and easier to make viral (a HackerOne leaderboard screenshot is a better pitch-deck slide than a reduced false-positive rate). That is an uncomfortable asymmetry for anyone building defense, not a reason to avoid it.

#What this means for us

  • The deterministic-core incumbents (Pentera, Horizon3.ai) prove that "keep the risky part out of the LLM call" is a viable, fundable architecture — worth studying even for a defensive product, since the same logic (deterministic core, LLM for reasoning/summarization only) reduces your own liability surface.
  • XBOW's HackerOne #1 claim is real but far more qualified than the marketing implies — see XBOW for the specific criticisms before citing it to anyone.
  • No independent, standardized benchmark exists comparing the LLM-agent pure-plays against each other. That gap is itself an opportunity: publishing a credible third-party evaluation methodology (see Cyber benchmarks and evals) is currently rare enough in this space to earn outsized attention when done well.
  • The frontier labs (Anthropic, Google) are the actual ceiling on offensive capability and they gate access rather than sell products — plan around a world where the most dangerous offensive AI never reaches the open market, but where a leaked or independently-replicated fraction of it eventually will. See Dual-use risk and what it costs you and Germany: §202c and the Berlin question.
  • Offensive vendors treat model-provider identity as commercially sensitive and mostly don't disclose it — Horizon3.ai's Claude/LLaMA/Mistral-via-Bedrock statement is the exception, not the norm. Don't assume silence means a hard blocking conflict with a model provider; it more likely means competitive secrecy.
  • The capital asymmetry between offense and defense is a real signal about where investor attention is, not a verdict on which is more valuable to build — see What could actually be defensible and Where the gaps actually are for where the defensive gaps actually sit.