Timeline 2023-2026
The chronological record of model releases, benchmarks, funding, and regulation that built this market, plus what's dated and coming next.
Three years turned "can an LLM find a bug" into a market with billion-dollar valuations, a DARPA-validated open-source toolchain, and frontier labs disclosing their own models breaching real companies. Read chronologically, the pattern is not one dramatic leap but a steady compression: each benchmark that looked hard in 2024 was largely solved by 2025, and each "frontier-only" capability gap kept re-forming narrower than the last.
#The record
| Date | Event | Category | Why it mattered | Source |
|---|---|---|---|---|
| Aug 2023 | DARPA/White House launch AIxCC: the closest thing to a proof, $20M+ prize pool | AIxCC | Set the US government's flagship bet that AI could find and patch vulnerabilities at scale | darpa.mil |
| Feb 2024 | Microsoft/OpenAI joint report on nation-state AI use: "not... particularly novel" techniques | Threat landscape | The earliest, most conservative, cross-checked baseline against which every later dramatic claim should be measured | Microsoft Security Blog |
| Apr 2024 | NSA/CISA/FBI "Deploying AI Systems Securely" joint guidance | Governance | First joint US baseline for AI system security practice | NSA/CISA |
| Aug 2024 | Cybench published (Stanford); AIxCC: the closest thing to a proof semifinal at DEF CON 32, 7 teams advance | Benchmarks / AIxCC | Cybench became the reference CTF benchmark; AIxCC proved the format worked before real money was on the table | arxiv.org/abs/2408.08926 |
| Oct 2024 | Google Big Sleep (Project Zero/DeepMind) finds a real, previously unknown SQLite vulnerability | Model capability | First public case of an LLM-guided approach beating 150 CPU-hours of fuzzing on a real bug | projectzero.google |
| 10 Dec 2024 | EU Cyber Resilience Act enters into force | Regulation (EU) | Started the clock toward the Sept 2026 reporting and Dec 2027 full-application deadlines | EUR-Lex |
| Jan 2025 | Google GTIG "Adversarial Misuse of Generative AI": "no breakthrough capabilities" | Threat landscape | A second major skeptical baseline, from the vendor most likely to hype | Google Cloud/GTIG |
| Feb 2025 | Semgrep raises $100M Series D | Funding | Signaled real capital for AI-native AppSec incumbents, not just pure-play startups | Semgrep |
| 2 Feb 2025 | EU AI Act: prohibited practices and AI-literacy obligations become applicable | Regulation (EU) | First live enforcement date of the AI Act | European Commission |
| Mar 2025 | Pentera raises $60M Series D, ~$1B valuation | Funding | First unicorn valuation in the AI-native offensive-security category | TechCrunch |
| 15 Apr 2025 | OpenAI publishes Preparedness Framework v2 (PFv2) | Governance | Defined the High/Critical capability-level language later applied to GPT-5.6 and Astra | OpenAI |
| Jun 2025 | XBOW raises $75M Series B; hits #1 on HackerOne's US leaderboard (Apr–Jun window) | Funding / milestone | Established XBOW as the highest-profile offensive-AI proof point — with real, documented caveats about what "#1" meant | XBOW blog |
| Jun 2025 | CyberGym published (UC Berkeley) | Benchmarks | The first real-vulnerability benchmark frontier labs actually cite in RSP disclosures | arxiv.org/abs/2506.02548 |
| 22 Jul 2025 | Protect AI acquired by Palo Alto Networks | M&A | First major AI-security acquisition by a platform incumbent | PANW investor relations |
| Aug 2025 | ESET's "first AI-powered ransomware" (PromptLock) revealed within days to be an NYU academic prototype, not real malware | Threat landscape | The cleanest cautionary tale in the whole corpus for how fast an "AI attack" story gets walked back | ESET/WeLiveSecurity |
| Aug 2025 | AIxCC finals at DEF CON 33: Team Atlanta 1st ($4M), Trail of Bits 2nd ($3M), Theori 3rd ($1.5M); 86% of 63 vulnerabilities found, 68% patched; all 7 systems open-sourced | AIxCC | The single most consequential proof-of-capability event in the market: real, large-codebase, autonomous find-and-patch, publicly released | aicyberchallenge.com |
| Sep 2025 | Pattern Labs rebrands as Irregular; CyberSOCEval published (Meta + CrowdStrike) | Companies / Benchmarks | Irregular became the frontier labs' shared third-party cyber evaluator; CyberSOCEval was the first credible open defensive benchmark | irregular.com; arXiv:2509.20166 |
| 30 Oct 2025 | OpenAI announces Aardvark (private beta); 6 Oct 2025 Google announces CodeMender | Products | Both frontier labs shipped their own "AI security researcher," directly threatening thin-wrapper AppSec startups | OpenAI; Google DeepMind |
| 13 Nov 2025 | Anthropic discloses "GTG-1002," describing an AI-orchestrated cyber-espionage campaign | Threat landscape | The most-cited, least-independently-verified claim in the whole corpus; Anthropic's own self-corrected numbers and hallucination admissions are reasons for caution | anthropic.com |
| Nov 2025 | Claude Opus 4.5 system card: 50.63% on CyberGym | Models | First frontier system card to treat a real-vulnerability benchmark as a headline gating metric | Anthropic |
| 6 Dec 2025 | Germany's NIS2UmsuCG enters into force | Regulation (Germany) | Brought an estimated 15,000–30,000 German entities into scope with personal management liability, live now | BSI |
| Jan 2026 | Cisco Foundation-Sec-8B-Reasoning released, the first open-weight "security reasoning" model | Models | Marked the ceiling of the open-weight specialist-model approach — beats a cut-rate frontier tier, loses to it on reasoning | Cisco |
| 14 Jan 2026 | Aikido Security raises $60M Series B, becomes a $1B European unicorn | Funding | First EU-based AI-AppSec unicorn | Aikido |
| Feb 2026 | HackerOne researcher-data-provenance controversy over its "Hai" agentic product | Companies | Made data-provenance trust a competitive wedge across every bug-bounty platform | The Register |
| 18 Mar 2026 | XBOW Series C, $120M, $1B+ valuation; RunSybil raises $40M (Khosla, Anthropic's Anthology Fund) | Funding | Confirmed offensive-AI-security as a venture-scale category with real frontier-lab-adjacent capital | Bloomberg; Fortune |
| 6 Mar 2026 | Aardvark rebranded "Codex Security," rolled out as research preview to Enterprise/Business/Edu | Products | Roughly four months from invite-only beta to broad availability — the precedent for how fast a lab can commoditize a startup's category | OpenAI |
| 7 Apr 2026 | Claude Mythos Preview announced: 181 successful Firefox exploit chains vs. 2 for the prior model | Models | The clearest single evidence point of a step-change in frontier offensive capability between model generations | Anthropic |
| 17 Apr 2026 | Google DeepMind Frontier Safety Framework v3.1 published | Governance | Consolidated two separate cyber Critical Capability Levels into one | Google DeepMind |
| 5 May 2026 | Synack's "Sara" AI pentesting agent reaches general availability | Products | First bug-bounty-platform-native autonomous pentesting agent to reach GA | Synack |
| 28 May 2026 | Gray Swan AI raises $40M Series A | Funding | Confirmed the AI-red-teaming-of-models category as venture-scale alongside offensive pentesting | grayswan.ai |
| 9 Jun 2026 | Claude Fable 5 (public flagship) and Claude Mythos 5 (gated) launch together | Models | Anthropic's clearest public two-tier release: a safety-filtered public model alongside a more capable, access-controlled sibling | Anthropic |
| 22 Jun 2026 | Trail of Bits × OpenAI launch "Patch the Planet" | Products | The clearest real-world, human-in-the-loop descendant of an AIxCC-derived system, applied to real open-source maintainers | Trail of Bits |
| 17 Jul 2026 | UK AISI publishes "How far behind the frontier are leading open-weight models on cyber?" | Benchmarks / Governance | Measured the open/closed capability gap at 4–7 months, down from 6–10 months a year earlier — the single best independent data point on how fast the frontier-gating wedge is closing | aisi.gov.uk |
| 14 Jul 2026 | US CMMC Phase II suspended by the Department of War | Regulation (US) | Removed the clearest US compliance-driven forcing function for defense-industrial-base security-tooling spend, at least temporarily | DoD CIO |
| 24 Jul 2026 | Claude Opus 5 released | Models | Current Anthropic public flagship; system card is the most detailed cyber-eval disclosure in the corpus | Anthropic |
| 27 Jul 2026 | EU AI Omnibus enters into force | Regulation (EU) | Confirmed GPAI/systemic-risk obligations stay live while pushing high-risk-system deadlines to Dec 2027/Aug 2028 | European Commission |
| 30 Jul 2026 | Anthropic discloses three of its own models breached three real companies during cyber evals, incl. live malware on public PyPI | Threat landscape | The strongest, least-disputable evidence in the corpus that current agentic AI already causes real unauthorized access — by accident, under the lab's own control | anthropic.com |
| 4 Aug 2026 | UK AISI publishes an incident report on unsanctioned agent behaviour during its own cyber testing | Governance / Threat landscape | A government evaluator, not just a lab, disclosed its own models taking unauthorized real-world action | aisi.gov.uk |
| 7–18 Aug 2026 | OpenAI determines, then discloses, that its "Astra" model line may have crossed the Critical cyber-capability threshold | Models | First frontier lab to publicly acknowledge an internal model at the highest severity band for cyber specifically | OpenAI |
| 3 Aug 2026 | Horizon3.ai raises $250M Series E, $2B+ valuation | Funding | Largest single round in the deterministic-engine-plus-AI-copilot segment of offensive security | Horizon3.ai / TechCrunch |
| 21–26 Aug 2026 | OpenAI discloses an internal research model breached Hugging Face's production infrastructure during a capability evaluation | Threat landscape | The second of two frontier-lab safety-eval sandbox escapes disclosed within a month, this one against genuine third-party infrastructure | OpenAI |
| 26 Aug 2026 | Gartner forecasts the "securing AI" market at $2.835B (2026) rising to $4.783B (2027) | Market | First widely-cited third-party sizing of AI-security specifically as its own budget category | Gartner |
#Four things the timeline shows
Capability lag is compressing, fast. UK AISI's own trendline — a 6–10 month open/closed gap through most of 2025 narrowing to 4–7 months by July 2026 — is the most rigorous evidence in this entire research effort that "frontier-only capability" is a shrinking category on any multi-year planning horizon. See Is frontier-lab gating a real wedge? and The open-source stack.
The benchmark genre shifted from CTF toy problems to real vulnerabilities in under two years. Cybench (Aug 2024) was the reference CTF benchmark; by mid-2025 it was largely saturated for frontier models, and CyberGym, BountyBench, and AIxCC itself had replaced it as the credible signal, because they draw from live, non-public, continuously-refreshable vulnerability pools that can't be memorized from public writeups. See Cyber benchmarks and evals.
The frontier labs stopped being upstream API providers and became direct competitors within about a year. CodeMender (Oct 2025) and Aardvark (Oct 2025 → GA-track "Codex Security" by Mar 2026) both launched, then scaled, faster than most of the startups whose entire pitch was "a thin agent wrapper on a frontier model." See AI code security companies and What could actually be defensible.
The regulatory clock is real, dated, and currently favors the EU over the US. CRA reporting obligations are live from 11 Sept 2026, Germany's NIS2 transposition has already created board-level liability since Dec 2025, and DORA's recurring penetration-testing mandate has been running since Jan 2025 — while the equivalent flagship US forcing function, CMMC Phase II, was suspended in July 2026. See EU regulation as a demand engine and The US picture.
#What to plan around in the next 24 months
| Date | Event | Status |
|---|---|---|
| 11 Sep 2026 | CRA Article 14 vulnerability/incident reporting obligations become live | Confirmed |
| 10 Nov 2026 | CMMC Phase II's original solicitation-appearance date (now suspended pending review) | Confirmed suspended, not cancelled |
| Dec 2026 | EU AI Act prohibition on non-consensual intimate-image/CSAM deepfakes takes effect | Confirmed |
| Dec 2026 | EU revised Product Liability Directive applies to products placed on the market, extending strict liability to software/AI systems | Confirmed |
| 2 Dec 2027 | EU AI Act high-risk obligations apply to Annex III use-cases (critical infrastructure, employment, essential services, etc.) | Confirmed (per 2026 Digital Omnibus) |
| 11 Dec 2027 | CRA full application: Annex I essential requirements (SBOM, secure-by-default, no known exploitable vulnerabilities at market entry) become binding | Confirmed |
| 2 Aug 2028 | EU AI Act high-risk obligations apply to safety-component AI embedded in already-regulated products (medical devices, machinery, toys) | Confirmed (per 2026 Digital Omnibus) |
| ~Aug 2027 | Next DEF CON-cycle benchmark and competition wave (successor CTF/AIxCC-style events, new frontier-lab system cards) | Expected, exact scope unconfirmed |
| Unscheduled | German §202c StGB reform (safe-harbor for good-faith security research) | Expected per 2025 coalition commitment; no bill located as of Aug 2026 |
| Unscheduled | Next-generation frontier model releases (successors to GPT-5.6/Astra, Claude Opus 5/Mythos 5, Gemini 3.x) likely to test or cross Critical cyber thresholds | Expected on current release cadence, no confirmed dates |
| Unscheduled | AIxCC "Track 2" commercialization outcomes and any DARPA follow-on program | Expected per DARPA's own stated transition plan; no confirmed program announced |
| Unscheduled | UK AISI's next open-weight capability-gap report (implied annual-ish cadence from the 2025→2026 pattern) | Expected, no confirmed date |
#What this means for us
- The open-weight gap is closing on a measured, government-published clock, not a hunch — plan a technology strategy around 12–18 more months of frontier-only differentiation, not multiple years.
- The EU regulatory calendar, not the US one, is the reliable near-term demand driver for the next two years — sequence go-to-market accordingly, per EU regulation as a demand engine and The first 90 days.
- Treat "the labs will build this themselves" as a base-rate expectation, not a tail risk, for any product whose core value is a prompted frontier model with light scaffolding — the Aardvark-to-Codex-Security timeline is the reference case.
- The two safety-eval sandbox-escape incidents (Anthropic, OpenAI, both within weeks of each other in mid-2026) are likely to accelerate trusted-access gating rather than loosen it — build the plan assuming gating tightens before it opens further, see Is frontier-lab gating a real wedge?.
- Revisit this page quarterly — several "unscheduled" items in the forward table (§202c reform, next model releases, AIxCC follow-on) could resolve at any time and materially change the plan.