Start here

Executive summary

The gating thesis is half right and closing. The defensible position is verification, not detection, sold into the EU compliance clock.

evidence: medium9 minupd 2026-08-29thesisstrategydecision

Three weeks of desk research across competitors, models, benchmarks, capital and law produces one uncomfortable finding and one genuine opening.

The uncomfortable finding: the premise this project started from — the frontier labs guard their cyber models, so getting on the right side of that gate is a wedge — is half right, and the half that is right is closing on a published clock. The genuine opening: almost every product in this market reports findings, and almost none of them report verified findings. That gap is technical, it is expensive to close, it is not something a frontier lab ships as a checkbox, and in Europe there is a regulation with a date on it that turns verified findings into a compliance artifact somebody has to buy.

So what

The market has an abundance of detection and a scarcity of proof. Every buyer complaint in Who buys, and what they pay, every benchmark gap in Cyber benchmarks and evals, and the single most striking result in AIxCC: the closest thing to a proof all point at the same missing thing: a system that can say "this is real, here is the reproduction, here is the fix, here is the evidence the fix works" and be right often enough to be trusted.

#What the research actually found

On the gating thesis. There is a hard block, and it is narrower than the premise assumes. Anthropic's Mythos-class and OpenAI's Astra-class models are genuinely unavailable to anyone through any self-serve channel — that part is real. But most of what founders call "gating" turns out to be over-refusal (thinly evidenced), enterprise agreements that already solve it, or cost. Meanwhile UK AISI's own measurements put open-weight models 4–7 months behind the frontier on cyber, down from 6–10 months a year earlier, at up to 45x lower cost. RunSybil reproduced top AIxCC results on standard, non-gated models for roughly $600. A business whose defensibility rests on the gate staying shut is underwriting a trend line that is moving the wrong way. Full argument in Is frontier-lab gating a real wedge?.

On the competitive field. Offensive AI security is well funded and loud — XBOW past $1B, Horizon3 raising at $2B+ — and it is not where a Berlin-based defensive lab should fight (Offensive AI security companies). Defensive code security is crowded at the surface and empty underneath: PR-comment review is now free from GitHub and Cursor and cheap from CodeRabbit and Greptile, while exploitability validation, whole-repo reasoning and trustworthy autofix remain unsolved by everyone including OpenAI and Google (AI code security companies). The AI SOC segment is capital-rich, integration-heavy, and the worst possible fit for a research-capability team (AI SOC and detection companies).

On what autonomous find-and-fix can actually do. AIxCC is the best public evidence that exists, and it contains one fact that should reshape any roadmap: across the finals, systems found 86% of 63 synthetic vulnerabilities and patched 68% of them — but of the real-world bugs they found, every single patch that worked was in Java, and zero C bugs were successfully patched. Memory-safe languages are tractable today. C and C++ memory-safety remediation is not. See AIxCC: the closest thing to a proof.

On post-training. Nobody has publicly shown a post-trained model beating a frontier model plus good scaffolding on agentic security tasks — not Cisco, not Trend Micro, not anyone on Hugging Face. Cisco's own Foundation-Sec-8B-Reasoning loses to GPT-5-Nano on Cisco's own benchmark. What the download data does show is that the market buys classifiers and safety layers, not monolithic "cyber brain" chatbots: SecBERT and Llama Guard 4 out-download every generative cyber model by an order of magnitude (Open-weight security models). Post-training wins on deployment constraints and unit economics at volume — not on capability (Post-training playbook, Unit economics and the compute bill).

On the law, which is the most underrated finding here. The EU Cyber Resilience Act starts biting on 11 September 2026 — thirteen days from this writing — with Article 14 vulnerability and incident reporting, and applies in full on 11 December 2027. It makes vulnerability handling, SBOMs and security updates mandatory for products with digital elements. That is a dated, enforceable, budget-creating obligation, and it is the single strongest structural demand driver in this entire research (EU regulation as a demand engine). Against that, Germany's §202c StGB — the Hackerparagraf — remains unreformed as of August 2026 despite a coalition pledge, which is a real constraint on where offensive capability can legally live (Germany: §202c and the Berlin question). In the US, CMMC Phase II was suspended on 14 July 2026, removing the clearest near-term US compliance deadline (The US picture).

Key numbers

#The three ideas, judged

Short form. The reasoning, and the conditions under which each judgement flips, is in The three ideas, judged.

Idea Verdict One-line reason
Post-train a defensive cyber model and win on benchmarks Not first No one has beaten frontier+scaffolding on agentic security tasks; the win is deployment and cost at volume you do not yet have
A security review agent on every PR Yes, but one layer down Layer-1 PR comments are commoditized and free; the unsolved layer is exploitability validation and verified fixes
Release your own benchmark / eval Yes as an artifact, no as a business Benchmarks generate citations and lab relationships, never revenue — make it narrow and attached to a product claim

Build a system whose output is a verified defect, not a finding: a reproduced vulnerability with a working proof, a patch, and machine-checkable evidence that the patch closes it and breaks nothing. Start with memory-safe languages, where AIxCC shows the patch loop actually converges. Sell it into EU companies facing CRA Article 14 obligations, where the same artifact doubles as compliance evidence. Use an open, reproducible benchmark and published third-party evaluations as the credibility engine — the one thing this market conspicuously lacks and that a small European lab can credibly own (Where the gaps actually are, What could actually be defensible).

Post-training comes later, when review volume makes it an economic argument rather than a capability claim, and when the verification harness has already produced the training data that makes it possible (Data, and whether a moat is possible).

Caution

This entire document was assembled by parallel research agents, several of which ran with an exhausted web-search budget. An adversarial fact-check found multiple confidently-stated claims to be false, including an invented accusation against a named company and a non-existent government institute. Every correction is recorded in Verification ledger. Read it before quoting anything here to a third party.

#What this means for us

  • The decision is not "which of my three ideas" — it is that idea two, moved one layer down and pointed at the CRA clock, is the business, and ideas one and three are inputs to it rather than alternatives.
  • The first ninety days should produce a public artifact, not a product: an evaluation nobody else has published, on a gap nobody else has measured. See The first 90 days.
  • Berlin is an asset for selling and a liability for building offensive capability. That argues for a specific corporate structure, not for moving. See Germany: §202c and the Berlin question.
  • Every number in this hub is a starting point for a conversation with a real practitioner, not a substitute for one. The highest-value unknowns are in Open questions and the research backlog, and several are one phone call away.