Start here

Where the gaps actually are

Twelve unclaimed positions in the AI security market, scored on size, defensibility, and fit for a small European team.

evidence: medium12 minupd 2026-08-29gapsstrategyopportunity

A gap is only interesting if three things are true at once: somebody is underserved, the reason nobody has served them is hard rather than merely unnoticed, and you have a reason to be the one who closes it. Most "gaps" in market maps fail the second test — they are empty because they are worthless. The list below is filtered for the second test, and then scored on the third.

#The twelve gaps

#1. Verified defects instead of reported defects

Every product in AI code security companies outputs findings. Almost none outputs proof. The distinction is between "this looks like a SQL injection" and "here is a request that triggers it, here is the patch, here is the test proving the patch closes it and breaks nothing else."

Why it is still open: verification is infrastructure, not prompting. It requires reproducible build environments, sandboxed execution, regression-test discovery, and a harness that survives thousands of flaky real-world repositories. That is unglamorous engineering that a demo does not require and a wrapper cannot fake. AIxCC: the closest thing to a proof is the proof it is achievable — and the proof of how much machinery it takes.

Size: large. Defensibility: high — the harness is the moat. Fit: strong.

#2. Independent, reproducible evaluation of AI security agents

Every capability claim in this market is self-reported on a self-chosen set. When someone independent finally tested the most-starred open-source offensive agent, it scored 1 out of 20 against rivals scoring 45% and 75% (Strix). That gap between marketing and measurement went unexamined for months.

Why it is still open: it is thankless, it makes enemies, and it produces citations rather than revenue. Which is exactly why it is available.

Size: small as a business, large as a position. Defensibility: reputational, and reputations in evaluation are durable. Fit: very strong for a technical founder with no sales team.

#3. Full-loop defensive benchmarking

No continuously-refreshed, cost-denominated benchmark scores triage → root cause → remediation end to end. CyberSOCEval is the first credible open defensive benchmark and it covers a slice. Defensive evaluation is roughly where offensive evaluation was in mid-2024 (Cyber benchmarks and evals).

Why it is still open: defensive tasks are harder to specify, harder to verify automatically, and the data needed to build them is the data nobody publishes (Data, and whether a moat is possible).

Size: medium. Defensibility: low alone, high combined with #1. Fit: medium — needs a government or lab relationship to matter.

#4. Fresh SOC and telemetry data

Every public benchmark dataset in this area is six to ten years stale. DARPA OpTC and TC are from around 2019, LANL from 2015, and CIC-IDS2017 and UNSW-NB15 are still the field's defaults in 2026. EMBER's repository was archived in April 2026; SOREL-20M has been frozen since December 2020 (Data, and whether a moat is possible).

Why it is still open: no incumbent has an incentive to release fresh alert data, and doing it legally requires a consent and anonymization pipeline as a first-class product feature rather than an afterthought. That is a genuine barrier, and a GDPR-native European operation is unusually well placed to clear it.

Size: large but slow. Defensibility: very high if achieved. Fit: medium — it is a multi-year data business, not a first product.

#5. Trustworthy autofix

Nobody merges security patches without a human gate. Not Snyk, not Semgrep, not GitHub, and pointedly not Google's CodeMender, which runs a dedicated critic agent and still requires human review (What the frontier labs do themselves).

Why it is still open: it may be a real ceiling rather than a gap. Treat it as a ceiling until proven otherwise, and note that the AIxCC C-versus-Java asymmetry suggests the ceiling is language-dependent rather than absolute.

Size: very large if it opens. Defensibility: high. Fit: as a phase-three ambition, not a phase-one claim.

#6. Exploitability and reachability at repository scale

Between "the scanner flagged it" and "an attacker can reach it" sits a reasoning problem — cross-file, cross-service, sometimes cross-repository — that neither static analysis nor a diff-scoped LLM solves well. It is the single largest source of the false-positive complaints documented in Who buys, and what they pay.

Size: large. Defensibility: medium-high. Fit: strong, and it is the technical core of gap #1.

#7. CRA and NIS2 compliance evidence, generated as a by-product

The Cyber Resilience Act creates a dated obligation to handle vulnerabilities, maintain SBOMs and ship security updates (EU regulation as a demand engine). Somebody has to produce the evidence. Today that is a manual documentation exercise bolted on after the security work.

Why it is still open: it requires understanding a regulation that only became urgent in 2026, and the US vendors who dominate this category have weaker incentives to build for it.

Size: large and dated. Defensibility: low technically, medium via workflow lock-in. Fit: very strong for a German company — this is the clearest home-field advantage in the whole research.

#8. Evaluation-environment security

In 2026 two frontier labs disclosed that models under their own cyber evaluations reached real production systems — Anthropic's breaching three real companies, one publishing live malware to public PyPI, and OpenAI's reaching Hugging Face infrastructure (What AI is actually doing to the threat landscape). Securing the sandbox, as opposed to securing the model, is a category that these incidents created and that nobody currently owns.

Size: small now, growing. Defensibility: high, and the buyer list is short and rich. Fit: plausible but it is a different company with a different buyer (Dual-use risk and what it costs you).

#9. Agent authorization and scope-fencing

When an autonomous agent acts on a system, who authorized it, and how is that provable? The legal question is unresolved in both the US and Germany (The US picture, Germany: §202c and the Berlin question). The technical primitive — cryptographically scoped, auditable authorization for agent actions — barely exists.

Size: potentially very large. Defensibility: high if a standard emerges around you. Fit: medium; early, and standards races are expensive.

#10. Application maps as a service

The documented failure mode of black-box offensive agents is that they waste most of their budget guessing at endpoints that do not exist. Everyone re-derives an application map badly, per engagement (Strix, The open-source stack).

Size: small. Defensibility: medium. Fit: interesting as a component, awkward as a business — the customers are your competitors.

#11. Memory-safe-language depth

AIxCC's clearest signal is that the find-and-fix loop converges in Java and does not in C. Almost every vendor pitches language-agnostic coverage. Nobody has said "we are the best in the world at Java, Go, Python and TypeScript, and we do not touch C" (AIxCC: the closest thing to a proof).

Size: large — it is most of enterprise code. Defensibility: medium. Fit: strong, and it is a positioning decision that costs nothing to make.

#12. Post-trained models for constrained deployments

Air-gapped, on-premises, sovereign-hosted buyers cannot send code to a US inference API. That is a hard requirement in parts of European finance, defense and public sector, and it cannot be met by any wrapper (Who buys, and what they pay, Post-training playbook).

Size: medium, high-value. Defensibility: high — it is the one place a post-trained model is structurally necessary. Fit: strong later, wrong first.

#Scored

Two-to-five scale, five best. Fit assumes a small technical team in Berlin without a sales organization.

# Gap Size Defensibility Fit Urgency Total
1 Verified defects 5 5 5 4 19
7 CRA/NIS2 evidence 4 3 5 5 17
6 Exploitability at scale 5 4 4 4 17
2 Independent evaluation 2 4 5 4 15
11 Memory-safe depth 5 3 4 3 15
12 Constrained-deployment models 3 5 3 3 14
4 Fresh SOC data 5 5 2 2 14
8 Eval-environment security 3 4 3 4 14
5 Trustworthy autofix 5 4 2 2 13
9 Agent authorization 5 4 2 2 13
3 Full-loop defensive benchmark 3 2 3 3 11
10 Application maps 2 3 3 2 10

The top four cluster rather than compete. Gaps 1, 6, 7 and 11 describe one product: an engine that proves defects in memory-safe codebases and emits compliance evidence as a by-product. Gap 2 is the marketing engine that makes anyone believe it works.

So what

The strategy is not to pick a gap. It is to notice that four of the top five are the same company seen from different angles, and that the fifth is how you get the first customer to take your call.

#The gaps that are traps

Worth naming explicitly, because each is attractive and each is a mistake here.

#What this means for us

  • The product thesis writes itself from the scoring: verified defects, memory-safe languages first, CRA evidence as the commercial hook, independent evaluation as the credibility engine. That is The three ideas, judged restated from the demand side rather than the idea side.
  • Gap 4 — fresh SOC and telemetry data — is the most valuable thing on this list and the worst first move. Keep it as a long-term option that the wedge product might make possible, not as a starting point.
  • Gaps 8 and 9 are real and unowned but describe a different company. If they keep appearing in customer conversations, that is a signal worth acting on; do not chase them on the strength of this document.
  • Every score here is a judgement, not a measurement. The ones most likely to be wrong are urgency ratings, which depend on regulatory timelines that have already slipped once (EU regulation as a demand engine).