Open questions and the research backlog
A prioritised backlog of what desk research could not settle, with a cheap next step and a decision each answer would change.
Desk research answers "what happened." It cannot answer "what happens next," "would this specific customer actually pay," or "what would a prosecutor actually do." This page separates the two kinds of unknown: things a cheap follow-up search or fetch could resolve, and things that require a phone call, a pilot, or a shipped product. Every item below states what decision hinges on the answer, because an open question with no attached decision is trivia, not a backlog.
#Market questions
Is the "gated frontier model" wedge a durable moat, or a relationship-and-timing advantage that evaporates once you're inside? Is frontier-lab gating a real wedge? found that Anthropic's Mythos-tier access and OpenAI's Trusted Access for Cyber both exist and are real, but neither has a documented, open-to-all-qualified-startups application process — the clearest case (XBOW) got in via direct invitation, not a form. Why it matters: if the moat is "become a trusted lab partner," the whole go-to-market plan should be a relationship-building campaign aimed at frontier-lab BD teams, not a product roadmap. How to check it cheaply: this is not desk-research-answerable (see below) — it requires actually applying.
Will the frontier labs' own security products (Codex Security, Claude Security, CodeMender) commoditize the entire "AI reviews your PR for vulnerabilities" category before a startup can build distribution? Aardvark went from invite-only beta to broad enterprise availability in about four months (Is frontier-lab gating a real wedge?). Why it matters: determines whether the defensive-code layer is investable at all, or only investable one level up the stack (orchestration, precision-tuning, compliance-specific workflows the labs won't build). Cheap check: track Codex Security's and Claude Security's public pricing and feature announcements monthly; the moment either ships autofix-with-auto-merge at enterprise scale, the thin-wrapper part of this market is closed.
What is a realistic willingness-to-pay at 500–1,000 developers, where none of the vendors publish real pricing? Who buys, and what they pay has solid data at 50–100 developers (roughly $45–100K/year, Snyk and Semgrep converge on a similar band) but nothing credible above that — enterprise deals go custom and undisclosed. Why it matters: a business plan's TAM calculation is only as good as its largest-account assumption. Cheap check: pull three to five real RFP responses or Vendr-style deal data for GitHub Advanced Security or a comparable platform at 1,000+ seats.
Is there a genuine Chinese offensive-AI-security startup sector, or does that activity simply not surface in English-language sources? Offensive AI security companies found essentially no VC-funded Chinese pentest-AI startups despite deliberate searching — a real gap, not confirmed absence. Why it matters: changes the competitive-threat model for anyone selling into markets where Chinese tooling might undercut on price. Cheap check: a Chinese-language search pass (Baidu, WeChat public accounts, 36Kr) specifically for "AI 渗透测试" (AI pentesting) funding announcements — genuinely five minutes of work this research pass did not do.
Does a defensive/blue-team AI product category actually exist yet, commercially, or is it pre-product-market-fit dressed as a market? Cyber benchmarks and evals found offense has six mature benchmarks and five frontier-lab internal suites; defense has essentially one (CyberSOCEval, published Sept 2025) and a scattering of vendor marketing claims. Why it matters: "biggest hole in the market" and "no market yet" are different diagnoses requiring opposite strategies. Cheap check: count actual paying-customer logos (not pilot logos) across the AI-SOC vendors in AI SOC and detection companies with a public case study naming a Fortune 500 buyer — a low but real bar.
#Technical questions
Does post-training an open model on defensive-security tasks ever beat prompting a frontier model with good scaffolding, at the volumes a real product would run? Post-training playbook found scaffold choice alone produces up to 2.6x variance on identical models (CAIBench), suggesting the first engineering dollar should go to harness quality, not weight updates — but nobody has published a controlled head-to-head at production inference volume. Why it matters: determines whether "train our own model" is ever the right early move, or purely a later cost-optimization once volume justifies it. Cheap check: run the same eval suite (CyberGym or Repair-CVE-Bench) against a well-scaffolded frontier model and a naively-prompted one, holding scaffold constant — this alone would produce a genuinely new data point nobody in the research corpus has published.
Is there a clean, non-gameable reward signal for defensive tasks, the way "PoC crashes the target, patched build doesn't" works for offense? Post-training playbook's own verifier table shows every offense-side signal (sanitizer fires, fuzzer crash reproduces) has a known gaming failure mode that's manageable with careful design — but no equivalent table exists for triage, root-cause attribution, or IR containment, because nobody has built one at scale. Why it matters: without a clean verifier, RLVR-style post-training for a defensive agent is not currently a solved engineering problem, only an offense one is. This is the single most concrete unclaimed research contribution surfaced anywhere in this corpus.
How much of the "black-box pentest agent" failure mode (Strix finding 1/20 vulnerabilities on a bare URL vs. 15/20 for a commercial competitor with source access) generalizes across the whole category, or is it Strix-specific? The open-source stack found this from a single independent benchmark (Escape.tech). Why it matters: if it generalizes, "point our agent at a URL with no source access" is not a credible product claim for anyone in Offensive AI security companies, and marketing claims across the sector need serious discounting. Cheap check: rerun the same Escape-style benchmark against two more agents (PentAGI, one commercial vendor) under identical black-box conditions.
What is the real, independently-measured false-positive rate of any AI code-review or pentest tool, at scale? Every number in AI code security companies (Semgrep's 96% agreement, Greptile's 82% catch rate, Aardvark's 92% recall) is self-reported on the vendor's own benchmark. The one independent number found — the disputed arXiv code-review study — could not even be confirmed to exist (see Verification ledger). Why it matters: a company's entire differentiation claim ("we have fewer false positives") is currently unverifiable industry-wide, including for a competitor's claims. This can only be answered by building the SWE-bench-equivalent standardized benchmark nobody in the space has built yet (see Cyber benchmarks and evals Q6).
#Legal questions
Will §202c StGB actually be reformed, and on what timeline? Germany: §202c and the Berlin question confirmed the statute is unreformed as of Aug 2026 despite a 2025 coalition-agreement commitment; no Referentenentwurf could be located. Why it matters: directly determines whether a Berlin-incorporated entity can safely build or distribute exploit-generating capability, or must route that product line through a US entity from day one. Cheap check: this needs a person, not a search — see below.
Does the EU AI Act's systemic-risk designation actually apply to a fine-tuned/post-trained cyber-offense model that doesn't cross the 10^25 FLOP compute threshold on its own training run? EU regulation as a demand engine found this is a genuinely contested interpretation — the Commission retains discretionary designation power below the compute threshold, and a model unusually effective at autonomous exploitation is a plausible candidate. Why it matters: determines whether Article 55's adversarial-testing and incident-reporting obligations apply to a defensive lab's own red-team tooling, not just to frontier labs. This needs EU counsel, not more search.
Who bears CFAA liability when an autonomous agent, not a human, drifts outside an authorized scope of work? The US picture confirmed no appellate authority resolves this post-Van Buren. Why it matters: is the single biggest reason to over-invest in hard technical scope-fencing (not prompt-level instructions) for any offensive-testing product — the legal answer isn't coming in time to inform the engineering decision, so the engineering has to assume the worst case.
Does publishing a trained-on-exploits model's weights, or a benchmark containing working exploit code, trigger EU or US dual-use export controls? Flagged but not resolved in EU regulation as a demand engine and Data, and whether a moat is possible. Why it matters: affects the release strategy for any open research artifact the lab wants to publish for credibility (see The open-source stack on garak/Opengrep as a credibility play) — a release that seemed like a pure reputation win could carry real compliance exposure.
#Questions only reachable by talking to people
Desk research hit a hard wall on anything requiring a first-hand account rather than a public document. The clearest gap: no on-the-record account exists anywhere in this corpus of a security startup saying "a frontier lab's product killed our business" or "our account got flagged/rate-limited by Anthropic or OpenAI" — not because it doesn't happen, but because founders don't want to publicly antagonize a platform they depend on, and because the research tooling used here (WebSearch exhausted, Reddit blocked) could not surface private complaints.
| Who | What to ask |
|---|---|
| XBOW, RunSybil, Strix maintainers, MindFort | How did you actually get Mythos/Astra-tier access — application, invitation, or relationship? Has a frontier lab's own product ever displaced a deal you were closing? |
| Anthropic/OpenAI BD or partnerships teams | What's the real eligibility bar for Glasswing / Trusted Access for Cyber for a company that isn't critical infrastructure or a government? |
| Working CISOs / heads of AppSec at 500+ dev companies | What would you actually pay for autonomous vuln finding-and-fixing at your scale? What's blocking you from trusting an autonomous fix today? |
| Daniel Stenberg (curl), or another maintainer living the AI-slop bug-bounty problem | Has the noise/signal ratio actually stabilized, or is the 2025 "closed the bounty temporarily" episode going to repeat? |
| A German cyber-law practitioner (e.g., referencing Kipker's published scholarship) | Is a §202c reform bill actually moving, or is the 2025 coalition commitment stalled? What would a defensible corporate structure look like today? |
| HackerOne, Bugcrowd program leads | Is the researcher-data-provenance backlash (Feb 2026, per Offensive AI security companies) actually costing platforms researcher trust, or did it blow over? |
| A Series A investor active in this space (e.g., a fund from Who funds this and at what price) | What due diligence do you actually run on a "we post-train our own model" claim, given how unverifiable most of these claims turned out to be in this research? |
#What cannot be answered by desk research — must be built
Four things in this backlog are not resolvable by more searching, no matter how much web-search budget is restored:
- Whether a defensive verifier design actually works. The only way to know if a given reward signal for triage/IR is gameable is to build it, post-train against it, and see what the model learns to exploit — exactly the RLVR crux Post-training playbook identifies as unsolved.
- Whether a real customer will trust an autonomous fix without human review. Every vendor in AI code security companies currently keeps a human merge gate; nobody has published evidence customers would accept otherwise. Only a live pilot answers this.
- What frontier-lab access actually looks like from the inside. No public document describes the eligibility bar for trusted-access programs with precision — this can only be learned by applying and seeing what happens.
- The real economics of post-training vs. scaffolding at production volume. Post-training playbook's cost model ($100–200K to a released checkpoint) is a budget estimate, not a demonstrated ROI — only running both paths against the same workload for months produces a real number.
#The three questions that would most change the plan
- Does the Glasswing/Trusted-Access tier ever open to funded startups outside critical infrastructure and governments, and on what basis? A yes reframes the entire strategy around winning that relationship early; a no means competing purely on the standard API tier, where Is frontier-lab gating a real wedge?'s evidence says the gap is already closing fast.
- Is the open-weight cyber capability gap (currently 4–7 months per UK AISI) closing fast enough that betting on open models is viable within a 2-year building window, or does it plateau? This determines whether to build model-agnostic from day one or make an early, possibly-premature bet on open weights for cost and deployability reasons.
- Does §202c StGB get reformed, and on what timeline? This is the one legal question with a clean binary answer that directly determines corporate structure, product-line sequencing, and whether "build in Berlin" is compatible with an offense-adjacent product line at all — see The first 90 days.
#What this means for us
- Treat "we couldn't find it" as a task, not a conclusion — several of the highest-value answers here (Glasswing eligibility, §202c's real trajectory) are one phone call away, not one search query away.
- Do not finalize a corporate structure decision until the §202c question is answered by counsel, not inferred from a coalition-agreement press release.
- Build the black-box benchmark rerun and the verifier-design experiment early — both are cheap, both produce genuinely novel data nobody else in this market has published, and both directly inform whether to compete on offense or defense first.
- Prioritize outreach to XBOW/RunSybil/Strix and to a frontier-lab BD contact in the first 90 days — the single most decision-relevant unknown (real Glasswing-tier eligibility) cannot be resolved any other way.
- Revisit this backlog every quarter; several entries (open-weight gap, frontier-lab product commoditization) are moving targets whose answer will change under you.