Reference

Open questions and the research backlog

A prioritised backlog of what desk research could not settle, with a cheap next step and a decision each answer would change.

evidence: medium9 minupd 2026-08-29backlogopen-questionsdiligence

Desk research answers "what happened." It cannot answer "what happens next," "would this specific customer actually pay," or "what would a prosecutor actually do." This page separates the two kinds of unknown: things a cheap follow-up search or fetch could resolve, and things that require a phone call, a pilot, or a shipped product. Every item below states what decision hinges on the answer, because an open question with no attached decision is trivia, not a backlog.

#Market questions

Is the "gated frontier model" wedge a durable moat, or a relationship-and-timing advantage that evaporates once you're inside? Is frontier-lab gating a real wedge? found that Anthropic's Mythos-tier access and OpenAI's Trusted Access for Cyber both exist and are real, but neither has a documented, open-to-all-qualified-startups application process — the clearest case (XBOW) got in via direct invitation, not a form. Why it matters: if the moat is "become a trusted lab partner," the whole go-to-market plan should be a relationship-building campaign aimed at frontier-lab BD teams, not a product roadmap. How to check it cheaply: this is not desk-research-answerable (see below) — it requires actually applying.

Will the frontier labs' own security products (Codex Security, Claude Security, CodeMender) commoditize the entire "AI reviews your PR for vulnerabilities" category before a startup can build distribution? Aardvark went from invite-only beta to broad enterprise availability in about four months (Is frontier-lab gating a real wedge?). Why it matters: determines whether the defensive-code layer is investable at all, or only investable one level up the stack (orchestration, precision-tuning, compliance-specific workflows the labs won't build). Cheap check: track Codex Security's and Claude Security's public pricing and feature announcements monthly; the moment either ships autofix-with-auto-merge at enterprise scale, the thin-wrapper part of this market is closed.

What is a realistic willingness-to-pay at 500–1,000 developers, where none of the vendors publish real pricing? Who buys, and what they pay has solid data at 50–100 developers (roughly $45–100K/year, Snyk and Semgrep converge on a similar band) but nothing credible above that — enterprise deals go custom and undisclosed. Why it matters: a business plan's TAM calculation is only as good as its largest-account assumption. Cheap check: pull three to five real RFP responses or Vendr-style deal data for GitHub Advanced Security or a comparable platform at 1,000+ seats.

Is there a genuine Chinese offensive-AI-security startup sector, or does that activity simply not surface in English-language sources? Offensive AI security companies found essentially no VC-funded Chinese pentest-AI startups despite deliberate searching — a real gap, not confirmed absence. Why it matters: changes the competitive-threat model for anyone selling into markets where Chinese tooling might undercut on price. Cheap check: a Chinese-language search pass (Baidu, WeChat public accounts, 36Kr) specifically for "AI 渗透测试" (AI pentesting) funding announcements — genuinely five minutes of work this research pass did not do.

Does a defensive/blue-team AI product category actually exist yet, commercially, or is it pre-product-market-fit dressed as a market? Cyber benchmarks and evals found offense has six mature benchmarks and five frontier-lab internal suites; defense has essentially one (CyberSOCEval, published Sept 2025) and a scattering of vendor marketing claims. Why it matters: "biggest hole in the market" and "no market yet" are different diagnoses requiring opposite strategies. Cheap check: count actual paying-customer logos (not pilot logos) across the AI-SOC vendors in AI SOC and detection companies with a public case study naming a Fortune 500 buyer — a low but real bar.

#Technical questions

Does post-training an open model on defensive-security tasks ever beat prompting a frontier model with good scaffolding, at the volumes a real product would run? Post-training playbook found scaffold choice alone produces up to 2.6x variance on identical models (CAIBench), suggesting the first engineering dollar should go to harness quality, not weight updates — but nobody has published a controlled head-to-head at production inference volume. Why it matters: determines whether "train our own model" is ever the right early move, or purely a later cost-optimization once volume justifies it. Cheap check: run the same eval suite (CyberGym or Repair-CVE-Bench) against a well-scaffolded frontier model and a naively-prompted one, holding scaffold constant — this alone would produce a genuinely new data point nobody in the research corpus has published.

Is there a clean, non-gameable reward signal for defensive tasks, the way "PoC crashes the target, patched build doesn't" works for offense? Post-training playbook's own verifier table shows every offense-side signal (sanitizer fires, fuzzer crash reproduces) has a known gaming failure mode that's manageable with careful design — but no equivalent table exists for triage, root-cause attribution, or IR containment, because nobody has built one at scale. Why it matters: without a clean verifier, RLVR-style post-training for a defensive agent is not currently a solved engineering problem, only an offense one is. This is the single most concrete unclaimed research contribution surfaced anywhere in this corpus.

How much of the "black-box pentest agent" failure mode (Strix finding 1/20 vulnerabilities on a bare URL vs. 15/20 for a commercial competitor with source access) generalizes across the whole category, or is it Strix-specific? The open-source stack found this from a single independent benchmark (Escape.tech). Why it matters: if it generalizes, "point our agent at a URL with no source access" is not a credible product claim for anyone in Offensive AI security companies, and marketing claims across the sector need serious discounting. Cheap check: rerun the same Escape-style benchmark against two more agents (PentAGI, one commercial vendor) under identical black-box conditions.

What is the real, independently-measured false-positive rate of any AI code-review or pentest tool, at scale? Every number in AI code security companies (Semgrep's 96% agreement, Greptile's 82% catch rate, Aardvark's 92% recall) is self-reported on the vendor's own benchmark. The one independent number found — the disputed arXiv code-review study — could not even be confirmed to exist (see Verification ledger). Why it matters: a company's entire differentiation claim ("we have fewer false positives") is currently unverifiable industry-wide, including for a competitor's claims. This can only be answered by building the SWE-bench-equivalent standardized benchmark nobody in the space has built yet (see Cyber benchmarks and evals Q6).

Will §202c StGB actually be reformed, and on what timeline? Germany: §202c and the Berlin question confirmed the statute is unreformed as of Aug 2026 despite a 2025 coalition-agreement commitment; no Referentenentwurf could be located. Why it matters: directly determines whether a Berlin-incorporated entity can safely build or distribute exploit-generating capability, or must route that product line through a US entity from day one. Cheap check: this needs a person, not a search — see below.

Does the EU AI Act's systemic-risk designation actually apply to a fine-tuned/post-trained cyber-offense model that doesn't cross the 10^25 FLOP compute threshold on its own training run? EU regulation as a demand engine found this is a genuinely contested interpretation — the Commission retains discretionary designation power below the compute threshold, and a model unusually effective at autonomous exploitation is a plausible candidate. Why it matters: determines whether Article 55's adversarial-testing and incident-reporting obligations apply to a defensive lab's own red-team tooling, not just to frontier labs. This needs EU counsel, not more search.

Who bears CFAA liability when an autonomous agent, not a human, drifts outside an authorized scope of work? The US picture confirmed no appellate authority resolves this post-Van Buren. Why it matters: is the single biggest reason to over-invest in hard technical scope-fencing (not prompt-level instructions) for any offensive-testing product — the legal answer isn't coming in time to inform the engineering decision, so the engineering has to assume the worst case.

Does publishing a trained-on-exploits model's weights, or a benchmark containing working exploit code, trigger EU or US dual-use export controls? Flagged but not resolved in EU regulation as a demand engine and Data, and whether a moat is possible. Why it matters: affects the release strategy for any open research artifact the lab wants to publish for credibility (see The open-source stack on garak/Opengrep as a credibility play) — a release that seemed like a pure reputation win could carry real compliance exposure.

#Questions only reachable by talking to people

Desk research hit a hard wall on anything requiring a first-hand account rather than a public document. The clearest gap: no on-the-record account exists anywhere in this corpus of a security startup saying "a frontier lab's product killed our business" or "our account got flagged/rate-limited by Anthropic or OpenAI" — not because it doesn't happen, but because founders don't want to publicly antagonize a platform they depend on, and because the research tooling used here (WebSearch exhausted, Reddit blocked) could not surface private complaints.

Who What to ask
XBOW, RunSybil, Strix maintainers, MindFort How did you actually get Mythos/Astra-tier access — application, invitation, or relationship? Has a frontier lab's own product ever displaced a deal you were closing?
Anthropic/OpenAI BD or partnerships teams What's the real eligibility bar for Glasswing / Trusted Access for Cyber for a company that isn't critical infrastructure or a government?
Working CISOs / heads of AppSec at 500+ dev companies What would you actually pay for autonomous vuln finding-and-fixing at your scale? What's blocking you from trusting an autonomous fix today?
Daniel Stenberg (curl), or another maintainer living the AI-slop bug-bounty problem Has the noise/signal ratio actually stabilized, or is the 2025 "closed the bounty temporarily" episode going to repeat?
A German cyber-law practitioner (e.g., referencing Kipker's published scholarship) Is a §202c reform bill actually moving, or is the 2025 coalition commitment stalled? What would a defensible corporate structure look like today?
HackerOne, Bugcrowd program leads Is the researcher-data-provenance backlash (Feb 2026, per Offensive AI security companies) actually costing platforms researcher trust, or did it blow over?
A Series A investor active in this space (e.g., a fund from Who funds this and at what price) What due diligence do you actually run on a "we post-train our own model" claim, given how unverifiable most of these claims turned out to be in this research?

#What cannot be answered by desk research — must be built

Four things in this backlog are not resolvable by more searching, no matter how much web-search budget is restored:

  1. Whether a defensive verifier design actually works. The only way to know if a given reward signal for triage/IR is gameable is to build it, post-train against it, and see what the model learns to exploit — exactly the RLVR crux Post-training playbook identifies as unsolved.
  2. Whether a real customer will trust an autonomous fix without human review. Every vendor in AI code security companies currently keeps a human merge gate; nobody has published evidence customers would accept otherwise. Only a live pilot answers this.
  3. What frontier-lab access actually looks like from the inside. No public document describes the eligibility bar for trusted-access programs with precision — this can only be learned by applying and seeing what happens.
  4. The real economics of post-training vs. scaffolding at production volume. Post-training playbook's cost model ($100–200K to a released checkpoint) is a budget estimate, not a demonstrated ROI — only running both paths against the same workload for months produces a real number.

#The three questions that would most change the plan

  1. Does the Glasswing/Trusted-Access tier ever open to funded startups outside critical infrastructure and governments, and on what basis? A yes reframes the entire strategy around winning that relationship early; a no means competing purely on the standard API tier, where Is frontier-lab gating a real wedge?'s evidence says the gap is already closing fast.
  2. Is the open-weight cyber capability gap (currently 4–7 months per UK AISI) closing fast enough that betting on open models is viable within a 2-year building window, or does it plateau? This determines whether to build model-agnostic from day one or make an early, possibly-premature bet on open weights for cost and deployability reasons.
  3. Does §202c StGB get reformed, and on what timeline? This is the one legal question with a clean binary answer that directly determines corporate structure, product-line sequencing, and whether "build in Berlin" is compatible with an offense-adjacent product line at all — see The first 90 days.

#What this means for us

  • Treat "we couldn't find it" as a task, not a conclusion — several of the highest-value answers here (Glasswing eligibility, §202c's real trajectory) are one phone call away, not one search query away.
  • Do not finalize a corporate structure decision until the §202c question is answered by counsel, not inferred from a coalition-agreement press release.
  • Build the black-box benchmark rerun and the verifier-design experiment early — both are cheap, both produce genuinely novel data nobody else in this market has published, and both directly inform whether to compete on offense or defense first.
  • Prioritize outreach to XBOW/RunSybil/Strix and to a frontier-lab BD contact in the first 90 days — the single most decision-relevant unknown (real Glasswing-tier eligibility) cannot be resolved any other way.
  • Revisit this backlog every quarter; several entries (open-weight gap, frontier-lab product commoditization) are moving targets whose answer will change under you.