The first 90 days
A week-by-week plan to produce one public artifact, three design partner conversations, and enough evidence to decide.
This plan is built to resolve uncertainty, not to build a product. At day 90 the goal is not revenue or a launch — it is having replaced the four assumptions this entire research rests on with evidence, and having produced one public artifact that makes the next conversation easier.
Confidence on this page is deliberately marked low. It is a proposal derived from research, not a validated plan, and the Open questions and the research backlog backlog exists precisely because several inputs to it are unverified.
#The four things to learn
Everything below serves one of these:
- Can you validate exploitability materially better than the obvious baseline? If not, idea two has no product underneath it.
- Does the CRA deadline actually move budget, or do people nod and do nothing? The strongest demand signal in the research is a legal deadline, and legal deadlines routinely fail to convert.
- What is real frontier-lab access eligibility? Not the policy documents — the actual answer from someone who grants it (Is frontier-lab gating a real wedge?).
- Is §202c a manageable constraint or a structural one? From counsel, not from a coalition agreement (Germany: §202c and the Berlin question).
#Weeks 1–2 — Baseline and legal
Build the baseline before building anything. Assemble a held-out evaluation set of real vulnerabilities in memory-safe codebases — ARVO is the obvious starting substrate, and its ~6,100 reproducible OSS-Fuzz cases are already reproducible by construction (Data, and whether a moat is possible). Filter to Java, Go, Python and TypeScript per the AIxCC language finding (AIxCC: the closest thing to a proof). Aim for 100–200 cases with known ground truth, deliberately including a control group of non-vulnerabilities, because precision is the thing being measured.
Then measure the baseline you have to beat: a well-configured Semgrep plus a frontier model doing diff review. Write down its precision and recall. This number is the reference point for every claim you will make for the next two years, and almost nobody in this market has one.
In parallel, book the lawyer. A German technology lawyer with §202c experience, with a specific brief: what can a German entity legally build, publish and sell in the area of exploit-generating tooling, and does structuring offensive capability under a non-German entity actually change the analysis. Do not proceed on inference from press coverage (Germany: §202c and the Berlin question).
Also start the slow things now because they gate everything later: SOC 2 Type II scoping, and an EU-region or self-hostable deployment story (Who buys, and what they pay).
#Weeks 3–6 — The verification harness
This is the technical core and the actual moat (Where the gaps actually are). Build the thin version:
- Reproducible environment per target repository
- Sandboxed execution with real isolation — treat this as the most consequential engineering decision you will make, given that two frontier labs failed at it in 2026 with far more resources (Dual-use risk and what it costs you)
- Candidate generation from a frontier model with your own scaffolding — no post-training yet, per The three ideas, judged
- Verification: does the proof-of-concept reproduce, does a patch close it, does the regression suite stay green
- Structured output: the verified-defect artifact, designed from the start to double as CRA Article 14 evidence (EU regulation as a demand engine)
Use the AIxCC systems as architecture references rather than forks — they are open-sourced and mostly archived, but ATLANTIS and Buttercup encode a great deal of hard-won design that almost nobody outside the competition has looked at (The open-source stack).
Run it against the held-out set weekly and plot precision against the baseline. If the line does not separate by week 6, that is information, and it arrives cheaply.
#Weeks 3–8 — Design partner conversations, running in parallel
Twenty conversations, targeting three signed design partners. Filter hard for CRA, NIS2 or DORA exposure: German Mittelstand manufacturers shipping products with digital elements, financial-infrastructure firms under DORA, and medium-sized software vendors selling into the EU (Go to market).
The question to ask is not "would you use this." It is: who owns CRA Article 14 compliance in your organization today, what are they doing about it, and what is the budget line it comes out of? If nobody can name the owner, the demand signal is narrative rather than budget, and that finding is worth more than a polite yes.
Two questions to ask every single time, because they falsify the two most load-bearing assumptions in Who buys, and what they pay: what do you currently pay per developer for security tooling, and is self-hosted or bring-your-own-key a hard requirement.
#Weeks 6–10 — The public artifact
The credibility ladder in The AI cyber lab category runs on public artifacts, not headcount, and the cheapest available rung is the independent evaluation nobody else is publishing (Where the gaps actually are).
Publish a rigorous, reproducible evaluation of AI security agents on validated exploitability — your held-out set, your harness, open-sourced, with the methodology and the failure cases included. Include your own system in it and do not flatter it. The Escape.tech evaluation of Strix is the template for the genre and the demonstration of how much attention a single honest measurement attracts in a market where every number is self-reported (Strix, Cyber benchmarks and evals).
Keep it narrow and attached to a real claim. This is not a bid to become a benchmark organization — it is a demonstration that you measure things other people assert.
Alongside it, disclose real vulnerabilities found by the harness through proper responsible-disclosure channels. Verified findings in real software are the currency this market actually respects.
#Weeks 8–12 — Access, capital, and the decision
Frontier lab conversations. Approach with the published artifact in hand, because that is what changes the conversation from a cold ask to a credible one. The specific question to resolve: what does eligibility for the gated tier actually require, and is it available to a European company at this stage. Anthropic's Anthology Fund has already invested in a company in this exact category, which makes an investment or partnership conversation a more realistic ask than raw API access (Who funds this and at what price, Is frontier-lab gating a real wedge?).
Non-dilutive capital, started now because the timelines are long. Cyberagentur, SPRIND, EIC Accelerator, HTGF. These are slow and underused, and they price-in nothing (Who funds this and at what price).
The talent conversation. One hire matters more than the rest: the person who can do both security research and post-training. AIxCC finalist alumni are a small, public, and almost entirely un-recruited pool outside the US offensive-security cluster. Direct outreach referencing specific commits beats any job posting (Talent: the actual constraint).
#The day-90 decision gate
Check these honestly. The point of writing them down now is that they are much harder to rationalize away later.
| Signal | Threshold | If missed |
|---|---|---|
| Validated-finding precision vs. baseline | ≥2x on the held-out set | The technical premise of idea two is unproven — extend 6 weeks, then stop |
| Signed design partners | ≥3, at least 2 citing CRA/NIS2/DORA | Demand is narrative; re-test with a different buyer segment before building further |
| Public artifact | Published, with ≥1 external group running the harness | Credibility strategy is not working; reconsider the Go to market motion |
| §202c legal position | Written opinion in hand | Do not incorporate or publish offensive-capable tooling until you have it |
| Frontier access eligibility | A real answer, yes or no | A "no" is fine and useful — it makes Post-training playbook a live question sooner |
Notice what is not in this plan: no post-training run, no model release, no benchmark organization, and no product launch. Each of those is a month-six-or-later decision that depends on evidence this plan produces. Starting any of them in the first 90 days spends capital on a bet the evidence has not yet justified (The three ideas, judged).
#Rough cost
| Item | 90-day cost |
|---|---|
| Two people (founder + one engineer), Berlin | €55–75K [estimate] |
| Frontier model inference for harness development and evaluation runs | €8–20K [estimate] |
| Compute and sandbox infrastructure | €3–6K [estimate] |
| German legal opinion on §202c and product structure | €5–15K [estimate] |
| SOC 2 Type II scoping and readiness | €10–20K [estimate] |
| Total | €80–135K |
Every figure is an estimate and should be rebuilt with real quotes. The point of the table is the order of magnitude: this is a sub-€150K experiment that resolves the four assumptions the entire thesis rests on. That is a good trade at any plausible outcome.
#What this means for us
- The first 90 days buy information, not progress. Resisting the urge to build the product early is the discipline this plan is actually enforcing.
- The verification harness is simultaneously the product core, the benchmark infrastructure, the training-data generator for any later post-training, and the thing that makes the public artifact credible. Everything routes through it — which is why it starts in week three and nothing else does.
- Two of the four key unknowns are resolved by talking to people, not by building. Book those conversations in week one.
- Re-read Open questions and the research backlog at day 45 and day 90. Several entries will have moved, and at least one will have invalidated something on this page.