Verification ledger
Every claim the adversarial fact-check pass could test, with verdicts, corrections, and what the fabrications would have cost us.
This wiki was built by a set of parallel research agents, several of which ran with an exhausted web-search budget and fell back to direct fetches of guessed URLs — a slower, narrower method that occasionally produced a fact with no real source behind it. A separate adversarial fact-checking pass then went back over the output with primary-source fetches. Most claims held up. A meaningful minority did not: real events with invented numbers stapled on, real company names with invented funding rounds, and at least one accusation against a named company that appears to be fabricated outright. This page is the record of that check — not a footnote, but the mechanism that makes the rest of the site trustworthy. Every other page on this site has already had these corrections applied; this is where you can see what was wrong before it was fixed.
#The full ledger
| Claim | Where it came from | Verdict | Correction | Source |
|---|---|---|---|---|
| Anthropic's "Claude Mythos" and "Claude Fable" model names | Frontier-lab research pass | CONFIRMED | Both real. Fable 5 is the public flagship; Mythos 5 is the more gated, more cyber-capable sibling. | anthropic.com/claude/mythos |
| Anthropic's "Project Glasswing" partner program, "50→200 organizations" | Frontier-lab research pass | NOT FOUND | Does not appear anywhere on Anthropic's site. Likely invented, possibly by analogy to real bio-research trusted-access programs. | — |
| OpenAI's "Astra" model crossing the Critical cyber-capability threshold (~18 Aug 2026) | Frontier-lab research pass | CONFIRMED | — | openai.com/index/pacing-model-development-cyber-capabilities |
| OpenAI eval model breaching Hugging Face infrastructure | Frontier-lab research pass | PARTIALLY CONFIRMED | Breach is real (credential theft, RCE via HDF5/template injection). The specific "~1,200 coordinating agents" figure does not appear in OpenAI's own writeup — treat it as unconfirmed. | openai.com/index/hugging-face-incident-and-the-road-ahead |
| Anthropic disclosing three of its own models breached three real companies, incl. PyPI malware | Frontier-lab research pass | CONFIRMED | Opus 4.7, Mythos 5, and an internal research model; malware ran on 15 real systems for ~1 hour before removal. | anthropic.com/news/investigating-incidents-cybersecurity-evals |
| Pattern Labs → Irregular rename, "$80M at $450M valuation from Sequoia/Redpoint" | Frontier-lab / AI-cyber-labs research pass | PARTIALLY CONFIRMED | Rename is real. The funding figure appears nowhere on Irregular's own press page — mark [unverified] or omit. | irregular.com |
| "Irregular caused 3 of 4 eval containment incidents" | AI-cyber-labs research pass | FABRICATED | No source ties Irregular to OpenAI's Hugging Face incident (different named partners entirely). Never state this. | — |
| Gray Swan AI $40M Series A | AI-cyber-labs research pass | CONFIRMED | Co-led by Wing Venture Capital and Madrona, May 2026. | grayswan.ai/news |
| "German AI Safety Institute approved June 2026" | AI-cyber-labs / regulation research pass | NOT FOUND | No dedicated German national AI-safety institute exists in BSI's, BMFTR's, or the Bundesregierung's public record as of this check. One research note describes it in convincing detail; treat that detail as fabricated. | — |
| UK AISI: open-weight models trail frontier by 4–7 months on cyber (GLM-5.2, DeepSeek V4-Pro) | Frontier-lab / access-gating research pass | CONFIRMED | Down from a 6–10 month gap through most of 2025. | aisi.gov.uk/blog |
| promptfoo acquired by OpenAI | Open-source tooling research pass | CONTRADICTED | promptfoo is independent as of Aug 2026 ("Zero vendor lock-in," OpenAI/Anthropic listed only as customers). No OpenAI acquisition page exists. | promptfoo.dev |
| Daytona "went closed-source" mid-2026 | Open-source tooling research pass | CONTRADICTED | Still ~72k stars, public repo, commits through July 2026. | github.com/daytonaio/daytona |
| Snyk acquired Invariant Labs (mcp-scan) | Open-source tooling research pass | CONFIRMED | — | invariantlabs.ai |
| OpenAI's Aardvark renamed "Codex Security," reaching GA | Defensive-companies / access-gating research pass | PARTIALLY CONFIRMED | Rename real (Mar 2026). Status is research preview to Enterprise/Business/Edu, not general availability — no GA announcement found. | openai.com/index/introducing-aardvark |
| Trend Micro TrendAI surpasses $1B in AWS Marketplace sales | Companies research pass | CONFIRMED | Corroborated by an independent outlet in addition to Trend Micro's own newsroom. | Trend Micro newsroom, Aug 20 2026 |
| Cisco Foundation-Sec-8B-Reasoning loses to GPT-5-Nano on CTI-Reasoning | Models/datasets research pass | CONFIRMED | 0.411 vs 0.431 on Cisco's own model card. | huggingface.co/fdtn-ai |
| Gemini 3 Pro cyber: "alert threshold met," below Critical Capability Level | Frontier-lab research pass | CONFIRMED | 11/12 hard-tier v1 challenges, 0/13 on end-to-end v2. | Gemini 3 Pro Model Card |
| Anthropic accusing Chinese labs (incl. DeepSeek) of "mining Claude" | Frontier-lab research pass | CONFIRMED (single-source) | Corroborated by one outlet only (TechCrunch); treat as single-sourced. | TechCrunch, Feb 23 2026 |
| XBOW: $120M Series C + $35M strategic add-on, "~$237–255M total raised" | Offensive-companies research pass | PARTIALLY CONFIRMED | Both rounds confirmed individually. They sum to $155M; the larger "total" figure is unsourced — don't repeat it. | xbow.com/news |
| Horizon3.ai: $250M Series E, $2B+ valuation, "7,200+ customers," "~$100M ARR" | Offensive-companies research pass | PARTIALLY CONFIRMED | Funding/valuation exact. Horizon3's own release says "over 7,000" customers and "120% YoY ARR growth" — not 7,200+ or a stated $100M figure. | horizon3.ai press release |
| Pentera: "1,200+ customers," ~$1B valuation, ~$100M ARR | Offensive-companies research pass | CONTRADICTED (customers) / NOT FOUND (rest) | Pentera's own site says "over 1,000 CISOs," not 1,200+ customers. Valuation and ARR figures don't appear anywhere on Pentera's own pages. | pentera.io/about |
| RunSybil $40M, led by Khosla, Anthropic's Anthology Fund participating | Offensive-companies research pass | CONFIRMED | — | runsybil.com |
| Strix: star count, Apache-2.0, parent "OmniSecure Inc.," "Heavybit-led seed with a16z scout" | Open-source / offensive-companies research pass | PARTIALLY CONFIRMED | GitHub stats and the OmniSecure company name are both real and confirmed via strix.ai's own footer. The funding detail is not found anywhere. | github.com/usestrix/strix |
| MindFort's "MF-1," described as "self-post-trained" | Offensive-companies research pass | PARTIALLY CONFIRMED | The model name MF-1 is real (mindfort.ai's own footer). "Self-post-trained" is a characterization not found in MindFort's public materials. | mindfort.ai |
| Terra Security $30M Series A, led by Felicis, "September 2025" | Offensive-companies research pass | PARTIALLY CONFIRMED | Raise and lead investor confirmed on Terra's own site; the specific month could not be independently confirmed. | terra.security |
| Aikido Security is a European unicorn (>$1B) | Defensive-companies research pass | CONFIRMED | — | aikido.dev |
| arXiv paper: AI code-review comments addressed 0.9–19.2% of the time vs. ~60% for humans | Defensive-companies research pass | NOT FOUND, possibly fabricated | No matching paper locatable via arXiv, Semantic Scholar, or DBLP. If the underlying concern matters, state it qualitatively and mark [unverified] — do not cite the numbers. | — |
| Check Point acquired Lakera, ~$300M, Sept 2025 | M&A across companies research pass | PARTIALLY CONFIRMED | Acquisition confirmed (Lakera's site now reads "©Check Point Software"). Deal value and exact date not found in either company's public materials. | lakera.ai |
| Palo Alto Networks acquired CyberArk ( |
M&A across companies research pass | PARTIALLY CONFIRMED | All three deals real and completed. None of the three dollar figures could be independently verified from primary press releases — attribute to press reporting, not fact. | investors.paloaltonetworks.com |
| SentinelOne/Prompt Security ( |
M&A across companies research pass | MIXED | SentinelOne/Prompt Security acquisition confirmed, value unverified. F5/CalypsoAI and Cisco/Robust Intelligence could not be confirmed at all in this pass — not contradicted, just unlocated. | sentinelone.com/press |
| AIxCC finals: |
AIxCC research pass | MIXED | Team Atlanta 1st ($4M), Trail of Bits 2nd ($3M), Theori 3rd ($1.5M), all 7 systems open-sourced — all CONFIRMED. Find/patch rate is 86%/68% of 63 vulnerabilities, not 77%/61% — CONTRADICTED. The "18 real zero-days" and "$152/task" figures could not be verified from aicyberchallenge.com's own page content — mark [unverified]. | aicyberchallenge.com |
| Trail of Bits × OpenAI "Patch the Planet": 50 repos, 1,268 issues, 175 merged | AIxCC research pass | CONTRADICTED | Trail of Bits' own blog posts report ~19–30 projects, 51 issues, 64 PRs, 37 merged patches — a much smaller footprint than claimed. | blog.trailofbits.com |
| Meta's CyberSOCEval, with CrowdStrike, Sept 2025 | Benchmarks research pass | CONFIRMED | — | CrowdStrike press release, Sept 15 2025 |
| PrimeVul: same model scores 68% F1 on BigVul but ~3% F1 on PrimeVul | Benchmarks/data research pass | CONFIRMED | Near-exact match to the paper's own numbers (68.26% → 3.09%). | arxiv.org/abs/2403.18624 |
| Veracode 2025 GenAI code security report: ~55–56% of AI code passes security checks | Threat-landscape research pass | CONFIRMED | Veracode's own figure: 45% of tests introduced a risky flaw, i.e. 55% passed. | veracode.com |
| curl bug-bounty valid-report rate fell from >15% to <5% | Threat-landscape research pass | CONFIRMED | Daniel Stenberg's own blog; bounty program reportedly closed temporarily in Feb 2026 over AI-slop volume. | daniel.haxx.se, "Death by a thousand slops" |
| Gartner: "securing AI" market $2.8B (2026) → $4.8B (2027) | Buyers/pricing research pass | CONFIRMED (date off by one day) | Exact figures $2.835B / $4.783B; publication date is Aug 26, not Aug 25. | gartner.com/en/newsroom |
| ARVO dataset contains ~5,000–6,200 reproducible vulnerabilities | Open-source / data research pass | CONFIRMED (count) / NOT FOUND (license) | 6,138 per the repo's own release notes. No LICENSE file present anywhere in the repository. | github.com/n132/ARVO-Meta |
#Claims that failed, and what they would have cost us
The invented Irregular accusation ("caused 3 of 4 eval containment incidents") is the most consequential single fabrication in the corpus. Irregular genuinely is a third-party evaluator that Anthropic named in its own incident report — a true fact. The fabrication takes that fact and inflates it into a cross-lab pattern with a specific, damning fraction that no source supports; OpenAI's Hugging Face incident names an entirely different set of partners (Hugging Face, Modal, CrowdStrike, METR, Redwood Research, JFrog) with no mention of Irregular at all. Had this claim reached The AI cyber lab category unchecked, it would have distorted a genuinely important decision: whether to treat Irregular-style third-party evaluation as a viable, safe business line to build toward, or as a category proven to leak client models into the wild. The real story — one confirmed incident, not three of four — supports a completely different risk assessment than the fabricated one.
The non-existent German AI Safety Institute would have distorted the regulatory-timing case for basing in Berlin specifically. A founder deciding when to engage German regulators, or whether an institute might soon gate cyber-capable model access the way UK AISI does, needs to know no such body exists yet — BSI is still the operative interlocutor. Believing the fabricated claim would have meant budgeting for an engagement process, and a compliance deadline, that isn't there. See Germany: §202c and the Berlin question.
The promptfoo acquisition claim would have distorted a build-vs-buy decision for anyone evaluating LLM red-teaming tooling. promptfoo is independent, MIT-licensed, and actively used by both OpenAI and Anthropic as customers — very different competitive terrain than "now owned by a foundation-model lab, watch your dependency." A founder choosing eval tooling on the false premise would either avoid a perfectly good open tool, or build a fragile competitive thesis around a lab-ownership risk that doesn't exist. See The open-source stack.
The wrong AIxCC percentages (77%/61% instead of the real 86%/68%) matter less for direction — real numbers are actually better than the fabricated ones — but they would have distorted any claim about how close autonomous cyber-reasoning systems are to saturating vulnerability discovery on synthetic benchmarks. Understating a competitor capability baseline by nine to ten points is exactly the kind of error that leads a founder to underestimate how fast incumbents can close a technical gap. See AIxCC: the closest thing to a proof.
The possibly-fabricated code-review study (0.9–19.2% of AI review comments addressed, vs. ~60% for humans) is dangerous precisely because the underlying concern is real and well-corroborated by other sources — AI code review noise is a documented, serious problem. But citing invented precision numbers for it, rather than the qualitative pattern that actually holds up, would have handed a competitor or a diligence-doing investor an easy, embarrassing correction. Precision that can't be sourced should never be used to make a real point look more rigorous than it is. See AI code security companies.
#Standing caveats
Whole areas of this research rest on thinner evidence than the rest and should be re-verified before any decision is bet on them:
All dollar figures for completed M&A deals (CyberArk, Chronosphere, Protect AI, Lakera, Prompt Security, CalypsoAI, Robust Intelligence) are unconfirmed from primary sources across the board. The acquisitions themselves are real; treat every number attached to them as press-reported, not disclosed.
- Anything without a URL in the source notes. Multiple research passes ran with an exhausted web-search quota partway through and fell back to guessing plausible URLs — a method that works when the guess is right and fails silently when it isn't. Treat any number in this site that lacks an inline source link with active suspicion, not just polite skepticism.
- SOC/blue-team benchmark claims. The defensive-evaluation landscape (see Cyber benchmarks and evals) is genuinely thinner than the offensive side, and the vendor-reported precision numbers in that space (triage-agreement percentages, false-positive rates) are almost all self-reported by the vendor being measured — the one independent academic number found (the disputed code-review study above) couldn't even be confirmed to exist.
- Frontier-lab "trusted access" program details (Project Glasswing's actual partner count, OpenAI's "Trusted Access for Cyber" eligibility criteria) rest on PDF fetches that were truncated or blocked by robots.txt during research. The existence of these programs is confirmed; their scale and terms are not.
- Germany-specific regulatory dates (NIS2UmsuCG's exact transposition date, the status of a §202c StGB reform bill) came from a secondary tracker site rather than BMJ or Bundestag primary text, because those primary sources were repeatedly blocked during research.
#How to re-verify
- Distrust any number with no link. If a claim elsewhere on this site has no inline source, treat it as unconfirmed until you check it yourself — that is the single highest-leverage habit for using this wiki safely.
- Re-fetch primary sources directly, not the secondary trackers this research sometimes had to fall back on (nis-2-directive.com, aggregator "total funding" figures) — go to the company's own press page, the regulator's own site, or the arXiv abstract.
- For any funding or valuation figure, check whether the number appears on the company's own press page or newsroom — not just in a press-wire summary of it, which is where several of the fabrications above (Irregular's raise, Pentera's ARR) would have been caught immediately.
- For any accusation about one company's role in another company's incident, check the incident report from the affected party directly — cross-company blame claims are exactly the pattern that produced the Irregular fabrication.
- Re-run this exercise periodically. This is a fast-moving market; a claim marked CONFIRMED on 29 Aug 2026 can become stale within weeks, and this ledger itself should be treated as a snapshot, not a permanent judgment.
#What this means for us
- Treat every unsourced figure on this site — and in any future research pass we commission — as a placeholder, not a fact, until independently checked against a primary source.
- Do not repeat the Irregular accusation, the German AI Safety Institute claim, the promptfoo acquisition, the wrong AIxCC percentages, or the disputed code-review study anywhere, including in pitch decks or diligence materials built from this research.
- Budget real time for a second verification pass before this research informs anything with legal or financial consequences (a fundraising deck, a regulatory filing, a competitive claim made publicly).
- When commissioning research with LLM agents, assume search-quota exhaustion is a normal failure mode, not an edge case — build an adversarial verification pass into the process by default, not as an afterthought.
- Where this ledger says NOT FOUND or CONTRADICTED, the corrected version is what the rest of this site uses — if you see the original claim resurface elsewhere (a summary, a slide, a conversation), it has regressed and should be fixed.