AI Safety & Security
Your Security Embargo Is 1,000x Too Slow Now
An OCaml maintainer saw automated probes 10 minutes after a public patch PR. Against a 7-day embargo ceiling, defenders are budgeting 1,000x too much time.
A Cambridge professor patched a path-traversal bug in an OCaml web library this week and watched the internet start probing for it ten minutes later. Anil Madhavapeddy’s account of the cohttp 6.3.0 fix is the cleanest field measurement the industry has of a number every security team is quietly getting wrong: how long a public hint about a bug stays un-weaponized. The answer is minutes, and the processes built around it still assume weeks.
Set that against policy. The Linux kernel’s own security-bug documentation defers disclosure by at most seven days, exceptionally fourteen — the tightest embargo ceiling in mainstream open source. Seven days is 10,080 minutes. Madhavapeddy’s probes arrived in ten. The most aggressive disclosure discipline the ecosystem has written down budgets roughly 1,008 times the window an automated watcher actually needed, and slower projects budget far more. That ratio, not any single CVE, is the operator problem.
The rumour is now the exploit
The mechanism is unglamorous and therefore durable. Madhavapeddy reports that before examining the patch, he pointed his own agent at the affected code with nothing more than a rough description of the bug class — path normalisation — and it independently surfaced related issues and produced a working local probe in under a minute. Claude Fable refused the request under its security guardrails; DeepSeek V4 Pro did not. Two facts fall out of that sentence. Capability is available, and refusal is not a control when a substitute is one API key away.
The refusal asymmetry deserves its own line. Madhavapeddy is a professor of computer science at Cambridge, a core maintainer of the OCaml compiler, and the person responsible for fixing the bug — exactly the profile a defensive-access programme exists to serve. The commercial model with the strongest safety posture blocked him; the open Chinese alternative did not. Guardrails calibrated to deny capability rather than to verify identity produce that outcome by design, and they impose the cost entirely on the defender, since attackers were never going to ask a vendor for permission.
That behaviour is not new, only cheap. Fang and colleagues showed in 2024 that a GPT-4 agent given a CVE description exploited 87% of a fifteen-vulnerability benchmark, versus 7% without the description. The description was the payload. Two years later the description does not even need to be accurate — a PR title, an odd commit in an orphan branch, or a question on a mailing list carries enough signal for an agent to reconstruct the rest.
Google’s frontline data says the crossover already happened at population scale. M-Trends 2026, built from more than 500,000 hours of Mandiant incident investigation in 2025, puts the mean time to exploit at an estimated -7 days: exploitation routinely precedes the patch. The same report records the collapse of the criminal hand-off window from more than eight hours in 2022 to 22 seconds in 2025, and median dwell time rising to 14 days as intruders get better at hiding. Speed at the front, patience at the back.
The trajectory matters more than the current value. Madhavapeddy traces the same metric at roughly 63 days in 2018-19, crossing zero in 2024 and landing at -7 days for 2025 — a defender’s lead time that has not merely shrunk but inverted inside seven years.
Exploitation now arrives before the patch does
Mean time from patch availability to first exploitation, in days
What the incident data does and does not say
The temptation here is to declare an AI apocalypse and stop reading the counts. The counts are more interesting. VulnCheck’s first-half 2026 exploitation analysis finds the median time from CVE publication to known-exploited status fell from 120 days in 2025 to 80 days, while the share of bugs already exploited on or before publication actually eased from 28.93% to 23.43%. Roughly 200 CVEs reached known-exploited status within 31 days — statistically flat against 196 in 2024 and 194 in 2025 — even as CVE issuance grew 45% in six months.
The AI-attribution number is the one to tape to the wall. Of 1,061 vulnerabilities credited to AI-assisted discovery across VulnCheck’s consolidated datasets, 14 — 1.3% — have been confirmed exploited in the wild, roughly the base rate for all vulnerabilities. Machine-found bugs are not, so far, disproportionately dangerous bugs.
Put the two datasets side by side and the honest reading is narrower than the headlines: mass exploitation has not scaled with model capability, but the lead time on any bug that becomes public has compressed hard. Individual cases carry it. Sysdig clocked the marimo notebook RCE from advisory to first exploitation attempt in under ten hours with no public proof-of-concept in existence, and Langflow’s CVE-2026-33017 at twenty hours. Ten minutes, ten hours, twenty hours. None of those numbers survive a two-week release cycle.
There is a second-order effect worth naming. Anthropic reported more than 23,000 findings through its own defensive programme, per VulnCheck’s tally — evidence that model-assisted discovery works at scale when it is pointed at defense. The same tooling that produces 23,000 findings for a funded programme produces 40 inbound reports a month for a volunteer maintainer with no triage staff. Capability distributed asymmetrically does not cancel out; it relocates the queue.
The defender’s bottleneck is not detection either. A May 2026 paper coining the term “bugonomics” argues the constraint has moved to defender remediation throughput — validation, prioritisation, and release capacity — not to finding flaws. rclone’s maintainer, quoted by Simon Willison, reports about 20 security disclosures across the project’s first ten years and more than 40 in the last month alone, with roughly 75% containing something real, while GitHub’s CVE assignment slipped from two to three days to three to four weeks. The reports got faster. The humans did not.
The ways this reading breaks
Three counterpoints deserve airtime before anyone rewrites a disclosure policy.
First, the ten-minute figure is one observation on one repository by one maintainer who was looking for it. Automated scanners have probed public commits for years; some fraction of that traffic is undirected background noise that would have arrived regardless. Madhavapeddy’s own framing — that a determined attacker “could easily be exploiting them within seconds” — is inference, not measurement.
Second, VulnCheck’s data actively resists the strong version of the thesis. If frontier agents had industrialised exploitation, early-lifecycle exploitation counts would be climbing, and they are flat. The KEV-to-CVE ratio has fallen from 2.7% in late 2023 to 1.4% in the first half of 2026. Capability is not the same as deployment, and attacker economics still favour the same tired content-management systems, which account for a third of new known-exploited bugs.
There is also a measurement caveat inside the numbers themselves. The -7 day figure is Mandiant’s estimated mean, drawn from investigated intrusions rather than from the full population of disclosed vulnerabilities, and means are hostage to outliers in a distribution this skewed. VulnCheck’s medians tell a milder story precisely because they discard the tail. Anyone quoting -7 days as a planning constant should treat it as directional evidence that lead time has inverted, not as a scheduling input.
Third, the asymmetry is partly self-inflicted and therefore reversible. Madhavapeddy notes that Western frontier models blocked his security work while Anthropic’s Project Glasswing has extended access to 150 organisations across 15 countries — critical infrastructure operators, cloud and financial providers, the Linux Foundation — and not to the volunteer maintainers who carry the dependency graph those organisations run on. Anthropic reported more than 23,000 findings through the programme, per VulnCheck’s tally. That is a distribution problem with a distribution fix, and it is the single highest-leverage lever in this story.
What would change the verdict: two consecutive VulnCheck halves showing early-lifecycle exploitation counts breaking out of the ~200 range, or a second and third maintainer reproducing sub-hour probe arrival on unrelated ecosystems. Either would move this from a lead-time problem to a volume problem, and the remediation calculus changes completely.
Rebuild the clock, not the secret
The practical shift is from secrecy to speed. Secrecy assumes the description is the scarce resource; it no longer is. Speed assumes the fix pipeline is the scarce resource, which the bugonomics argument and every maintainer’s inbox now confirm. Google ships Chrome security updates twice weekly with restart-free dynamic patching; the kernel defers at most seven days. Neither strategy depends on attackers staying ignorant.
The pattern repeats across the paper’s recent reporting on agent infrastructure. When 1.5% of vendor AI documentation points at unowned code, the exposure is not a secret either — it is published, machine-readable, and waiting. It repeats in the model layer too: Tencent’s new flagship, which today’s brief prices at 5.6 times GLM’s input rate for a 2.4% rating edge, ships with the model’s own claim that it optimized its inference stack, and capability that finds bottlenecks finds bugs. The liability arrives the same way as the litigation in Anthropic’s music-publisher exposure: created upstream, collected downstream, priced by whoever moves last. When OpenAI priced its own oversight at roughly 20% of the inference compute being monitored, it converted a governance posture into a budget line. Disclosure timing deserves the same treatment: a number in the plan, not a norm in a wiki.
For teams shipping this quarter:
- Measure your own arrival time. Instrument one public patch — log probe traffic against the merge timestamp. If the gap is under an hour, every embargo in your runbook is decorative, and you now have the internal evidence to say so.
- Pre-stage the mitigation, not just the fix. Madhavapeddy’s cohttp bug had a one-line mitigation — normalise percent-encoded path separators — deployable the minute the report arrived, while the real patch went through review and packaging. WAF rules, feature flags, and config toggles ship in minutes; releases ship in weeks. Budget for both paths on every advisory.
- Cost the triage surge, then staff it. A 40-reports-per-month inbox at 75% signal is a headcount decision, not a volunteer weekend. Teams already tracking what agent tooling costs per unit of work have the framework; apply it to security review throughput before the queue sets your release cadence for you.
- Buy your maintainers model access. If your production stack depends on a small open-source project, defensive frontier-model access for its maintainers is cheaper than the incident. This is the same governance logic the paper drew from the AISI evaluation incident: controls belong where the work happens, not where the budget sits.
- Escalate CVE-pending releases. With assignment running three to four weeks, shipping with
CVE-PENDINGis now normal. Make sure your ingestion tooling does not silently drop advisories that lack an identifier.
The uncomfortable synthesis is that both stories are true at once. Exploitation volume has not exploded, and the interval between “someone mentions a bug” and “someone probes for it” has effectively gone to zero. Policies calibrated to the first fact while ignoring the second will look prudent right up to the morning the probes arrive before the release notes do. The paper’s earlier work on runtime controls for agent harnesses made the same argument about a different layer: the control has to live where the speed is.
Sources
- Anil Madhavapeddy — Just a rumour of a bug is enough to find a security exploit these days
- Linux kernel documentation — Security bugs and disclosure timing
- Google Cloud — M-Trends 2026 frontline metrics
- VulnCheck — State of Exploitation, first half of 2026
- Sysdig — marimo notebook RCE exploited in under ten hours
- Sysdig — CVE-2026-33017 exploited against Langflow pipelines in 20 hours
- Fang et al. — LLM agents can exploit one-day vulnerabilities
- Pesoli et al. — Demystifying the Mythos or Disrupting Bugonomics?
- Simon Willison — maintainers on the disclosure surge
- Anthropic — Project Glasswing