skip to content
The Weighted Average

AI Safety & Security

1.5% of Vendor AI Docs Point at Unowned Code

Researchers scanned 8,265 llms.txt files and found 120 pointing at unregistered packages or domains — in files coding agents treat as vendor ground truth.

Fiber optic cables connected to a network switch in a server rack
Fiber optic cables connected to a network switch in a server rack. Photograph by Kirill Sh

A convention almost nobody governs has quietly become an input to production code. Israeli researchers scanned 6,214 live domains belonging to Fortune 500 companies, Big Tech, and defense contractors, found 8,265 llms.txt and llms-full.txt files, and identified 120 — one on each of 120 different sites — pointing at code packages or domain names that nobody had registered, Ars Technica reported this week. That is 1.5 percent of a file type that coding agents read as authoritative vendor guidance.

The researchers registered a handful of the unclaimed names and hosted packages that phoned home. Within an hour a Fortune 500 company’s infrastructure called back; over time, a few dozen more organizations followed. Parent-process chains recorded in those beacons named the agents involved: Claude, OpenAI’s Codex, and Nous Research’s Hermes. Anthropic, OpenAI, and Nous did not respond to Ars by publication. “The trust model is broken,” researcher Alon Hertz told the publication. “Agents treat vendor docs as ground truth and don’t question them — and neither do the humans supervising them.”

The file nobody owns

llms.txt is a community convention for publishing a machine-readable map of a site for language models — the AI-era analogue of robots.txt, now formalized enough that Google’s Lighthouse ships an audit for it. The crucial difference from robots.txt is directionality. Robots.txt tells a crawler what not to fetch; llms.txt tells a model what to believe, and increasingly what to install.

Scale makes the governance gap concrete. Cloudflare’s correctly configured llms.txt runs 183 lines with 88 external links; its llms-full.txt is roughly 157,000 characters carrying 389 links. Every one of those is a pointer an agent may follow without a human in the loop. Apply the researchers’ 1.5 percent stale-reference rate to a corpus that size and the expected number of dangling pointers per large vendor file is not zero — it is a maintenance obligation nobody has staffed, because these files are typically generated once by marketing or docs tooling and never audited again.

The convention’s own documentation is explicit that these files are written for machine consumption and that the format expects curated, high-signal links rather than a dump of a sitemap. Curation implies a curator. In practice the file is generated once, shipped, and forgotten, and the researchers’ 1.5 percent is what forgetting looks like measured at scale across the most security-conscious tier of the corporate internet.

The derived figure worth holding: at 8,265 files across 6,214 domains, roughly 1.33 files per domain, and 120 defective files means about one in 52 scanned organizations publishes machine-readable guidance that resolves to unowned infrastructure. Not one in 52 attacks — one in 52 opportunities, sitting in public, in the docs directory of companies that spend heavily on supply-chain security everywhere else.

What this changes for the people running agents

This is a provenance problem, not an exploit story, and the fixes are ordinary governance. Three of them are available this quarter.

First, treat vendor documentation as an untrusted input in the same category as a package registry. The reason agents installed unowned code is that the docs were never in anyone’s threat model. Any agent with network and install permissions needs an allowlist for what it may fetch and a record of what it fetched — the same auditability argument this paper made when Claude gained send-and-delete rights in Gmail and Drive and the approval log became the only audit trail.

Second, own your own file. If your company publishes llms.txt or llms-full.txt, the links in it are now statements your agents’ customers act on. Diff it on a schedule, resolve every reference, and fail the build when one 404s. A stale link in a human-facing doc is an annoyance; the same link in a machine-facing one is an instruction.

Third, capture the install chain. The only reason this research produced findings rather than speculation is that the beacons recorded the parent-process chain that spawned each install. Most organizations cannot answer “which agent, on whose behalf, added this dependency” from their own telemetry. That gap is the actual finding, and it applies well beyond one file convention — it is the same observability deficit behind the AISI’s push for standardized agent evaluation and governance.

A fourth move is worth considering for teams that publish developer documentation commercially: publish a checksum or a signature alongside anything an agent is told to install. The dangling-reference problem is a subset of the broader question of whether machine-readable vendor guidance carries any provenance at all, and today it carries none. Nothing in the convention distinguishes a link a vendor deliberately maintains from one that survived three site migrations.

What would change the verdict? A rate this low might be noise: 120 defective files is small, the researchers are a stealth startup with a commercial interest in the finding, and the phone-home count — a few dozen organizations — measures curiosity as much as exposure. A replication across a random sample of domains, rather than a high-value target list, would settle whether 1.5 percent is the internet’s number or the Fortune 500’s. Until then the cheap move stands: audit the file, log the installs. The infrastructure economics in today’s lead on the chip-tariff exemption determine what agents cost to run; this determines what they are allowed to bring home.

Sources