What Server Logs Reveal About AI Crawlers and Opt-Out Files
Summary
A website operator used an edge sensor to measure whether automated agents that identify themselves as AI-training crawlers read the machine-readable rights notices published on the site. The site exposes a W3C TDM Reservation Protocol file at /.well-known/tdmrep.json, a machine-readable license, a policy page, HTML metadata, response headers, and robots.txt signals stating ai-train=no. From July 28 through August 10, the sensor recorded 17,490 automated reads, including 1,534 fetches from crawlers identifying as AI-training agents. OpenAI’s GPTBot accounted for hundreds of page reads, with 841 AI-agent fetches recorded on August 8 alone, while the median daily total was 15. Only one of the 1,534 fetches touched a rights file: it came from OAI-SearchBot, which OpenAI says is not used for training. GPTBot never requested the terms attached to the reservation. The sensor counted 80 rights-file fetches overall, but the identifiable requests came from the site’s own monitoring; much of the tooling did not identify itself as a machine. The author says the edge instrumentation records each visit before serving the file and reports an error if monitoring fails, making the result a measured zero rather than an unobserved gap. The author also stresses that one site over 14 days is an anecdote, not a broad study. The conclusion is that, in this observation, publishing machine-readable rights declarations was complete, but the relevant training crawler did not appear to consult them.