Back to News
RSS feedai-visibility.lastminutedealshq.com

368 of the Top 5,000 Sites Block at Least One AI Search Crawler

Summary

A census dated September 7, 2026 examined the robots.txt files served by sites in the Tranco top 5,000. Of the 2,771 domains that served a robots.txt, 368 blocked at least one of OAI-SearchBot, PerplexityBot, or Claude-SearchBot, the crawlers associated with citations in AI answers, and 200 blocked all three. The crawler-specific totals were 240 blocks for OAI-SearchBot, 358 for PerplexityBot, and 249 for Claude-SearchBot. A blocked crawler is not allowed to read the site root under the robots.txt result and therefore cannot cite that site through the relevant engine. The census deliberately measures citation access, not training access: rules blocking GPTBot, described as a training crawler, are excluded because they do not affect citations. The source also cautions that the Tranco ranking contains raw hostnames, including CDN and API endpoints such as gstatic.com and fbcdn.net, rather than only destinations people browse. Two of the 368 entries appear to be infrastructure rather than publisher sites, but they remain because their robots.txt results are genuine. Results were measured against each domain’s live robots.txt using the same matcher Google publishes, including wildcard and dollar-anchor behavior. Per-domain results are available on the census page and in an open dataset.