Studies / S2
Spoofed AI crawlers
How much traffic that claims to be an AI or search crawler is real?
Abstract
Site owners count "GPTBot", "ClaudeBot" or "ChatGPT-User" in their access logs as evidence that AI systems read or cite them. The User-Agent is free text. We checked every request claiming to be one of 23 AI or search crawlers on three small production websites (16,548 requests; one 68-day log, two 7–9-day logs, plus a 30-day Cloudflare edge view) against each operator's own published verification method — IP-range JSON files and forward-confirmed reverse DNS. 39.3% of the claims were spoofed (44.3% of those that could be tested), 49.5% verified, 11.2% unverifiable or indeterminate. User-initiated fetchers, the names most often read as an "AI citation" signal, were the worst: 67.8% spoofed, and Claude-User, Perplexity-User, MistralAI-User and DuckAssistBot were 0–4.2% genuine. Google-Extended, a token Google says is never sent as a User-Agent, appeared 371 times. All spoofing came from at most 54 addresses, 81% in cloud/hosting networks (one cloud provider's customer space alone: 66%); 97% of spoofed requests came from sources that also probed for credentials (.env, config.json, service-account.json) and 92% from sources that rotated through several operators' identities. On the main site, at the CDN edge, the spoofed volume was 3.1× what the origin logged; free-plan defaults blocked 0.7% of it. We release ai-bot-verify, a dependency-free verifier for any access log that outputs aggregates only and refuses to write IP addresses.
Cite this study
EasyxLab (2026). Spoofed AI crawlers. Study S2. EasyByte Hub S. Coop. Mad. https://github.com/easybytehub/easyxlab/tree/main/studies/s2-spoofed-ai-crawlers@techreport{easyxlab_s2,
title = {Spoofed AI crawlers},
author = {{EasyxLab}},
institution = {EasyByte Hub S. Coop. Mad.},
number = {S2},
year = {2026},
url = {https://github.com/easybytehub/easyxlab/tree/main/studies/s2-spoofed-ai-crawlers}
}