The Same Welcome Mat You're Told to Roll Out for AI Crawlers Can Be Used Against You
Updated: 6 days ago

I pulled a routine report on my own site's bot traffic recently, checking which AI tools were actually crawling my content. It's the kind of check I recommend to anyone working on AI search visibility, confirm ClaudeBot, GPTBot, and the others can actually reach you.
What I found instead was a list of requests for files like gcp-credentials.json, .azure/credentials, and .claude/settings.json. Not content. Credentials. Dozens of attempts to find exposed secrets, none of which existed on my site to find.
At first I assumed it was generic scanning, the background noise every public website gets. But I spent the first half of my career in security before moving into marketing, and that background doesn't fully switch off. Something about the pattern made me look closer instead of dismissing it, and I found a live security story I wasn't expecting.
AI Crawler Impersonation Exploits the Trust You're Told to Extend
Security firm GreyNoise published research in late August confirming exactly this pattern, a wave of AI crawler impersonation happening right now. Automated scanners, spread across hundreds of addresses, are forging the identities of real AI crawlers, ClaudeBot, GPTBot, Google-Extended, and others, to search websites for exposed secrets. None of the addresses involved matched the actual published IP ranges of the companies whose names they were using.
Here's the mechanism, and it's simpler than it sounds. A crawler identifies itself through one line in its request, the user agent. Anthropic's real crawler says ClaudeBot. But that string is just typed in by whoever's sending the request. Nothing about it proves it's true. Researchers caught the fakes because a real crawler checks robots.txt first, the file that tells it what it's allowed to access. The impersonators never did. They went straight for credentials instead.
I've spent months telling readers to welcome AI crawlers, check that robots.txt isn't blocking them, make sure the door is open. This is the necessary second half of that advice, and I hadn't given it yet: welcoming AI crawlers and trusting anything that claims to be one are not the same thing.
Why This Matters More Than It Might Seem
A lot of businesses, mine included, have started treating "known AI crawler" as a reason to relax scrutiny. Skip the rate limit. Don't flag it as suspicious. That's a reasonable instinct if the name can actually be trusted. The problem is that a name typed into a request header proves nothing on its own, and attackers realized that extending trust based on a name alone is exactly the gap worth exploiting.
The fix isn't complicated, and it's the one concrete thing worth taking from this. Any rule that treats a known AI crawler differently should also check that the request is actually coming from that company's published IP range, not the name by itself. That's a real, specific, one-time check most site owners have never thought to make, and it typically requires server access logs, a firewall tool, or someone on your team who can look, not something available on every platform. If you're running real infrastructure behind your site, this is the one thing worth going and confirming. I can't run it myself, my own site sits on Wix with no server-level access to check, which is worth knowing too: plenty of small business sites are in the same position, unable to verify this even if they wanted to.
Where This Fits With Everything Else I've Written
This connects directly to the technical foundation layer I've talked about throughout this whole series, structured data, entity clarity, and yes, a correctly configured robots.txt. That file and the trust it represents work both ways. Get it right, and you're genuinely more discoverable. Misunderstand what it's actually granting, and it becomes something else entirely.
One honest note, since I don't want to overstate my own stakes here. My site runs on Wix, no backend server, nothing that could actually be exposed by any of this. Every request in my own data came back at nothing, because there was nothing there to find. The real risk in this story belongs to anyone running actual infrastructure behind their site, a custom backend, exposed cloud services, self-hosted tools. If that's you, this is worth more than a passing read.
I didn't go looking for a security story. I went looking for AI crawler activity and found one anyway, which is probably the most honest way to say how much this space keeps overlapping with itself. There's been a lot of big, abstract conversation lately about AI safety, the kind that reaches for War Games and machines making decisions no one fully understands. This is a smaller, more mundane version of the same worry, not a rogue system deciding anything, just people exploiting the fact that we've started trusting a name in a request header the way we used to trust a name on a badge. The stakes here are lower than a war room. The lesson about misplaced trust isn't.



Comments