GreyNoise’s sensors caught a polite visitor in late August. The User-Agent said ClaudeBot, or GPTBot, or Google-Extended, names most content teams would welcome. The paths were uglier: .env files, cloud credential dumps, backup archives, leftover configs sitting on web roots. Help Net Security relayed the finding. Attackers are posing as AI crawlers from OpenAI, Anthropic, Google, Perplexity, and their peers while they hunt exposed secrets. If your cybersecurity program treats a famous User-Agent as a hall pass, you already built their on-ramp.

Nothing in the Request Proves the Name
The GreyNoise observation is the kind of finding that should change a rule, not a slide. Help Net Security quoted the researchers on the failure that makes this work:
Every program that visits a website announces itself in one line of the request. Chrome says it is Chrome. Googlebot says it is Googlebot. Anthropic’s crawler says it is ClaudeBot. Nothing in the request itself proves any of it is true.
Tape that to the WAF console.
Operators have spent two years carving exceptions for AI crawlers. Legal wants the brand in the training data, or kept out of it, with equal confidence. Marketing wants citations. Engineering files the ticket you have already seen: please allow GPTBot, we are catching 403s. Nobody asked whether those 403s came from the vendor’s published ranges or from a fresh proxy in a country you do not do business with.
Spoofed crawlers do not need to win a brute-force fight against your SSH banner. They need one mislaid file. A staging host that inherited production secrets. A CI artifact under /downloads. An open .git. Volume can stay low, which is how the activity slips past threat detection tuned for noisy scanners. Your firewall still classifies the session as permitted crawler traffic. The log looks like the internet working as designed.
Verification exists, and it has existed for years. Major crawler operators publish IP ranges and often document reverse-DNS conventions. You can pin those ranges at the edge and treat every other claimant as unknown. Unknown can read public pages at a human pace. Unknown does not get the crawler fast-lane, and unknown does not get a free look at paths you never linked from the homepage.
When the costume works, incident response rarely starts with a dramatic endpoint popup. It starts with a cloud trail showing an access key that was born in a web-accessible file. By then the User-Agent has left the stage, and you are reconstructing which “ClaudeBot” actually walked off with the secret.
Adware Played the Same Trick
A different costume landed in download portals the same week. Kaspersky’s Securelist team unpacked ValleyRAT arriving behind an installer that presented as adware. People click through adware. They have been trained that it is a toolbar, a lecture from the browser, a nuisance you dismiss so you can get back to work. The payload was a backdoor. Trust the wrapper, skip the scrutiny. That is the same bet the fake crawler is making against your edge policy.

Nightmare Eclipse then published HardBreacher, an exploit aimed at Kaspersky Endpoint Security. Kaspersky told SecurityWeek it had patched the vulnerability. If you run that agent, read the operational fact and ignore the vendor theater: the threat-protection process on the box is software you have to patch and monitor. Defense in depth means you assume the agent can be abused, you treat its update SLA like OS patching, and you watch for odd handles and injections around it.
Two campaigns, one habit. Attackers wrap the hostile thing in a label your environment is socialized to forgive. Cyber security programs run on labels because labels scale. Allow this bot. Ignore this adware. Trust this process because its name matches the security suite. Labels are also how you get walked past the front desk.
Stop Treating User-Agents as Cybersecurity Controls
Do this week, without buying a product. Pull 30 days of web and CDN logs for User-Agents matching GPTBot, ClaudeBot, Claude-Web, Google-Extended, Bytespider, PerplexityBot, Amazonbot, and the rest of the public crawler roster. Join those rows to each vendor’s published source ranges. Every name-and-range mismatch is reconnaissance. Then read the URLs those IPs requested. If they touched credential-shaped paths, rotate those secrets the same day. Do not wait for a tidy change window. A key that sat in a web root is a key you should assume has been copied.
If you granted WAF or CDN bypasses so “AI crawlers” can index docs, put those bypasses behind vendor IP allowlists. String matches on User-Agent are the entire trick. A header is a claim. An authenticated source range is a control.
Ongoing security hardening is boring on purpose. Keep credential material off web roots, including the “temporary” zip a developer left in /backup. Default-deny hidden path enumeration; a client that walks .env, wp-config.php, .aws/credentials, and docker-compose.yml in one session has told you what it is. Rate-limit unknowns even when they smile in the header. Feed confirmed spoofed crawler IPs into edge blocks the same way you would any other scanner. Tune detections for path shape plus identity mismatch, not for the word “bot” in a log field.
For the adware path, treat unexpected installers as execution events in EDR policy. First-seen publishers and “optimizer” downloads from the browser should not get a courtesy pass because they look low-grade. For endpoint suites, patch on the vendor’s timeline and alert on processes that tamper with the security product itself. HardBreacher is public because someone published the research. The next exploit against your agent will not send a courtesy note first.
The next costume will use a different famous name. Your job is to stop granting power to names.
Sources
- Threat actors are posing as AI crawlers to hunt for exposed credentials
- ValleyRAT masquerading as adware
- Nightmare Eclipse Drops ‘HardBreacher’ Kaspersky Product Exploit
Take Control of Your Server Security
Don't let brute-force attacks slow you down. Try IPBan Pro risk-free for 30 days.
Secure. Automated. Lightweight.
Take Control of Your Server Security
Don't let brute-force attacks slow you down. Try IPBan Pro risk-free for 30 days.
Secure. Automated. Lightweight.
