In this article
The hosting bill has gone up and the site slows down at odd hours, but sales and enquiries have not moved. The access logs explain it: a large share of the requests are not people. The natural reaction is to block everything that looks like a bot. That is how websites lose their search traffic, because the crawler that brings customers and the scraper that wastes bandwidth can look alike in a log file.
This guide shows how to tell them apart, which controls work on which kind of bot, and how to check that a change has not cost you visibility. It covers search crawlers, the newer AI crawlers, rate limits and firewall rules, and it names the official documentation behind each claim. It does not promise a ranking outcome. Nobody can.
Sort the traffic into groups before you touch a setting
Bot traffic is not one thing. Each group below wants something different from your site, identifies itself differently, and responds to different controls.
| Traffic | Examples | Default action | Why |
|---|---|---|---|
| Verified search crawlers | Googlebot, Bingbot | Allow. Never challenge. | They are how customers find you. Verify the identity; do not trust the name. |
| Declared AI training crawlers | GPTBot, ClaudeBot | Your decision, made in robots.txt | They identify themselves and say they honour robots.txt. |
| Declared AI search crawlers | OAI-SearchBot, Claude-SearchBot, PerplexityBot | Allow, if you want to appear in AI answers | Blocking them removes you from those results. |
| User-triggered fetchers | ChatGPT-User, Claude-User, Perplexity-User | Allow, or rate-limit | A person asked an assistant to read your page. These may not follow robots.txt. |
| Tools you invited | Uptime monitors, SEO auditors, payment gateway callbacks | Allow by IP address or token | You depend on them. Blocking them breaks your own operations. |
| Unidentified automation | Scrapers, vulnerability scanners, login attacks, fake Googlebot | Rate-limit, challenge or block | No identity and no benefit to you. |
Read the logs before changing anything
A week of access logs, or the equivalent report from your CDN, answers most of the questions that matter. You are looking for what costs you resources, not for a list of every bot.
- Which user agents and IP ranges send the most requests, and transfer the most data?
- Which URLs do they request? Internal search, filters, calendars and cart pages are expensive to generate and can produce an endless number of URLs.
- What does your server answer? A high rate of errors or timeouts to crawlers is a visibility problem in its own right.
- When are the peaks, and do they line up with the slow periods your customers notice?
Two free reports add to the raw logs. Google Search Console's Crawl Stats report shows how many requests Google made, the response codes it received, and whether it could reach your robots.txt. If you use Cloudflare, AI Crawl Control lists the AI crawlers visiting your site, how often, and whether they follow your robots.txt.
Often the finding is one crawler caught in a loop of filter combinations, or one scraper hitting a search page. That is a small, specific problem, and it needs a small, specific fix.
Verify the crawler. The name is only a claim.
The user agent is a line of text the visitor sends. Anyone can send "Googlebot". Bad bots do, because they know site owners are afraid to block it. The search engines publish ways to check.
| Operator | Verification they document |
|---|---|
| Reverse DNS lookup on the IP address, confirm the host name ends in googlebot.com, google.com or googleusercontent.com, then a forward lookup that returns the same IP. Or match the IP against Google's published range files. | |
| Bing | Reverse DNS lookup ending in search.msn.com, then a forward lookup that matches. Bing also offers a verification tool and a published IP list. |
| OpenAI, Anthropic, Perplexity | Published IP address lists for each of their crawlers. |
Most businesses should not build this themselves. A CDN or web application firewall that classifies verified bots does the same checks continuously. Cloudflare maintains a verified bots list, and AWS WAF Bot Control labels verified bots so that your rules can treat them differently. If you maintain your own allow-list, note that Google moved and renamed its range files in February 2026. A list copied from an old tutorial may be out of date.
A request that claims to be Googlebot and fails verification is the easiest decision in this guide. Block it.
robots.txt is a request, not a lock
Much of the confusion about bots comes from expecting robots.txt to do a job it was never designed for. The standard that defines it, RFC 9309, says outright that its rules are not a form of access authorisation. Google's documentation says the same: robots.txt manages crawler traffic, and it is not a way to keep a page out of search results.
- Tells well-behaved crawlers which paths not to crawl.
- Depends entirely on the crawler choosing to comply.
- Is public. Anyone can read which paths you listed.
- Does not remove a page from search results. Use noindex for that.
- Right for steering crawlers away from filters and internal search, and for opting out of declared AI training crawlers.
- Refuses, slows or challenges a request whatever the visitor intends.
- Works on bots that ignore robots.txt.
- Can act on verified identity, IP range, request rate or behaviour.
- Can block real crawlers and real customers if the rule is careless.
- Right for scrapers, abusive request rates, login attacks and anything private.
- Do not list private paths in robots.txt to hide them. Protect them with a login.
- A page blocked in robots.txt cannot show Google its noindex tag, because Google is not allowed to fetch the page. To remove a page from results, allow the crawl and use noindex.
- Google does not support the crawl-delay directive. Bing does, and Anthropic says its crawler does. Check each operator.
- Keep robots.txt itself reachable. Google's documentation says that if the file returns a server error, it stops crawling the site for a period. Exclude robots.txt and your sitemap from firewall challenges and rate limits.
AI crawlers do three different jobs
The companies behind the main AI assistants now publish separate crawlers for separate purposes. Treating them as one group is the commonest mistake, because the decision is different for each.
| Operator | Model training | Search and answers | On a user's request |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Perplexity | Says PerplexityBot is not used for training | PerplexityBot | Perplexity-User |
| Google-Extended, a robots.txt token and not a separate crawler | Googlebot | Google's user-triggered fetchers |
- Training crawlers collect content to train models. Whether to allow them is a business decision about how your content is used. These operators say their training crawlers honour robots.txt.
- Search crawlers build the index an assistant draws on when it answers with links. OpenAI states that a site which opts out of OAI-SearchBot will not be shown in ChatGPT search answers. If customers ask assistants for recommendations in your market, blocking these has a cost.
- User-triggered fetchers act when a person asks an assistant to open a page. OpenAI and Perplexity both say robots.txt rules may not apply to these, on the reasoning that a person made the request. If you need to stop them, that takes a firewall rule.
Does blocking AI crawlers affect Google rankings?
Here is what Google has put in writing. Google-Extended, the token that controls whether your content is used for Gemini training and grounding, does not affect a site's inclusion in Google Search and is not a ranking signal. Separately, Search Console now has a Search generative AI setting, available to all sites since August 2026, that lets you exclude your site from AI Overviews and AI Mode. Google says that setting is not used as a ranking signal for other parts of Search either.
Google has published nothing about other companies' crawlers. We found no official statement that blocking GPTBot or ClaudeBot changes a Google ranking in either direction, and there is no obvious mechanism by which it would. Be cautious about confident claims on this point, whichever way they lean.
The real risk is collateral damage: a switch labelled "AI bots" that catches a search crawler as well. Two current examples, both from the vendors' own documentation:
- Cloudflare replaced its "Block AI bots" setting in September 2026 with separate Search, Training and Agent policies. Its announcement states that setting Training to Block also stops Googlebot, Bingbot and Applebot, because those crawlers serve both search and AI. A separate option, Disallow AI Training, keeps them for search.
- AWS WAF Bot Control does not block verified bots in general, but its AI category rule blocks AI bots whether they are verified or not.
Rate limits, challenges and firewall rules, used selectively
For traffic that does not identify itself or does not behave, robots.txt is irrelevant and enforcement is the only tool. The aim is to make abuse expensive without putting anything in front of customers or verified crawlers.
- Make pages cheap before you block anyoneCache pages for anonymous visitors at the CDN, and stop filter and sort combinations from generating unlimited URLs. A request served from cache costs you very little, whoever sent it.
- Rate-limit the expensive URLsSearch, filters, login, cart and API endpoints first. Limit by IP address or session for unverified clients, and exempt verified search crawlers. Cloudflare's free plan includes one rate limiting rule. AWS WAF rate-based rules count requests over a window of one to ten minutes.
- Challenge suspicious traffic on pages onlyA browser challenge suits traffic that is probably automated but not certainly. Never apply one to robots.txt, sitemaps, feeds, APIs or payment callbacks, because none of those clients can solve it.
- Block what is clearly hostileFake crawlers that fail verification, scanners probing for software you do not run, and sources that keep going after being rate-limited.
- Try a rule in log-only mode firstWhere your firewall can count or log matches without acting, run a new rule that way for a few days and read what it would have blocked.
One Cloudflare detail deserves a mention, because it surprises people. Bot Fight Mode on the free plan cannot be bypassed with custom rules. If you have monitoring tools, API clients or payment callbacks reaching your site, test them after enabling it.
If your server is overloaded and you need Google itself to slow down, Google documents a way: return a 500, 503 or 429 status to its crawler for a short period. The same page warns against doing so for longer than a day or two, because URLs that keep returning errors are eventually dropped from the index. It is an emergency brake.
After every change, watch crawling and indexing
A blocking rule that goes wrong does not announce itself. Traffic falls a few weeks later, and by then the cause is hard to trace. Change one thing at a time, write down the date, and check these afterwards.
Seven mistakes that cost search visibility
- Blocking every user agent that contains the word "bot". Googlebot and Bingbot contain it too.
- Challenging or blocking whole countries or data-centre networks without exempting verified crawlers, which operate from data centres.
- Putting robots.txt or the sitemap behind a firewall challenge.
- Using robots.txt to hide pages, then finding them in results with no description.
- Blocking the CSS and JavaScript files a crawler needs to render the page.
- Leaving an emergency 503 or a strict rate limit in place for weeks.
- Switching on an AI bot control without reading what it covers.
Block what you can identify as harmful, allow what you can verify as useful, and slow down everything in between.
A one-week plan
- Days 1 and 2: measureExport a week of logs. Rank user agents, IP ranges and URLs by requests and data transferred. Read Crawl Stats.
- Day 3: decide the policyWrite down which groups you allow, limit and block, and your position on AI training and AI search crawlers. Update robots.txt for the crawlers that honour it.
- Day 4: reduce the costAdd caching for anonymous pages and close the URL traps the logs revealed.
- Day 5: enforce, in log-only modeAdd rate limits on expensive URLs and a rule for fake crawlers. Read what they would have blocked.
- Days 6 and 7: switch on and watchEnable the rules one at a time and work through the monitoring checklist. Repeat the log review monthly, because crawler names and behaviour change.
Techimpace reviews websites for exactly this: what the automated traffic is, what it costs, and what can be limited without touching search visibility. The review covers logs, caching, robots.txt, firewall rules and the crawl reports, and ends with a short list of changes in the order to make them.
Frequently asked questions
Will blocking bots hurt my SEO?
Only if the rule catches verified search crawlers or the files they need. Blocking unidentified scrapers and fake crawlers does not affect search. Verify crawler identity, exempt verified search crawlers from challenges and rate limits, and check Search Console after each change.
Does robots.txt stop bots from accessing my website?
No. It is a published request that well-behaved crawlers choose to follow. It has no enforcement. To stop a bot that ignores it, you need a rule on your server, CDN or firewall.
How can I tell whether a request is really from Googlebot?
Do a reverse DNS lookup on the IP address, confirm the host name ends in googlebot.com, google.com or googleusercontent.com, then do a forward lookup and confirm it returns the same IP. You can also match the address against Google's published IP ranges, or rely on a CDN that classifies verified bots.
Should I block AI crawlers such as GPTBot and ClaudeBot?
It is a business decision. Training crawlers and AI search crawlers are separate, so you can opt out of training while staying visible in AI search results. Decide each one deliberately in robots.txt.
Does blocking AI crawlers affect Google rankings?
Google says its own AI control, Google-Extended, is not a ranking signal and does not affect inclusion in Search. Google has published nothing about other companies' AI crawlers and rankings. The practical risk is a setting that blocks a search crawler along with the AI ones.
How can I reduce bot load without blocking anything?
Cache pages for anonymous visitors, stop filters and internal search from generating unlimited URLs, and rate-limit expensive endpoints. These lower the cost of each automated request, which is often enough.
Can Techimpace review bot traffic on our website?
Yes. A review covers the access logs, caching, robots.txt, firewall rules and the crawl and indexing reports, and produces a prioritised list of changes. We do not guarantee ranking outcomes.
Paritosh Bag is a software engineer and the Founder & CEO of Techimpace Innovations Pvt Ltd. He has been building business software since 2010 and has led Techimpace since founding it in 2013, working across PHP and Laravel, JavaScript, cloud infrastructure, payment systems and AI automation.