Looking at your logs and seeing a name you don’t recognize? This page answers one question: what is this bot and what should you do about it. Check the operator, what the bot does, its robots.txt rules and the number of requests I recorded on my own site.
The “podrez.pl logs” column shows how many requests carrying a given bot name in the user-agent field I recorded on my site between 2 February and 5 August 2026. Read these numbers as a hint of what to expect, not as a market average.
Bot and user-agent list [searchable]
The table covers AI crawlers, classic search engines, SEO tools, link previews from messaging apps, monitoring services and plain HTTP libraries. Narrow the category or type a name – the table filters as you type, and when a single row matches you get a ready answer, plus a robots.txt entry wherever the bot has a usable token.
What do the bot categories mean?Eleven classes of bots, one sentence each.
- AI: training
- collects content so that models can learn from it
- AI: indexes and answers
- builds an index that a generative system quotes from
- User-triggered
- fetches a page because someone has just asked about something
- Tokens without a bot
- send no requests, they exist only as robots.txt entries
- Search engines
- classic search engines and their helper bots
- Link previews
- generate the link thumbnail in a messenger or on social media
- SEO and marketing
- SEO tools, audits and checks on ad landing pages
- Monitoring and security
- site availability, brand mentions, abuse scanning
- Research and archives
- academic corpora and page archiving
- Tools and libraries
- traffic triggered by a browser, plugin, app or someone’s own script
- Unknown or suspicious
- operator, purpose or user-agent authenticity has not been confirmed
The table scrolls sideways →
| Bot and operator | Purpose | robots.txt token | Honors it | Requests podrez.pl, 2 Feb – 5 Aug 2026 |
|---|---|---|---|---|
| GPTBotOpenAI | Model training | GPTBot | Yes | 9,367 |
| AmazonbotAmazon | Improving Amazon products; may also be used for model training | Amazonbot | Yes | 5,617 |
| ClaudeBotAnthropic | Model training | ClaudeBot | Yes | 4,398 |
| meta-externalagentMeta | Content collection, including for training | Meta-ExternalAgent | Stated by operator | 1,027 |
| BytespiderByteDance | Collects content for ByteDance models | Bytespider | No clear documentation | 619 |
| CCBotCommon Crawl | Open corpus of web pages, used among other things for training | CCBot | Yes | 483 |
| cohere-ai CohereBotCohere | Content collection | cohere-ai CohereBot | No clear documentation | 36 |
| MistralAI-TrainingMistral AI | Model training | MistralAI-Training | Yes | 0 |
| LinkupBotLinkup (linkup.so) | Stated index for a search API used by AI applications | LinkupBot | No clear documentation | 4,935 |
| OAI-SearchBotOpenAI | ChatGPT search index | OAI-SearchBot | Yes | 4,527 |
| ShapBotParallel | Parallel’s search index, feeds its API for AI agents | ShapBot | Yes | 2,990 |
| PerplexityBotPerplexity | Search index | PerplexityBot | Yes | 2,353 |
| YouBotYou.com | Search index | YouBot | Yes | 108 |
| Google-CloudVertexBotGoogle | Crawl ordered by the site owner for Vertex AI | Google-CloudVertexBot | Yes | 35 |
| Claude-SearchBotAnthropic | Search index | Claude-SearchBot | Yes | 28 |
| DuckAssistBotDuckDuckGo | Generative answers in DuckDuckGo | DuckAssistBot | Yes | 14 |
| meta-webindexerMeta | Index behind Meta AI search results | Meta-WebIndexer | No clear documentation | 13 |
| DiffbotDiffbot | Diffbot’s search engine and Knowledge Graph, not used for training | Diffbot | Yes, with possible contractual exceptions | 3 |
| MistralAI-IndexMistral AI | Search index for Vibe | MistralAI-Index | Yes | 0 |
| ChatGPT-UserOpenAI | Fetch on user request | OpenAI does not document it as a control token | May not apply | 6,005 |
| Google-GeminiNotebook formerly Google-NotebookLMGoogle | Fetches sources added by the user | none, server-side blocking only | Usually not | 3,310** |
| FeedFetcher-GoogleGoogle | Fetches RSS and Atom feeds on user request | none, user fetcher | Usually not | 1,322 |
| Claude-UserAnthropic | Fetch on user request | Claude-User | Yes | 572 |
| Google-Read-AloudGoogle | Reads a page aloud on user request | none, user fetcher | Usually not | 276 |
| Diffbot-UserDiffbot | Fetch on request of a tool user | Diffbot-User | Yes | 158 |
| MistralAI-UserMistral AI | Fetch on request of a Vibe user, with a link to the source | MistralAI-User | Yes | 148 |
| Perplexity-UserPerplexity | Fetch on user request | Perplexity-User | Usually not | 89 |
| Google-AgentGoogle | Google agents performing tasks on user request | none, server-side blocking only | Usually not | 0 |
| Amzn-UserAmazon | Fetch on user request, e.g. a question put to Alexa | Amzn-User | May not apply | 0 |
| meta-externalfetcherMeta | Fetch on user request | Meta-ExternalFetcher | Sometimes ignored | 0 |
| Google-PinpointGoogle | Fetches sources added by a Pinpoint user | none, user fetcher | Usually not | 0 |
| Google-CWSGoogle | Fetcher tied to the Chrome Web Store | none, user fetcher | Usually not | 0 |
| Shap-UserParallel | Fetches a page on request of a Parallel user | none – Parallel describes it as a visibility signal, not a control mechanism | Not applicable | 0 |
| Google-ExtendedGoogle | Consent for Gemini training and grounding. Every appearance of this name in your logs is an impersonation | Google-Extended | Control token, no user-agent | 31 impersonation only |
| Applebot-ExtendedApple | Consent for Apple Intelligence training. Every appearance of this name in your logs is an impersonation | Applebot-Extended | Control token, no user-agent | 5 impersonation only |
| GooglebotGoogle | Search engine, also feeds AI Overviews and AI Mode | Googlebot | Yes | 12,538 |
| bingbotMicrosoft | Bing search engine, also feeds Copilot | bingbot | Yes | 6,843 |
| ApplebotApple | Siri, Spotlight and Safari; also feeds Apple’s generative answers | Applebot | Yes | 3,627 |
| BaiduspiderBaidu | Chinese search engine; the -render variant renders JavaScript | Baiduspider | Yes | 2,546 |
| PetalBotHuawei | Petal Search engine | PetalBot | Yes | 2,524 |
| YandexBotYandex | Yandex search engine | YandexBot | Yes | 1,725 |
| Googlebot-ImageGoogle | Image indexing | Googlebot-Image Googlebot | Yes | 1,434 |
| SeznamBotSeznam | Czech search engine | SeznamBot | Yes | 1,146 |
| GoogleOtherGoogle | General-purpose crawler | GoogleOther | Yes | 681 |
| DuckDuckBotDuckDuckGo | Classic search results | DuckDuckBot | Yes | 666 |
| YandexRenderResourcesBot YandexImages YandexFavicons YaDirectFetcherYandex | JS rendering, images, favicons and landing pages for Yandex Direct ads | separate tokens e.g. YandexImages | Yes | 483 |
| ExabotExalead? – to verify | Historically the Exalead search engine. Check the full user-agent and the address it contains | Exabot | No clear documentation | 98 |
| BingPreview / bingbotMicrosoft | Page snapshots for previews in Bing. User-agent observed in my logs; current documentation shows this traffic as bingbot | no separate token; bingbot rules apply | Per bingbot rules | 88 |
| QwantbotQwant | French privacy-focused search engine | Qwantbot | Yes | 48 |
| Yahoo! SlurpYahoo | Yahoo search engine | Slurp | Yes | 9 |
| Y!J-DLCYahoo Japan? – to verify | Agent associated with Yahoo Japan, with no current confirmation from the operator | none | No clear documentation | 38 |
| Bravebot***Brave? – unconfirmed | User-agent attributed to the Brave Search index; the operator does not document it | Bravebot | No clear documentation | 42 |
| Storebot-GoogleGoogle | Google Shopping | Storebot-Google | Yes | 12 |
| SogouSogou | Chinese search engine | Sogou | Yes | 8 |
| Amzn-SearchBotAmazon | Search across Amazon products, including Alexa | Amzn-SearchBot | Yes | 5 |
| Googlebot-VideoGoogle | Video indexing | Googlebot-Video Googlebot | Yes | 2 |
| MojeekBotMojeek | Independent search engine with its own index | MojeekBot | Yes | 2 |
| BingVideoPreviewMicrosoft | Video previews in Bing | BingVideoPreview | Yes | 0 |
| YepBotAhrefs / Yep | The Yep search engine, currently only for IndexNow requests. The index itself is built by AhrefsBot | YepBot | Yes, also honors Crawl-delay | 0 |
| PinterestbotPinterest | Indexes content and products for Pins, updates prices and dead links | Pinterestbot | Yes; honors Crawl-delay up to a value of 1, treats higher values as 1 | 0 |
| GoogleOther-Image GoogleOther-VideoGoogle | GoogleOther variants for public images and video | GoogleOther-Image GoogleOther-Video GoogleOther | Yes | 0 |
| GoogleProducerGoogle | Feeds configured by a publisher in Google Publisher Center | none, user fetcher | Usually not | 0 |
| facebookexternalhitMeta | Preview of links shared across Meta services | facebookexternalhit | Sometimes ignored | 4,479 |
| Slackbot-LinkExpanding Slack-ImgProxy SlackbotSlack | Link previews, image fetching and Slack’s remaining requests | none | No – Slack states that it does not process robots.txt | 407 |
| SkypeUriPreviewMicrosoft | Link previews in Skype | none | No clear documentation | 327 |
| TwitterbotX (formerly Twitter) | Link previews in X cards | Twitterbot | No clear documentation | 273 |
| LinkedInBotLinkedIn | Preview of links shared on LinkedIn | LinkedInBot | No clear documentation | 270 |
| WhatsAppMeta | Preview of links sent in WhatsApp | none | No clear documentation | 166 |
| DiscordbotDiscord | Link previews on Discord | Discordbot | No clear documentation | 6 |
| MicrosoftPreviewMicrosoft | Page previews in Microsoft products | MicrosoftPreview | Yes | 0 |
| Google MessagesGoogle | Preview of links sent in Google Messages | none, user fetcher | Usually not | 0 |
| AhrefsSiteAuditAhrefs | Technical audit ordered by the site owner | AhrefsSiteAudit | Yes by default; the site owner can switch this off in the audit settings | 127,925 |
| Screaming Frog SEO Spiderwhoever runs the program | Manual crawl, usually your own or your agency’s | Screaming Frog SEO Spider | Depends on the settings | 7,787 |
| BarkrowlerBabbar | Link index | Barkrowler | Yes | 5,975 |
| AhrefsBotAhrefs | Index of links and content for Ahrefs and the Yep search engine | AhrefsBot | Yes | 5,347 |
| SemrushBotSemrush | Index of links and visibility data | SemrushBot | Yes | 4,639 |
| serpstatbotSerpstat | Backlink index | serpstatbot | Yes | 2,741 |
| MJ12botMajestic | Link index | MJ12bot | Yes | 1,543 |
| DataForSeoBotDataForSEO | Index of SEO data sold through an API | DataForSeoBot | Yes | 1,219 |
| SERankingBacklinksBotSE Ranking | Backlink index | SERankingBacklinksBot | Yes | 752 |
| SiteAuditBotSemrush | Technical audit on request | SiteAuditBot | Yes | 666 |
| SEBot-WASE Ranking | SE Ranking’s technical audit crawler | SEBot-WA | Yes by default; the site owner can change this | 528 |
| AdsBot-GoogleGoogle | Quality assessment of ad landing pages | AdsBot-Google | Yes, but ignores User-agent: * | 228 |
| Mediapartners-GoogleGoogle | Matching AdSense ads to page content | Mediapartners-Google | Yes, but ignores User-agent: * | 81 |
| AdsBot-Google-MobileGoogle | Quality assessment of mobile ad landing pages | AdsBot-Google-Mobile | Yes, but ignores User-agent: * | 39 |
| AdIdxBotMicrosoft | Quality control of landing pages for Microsoft Advertising | adidxbot | Yes | 0 |
| DotBotMoz | Crawler building the Moz Link Explorer link index | DotBot | No current clear documentation from the operator | 0 |
| rogerbotMoz | Audits of sites added to Moz Pro campaigns by their owners | rogerbot | No current clear documentation from the operator | 0 |
| OAI-AdsBotOpenAI | Verification of ad landing pages in ChatGPT | OAI-AdsBot | Yes* | 0 |
| UptimeRobotUptimeRobot | Site availability monitoring, usually set up by the owner | UptimeRobot | Not applicable | 51,781 |
| AwarioBotAwario | Brand mention monitoring | AwarioBot | Yes | 1,601 |
| NetcraftSurveyAgent CheckMarkNetwork StormIntelCrawler OI-Crawlersecurity research and domain inventory | Scanning the public surface of websites, domain statistics | separate tokens | No clear documentation | 64 |
| Buck (Hypefactors) RecordedFuturemedia monitoring | Tracking mentions of brands and topics | separate tokens | No clear documentation | 36 |
| Google-SafetyGoogle | Detecting malware and abuse behind publicly shared links | none | No – Google states that it ignores robots.txt | 0 |
| archive.org_botInternet Archive | Archiving pages in the Wayback Machine | archive.org_bot | No current clear documentation from the operator | 381 |
| ClueWeb-CrawlerCarnegie Mellon University | Building the public ClueWeb corpus for search research | ClueWeb-Crawler | Yes | 82 |
| Aranea Web-Crawled CorporaSlovak Academy of Sciences | Building public language corpora | none | No clear documentation | 79 |
| curl Wget python-requests Scrapy Go-http-client axios and relatedany script | Not bots, just libraries for fetching pages. Behind each one stands a person or someone else’s program | none | Not applicable – depends on the script, the library enforces nothing | 22,176 |
| Google-InspectionToolGoogle | Search Console testing tools | Google-InspectionTool Googlebot | Yes | 4,874 |
| RankMath Link CheckerRank Math | SEO plugin checking links. If you run Rank Math, this is your own traffic | none | Not applicable | 1,408 |
| Chrome Privacy Preserving Prefetch ProxyGoogle | Page prefetching for Chrome users. Looks like a bot, but a browser stands behind it | none | Not applicable | 1,234 |
| Inoreader FeedlyRSS readers | Fetching an RSS feed on behalf of a subscriber | separate tokens | Usually yes | 321 |
| WP Rocket Pre-fetchWP Rocket | Cache plugin prefetching links from your own site. Also your own traffic | none | Not applicable | 43 |
| datasets library (Hugging Face)any user | A Python library for building datasets. Someone was pulling content into a set of their own | none | Not applicable – depends on the script | 14 |
| APIs-GoogleGoogle | Delivering notifications from Google APIs, e.g. PubSubHubbub | APIs-Google | Yes, but ignores User-agent: * | 0 |
| Google-Site-VerificationGoogle | Checking site ownership verification in Search Console | none, user fetcher | Usually not | 0 |
| 2ip bot2ip.io | Service for checking IP addresses and websites | 2ip bot | No clear documentation | 37,237 |
| br-crawler crawler_eb_germany Hermes-SVF-static-crawler VelenPublicWebCrawler and othersnot established | No documentation and no way to establish the operator. Judge them by behavior | none | Unknown | 3,900 |
| BingSapphireMicrosoft? – to verify | Absent from Microsoft’s official crawler list. Check the IP address | none | No clear documentation | 76 |
| EmailCrawlerunknown | Harvesting email addresses. The one entry where blocking needs no thought | EmailCrawler | Unknown; if you want certainty, block it on the server | 21 |
Urgent if you block NotebookLM. On 16 July 2026 Google changed the user-agent name from Google-NotebookLM to Google-GeminiNotebook. The old name is supported only for a transition period indicated as August 2026. A rule in .htaccess or in your WAF written against the old name will soon stop catching anything. On my site this bot made 3,310 requests, so the problem is not theoretical.
How this table was built
I checked purposes and tokens in the operators' documentation everywhere such documentation exists. Entries with no official source are described on the basis of the full user-agent, the address given in its string and its behavior in the logs – a missing confirmation is clearly marked. The frequency data comes from over 1.8 million requests in the podrez.pl logs for the period from 2 February to 5 August 2026.
Some entries group several related user-agents from one operator, so there are more names than rows.
Compiled by Ewelina Podrez-Siama, working in SEO since 2009. Last verified: 6 August 2026.
Sources: bot operator documentationTwenty-two documents in which I checked tokens, purposes and rules. Check for yourself if anything raises a doubt.
OpenAI
Overview of OpenAI Crawlers – GPTBot, OAI-SearchBot, ChatGPT-User, OAI-AdsBot
Advertiser Guidance for Allowing OpenAI Web Crawlers – the OAI-AdsBot token and rule compliance
Anthropic
Does Anthropic crawl data from the web… – ClaudeBot, Claude-User, Claude-SearchBot
Perplexity
Perplexity Crawlers – PerplexityBot, Perplexity-User, IP address lists
Google
Google's common crawlers – Googlebot, GoogleOther, Google-Extended, Google-CloudVertexBot
Google special-case crawlers – AdsBot, Mediapartners, APIs-Google, Google-Safety
Google user-triggered fetchers – Gemini Notebook, Google-Agent, GoogleProducer
Verify requests from Google – three classes of bots, three JSON files, reverse DNS
Crawler documentation changelog – change dates, including 16.07.2026 and 20.03.2026
AI Features and Your Website – preview directives with regard to AI Overviews and AI Mode
Apple
About Applebot – Applebot, Applebot-Extended, inheriting rules from Googlebot
Amazon
About Amazonbot – Amazonbot, Amzn-SearchBot, Amzn-User, three IP address lists
Mistral AI
Mistral crawlers – MistralAI-Training, MistralAI-Index, MistralAI-User
Microsoft
Overview of Bing crawlers – bingbot, adidxbot, BingPreview, MicrosoftPreview, BingVideoPreview
Other operators
Common Crawl – CCBot
Parallel – ShapBot and Shap-User
Slack Robots – three agents and a statement about not honoring the rules
Yep – YepBot
Pinterest – Pinterestbot
Diffbot – Does Diffbot respect robots.txt?
Brave Search Crawler
Context
Cloudflare – Your site, your rules: new AI traffic options – the Search, Agent and Training categories and the change of default settings from 15.09.2026
All links checked on 6 August 2026. Entries marked in the table as "no clear documentation" have no counterpart on this list – I describe them solely on the basis of the user-agent and behavior in the logs.
* OAI-AdsBot. OpenAI's developer documentation lists only GPTBot and OAI-SearchBot as control tokens, but the help article for advertisers gives an explicit User-agent: OAI-AdsBot example and declares that the rules are honored. I go with the latter, because it is newer and more detailed.
** Gemini Notebook. All 3,310 requests arrived under the old name Google-NotebookLM. The new one has not appeared in my logs even once.
*** Bravebot. The one entry where the documentation contradicts observation. The Brave Search help page says its crawler does not identify itself with a separate user-agent, and yet requests signed Bravebot and pointing to search.brave.com are recorded by independent registries and by my logs. The documentation may be out of date, but without confirmation from the operator I can't rule out impersonation.
The "honors robots.txt" column. It is based on operator documentation, not on my own tests. "Stated by operator" means the operator says so, though reports of deviations circulate. "No clear documentation" means exactly what it says – I found no reliable source and I am not guessing.
Why doesn't robots.txt block every AI bot?
An entry in robots.txt does not work the same way on every bot. Crawlers that move around the web on their own initiative usually follow it. With fetchers triggered by a user things get complicated, because here every operator has a policy of its own.
Automated crawlers
They move around the web on the operator's initiative: GPTBot, ClaudeBot, PerplexityBot, Googlebot. They honor
robots.txtand this file was created with them in mind.User-triggered fetchers
They fetch a page because a specific user action required it. Operators differ here. Anthropic and Mistral provide separate tokens (
Claude-User,MistralAI-User) and declare that they honor them. Google, OpenAI, Perplexity and Amazon note that the rules may not apply, because the request is initiated by a user.Control tokens without a bot
Google-ExtendedandApplebot-Extendedsend no requests at all. You put them inrobots.txt, but you will never see them in your logs. Google-Extended covers training of Gemini models and grounding, meaning content passed to the model at the moment of answering. Applebot-Extended covers training of Apple's models.
The practical conclusion: before you block a fetcher triggered by a user, check the policy of that particular operator. A Disallow for Claude-User or MistralAI-User genuinely closes the door, because Anthropic and Mistral declare that they honor those tokens. With Google, OpenAI, Perplexity and Amazon the same entry is a request, so if you need certainty, add a rule at the server or WAF level. How to check this on your own site, I described in the article on server log analysis.
Suspicious user-agents: impossible names and names to verify
Five names turned up in my logs that cannot be reconciled with the documentation of their supposed operators. They are worth knowing, because they look credible and slip through a filter easily.
Two of these names cannot come from the operator, because its own documentation says those tokens do not crawl at all. Three more do not appear in the current documentation of their supposed operators – that does not prove impersonation, but it does mean you should check the IP address before treating them as genuine.
| Name in the log | Status | Why | What it was looking for on my site |
|---|---|---|---|
Google-Extended | Impossible | A control token in robots.txt; Google does not send requests with it | /.aws/credentials |
Applebot-Extended | Impossible | Apple states outright that this token does not crawl pages | /privatekey.key |
BingIndexCrawler | To verify | The name does not appear in Microsoft's current crawler documentation; Bing crawls as bingbot | ordinary subpages |
anthropic-ai | To verify | A name outside Anthropic's current documentation | /secrets.yml |
Claude-Web | To verify | A name outside Anthropic's current documentation | /_profiler/open, /info.php |
Addresses come from the podrez.pl logs, February–August 2026. Number of occurrences: Google-Extended 31, anthropic-ai 28, Claude-Web 20, BingIndexCrawler 10, Applebot-Extended 5.
The pattern is clear: suspicious requests reach for tokens that are not user-agents, and for names absent from operators' current documentation. It is worth distinguishing intent here. Requests signed as Google-Extended or anthropic-ai were looking for password and configuration files on my site, so that is a scanner. A name that does not match the documentation need not mean bad intent, though – which is why the IP address decides, not the name alone. The problem is not limited to my site either: Common Crawl explicitly warns in its documentation that it knows of crawlers falsely claiming to be CCBot, and recommends verifying the user-agent.
How to check whether a bot is real: verification by IP address
A user-agent can be faked in ten seconds, so the name alone proves nothing. Certainty comes only from matching the name with the IP address and the operator's official ranges, and where possible with a reverse and forward DNS check. Google is also testing cryptographic authentication of bots, that is Web Bot Auth.
| Operator | Source of addresses |
|---|---|
| OpenAI | openai.com/gptbot.json, openai.com/searchbot.json, openai.com/chatgpt-user.json, openai.com/adsbot.json |
| Anthropic | claude.com/crawling/bots.json |
| Perplexity | perplexity.com/perplexitybot.json, perplexity.com/perplexity-user.json |
| Google – common crawlers | common-crawlers.json; reverse DNS ends in googlebot.com |
| Google – special-case crawlers | special-crawlers.json; reverse DNS ends in google.com |
| Google – user fetchers and agents | separate JSON files; reverse DNS ends in google.com or gae.googleusercontent.com |
| Mistral | mistral.ai/mistralai-user-ips.json, mistral.ai/mistralai-index-ips.json; for MistralAI-Training the list is not published |
| Amazon | three separate lists: /ip-addresses/, /searchbot-ip-addresses/, /live-ip-addresses/ |
| Parallel | docs.parallel.ai/resources/shapbot.json |
| Microsoft | the Verify Bingbot tool in Bing Webmaster Tools; reverse DNS ends in search.msn.com |
| Common Crawl | index.commoncrawl.org/ccbot.json; reverse DNS ends in crawl.commoncrawl.org |
How to block AI bots in robots.txt? Ready-made examples.
Before you block anything: in most cases I do not recommend it. A block cuts you off from generative answers, and so from a channel where your brand can be mentioned and cited. For most companies, experts and shops that is a loss, not a gain.
The exceptions are real and concern publishers, portals and media above all – those who make a living from selling content or from traffic on their own site. There, a block can be a justified business decision – but still a decision, not a default setting.
So treat the configurations below as a starting point for an informed choice, not as a recommendation.
The first variant limits the main crawlers used for model training. The second blocks separately identified AI bots and fetchers, leaving access open to Googlebot, bingbot and Applebot. Both are a starting point, not a universal configuration for every site.
Variant 1 · No training
I protect my contentI let in everything that can cite me and cut off content collection for model training.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Bytespider
User-agent: Meta-ExternalAgent
User-agent: MistralAI-Training
User-agent: cohere-ai
User-agent: CohereBot
User-agent: Amazonbot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: DuckAssistBot
User-agent: MistralAI-Index
User-agent: MistralAI-User
User-agent: Amzn-SearchBot
User-agent: ShapBot
Allow: /
Variant 2 · Separate AI bots
Google and Apple stayI block separate training crawlers, AI indexes and user fetchers, but leave classic search engines open. This does not remove the site from AI Overviews, AI Mode or Apple's generative answers.
# search engines stay
User-agent: Googlebot
User-agent: bingbot
User-agent: Applebot
Allow: /
# model training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Bytespider
User-agent: Meta-ExternalAgent
User-agent: MistralAI-Training
User-agent: cohere-ai
User-agent: CohereBot
User-agent: Amazonbot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
# indexes behind generative answers
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: MistralAI-Index
User-agent: Meta-WebIndexer
User-agent: DuckAssistBot
User-agent: Amzn-SearchBot
User-agent: LinkupBot
User-agent: ShapBot
User-agent: YouBot
User-agent: Diffbot
Disallow: /
# fetches on user request
User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User
User-agent: MistralAI-User
User-agent: Meta-ExternalFetcher
User-agent: Amzn-User
Disallow: /
The second variant has a hole and it is better to know about it right away. The last group are fetchers triggered by a user, and some of them may ignore robots.txt. Gemini Notebook and Google-Agent do not even have a token of their own. If the block is to be airtight, it has to be added at the server level. An example for Apache and the .htaccess file:
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} (Google-GeminiNotebook|Google-NotebookLM|Google-Agent|GoogleProducer|ChatGPT-User|Perplexity-User|Amzn-User|Meta-ExternalFetcher|Shap-User) [NC]
RewriteRule ^ - [F,L]
One caveat that is easy to forget: robots.txt is not a tool for protecting data. The file is public, the directives are not access control, and a blocked URL can still end up in the index on the strength of external links. Genuinely private sections are secured with server-side authentication.
Watch out for groups and for the wildcard. If a more specific matching group exists for a bot, it takes precedence over User-agent: *. The wildcard applies only when nothing more specific matches. Fetchers triggered by a user are a separate case: some honor dedicated tokens, some may skip them and need a rule at the server or WAF level.
How to limit AI Overviews? On nosnippet, data-nosnippet and max-snippet
robots.txt tells compliant bots whether they should fetch a page. Meta directives, by contrast, control indexing and which fragments can reach the results and the AI features. These are two different layers, and with AI Overviews the second one is often the only one you have left.
The key point
There is no separate opt-out from AI Overviews and AI Mode that would leave your ordinary presence in Google untouched. They are fed by the same Googlebot and the same index, and Google-Extended covers only Gemini training and grounding, so blocking it changes nothing here.
You can, however, limit how much content reaches those features. Google lists four directives: nosnippet, data-nosnippet, max-snippet and noindex. Each of them works on ordinary results and on AI features at the same time – you cannot separate one from the other.
| Directive | What it does | What it costs you |
|---|---|---|
nosnippet | Blocks the display of a page fragment in every form of results, including AI Overviews and AI Mode | The ordinary description under the title also disappears, along with the chance of a direct answer |
data-nosnippet | The same, but only for the marked HTML fragment | The rest of the page works normally – the most precise tool in this set |
max-snippet:N | Limits the fragment length to N characters and with it the amount of content passed to the model | At a low value the effect comes close to nosnippet |
noindex | Removes the page from every form of Google results, including AI Overviews and AI Mode | The page also disappears from ordinary search results |
noarchive | At Amazon it means "do not use this page to train models" | Works only with operators that honor this directive |
Apple honors nosnippet as an opt-out from the use of content in generative answers based on current information. Amazon declares that it honors noarchive, noindex and none, but does not support crawl-delay.
The most precise tool is data-nosnippet. It excludes only the marked fragment of a page from previews and from direct use by AI features. Use it deliberately: the very paragraphs that would most easily replace a visit to your site – definitions, step lists, tables – are usually the ones that get you cited in the first place.
What AI bots take from your site: content or files?
The number of visits alone says little. What is more interesting is how many of them land on a page with content, and how many on CSS files, images and API endpoints.
| Bot | Successful requests | Share of hits on content |
|---|---|---|
| Perplexity-User | 57 | 91.2% |
| CCBot | 405 | 82.5% |
| ChatGPT-User | 5,870 | 73.9% |
| PerplexityBot | 2,127 | 61.1% |
| Claude-User | 389 | 60.2% |
| OAI-SearchBot | 3,157 | 33.7% |
| GPTBot | 5,283 | 32.0% |
| ClaudeBot | 2,244 | 28.6% |
| Google-GeminiNotebook | 2,945 | 20.0% |
| Applebot | 3,192 | 15.0% |
| Amazonbot | 3,337 | 14.6% |
podrez.pl logs, 2 Feb – 5 Aug 2026. Perplexity-User, ChatGPT-User and Claude-User land mostly straight on content, but the class of bot alone does not determine its behavior: CCBot has a high share of hits on content, and Gemini Notebook a low one.
Frequently asked questions
How do I check which bot visited my site?
The bot name sits in the user-agent field of your server's access log, usually in access.log or domain.com.log in the logs directory. Copy the name and search for it in the table above. If you want to be sure the bot is genuine, compare its IP address with the ranges published by the operator.
Does blocking GPTBot remove me from ChatGPT results?
No. GPTBot is responsible for collecting content for model training, while presence in ChatGPT search results is handled by OAI-SearchBot. These are two independent tokens and you can set them differently: block training, leave search open. Anthropic and Mistral apply an analogous three-way split. Amazon also divides its traffic across three agents, but its Amazonbot is a multi-purpose crawler that only may be used for training. Perplexity, by contrast, separates just the search index from fetches on user request and has no documented training crawler.
Why don't I see Google-Extended in my logs?
Because Google-Extended is not a bot. It is a control token in robots.txt with which you decide about the use of your content for training Gemini models and for grounding. The fetching itself is done by Google's ordinary robots. If you do see this name in your log, you are looking at an impersonator.
How do I opt out of AI Overviews?
Through robots.txt you can't. AI Overviews and AI Mode use Google's ordinary index, built by Googlebot, and the Google-Extended token covers only the training of Gemini models and grounding. The only available tool is the preview directives: nosnippet, data-nosnippet, max-snippet and noindex. Each of them also limits your ordinary search result, so leaving AI Overviews means partly leaving Google.
Will a robots.txt block stop every AI bot?
No. Fetchers triggered by a user, such as Gemini Notebook or ChatGPT-User, may skip robots.txt, because the request is initiated by a person, not a machine. To stop them you need a rule at the server level, in your WAF configuration or in the .htaccess file.
What changed with NotebookLM?
On 16 July 2026 Google updated its documentation and changed the user-agent name from Google-NotebookLM to Google-GeminiNotebook, after the product itself was renamed Gemini Notebook. The old name is supported only for a transition period indicated as August 2026, so rules written against the old name need updating.
The name in the log is only the beginning. Far more interesting is what that bot got in response – and whether it got anything at all. How to check that and what I found on my own site, I described in the article on server log analysis.
Not sure which bots your site lets in?
Log analysis covering AI bots, search engine crawlers, SEO tools and suspicious traffic is part of my audit. If you would rather start with a conversation, book a short call.
History of this page 06.08.2026 – first version.
Timeline of changes at the operators 16.07.2026 – Google changed the NotebookLM user-agent name to
Google-GeminiNotebook; the old name supported for a transition period.01.07.2026 – Cloudflare introduced a split of bots into the Search, Agent and Training categories.
07.04.2026 – Anthropic split its documentation into three bots, adding
Claude-SearchBot and Claude-User.20.03.2026 – Google documented the
Google-Agent user-agent.
