What bot is this? List of AI crawlers, search bots and user-agents

What bot is this? List of AI crawlers, search bots and user-agents

Looking at your logs and seeing a name you don’t recognize? This page answers one question: what is this bot and what should you do about it. Check the operator, what the bot does, its robots.txt rules and the number of requests I recorded on my own site.

The “podrez.pl logs” column shows how many requests carrying a given bot name in the user-agent field I recorded on my site between 2 February and 5 August 2026. Read these numbers as a hint of what to expect, not as a market average.

Bot and user-agent list [searchable]

The table covers AI crawlers, classic search engines, SEO tools, link previews from messaging apps, monitoring services and plain HTTP libraries. Narrow the category or type a name – the table filters as you type, and when a single row matches you get a ready answer, plus a robots.txt entry wherever the bot has a usable token.

1. Narrow the category

What do the bot categories mean?Eleven classes of bots, one sentence each.

AI: training
collects content so that models can learn from it
AI: indexes and answers
builds an index that a generative system quotes from
User-triggered
fetches a page because someone has just asked about something
Tokens without a bot
send no requests, they exist only as robots.txt entries
Search engines
classic search engines and their helper bots
Link previews
generate the link thumbnail in a messenger or on social media
SEO and marketing
SEO tools, audits and checks on ad landing pages
Monitoring and security
site availability, brand mentions, abuse scanning
Research and archives
academic corpora and page archiving
Tools and libraries
traffic triggered by a browser, plugin, app or someone’s own script
Unknown or suspicious
operator, purpose or user-agent authenticity has not been confirmed

The table scrolls sideways →

Bot and operatorPurposerobots.txt tokenHonors itRequests
podrez.pl, 2 Feb – 5 Aug 2026
GPTBotOpenAIModel trainingGPTBotYes9,367
AmazonbotAmazonImproving Amazon products; may also be used for model trainingAmazonbotYes5,617
ClaudeBotAnthropicModel trainingClaudeBotYes4,398
meta-externalagentMetaContent collection, including for trainingMeta-ExternalAgentStated by operator1,027
BytespiderByteDanceCollects content for ByteDance modelsBytespiderNo clear documentation619
CCBotCommon CrawlOpen corpus of web pages, used among other things for trainingCCBotYes483
cohere-ai
CohereBot
Cohere
Content collectioncohere-ai
CohereBot
No clear documentation36
MistralAI-TrainingMistral AIModel trainingMistralAI-TrainingYes0
LinkupBotLinkup (linkup.so)Stated index for a search API used by AI applicationsLinkupBotNo clear documentation4,935
OAI-SearchBotOpenAIChatGPT search indexOAI-SearchBotYes4,527
ShapBotParallelParallel’s search index, feeds its API for AI agentsShapBotYes2,990
PerplexityBotPerplexitySearch indexPerplexityBotYes2,353
YouBotYou.comSearch indexYouBotYes108
Google-CloudVertexBotGoogleCrawl ordered by the site owner for Vertex AIGoogle-CloudVertexBotYes35
Claude-SearchBotAnthropicSearch indexClaude-SearchBotYes28
DuckAssistBotDuckDuckGoGenerative answers in DuckDuckGoDuckAssistBotYes14
meta-webindexerMetaIndex behind Meta AI search resultsMeta-WebIndexerNo clear documentation13
DiffbotDiffbotDiffbot’s search engine and Knowledge Graph, not used for trainingDiffbotYes, with possible contractual exceptions3
MistralAI-IndexMistral AISearch index for VibeMistralAI-IndexYes0
ChatGPT-UserOpenAIFetch on user requestOpenAI does not document it as a control tokenMay not apply6,005
Google-GeminiNotebook
formerly Google-NotebookLM
Google
Fetches sources added by the usernone, server-side blocking onlyUsually not3,310**
FeedFetcher-GoogleGoogleFetches RSS and Atom feeds on user requestnone, user fetcherUsually not1,322
Claude-UserAnthropicFetch on user requestClaude-UserYes572
Google-Read-AloudGoogleReads a page aloud on user requestnone, user fetcherUsually not276
Diffbot-UserDiffbotFetch on request of a tool userDiffbot-UserYes158
MistralAI-UserMistral AIFetch on request of a Vibe user, with a link to the sourceMistralAI-UserYes148
Perplexity-UserPerplexityFetch on user requestPerplexity-UserUsually not89
Google-AgentGoogleGoogle agents performing tasks on user requestnone, server-side blocking onlyUsually not0
Amzn-UserAmazonFetch on user request, e.g. a question put to AlexaAmzn-UserMay not apply0
meta-externalfetcherMetaFetch on user requestMeta-ExternalFetcherSometimes ignored0
Google-PinpointGoogleFetches sources added by a Pinpoint usernone, user fetcherUsually not0
Google-CWSGoogleFetcher tied to the Chrome Web Storenone, user fetcherUsually not0
Shap-UserParallelFetches a page on request of a Parallel usernone – Parallel describes it as a visibility signal, not a control mechanismNot applicable0
Google-ExtendedGoogleConsent for Gemini training and grounding. Every appearance of this name in your logs is an impersonationGoogle-ExtendedControl token, no user-agent31
impersonation only
Applebot-ExtendedAppleConsent for Apple Intelligence training. Every appearance of this name in your logs is an impersonationApplebot-ExtendedControl token, no user-agent5
impersonation only
GooglebotGoogleSearch engine, also feeds AI Overviews and AI ModeGooglebotYes12,538
bingbotMicrosoftBing search engine, also feeds CopilotbingbotYes6,843
ApplebotAppleSiri, Spotlight and Safari; also feeds Apple’s generative answersApplebotYes3,627
BaiduspiderBaiduChinese search engine; the -render variant renders JavaScriptBaiduspiderYes2,546
PetalBotHuaweiPetal Search enginePetalBotYes2,524
YandexBotYandexYandex search engineYandexBotYes1,725
Googlebot-ImageGoogleImage indexingGooglebot-Image
Googlebot
Yes1,434
SeznamBotSeznamCzech search engineSeznamBotYes1,146
GoogleOtherGoogleGeneral-purpose crawlerGoogleOtherYes681
DuckDuckBotDuckDuckGoClassic search resultsDuckDuckBotYes666
YandexRenderResourcesBot
YandexImages
YandexFavicons
YaDirectFetcher
Yandex
JS rendering, images, favicons and landing pages for Yandex Direct adsseparate tokens
e.g. YandexImages
Yes483
ExabotExalead? – to verifyHistorically the Exalead search engine. Check the full user-agent and the address it containsExabotNo clear documentation98
BingPreview / bingbotMicrosoftPage snapshots for previews in Bing. User-agent observed in my logs; current documentation shows this traffic as bingbotno separate token; bingbot rules applyPer bingbot rules88
QwantbotQwantFrench privacy-focused search engineQwantbotYes48
Yahoo! SlurpYahooYahoo search engineSlurpYes9
Y!J-DLCYahoo Japan? – to verifyAgent associated with Yahoo Japan, with no current confirmation from the operatornoneNo clear documentation38
Bravebot***Brave? – unconfirmedUser-agent attributed to the Brave Search index; the operator does not document itBravebotNo clear documentation42
Storebot-GoogleGoogleGoogle ShoppingStorebot-GoogleYes12
SogouSogouChinese search engineSogouYes8
Amzn-SearchBotAmazonSearch across Amazon products, including AlexaAmzn-SearchBotYes5
Googlebot-VideoGoogleVideo indexingGooglebot-Video
Googlebot
Yes2
MojeekBotMojeekIndependent search engine with its own indexMojeekBotYes2
BingVideoPreviewMicrosoftVideo previews in BingBingVideoPreviewYes0
YepBotAhrefs / YepThe Yep search engine, currently only for IndexNow requests. The index itself is built by AhrefsBotYepBotYes, also honors Crawl-delay0
PinterestbotPinterestIndexes content and products for Pins, updates prices and dead linksPinterestbotYes; honors Crawl-delay up to a value of 1, treats higher values as 10
GoogleOther-Image
GoogleOther-Video
Google
GoogleOther variants for public images and videoGoogleOther-Image
GoogleOther-Video
GoogleOther
Yes0
GoogleProducerGoogleFeeds configured by a publisher in Google Publisher Centernone, user fetcherUsually not0
facebookexternalhitMetaPreview of links shared across Meta servicesfacebookexternalhitSometimes ignored4,479
Slackbot-LinkExpanding
Slack-ImgProxy
Slackbot
Slack
Link previews, image fetching and Slack’s remaining requestsnoneNo – Slack states that it does not process robots.txt407
SkypeUriPreviewMicrosoftLink previews in SkypenoneNo clear documentation327
TwitterbotX (formerly Twitter)Link previews in X cardsTwitterbotNo clear documentation273
LinkedInBotLinkedInPreview of links shared on LinkedInLinkedInBotNo clear documentation270
WhatsAppMetaPreview of links sent in WhatsAppnoneNo clear documentation166
DiscordbotDiscordLink previews on DiscordDiscordbotNo clear documentation6
MicrosoftPreviewMicrosoftPage previews in Microsoft productsMicrosoftPreviewYes0
Google MessagesGooglePreview of links sent in Google Messagesnone, user fetcherUsually not0
AhrefsSiteAuditAhrefsTechnical audit ordered by the site ownerAhrefsSiteAuditYes by default; the site owner can switch this off in the audit settings127,925
Screaming Frog SEO Spiderwhoever runs the programManual crawl, usually your own or your agency’sScreaming Frog SEO SpiderDepends on the settings7,787
BarkrowlerBabbarLink indexBarkrowlerYes5,975
AhrefsBotAhrefsIndex of links and content for Ahrefs and the Yep search engineAhrefsBotYes5,347
SemrushBotSemrushIndex of links and visibility dataSemrushBotYes4,639
serpstatbotSerpstatBacklink indexserpstatbotYes2,741
MJ12botMajesticLink indexMJ12botYes1,543
DataForSeoBotDataForSEOIndex of SEO data sold through an APIDataForSeoBotYes1,219
SERankingBacklinksBotSE RankingBacklink indexSERankingBacklinksBotYes752
SiteAuditBotSemrushTechnical audit on requestSiteAuditBotYes666
SEBot-WASE RankingSE Ranking’s technical audit crawlerSEBot-WAYes by default; the site owner can change this528
AdsBot-GoogleGoogleQuality assessment of ad landing pagesAdsBot-GoogleYes, but ignores User-agent: *228
Mediapartners-GoogleGoogleMatching AdSense ads to page contentMediapartners-GoogleYes, but ignores User-agent: *81
AdsBot-Google-MobileGoogleQuality assessment of mobile ad landing pagesAdsBot-Google-MobileYes, but ignores User-agent: *39
AdIdxBotMicrosoftQuality control of landing pages for Microsoft AdvertisingadidxbotYes0
DotBotMozCrawler building the Moz Link Explorer link indexDotBotNo current clear documentation from the operator0
rogerbotMozAudits of sites added to Moz Pro campaigns by their ownersrogerbotNo current clear documentation from the operator0
OAI-AdsBotOpenAIVerification of ad landing pages in ChatGPTOAI-AdsBotYes*0
UptimeRobotUptimeRobotSite availability monitoring, usually set up by the ownerUptimeRobotNot applicable51,781
AwarioBotAwarioBrand mention monitoringAwarioBotYes1,601
NetcraftSurveyAgent
CheckMarkNetwork
StormIntelCrawler
OI-Crawler
security research and domain inventory
Scanning the public surface of websites, domain statisticsseparate tokensNo clear documentation64
Buck (Hypefactors)
RecordedFuture
media monitoring
Tracking mentions of brands and topicsseparate tokensNo clear documentation36
Google-SafetyGoogleDetecting malware and abuse behind publicly shared linksnoneNo – Google states that it ignores robots.txt0
archive.org_botInternet ArchiveArchiving pages in the Wayback Machinearchive.org_botNo current clear documentation from the operator381
ClueWeb-CrawlerCarnegie Mellon UniversityBuilding the public ClueWeb corpus for search researchClueWeb-CrawlerYes82
Aranea Web-Crawled CorporaSlovak Academy of SciencesBuilding public language corporanoneNo clear documentation79
curl
Wget
python-requests
Scrapy
Go-http-client
axios and related
any script
Not bots, just libraries for fetching pages. Behind each one stands a person or someone else’s programnoneNot applicable – depends on the script, the library enforces nothing22,176
Google-InspectionToolGoogleSearch Console testing toolsGoogle-InspectionTool
Googlebot
Yes4,874
Chrome Privacy Preserving Prefetch ProxyGooglePage prefetching for Chrome users. Looks like a bot, but a browser stands behind itnoneNot applicable1,234
Inoreader
Feedly
RSS readers
Fetching an RSS feed on behalf of a subscriberseparate tokensUsually yes321
WP Rocket Pre-fetchWP RocketCache plugin prefetching links from your own site. Also your own trafficnoneNot applicable43
datasets library (Hugging Face)any userA Python library for building datasets. Someone was pulling content into a set of their ownnoneNot applicable – depends on the script14
APIs-GoogleGoogleDelivering notifications from Google APIs, e.g. PubSubHubbubAPIs-GoogleYes, but ignores User-agent: *0
Google-Site-VerificationGoogleChecking site ownership verification in Search Consolenone, user fetcherUsually not0
2ip bot2ip.ioService for checking IP addresses and websites2ip botNo clear documentation37,237
br-crawler
crawler_eb_germany
Hermes-SVF-static-crawler
VelenPublicWebCrawler and others
not established
No documentation and no way to establish the operator. Judge them by behaviornoneUnknown3,900
BingSapphireMicrosoft? – to verifyAbsent from Microsoft’s official crawler list. Check the IP addressnoneNo clear documentation76
EmailCrawlerunknownHarvesting email addresses. The one entry where blocking needs no thoughtEmailCrawlerUnknown; if you want certainty, block it on the server21

Urgent if you block NotebookLM. On 16 July 2026 Google changed the user-agent name from Google-NotebookLM to Google-GeminiNotebook. The old name is supported only for a transition period indicated as August 2026. A rule in .htaccess or in your WAF written against the old name will soon stop catching anything. On my site this bot made 3,310 requests, so the problem is not theoretical.

How this table was built

I checked purposes and tokens in the operators' documentation everywhere such documentation exists. Entries with no official source are described on the basis of the full user-agent, the address given in its string and its behavior in the logs – a missing confirmation is clearly marked. The frequency data comes from over 1.8 million requests in the podrez.pl logs for the period from 2 February to 5 August 2026.

Some entries group several related user-agents from one operator, so there are more names than rows.

Compiled by Ewelina Podrez-Siama, working in SEO since 2009. Last verified: 6 August 2026.

Sources: bot operator documentationTwenty-two documents in which I checked tokens, purposes and rules. Check for yourself if anything raises a doubt.

OpenAI
Overview of OpenAI Crawlers – GPTBot, OAI-SearchBot, ChatGPT-User, OAI-AdsBot
Advertiser Guidance for Allowing OpenAI Web Crawlers – the OAI-AdsBot token and rule compliance

Anthropic
Does Anthropic crawl data from the web… – ClaudeBot, Claude-User, Claude-SearchBot

Perplexity
Perplexity Crawlers – PerplexityBot, Perplexity-User, IP address lists

Google
Google's common crawlers – Googlebot, GoogleOther, Google-Extended, Google-CloudVertexBot
Google special-case crawlers – AdsBot, Mediapartners, APIs-Google, Google-Safety
Google user-triggered fetchers – Gemini Notebook, Google-Agent, GoogleProducer
Verify requests from Google – three classes of bots, three JSON files, reverse DNS
Crawler documentation changelog – change dates, including 16.07.2026 and 20.03.2026
AI Features and Your Website – preview directives with regard to AI Overviews and AI Mode

Apple
About Applebot – Applebot, Applebot-Extended, inheriting rules from Googlebot

Amazon
About Amazonbot – Amazonbot, Amzn-SearchBot, Amzn-User, three IP address lists

Mistral AI
Mistral crawlers – MistralAI-Training, MistralAI-Index, MistralAI-User

Microsoft
Overview of Bing crawlers – bingbot, adidxbot, BingPreview, MicrosoftPreview, BingVideoPreview

Other operators
Common Crawl – CCBot
Parallel – ShapBot and Shap-User
Slack Robots – three agents and a statement about not honoring the rules
Yep – YepBot
Pinterest – Pinterestbot
Diffbot – Does Diffbot respect robots.txt?
Brave Search Crawler

Context
Cloudflare – Your site, your rules: new AI traffic options – the Search, Agent and Training categories and the change of default settings from 15.09.2026

All links checked on 6 August 2026. Entries marked in the table as "no clear documentation" have no counterpart on this list – I describe them solely on the basis of the user-agent and behavior in the logs.

 

* OAI-AdsBot. OpenAI's developer documentation lists only GPTBot and OAI-SearchBot as control tokens, but the help article for advertisers gives an explicit User-agent: OAI-AdsBot example and declares that the rules are honored. I go with the latter, because it is newer and more detailed.

** Gemini Notebook. All 3,310 requests arrived under the old name Google-NotebookLM. The new one has not appeared in my logs even once.

*** Bravebot. The one entry where the documentation contradicts observation. The Brave Search help page says its crawler does not identify itself with a separate user-agent, and yet requests signed Bravebot and pointing to search.brave.com are recorded by independent registries and by my logs. The documentation may be out of date, but without confirmation from the operator I can't rule out impersonation.

The "honors robots.txt" column. It is based on operator documentation, not on my own tests. "Stated by operator" means the operator says so, though reports of deviations circulate. "No clear documentation" means exactly what it says – I found no reliable source and I am not guessing.

Why doesn't robots.txt block every AI bot?

An entry in robots.txt does not work the same way on every bot. Crawlers that move around the web on their own initiative usually follow it. With fetchers triggered by a user things get complicated, because here every operator has a policy of its own.

  • Automated crawlers

    They move around the web on the operator's initiative: GPTBot, ClaudeBot, PerplexityBot, Googlebot. They honor robots.txt and this file was created with them in mind.

  • User-triggered fetchers

    They fetch a page because a specific user action required it. Operators differ here. Anthropic and Mistral provide separate tokens (Claude-User, MistralAI-User) and declare that they honor them. Google, OpenAI, Perplexity and Amazon note that the rules may not apply, because the request is initiated by a user.

  • Control tokens without a bot

    Google-Extended and Applebot-Extended send no requests at all. You put them in robots.txt, but you will never see them in your logs. Google-Extended covers training of Gemini models and grounding, meaning content passed to the model at the moment of answering. Applebot-Extended covers training of Apple's models.

The practical conclusion: before you block a fetcher triggered by a user, check the policy of that particular operator. A Disallow for Claude-User or MistralAI-User genuinely closes the door, because Anthropic and Mistral declare that they honor those tokens. With Google, OpenAI, Perplexity and Amazon the same entry is a request, so if you need certainty, add a rule at the server or WAF level. How to check this on your own site, I described in the article on server log analysis.

Suspicious user-agents: impossible names and names to verify

Five names turned up in my logs that cannot be reconciled with the documentation of their supposed operators. They are worth knowing, because they look credible and slip through a filter easily.

Two of these names cannot come from the operator, because its own documentation says those tokens do not crawl at all. Three more do not appear in the current documentation of their supposed operators – that does not prove impersonation, but it does mean you should check the IP address before treating them as genuine.

Name in the logStatusWhyWhat it was looking for on my site
Google-ExtendedImpossibleA control token in robots.txt; Google does not send requests with it/.aws/credentials
Applebot-ExtendedImpossibleApple states outright that this token does not crawl pages/privatekey.key
BingIndexCrawlerTo verifyThe name does not appear in Microsoft's current crawler documentation; Bing crawls as bingbotordinary subpages
anthropic-aiTo verifyA name outside Anthropic's current documentation/secrets.yml
Claude-WebTo verifyA name outside Anthropic's current documentation/_profiler/open, /info.php

Addresses come from the podrez.pl logs, February–August 2026. Number of occurrences: Google-Extended 31, anthropic-ai 28, Claude-Web 20, BingIndexCrawler 10, Applebot-Extended 5.

The pattern is clear: suspicious requests reach for tokens that are not user-agents, and for names absent from operators' current documentation. It is worth distinguishing intent here. Requests signed as Google-Extended or anthropic-ai were looking for password and configuration files on my site, so that is a scanner. A name that does not match the documentation need not mean bad intent, though – which is why the IP address decides, not the name alone. The problem is not limited to my site either: Common Crawl explicitly warns in its documentation that it knows of crawlers falsely claiming to be CCBot, and recommends verifying the user-agent.

How to check whether a bot is real: verification by IP address

A user-agent can be faked in ten seconds, so the name alone proves nothing. Certainty comes only from matching the name with the IP address and the operator's official ranges, and where possible with a reverse and forward DNS check. Google is also testing cryptographic authentication of bots, that is Web Bot Auth.

OperatorSource of addresses
OpenAIopenai.com/gptbot.json, openai.com/searchbot.json, openai.com/chatgpt-user.json, openai.com/adsbot.json
Anthropicclaude.com/crawling/bots.json
Perplexityperplexity.com/perplexitybot.json, perplexity.com/perplexity-user.json
Google – common crawlerscommon-crawlers.json; reverse DNS ends in googlebot.com
Google – special-case crawlersspecial-crawlers.json; reverse DNS ends in google.com
Google – user fetchers and agentsseparate JSON files; reverse DNS ends in google.com or gae.googleusercontent.com
Mistralmistral.ai/mistralai-user-ips.json, mistral.ai/mistralai-index-ips.json; for MistralAI-Training the list is not published
Amazonthree separate lists: /ip-addresses/, /searchbot-ip-addresses/, /live-ip-addresses/
Paralleldocs.parallel.ai/resources/shapbot.json
Microsoftthe Verify Bingbot tool in Bing Webmaster Tools; reverse DNS ends in search.msn.com
Common Crawlindex.commoncrawl.org/ccbot.json; reverse DNS ends in crawl.commoncrawl.org

How to block AI bots in robots.txt? Ready-made examples.

Before you block anything: in most cases I do not recommend it. A block cuts you off from generative answers, and so from a channel where your brand can be mentioned and cited. For most companies, experts and shops that is a loss, not a gain.

The exceptions are real and concern publishers, portals and media above all – those who make a living from selling content or from traffic on their own site. There, a block can be a justified business decision – but still a decision, not a default setting.

So treat the configurations below as a starting point for an informed choice, not as a recommendation.

The first variant limits the main crawlers used for model training. The second blocks separately identified AI bots and fetchers, leaving access open to Googlebot, bingbot and Applebot. Both are a starting point, not a universal configuration for every site.

Variant 1 · No training

I protect my content

I let in everything that can cite me and cut off content collection for model training.

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Bytespider
User-agent: Meta-ExternalAgent
User-agent: MistralAI-Training
User-agent: cohere-ai
User-agent: CohereBot
User-agent: Amazonbot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /

User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: DuckAssistBot
User-agent: MistralAI-Index
User-agent: MistralAI-User
User-agent: Amzn-SearchBot
User-agent: ShapBot
Allow: /
The cost: with multi-purpose crawlers such as Amazonbot or CCBot, a block also switches off uses other than training.

Variant 2 · Separate AI bots

Google and Apple stay

I block separate training crawlers, AI indexes and user fetchers, but leave classic search engines open. This does not remove the site from AI Overviews, AI Mode or Apple's generative answers.

# search engines stay
User-agent: Googlebot
User-agent: bingbot
User-agent: Applebot
Allow: /

# model training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Bytespider
User-agent: Meta-ExternalAgent
User-agent: MistralAI-Training
User-agent: cohere-ai
User-agent: CohereBot
User-agent: Amazonbot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /

# indexes behind generative answers
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: MistralAI-Index
User-agent: Meta-WebIndexer
User-agent: DuckAssistBot
User-agent: Amzn-SearchBot
User-agent: LinkupBot
User-agent: ShapBot
User-agent: YouBot
User-agent: Diffbot
Disallow: /

# fetches on user request
User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User
User-agent: MistralAI-User
User-agent: Meta-ExternalFetcher
User-agent: Amzn-User
Disallow: /
The effect: you limit access for systems that have bots of their own – ChatGPT, Claude, Perplexity, Mistral. This is not a full exit from generative answers: AI Overviews and AI Mode are fed by Googlebot, Copilot by bingbot, and Applebot can be a source of context for Apple's answers.

The second variant has a hole and it is better to know about it right away. The last group are fetchers triggered by a user, and some of them may ignore robots.txt. Gemini Notebook and Google-Agent do not even have a token of their own. If the block is to be airtight, it has to be added at the server level. An example for Apache and the .htaccess file:

RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} (Google-GeminiNotebook|Google-NotebookLM|Google-Agent|GoogleProducer|ChatGPT-User|Perplexity-User|Amzn-User|Meta-ExternalFetcher|Shap-User) [NC]
RewriteRule ^ - [F,L]

One caveat that is easy to forget: robots.txt is not a tool for protecting data. The file is public, the directives are not access control, and a blocked URL can still end up in the index on the strength of external links. Genuinely private sections are secured with server-side authentication.

Watch out for groups and for the wildcard. If a more specific matching group exists for a bot, it takes precedence over User-agent: *. The wildcard applies only when nothing more specific matches. Fetchers triggered by a user are a separate case: some honor dedicated tokens, some may skip them and need a rule at the server or WAF level.

How to limit AI Overviews? On nosnippet, data-nosnippet and max-snippet

robots.txt tells compliant bots whether they should fetch a page. Meta directives, by contrast, control indexing and which fragments can reach the results and the AI features. These are two different layers, and with AI Overviews the second one is often the only one you have left.

The key point

There is no separate opt-out from AI Overviews and AI Mode that would leave your ordinary presence in Google untouched. They are fed by the same Googlebot and the same index, and Google-Extended covers only Gemini training and grounding, so blocking it changes nothing here.

You can, however, limit how much content reaches those features. Google lists four directives: nosnippet, data-nosnippet, max-snippet and noindex. Each of them works on ordinary results and on AI features at the same time – you cannot separate one from the other.

DirectiveWhat it doesWhat it costs you
nosnippetBlocks the display of a page fragment in every form of results, including AI Overviews and AI ModeThe ordinary description under the title also disappears, along with the chance of a direct answer
data-nosnippetThe same, but only for the marked HTML fragmentThe rest of the page works normally – the most precise tool in this set
max-snippet:NLimits the fragment length to N characters and with it the amount of content passed to the modelAt a low value the effect comes close to nosnippet
noindexRemoves the page from every form of Google results, including AI Overviews and AI ModeThe page also disappears from ordinary search results
noarchiveAt Amazon it means "do not use this page to train models"Works only with operators that honor this directive

Apple honors nosnippet as an opt-out from the use of content in generative answers based on current information. Amazon declares that it honors noarchive, noindex and none, but does not support crawl-delay.

The most precise tool is data-nosnippet. It excludes only the marked fragment of a page from previews and from direct use by AI features. Use it deliberately: the very paragraphs that would most easily replace a visit to your site – definitions, step lists, tables – are usually the ones that get you cited in the first place.

What AI bots take from your site: content or files?

The number of visits alone says little. What is more interesting is how many of them land on a page with content, and how many on CSS files, images and API endpoints.

BotSuccessful requestsShare of hits on content
Perplexity-User5791.2%
CCBot40582.5%
ChatGPT-User5,87073.9%
PerplexityBot2,12761.1%
Claude-User38960.2%
OAI-SearchBot3,15733.7%
GPTBot5,28332.0%
ClaudeBot2,24428.6%
Google-GeminiNotebook2,94520.0%
Applebot3,19215.0%
Amazonbot3,33714.6%

podrez.pl logs, 2 Feb – 5 Aug 2026. Perplexity-User, ChatGPT-User and Claude-User land mostly straight on content, but the class of bot alone does not determine its behavior: CCBot has a high share of hits on content, and Gemini Notebook a low one.

Frequently asked questions

How do I check which bot visited my site?

The bot name sits in the user-agent field of your server's access log, usually in access.log or domain.com.log in the logs directory. Copy the name and search for it in the table above. If you want to be sure the bot is genuine, compare its IP address with the ranges published by the operator.

Does blocking GPTBot remove me from ChatGPT results?

No. GPTBot is responsible for collecting content for model training, while presence in ChatGPT search results is handled by OAI-SearchBot. These are two independent tokens and you can set them differently: block training, leave search open. Anthropic and Mistral apply an analogous three-way split. Amazon also divides its traffic across three agents, but its Amazonbot is a multi-purpose crawler that only may be used for training. Perplexity, by contrast, separates just the search index from fetches on user request and has no documented training crawler.

Why don't I see Google-Extended in my logs?

Because Google-Extended is not a bot. It is a control token in robots.txt with which you decide about the use of your content for training Gemini models and for grounding. The fetching itself is done by Google's ordinary robots. If you do see this name in your log, you are looking at an impersonator.

How do I opt out of AI Overviews?

Through robots.txt you can't. AI Overviews and AI Mode use Google's ordinary index, built by Googlebot, and the Google-Extended token covers only the training of Gemini models and grounding. The only available tool is the preview directives: nosnippet, data-nosnippet, max-snippet and noindex. Each of them also limits your ordinary search result, so leaving AI Overviews means partly leaving Google.

Will a robots.txt block stop every AI bot?

No. Fetchers triggered by a user, such as Gemini Notebook or ChatGPT-User, may skip robots.txt, because the request is initiated by a person, not a machine. To stop them you need a rule at the server level, in your WAF configuration or in the .htaccess file.

What changed with NotebookLM?

On 16 July 2026 Google updated its documentation and changed the user-agent name from Google-NotebookLM to Google-GeminiNotebook, after the product itself was renamed Gemini Notebook. The old name is supported only for a transition period indicated as August 2026, so rules written against the old name need updating.

The name in the log is only the beginning. Far more interesting is what that bot got in response – and whether it got anything at all. How to check that and what I found on my own site, I described in the article on server log analysis.

Not sure which bots your site lets in?

Log analysis covering AI bots, search engine crawlers, SEO tools and suspicious traffic is part of my audit. If you would rather start with a conversation, book a short call.

Last verified: 6 August 2026 Bot tokens and purposes checked in the operators' documentation: OpenAI, Anthropic, Perplexity, Google, Apple, Amazon, Common Crawl, Mistral, Diffbot, Parallel, Slack, Yep, Pinterest. Entries marked as "no clear documentation" have no confirmation from the operator and are described solely on the basis of observation. Brave and Linkup: purposes assessed on the basis of product documentation, user-agents and my logs, without current documentation of the crawler itself. Observational data comes from the podrez.pl access logs for the period 2 Feb – 5 Aug 2026, over 1.8 million requests.

History of this page 06.08.2026 – first version.

Timeline of changes at the operators 16.07.2026 – Google changed the NotebookLM user-agent name to Google-GeminiNotebook; the old name supported for a transition period.
01.07.2026 – Cloudflare introduced a split of bots into the Search, Agent and Training categories.
07.04.2026 – Anthropic split its documentation into three bots, adding Claude-SearchBot and Claude-User.
20.03.2026 – Google documented the Google-Agent user-agent.