- Every engine cites from its own retrieval stack. Blocking OAI-SearchBot removes a site from ChatGPT search answers, while blocking GPTBot only affects training1.
- In a 1,000-question test run on 3 August 2026, Perplexity returned 19.08 links per answer, Claude 12.49, Google AI Overviews 9.42 and ChatGPT 3.1017.
- Google AI Overviews and AI Mode cited the same URLs only 13.7% of the time across 540,000 query pairs, even though their answers were 86% semantically similar16.
- In Profound's 680 million citation dataset, Wikipedia was 47.9% of ChatGPT's top 10 sources, and Reddit was 46.7% of Perplexity's15.
- Bing Webmaster Tools' AI Performance report, in public preview since February 2026, lists which of your URLs Copilot and Bing's AI summaries cited; no other major engine offers a per-URL citation report24.
How do AI engines decide which sources to cite?
An engine can only cite what its search step hands the model. The model doesn't browse the open web. It sends one or more queries to an index, reads a shortlist of results, writes the answer and attaches links to the passages it used. That loop is retrieval-augmented generation, and tying the answer to fetched pages is called grounding.
So there are three gates, and you can fail any of them. Your page has to be in that engine's index, which depends on which crawler you let in and whether a CDN rule quietly blocks it. It has to match the query the engine actually sends, which is often a rewritten or split version of what the user typed. And it has to contain a passage the model finds worth quoting.
The engines differ on the first gate more than people expect. That's why a page can be cited by Perplexity every day and never once by ChatGPT.
Where does ChatGPT get its sources?
ChatGPT search pulls sources through OpenAI's own crawler, OAI-SearchBot, combined with third-party search providers and content supplied by publishing partners. OpenAI's launch post for ChatGPT search on 31 October 2024 described it that way2, and Microsoft had announced at Build in May 2023 that Bing would be ChatGPT's default search experience4. OpenAI has never said how much of today's results come from each provider.
OpenAI's crawler documentation is explicit about the controls. Sites that opt out of OAI-SearchBot "will not be shown in ChatGPT search answers," though they can still appear as navigational links; GPTBot is a separate training crawler with an independent setting; and ChatGPT-User handles visits a user triggers and is not used to decide what appears in search1. The same page says robots.txt changes take about 24 hours to reach search results1. The publisher FAQ adds that ChatGPT tags referral links with utm_source=chatgpt.com3.
ChatGPT shows few links. In Slate's August 2026 test of 1,000 B2B software questions it returned 3.10 links per answer, and 372 of the 1,000 answers had no link at all17. Those answers named no source, so whatever they said about a brand most likely came from training data, which is where GPTBot still matters. When it does cite, it favors reference sites: Profound's analysis of 680 million citations from August 2024 to June 2025 found Wikipedia was 7.8% of all ChatGPT citations and 47.9% of its top 10 cited domains15.
If you want ChatGPT to cite you
- Grep your access logs for
OAI-SearchBotand check the IPs against OpenAI's published searchbot.json. No verified hits in 30 days usually means a firewall or bot-management rule is blocking it, even when robots.txt says allow. - Verify the site in Bing Webmaster Tools and submit your sitemap. OpenAI uses unnamed third-party providers and Bing was the announced default, so a page missing from Bing is a page you are betting OpenAI found on its own.
- Because only about three links make the cut, run your ten most valuable prompts and list every domain ChatGPT cites. Those domains, often Wikipedia, review sites and a couple of trade publications, are your outreach list. Getting your product named on them beats rewriting your own page for the fourth time.
- Leave GPTBot allowed if you sell software. 37% of ChatGPT's answers in the Slate test carried no link, so what the model learned in training is the only thing those answers can say about you.
- Build a GA4 exploration filtered to
session source = chatgpt.comand watch which landing pages get the clicks. Those are the pages ChatGPT already trusts; extend them before you write new ones.
Where does Perplexity get its sources?
Perplexity cites from its own index, built by PerplexityBot. When it opened that index to developers as the Search API in September 2025, it described hundreds of billions of webpages, the same infrastructure behind its answer engine6. Its crawler documentation says PerplexityBot surfaces and links websites and is not used to train foundation models5. A second agent, Perplexity-User, fetches pages when a user's question needs them and "generally ignores robots.txt rules," per the same docs5.
It is the most citation-heavy engine in every recent test: 19.08 links per answer in the Slate study, with a link on all 1,000 questions17. The mix tilts toward community content. Reddit was 6.6% of all Perplexity citations in Profound's dataset and 46.7% of its top 10 cited domains15. Perplexity also ranks passages inside pages, which its launch post called sub-document precision, so a single strong paragraph on a long page can win the citation6.
If you want Perplexity to cite you
- Allow PerplexityBot in robots.txt and at the firewall. Perplexity's docs tell Cloudflare users to create a rule with the action set to Allow for its published IP ranges, so it bypasses other security rules5.
- Make each H2 section stand on its own: the question as the heading, the answer in the first sentence, the number and its source in the next two. Passage-level ranking means a section that needs the previous section to make sense is a weak candidate.
- Search your category on Reddit, find the three or four threads Perplexity already cites for your prompts, and contribute where you have something specific to add. With Reddit at nearly half of Perplexity's top 10 domains, that thread is often the real competition.
- Open the Sources panel for each tracked prompt and note which of your competitors' pages appear. Perplexity shows enough links per answer that you can usually see exactly which page format it prefers for a query.
How do Gemini, AI Overviews and AI Mode pick sources?
Google's AI features cite pages from the Google Search index, found through what Google calls query fan-out. Its AI features documentation says AI Overviews and AI Mode may issue multiple related searches across subtopics and data sources, and keep finding supporting pages while the response is generated7. The May 2025 AI Mode announcement said Deep Search can issue hundreds of searches8. Our glossary entry on query fan-out covers the mechanics.
The eligibility rule is short. A supporting link "must be indexed and eligible to be shown in Google Search with a snippet," and Google says there are no additional requirements or special optimizations7. Googlebot does the crawling. Google-Extended, the token many publishers block, has no effect on Search; it governs Gemini training and grounding in the Gemini app and Vertex AI9. The Gemini API's grounding documentation says the model decides whether a search would help, runs one or more queries, and returns citation annotations tied to text spans10.
Fan-out makes Google's AI results unstable. Ahrefs compared AI Overviews and AI Mode on the same September 2025 US queries and found only 13.7% of cited URLs overlapped (16.3% for the top three), though the answers averaged 86% semantic similarity16. Wikipedia was in 28.9% of AI Mode citations against 18.1% for AI Overviews16. In the Slate test, AI Overviews showed 9.42 links per answer, while Gemini 2.5 Flash through the API returned links on just 16 of 1,000 answers17.
If you want Google's AI features to cite you
- Run your money pages through URL Inspection in Search Console. "URL is on Google" plus no
nosnippetis the whole technical requirement; anything else is a blocker you can fix today. - Search your templates for
data-nosnippetandmax-snippet. Google won't use a passage it isn't allowed to show as a snippet, and these often get added for one widget and then copied site-wide. - Pull the "People also ask" questions and the related searches for your head term and give each one its own H2 or H3 with a one-sentence answer. That's the closest public proxy for the sub-queries fan-out sends.
- Check that any structured data matches the visible text. Google's AI features page asks for exactly that and says no special schema is required7.
- Don't wait for a separate AI report. Google counts AI Overviews and AI Mode traffic inside the normal Performance report under the Web search type7, so track the queries where your clicks fall while impressions hold.
Where does Claude get its sources?
Claude cites pages returned by Anthropic's web search tool, which searches when a request depends on current information and answers from training data when it doesn't11. Citations are always on for web search results, and Anthropic's help center says every response that uses web search includes them1122.
Anthropic doesn't name its search provider in product docs. Brave Search was added to its subprocessor list in March 2025, and Simon Willison found that Claude's cited results for a test query matched Brave's for the same search13. That matters for robots.txt, because Brave's crawler help page says it uses no distinct user agent and won't crawl a page Googlebot isn't allowed to crawl26. Anthropic also runs its own crawlers: Claude-SearchBot indexes content for search quality (blocking it may reduce visibility in Claude's results), Claude-User fetches pages for a user's question, and ClaudeBot collects training data12.
Claude cites more than its reputation suggests: 12.49 links per answer from Claude Sonnet 4.6 in the Slate test, with links on 978 of 1,000 answers17. Claude and Perplexity shared 13.6% of cited pages, the highest overlap of any pair in that study17.
If you want Claude to cite you
- Search your target queries on search.brave.com. If your page isn't in Brave's top results, Claude is unlikely to see it. Brave's help page has a form to request a re-fetch of a page26.
- Keep Googlebot allowed everywhere you want Claude to reach, since Brave follows your Googlebot rules26. A leftover
Disallowfor Googlebot on a docs or help subdomain would hide it from both. - Allow Claude-SearchBot and Claude-User, and confirm hits against Anthropic's IP list12.
- Reuse your Perplexity work. With the highest overlap of any engine pair, the same answer-first, well-sourced sections tend to help in both.
Where does Microsoft Copilot get its sources?
Copilot grounds web answers in Bing. Microsoft's documentation for Microsoft Copilot and Copilot Chat says it generates a short search query from the prompt, sends it to Bing, and writes the response from what comes back14. Bingbot access and Bing indexing are the gate.
Bing has page-level controls for this. Content tagged NOARCHIVE is left out of chat answers and not used to train Microsoft's models, while NOCACHE allows only URLs, titles and snippets; both still appear in normal Bing results, per Bing's 2023 announcement21. In February 2026 Bing added an AI Performance report to Webmaster Tools in public preview. It shows total citations, which URLs were cited across Copilot, Bing's AI summaries and select partners, and a sample of the grounding queries behind them24.
If you want Copilot to cite you
- Open the AI Performance report and export the grounding queries. They are the closest thing any engine publishes to the real search queries an assistant sent for your pages.
- Set up IndexNow so Bing hears about new and updated URLs when you publish; Bing is one of the participating engines25.
- Search your templates for
NOCACHEandNOARCHIVErobots values. Either one limits what Copilot can show. - Verify bingbot hits with a reverse DNS lookup (the host should end in
search.msn.com) before you rate-limit anything that claims to be Bing.
How do the engines compare?
The table sums up the retrieval source, the crawler you need to allow, and citation behavior from the studies above. Link counts come from one API-based test, so use them to compare engines and avoid quoting them as fixed figures.
| Engine | Retrieval source | Crawler to allow | Links per answer (Slate, Aug 2026) | Source tilt |
|---|---|---|---|---|
| ChatGPT search | OAI-SearchBot crawl plus third-party search providers and partner content | OAI-SearchBot | 3.10 | Wikipedia leads its top 10 (47.9%) |
| Perplexity | Own index of hundreds of billions of pages | PerplexityBot | 19.08 | Reddit leads its top 10 (46.7%) |
| Google AI Overviews | Google Search index, query fan-out | Googlebot | 9.42 | Reddit leads its top 10 (21.0%) |
| Google AI Mode | Google Search index, query fan-out | Googlebot | Not measured | Wikipedia in 28.9% of citations |
| Gemini (app and API) | Grounding with Google Search | Googlebot; Google-Extended governs app grounding | 0.07 (API, Gemini 2.5 Flash) | Too few links to measure |
| Claude | Web search tool; Brave Search listed as subprocessor | Claude-SearchBot, Claude-User, and Googlebot rules for Brave | 12.49 | Closest overlap with Perplexity (13.6%) |
| Microsoft Copilot | Bing search service | bingbot | Not measured | Follows Bing indexing |
A worked example: who is even eligible for "how to block GPTBot"?
We can't reproduce engine answers here without fabricating them, so this example checks the part you can verify from public files: which top-ranking pages each engine is allowed to retrieve. We ran the query how to block GPTBot on Brave Search (the index Claude appears to use) while researching this guide, took the top 10 organic results, and parsed each site's live robots.txt for six user agents.
- Seven of the ten (Vercel's knowledge base, Mersel AI, Search Logistics, Krasamo, cside, iSocialWeb and Otterly.AI) allowed all six: GPTBot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and bingbot.
- Raptive blocked GPTBot only. Its page stays eligible for every AI search engine while opting out of training, which is the split OpenAI's docs describe.
- Mashable names OAI-SearchBot, Claude-SearchBot and PerplexityBot in disallow groups but lets Googlebot and bingbot in. Its article can support an AI Overview or a Copilot answer, and per OpenAI's docs it won't be shown in ChatGPT search answers. Brave still ranked it ninth, because Brave follows Googlebot rules, so whether Claude cites it depends on how Anthropic treats Brave results for a site that blocks Claude-SearchBot. Anthropic hasn't documented that.
- The Reddit thread in tenth place sits on a domain whose robots.txt, as served to our request, is two lines:
User-agent: *andDisallow: /. Reddit made that change in mid-2024, and Microsoft confirmed Bing stopped crawling it27. Reddit content still reaches Google through a data licensing deal28, a reminder that robots.txt isn't the only door into an index.
For anyone writing on this topic, 8 of the 10 competing pages are open to every AI search crawler, so your edge on this query has to come from passage quality and third-party mentions. News queries are different. In our robots.txt study of 261 major websites, 33 of 48 news sites blocked at least one AI search crawler, against 1 of 95 B2B SaaS sites, so on a news query simply being retrievable can put you in ChatGPT's three links.
To test real engine output, run the same prompt set in each engine on the same day, logged out or in a fresh session, and record what you see in a sheet like this:
prompt | engine | date run | answered with search? (y/n) | cited URLs (in order) | our domain cited? | our brand named in text? | competitors named | robots.txt allows engine's bot? (y/n)
Run each prompt three times per engine, because answers vary run to run, and count a citation only if it shows up in at least two of three. Twenty prompts across five engines is 300 runs, about a day of work by hand. Our guide to measuring AI visibility covers sampling and share of voice in more detail, and the free Arobis AI visibility checker runs a buyer-prompt set across engines if you'd rather not do it by hand. We compare it with nine other free options in our roundup of free AI visibility checkers, including which ones show actual citations and which only score readiness.
How little do the engines agree with each other?
Overlap is low in every study that measures it. Slate's same-page overlap ranged from 0.5% for ChatGPT and Claude to 13.6% for Perplexity and Claude, and 74.1% of pages cited by the three link-showing engines weren't in Google's top 10 or AI Overview data for the same questions17. The Digital Bloom's 2025 report puts the share of domains cited by both ChatGPT and Perplexity at 11%, though it doesn't publish its method18.
Watch how source shares get quoted. Profound reports two different metrics: Wikipedia's 47.9% for ChatGPT and Reddit's 46.7% for Perplexity are shares of each platform's top 10 domains, while their shares of all citations were 7.8% and 6.6%15. Plenty of articles mix the two up and end up claiming half of all ChatGPT citations are Wikipedia.
Referral traffic is shifting as well. Goodie's report of 21 May 2026 found ChatGPT's share of AI referral sessions on a mostly B2B panel fell from 89.1% in May to August 2025 to 62.6% in March and April 2026. Claude rose from 1.4% to 18.5%, Gemini from 2.4% to 10.6%, Perplexity from 3.1% to 7.3% and Copilot from 3.2% to 4.0%19. For B2B sites, optimizing for ChatGPT alone now ignores more than a third of AI referrals.
Citations aren't always right, either. The Tow Center at Columbia tested eight AI search tools on 1,600 queries in March 2025 and found they misattributed the source of news excerpts more than 60% of the time; Perplexity was wrong on 37% and Grok 3 on 94%20. Some tools also answered correctly about publishers whose robots.txt should have kept them out20. More figures like these are on our AI search statistics page.
Why one page wins in Perplexity and never shows up in ChatGPT
Usually it's the first gate. A page indexed by Google but blocked for OAI-SearchBot (directly, or by a CDN rule that challenges unknown bots) can sit in AI Overviews for months and never appear in ChatGPT search. A page Bing never indexed can't ground Copilot, however well it ranks on Google.
When access is fine, the difference is usually link budget. Perplexity shows around 19 links, so a decent page on a niche question gets in. ChatGPT shows about three, so it takes the most established sources, and for many commercial queries that means Wikipedia, a big review site and a trade publication. If your page wins in Perplexity and not ChatGPT, look at the pages ChatGPT does cite for that prompt and work on getting named there before you touch your own copy again. The original GEO paper found that adding statistics, quotations and cited sources lifted visibility in generative answers by up to 40%23, which helps once you're in the shortlist. Our generative engine optimization guide covers those on-page tactics, our AI crawler user agent list has the robots.txt rules, and the citability grader scores a page's answer-first structure and sourcing.
Frequently asked questions
Does blocking GPTBot remove my site from ChatGPT search?
No. OpenAI runs GPTBot and OAI-SearchBot as separate crawlers with independent settings. GPTBot collects content that may train future models; OAI-SearchBot is what puts pages into ChatGPT search answers. You can disallow GPTBot and keep OAI-SearchBot allowed, and OpenAI says robots.txt changes for search take about 24 hours to apply.
Does ChatGPT use Bing to find sources?
Partly, as far as anyone outside OpenAI can tell. Microsoft announced Bing as ChatGPT's default search in May 2023, and OpenAI's 2024 search launch said it uses third-party search providers plus partner content. OpenAI also crawls with OAI-SearchBot and has never published the split. Being indexed in Bing is cheap insurance either way.
How do you get mentioned by ChatGPT when it doesn't link a source?
An answer with no link most likely draws on what the model learned in training, so keep GPTBot allowed if you want new content to reach future models, and get your brand named on the Wikipedia pages, review sites and trade publications ChatGPT already cites for your prompts. In Slate's 1,000-question test, 372 ChatGPT answers carried no link at all.
Why does Perplexity cite so many more sources than ChatGPT?
Perplexity searches its own index on almost every prompt and shows the results as a source list, while ChatGPT often answers from the model without searching. In Slate's 1,000-question test from August 2026, Perplexity returned 19.08 links per answer and linked a source on every question; ChatGPT returned 3.10 links per answer and none at all on 372 answers.
Does Google-Extended affect AI Overviews?
No. Google says Google-Extended does not affect inclusion in Google Search and is not a ranking signal. AI Overviews and AI Mode are Search features crawled by Googlebot. Google-Extended only controls whether your content trains Gemini models and grounds answers in the Gemini app and Vertex AI, so blocking it costs you Gemini app citations, not AI Overviews.
Which search engine does Claude use?
Anthropic's product docs do not name a provider. Brave Search was added to Anthropic's subprocessor list in March 2025, and Simon Willison found Claude's cited results matched Brave's for the same query. Brave says its crawler follows your Googlebot rules, so blocking Googlebot likely hides you from Claude too. Anthropic also runs Claude-SearchBot for its own index.
How much do AI Overviews and AI Mode citations overlap?
Very little. Ahrefs compared 540,000 US query pairs from September 2025 and found the two features cited the same URLs only 13.7% of the time, or 16.3% for the top three citations. The answers themselves averaged 86% semantic similarity, so Google said much the same thing in both places while pulling it from different pages.
What should I fix first to get cited by AI engines?
Start with indexing and access, since no amount of writing helps a page the engine cannot retrieve. Confirm Google and Bing both index your key pages, then check that OAI-SearchBot, PerplexityBot and Claude-SearchBot get 200 responses in your logs and are not blocked by a CDN rule. Only then rewrite pages answer-first and work on third-party mentions.
Sources
- OpenAI, Overview of OpenAI crawlers
- OpenAI, Introducing ChatGPT search (31 October 2024)
- OpenAI Help Center, Publishers and developers FAQ
- CNBC, Microsoft says Bing can be default search engine for ChatGPT users (23 May 2023)
- Perplexity, Perplexity crawlers documentation
- Perplexity, Introducing the Perplexity Search API (September 2025)
- Google Search Central, AI features and your website
- Google, AI Mode in Google Search: updates from Google I/O 2025 (20 May 2025)
- Google Search Central, Google's common crawlers (Google-Extended)
- Google AI for Developers, Grounding with Google Search
- Anthropic, Web search tool documentation
- Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Simon Willison, Anthropic appears to be using Brave to power web search for Claude (21 March 2025)
- Microsoft Learn, Data, privacy, and security for web search in Microsoft Copilot
- Profound, AI platform citation patterns (680 million citations, Aug 2024 to Jun 2025)
- Ahrefs, AI Mode and AI Overviews cite different URLs (15 December 2025)
- Slate, How to get cited by ChatGPT, Perplexity, Claude and Gemini (1,000-question study, 27 August 2026)
- The Digital Bloom, 2025 AI citation and LLM visibility report
- Goodie (Higoodie), AI search traffic report 2026 (21 May 2026)
- Columbia Journalism Review, Tow Center: AI search has a citation problem (6 March 2025)
- Bing Webmaster Blog, New options for webmasters to control usage of their content in Bing Chat (September 2023)
- Claude Help Center, Enabling and using web search
- Aggarwal et al., GEO: Generative Engine Optimization (KDD 2024)
- Bing Webmaster Blog, Introducing AI Performance in Bing Webmaster Tools Public Preview (10 February 2026)
- IndexNow, protocol site and participating search engines
- Brave Search Help, Brave Search crawler
- Search Engine Land, Microsoft confirms Reddit blocked Bing search (July 2024)
- Google, Google expands partnership with Reddit (22 February 2024)
AI Ranked Editorial. "How to get cited by ChatGPT and other AI engines." AI Ranked, October 11, 2026. https://airanked.ai/guides/how-ai-engines-choose-sources