- Five of the ten organic US Google results we pulled for "llm optimization" were about engineering (inference speed, model accuracy). This guide covers the marketing meaning, which most practitioners call LLM SEO.
- Only about 12% of URLs cited by ChatGPT, Gemini, Copilot and Perplexity ranked in Google's top 10 for the same prompt, across 15,000 prompts in an Ahrefs study.
- Branded web mentions correlated with AI Overview brand visibility at 0.664 in Ahrefs' 75,000-brand study. Backlink count came in at 0.218.
- In the GEO paper's 10,000-query benchmark, adding quotations raised the main visibility score from 19.5 to 27.8, while keyword stuffing dropped it to 17.8.
- Blocking GPTBot or ClaudeBot opts you out of training. Blocking OAI-SearchBot or Claude-SearchBot is what pulls you out of their search answers.
Which "LLM optimization" is this page about?
This page is about LLM optimization for brands and content: making large language models mention, recommend and link to you. The phrase has a second, older meaning in machine learning, and Google mixes the two. When we pulled the US top 10 for "llm optimization" in Ahrefs, five organic results were engineering pages (NVIDIA, Mirantis, Iguazio, a Medium explainer and OpenAI's own accuracy guide) and three were marketing pages (Semrush, Conductor and an r/SEO thread).
If you came here to make a model faster or cheaper to run, you want the engineering sense, and NVIDIA's guide to inference optimization is a good start.
| LLM optimization for brands (LLM SEO) | LLM optimization in engineering | |
|---|---|---|
| Goal | Get named and cited in AI answers | Make a model faster, cheaper or more accurate |
| Who does it | SEO, content and PR teams | ML engineers, platform teams |
| Typical work | Crawler access, citable pages, third-party mentions, prompt tracking | Quantization, batching, KV caching, prompt engineering, fine-tuning, RAG pipelines |
| Success metric | Mention rate and citation rate across a prompt set | Latency, cost per token, eval accuracy |
What is LLM SEO, and how is it different from GEO and AEO?
LLM SEO is the same job as generative engine optimization (GEO) and answer engine optimization (AEO), under a name that points at the model instead of the interface. GEO comes from a 2023 Princeton-led paper10; LLM SEO is the practitioners' term for the same tactics. Our GEO vs AEO vs SEO comparison goes through where the edges actually differ.
Google's position is blunt. Its guide to generative AI features says that from Search's perspective, AEO and GEO are SEO, and that AI Overviews and AI Mode "are rooted in our core Search ranking and quality systems"7 (Google Search Central). For Google's own surfaces, that's largely right. For ChatGPT, Claude and Perplexity, which run their own retrieval and their own crawlers, a good Google ranking is a weaker predictor than most people assume (more on that below).
How do LLMs decide which sources and brands to mention?
An LLM answers from two places: what it absorbed during training (its parametric memory) and pages it retrieves at answer time. The second pattern is retrieval-augmented generation, described in a 2020 paper by Lewis and colleagues that paired a language model with a searchable index of passages8 (arXiv). Every consumer AI assistant now does some version of this when a question needs current or specific facts.
The training path
When a model answers without searching, it draws on training data collected before its knowledge cutoff. Brands that were written about often, and consistently, in the years before that cutoff come out of the model's memory first. You can't edit that memory. You can only add to what the next model learns, by being described in more places, and you can control whether your own pages are part of it through the training crawlers: GPTBot for OpenAI1, ClaudeBot for Anthropic3, and the Google-Extended token for Gemini5.
The retrieval path
When a model searches, it rewrites the question into one or more shorter queries, sends them to a search index, reads a handful of results and writes from them. OpenAI's help center describes exactly this for ChatGPT: it typically rewrites your prompt into targeted queries for third-party search providers and may follow up with narrower ones2 (OpenAI). Google calls its version query fan-out, "issuing multiple related searches across subtopics and data sources"6. Which index gets searched depends on the engine:
| Engine | Search crawler (blocking it hides you from answers) | Training crawler or token |
|---|---|---|
| ChatGPT | OAI-SearchBot; opted-out sites "will not be shown in ChatGPT search answers" except as navigational links1 | GPTBot |
| Claude | Claude-SearchBot; blocking it may reduce visibility in Claude's search results3 | ClaudeBot |
| Perplexity | PerplexityBot, which Perplexity says is not used to crawl content for foundation models4 | None declared |
| Google AI Overviews and AI Mode | Googlebot. Pages must be indexed and eligible for a snippet6 | Google-Extended, which "does not impact a site's inclusion in Google Search"5 |
Two numbers show how far this is from classic SEO. Ahrefs found that only about 12% of URLs cited by ChatGPT, Gemini, Copilot and Perplexity ranked in Google's top 10 for the same prompt; Perplexity was the closest to Google at 28.6%11 (Ahrefs). Even Google's own AI Overviews drift from its blue links: in a March 2026 study of about 4 million cited URLs, 37.9% also appeared in the top 10 for the query, down from about 76% in Ahrefs' earlier analysis12 (Ahrefs). Fan-out explains most of the gap: the engine is searching sub-questions you never typed, and the pages that answer those sub-questions get cited. Our explainer on how AI engines choose their sources goes engine by engine.
What happens to your page after retrieval
Research on long contexts found that models answer best when the relevant information sits near the start or end of the input and do worse when it's buried in the middle9 (Liu et al.). That study tested question answering, not AI search, so treat this as our reading of it: a page whose answer is stated plainly near the top, in a sentence that makes sense on its own, gives the model less to dig through.
The five levers that move LLM visibility
Everything that reliably moves AI answers falls under one of five levers. They're listed in the order we'd fix them, because each one depends on the one before it.
1. Crawl access: separate search bots from training bots
Most sites that block AI crawlers block the wrong ones. In our study of AI crawler access on 261 major sites, 31 of the 59 sites that disallow GPTBot still allow OAI-SearchBot, which is a deliberate choice to stay out of training while staying in ChatGPT search. Accidental blocks usually sit in the CDN: Cloudflare announced default blocking of AI crawlers on 1 July 202519 (Cloudflare), so a site can say "allow" in robots.txt and still return 403s.
Our default for a business that wants to be recommended: allow every search crawler, and allow training crawlers unless you sell the content itself (news, research, courses). OpenAI says robots.txt changes reach its search systems in about 24 hours1. Here's the starting file:
# AI search crawlers: keep these open if you want to appear in answers
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
# Training crawlers: block only if your content is the product
User-agent: GPTBot
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: Google-Extended
Allow: /
Then check what the bots actually get. On a standard combined log format, this counts status codes per AI user agent for the last log file:
grep -E "OAI-SearchBot|Claude-SearchBot|PerplexityBot|GPTBot|ClaudeBot" access.log \
| awk '{print $9}' | sort | uniq -c
Done when: every search crawler gets 200 responses on your pricing, product and comparison pages, and no CDN or WAF rule returns 403 or a challenge page to them.
2. Citable passages: answers that survive being lifted out
The GEO paper tested nine rewriting tactics on a 10,000-query benchmark10 (Aggarwal et al., KDD 2024). On its main visibility metric, a baseline of 19.5 rose to 27.8 with quotations added, 25.9 with statistics and 24.9 with cited sources. Keyword stuffing fell to 17.8, worse than doing nothing. The more interesting result is who benefits: citing sources raised visibility 115% for pages ranked fifth and lowered it about 30% for pages ranked first, so these edits help challengers far more than leaders.
A SaaS pricing paragraph, before and after (the product is hypothetical):
Before: Our flexible pricing scales with your business. Whether you're a startup or an enterprise, we have a plan that fits your needs, with powerful features and world-class support included.
After: Ledgerly costs $8 per active user per month on the Team plan and $14 on Business, billed annually, with a 14-day free trial. Business adds multi-entity accounting and NetSuite sync. Companies under 50 users usually pick Team; the break-even for Business is the NetSuite integration, which Team doesn't include.
The second version can be quoted in an answer to "how much does Ledgerly cost," "does Ledgerly integrate with NetSuite" and "Ledgerly Team vs Business." The first can't be quoted for anything. Google says you don't need to break content into tiny pieces for AI7, and you shouldn't: keep the page whole and make each paragraph carry one checkable fact. Run drafts through the citability grader to catch the vague ones.
Done when: each money page opens with a sentence that names the product, the category and who it's for, and every claim on the page has a number, a name or a link behind it.
3. Entity consistency: one version of the facts everywhere
Models learn what you are from every place you're described. When your About page says "revenue intelligence platform," G2 lists you under "sales analytics," and Crunchbase still carries your 2021 tagline, a model has three entity descriptions to choose between and no reason to prefer yours. Write a one-page fact sheet and push it out to every profile you control:
Field | Value (use word for word)
Category label | expense management software for mid-size companies
One-line pitch | Ledgerly automates expenses and cards for 100 to 1,000 employee companies
Founded / HQ | 2019 / Austin, Texas
Pricing model | per active user, public pricing page
Key integrations | NetSuite, QuickBooks, Xero
Profiles to sync | About page, LinkedIn, Crunchbase, G2, Capterra, Wikidata (if notable)
Use Organization schema whose sameAs list points to each of those profiles. Google says no special schema is needed for its AI features but asks that structured data match the visible text6, so treat schema as a consistency check. Nobody has shown it works as a ranking trick.
Done when: the category label and core facts read the same on every profile in the list, and the schema matches the page.
4. Third-party mentions: be named where the model reads
This lever usually decides "best X" prompts. In Ahrefs' study of 75,000 brands, branded web mentions correlated with AI Overview visibility at 0.664, branded anchors at 0.527, Domain Rating at 0.326 and backlink count at 0.21813 (Ahrefs). Correlation isn't cause, and big brands have more of everything, but the direction matches what the retrieval path predicts: the model repeats what the pages it reads agree on.
Which pages those are varies by engine and shifts fast. Profound's analysis of 680 million citations found Wikipedia made up 47.9% of ChatGPT's top 10 cited sources and Reddit 46.7% of Perplexity's15 (Profound). Semrush then watched the share of ChatGPT responses citing Reddit fall from close to 60% in early August 2025 to around 10% by mid-September16 (Semrush). So don't build a plan on someone's list of "domains AI loves." Pull the cited URLs for your own prompts and work that list: for each of the top 20, note whether you're named, who owns the page and what would make them add you.
Google warns that seeking inauthentic mentions "isn't as helpful as it might seem"7. Paid placements on low-quality listicles are the obvious case. Earned inclusion in reviews, comparisons and community threads where you actually fit is the version that holds up.
Done when: you're named on at least half of the 20 URLs most often cited for your prompt set.
5. Freshness: update facts, not dates
AI assistants cite newer pages than Google's organic results do. Ahrefs analyzed about 17 million cited URLs and found they averaged 1,064 days since publication against 1,432 for organic results, 25.7% fresher; ChatGPT's citations were the newest, at 458 days newer than organic. Google's AI Overviews were the exception, citing pages about 16 days older14 (Ahrefs).
Put your money pages and most-cited guides on a quarterly review and change something real each time: new prices, a new comparison row, a new data point. Bumping a date with no new content gives a model nothing new to cite.
What's hype in LLM optimization?
Plenty of LLM SEO advice is unproven or contradicted by the engines' own documentation:
| Claim | What the evidence says | Verdict |
|---|---|---|
| "Add an llms.txt file to rank in AI" | Google says Search doesn't use llms.txt and that it "will neither harm nor help" visibility7. No major AI search engine has said it ranks on it. Details in our llms.txt explainer. | Optional, low impact |
| "Chunk every page into short blocks for AI" | Google: no requirement to break content into tiny pieces7. | Skip; write clear paragraphs instead |
| "Stuff the prompt wording into your page" | Keyword stuffing scored below the do-nothing baseline in the GEO paper10. | Counterproductive |
| "Schema markup gets you into ChatGPT" | OpenAI hasn't said it uses structured data to choose sources. Google says no special schema is needed for its AI features6. | Housekeeping, not a lever |
| "We rank #1 in ChatGPT" (one screenshot) | SparkToro and Gumshoe found ChatGPT and Google's AI returned the same brand list for a prompt less than 1 time in 100, and the same order under 1 time in 1,00017. | Noise; ask for a rate over many runs |
| "Block GPTBot to protect your content" | GPTBot is training only. Blocking it doesn't remove you from ChatGPT search, which is OAI-SearchBot's job1. | Fine for publishers, usually wrong for brands |
How do you measure LLM optimization?
Measure it on a fixed set of prompts, run several times per engine, because single answers vary too much to read. The 12-prompt SparkToro study ran each prompt 60 to 100 times per platform and found that lists almost never repeat, while a brand's appearance rate across many runs held far steadier than its position17 (Search Engine Journal). That gives you the core metrics:
- Mention rate, the share of runs that name you. The headline number.
- Citation rate, the share of runs that link one of your URLs.
- Share of voice, your mentions divided by all brand mentions across the set.
- Accuracy, the share of answers that describe your pricing, category and audience correctly.
First-party data fills in the rest. For Google, Search Console's generative AI performance report counts impressions in AI Overviews and AI Mode, and Google says it reached all sites on 31 August 202620. Bing Webmaster Tools' AI Performance report, in public preview since February 2026, counts how often Copilot and Bing's AI summaries cite your pages and samples the grounding queries behind them21. ChatGPT adds utm_source=chatgpt.com to referral links22, so GA4 shows which pages it sends people to. Clicks will understate the effect: Pew found people clicked a link inside a Google AI summary on about 1% of visits where one appeared18.
Our guide to measuring AI visibility has the prompt-set template and sample-size math.
A six-step LLM optimization plan for the first 60 days
1. Write 30 to 50 prompts your buyers would type (days 1 to 3)
Mix category ("best expense software for a 300-person company"), comparison, alternatives, problem and branded prompts, with no more than a fifth branded. Pull the wording from sales calls and the "best" and "vs" queries in Search Console.
2. Baseline it, three runs per prompt per engine (days 3 to 7)
Logged out or in a temporary chat, so memory doesn't flatter you. Log mention, position, cited URLs and any wrong facts.
3. Fix access (week 2)
Apply lever 1. Check robots.txt with our robots.txt AI checker, then the logs and the CDN.
4. Rewrite the five pages that should be cited (weeks 2 to 4)
Usually the homepage, pricing, the main product page and your two most-searched comparisons. Apply levers 2 and 3.
5. Work the citation map (weeks 3 to 8 and ongoing)
Lever 4. Start with the cited URLs that appear on the most prompts. Expect weeks per placement, since you're on other editors' schedules.
6. Re-run the same prompts at day 60
Same prompts, same engines, same method. Compare mention rate by prompt type, then look at which cited URLs changed. If you want this turned into a quarter-long program with owners and a scoring rubric, our 90-day GEO strategy framework picks up from here.
Set expectations by lever. Access fixes can show up within days, given OpenAI's 24-hour figure. Rewritten pages count once they've been recrawled and retrieved, so plan on weeks. Third-party mentions depend on when other people update their pages, so plan on one to three months. Answers that come from training data move only when a new model ships, so don't judge the program on those.
Frequently asked questions
What is LLM optimization in SEO?
LLM optimization in SEO, often shortened to LLMO or called LLM SEO, means shaping how large language models describe and cite your brand. You make sure AI search crawlers can reach your pages, write passages a model can quote, keep your facts consistent across profiles, and earn mentions on the pages models retrieve. Progress is measured by how often you're named across a fixed set of prompts.
How much of LLM SEO is regular SEO?
Partly. Google says AI Overviews and AI Mode run on its core ranking systems, so good SEO covers much of Google's AI. ChatGPT, Claude and Perplexity use their own retrieval, and Ahrefs found only about 12% of the URLs AI assistants cite rank in Google's top 10 for the same prompt. LLM SEO adds AI crawler access, third-party mentions and prompt-based measurement.
Can you get your brand into ChatGPT's training data?
Not directly, and not on a timeline you control. Allowing GPTBot makes your public pages eligible for future OpenAI training, and being described consistently on many sites raises the odds a future model knows you. Those changes appear only when a new model ships. The faster route is ChatGPT search: allow OAI-SearchBot and get named on the pages it retrieves.
Does llms.txt help with LLM optimization?
There's no evidence that it improves visibility. Google says its Search doesn't use llms.txt and that the file will neither harm nor help rankings, and no major AI search engine has said it uses the file to pick sources. It can help coding agents read developer docs. Add one if it takes an hour, but don't expect it to change what ChatGPT says about you.
How long does LLM optimization take to work?
It depends on the lever. Crawler fixes are the fastest: OpenAI says ChatGPT search picks up a robots.txt change in roughly a day. Rewritten pages need to be recrawled and retrieved, which takes weeks. Mentions on third-party pages take one to three months because other editors set the schedule. Answers drawn from training data change only with new model releases.
Should I block GPTBot or ClaudeBot?
Only if your content is what you sell, such as news, research or paid courses. GPTBot and ClaudeBot collect training data. Blocking them doesn't remove you from ChatGPT or Claude search, which rely on OAI-SearchBot and Claude-SearchBot. For a business that wants to be recommended, appearing in future training data is usually worth more than the protection, so leave both allowed.
Sources
- OpenAI. "Overview of OpenAI crawlers."
- OpenAI Help Center. "ChatGPT search."
- Anthropic. "Does Anthropic crawl data from the web, and how can site owners block the crawler?"
- Perplexity. "Perplexity crawlers."
- Google Search Central. "Google's common crawlers" (Google-Extended).
- Google Search Central. "AI features and your website."
- Google Search Central. "Google's guide to optimizing for generative AI features on Google Search."
- Lewis et al. "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." NeurIPS 2020.
- Liu et al. "Lost in the Middle: How Language Models Use Long Contexts." TACL, 2023.
- Aggarwal et al. "GEO: Generative Engine Optimization." KDD 2024.
- Ahrefs. "Only 12% of AI cited URLs rank in Google's top 10 for the original prompt" (15,000 prompts). 11 August 2025.
- Ahrefs. "Update: 38% of AI Overview citations pull from the top 10" (863K SERPs, about 4M cited URLs). 2 March 2026.
- Ahrefs. "An analysis of AI Overview brand visibility factors (75K brands studied)." 26 May 2025.
- Ahrefs. "AI assistants prefer to cite fresher content" (about 17 million citations). 28 July 2025.
- Profound. "AI platform citation patterns" (680 million citations, August 2024 to June 2025).
- Semrush. "The most-cited domains in AI: a 3-month study" (230,000+ prompts, July to October 2025).
- Search Engine Journal. "AI recommendations change with nearly every query: SparkToro" (SparkToro and Gumshoe study).
- Pew Research Center. "Google users are less likely to click on links when an AI summary appears in the results." 22 July 2025.
- Cloudflare. "Cloudflare just changed how AI crawlers scrape the internet-at-large." 1 July 2025.
- Google Search Central Blog. "Introducing Search generative AI performance reports in Search Console." 3 June 2026.
- Bing Webmaster Blog. "Introducing AI Performance in Bing Webmaster Tools public preview." 10 February 2026.
- OpenAI Help Center. "Publishers and developers FAQ."
AI Ranked Editorial. "LLM optimization (LLM SEO): how to get your brand into AI answers." AI Ranked, October 11, 2026. https://airanked.ai/guides/llm-optimization