GEO glossary: 40 AI search and AEO terms defined
The vocabulary of AI search in plain English. Each entry starts with a definition you can quote. Where a worked example or a common mistake teaches more than the definition, we include one.
- AI ModeGoogle Search feature
AI Mode is Google's conversational search mode. It answers harder questions with a longer AI-written response, runs many related searches behind the scenes (query fan-out) and links the pages it used.
Google tested it in Labs from March 2025, opened it to all U.S. users that May, and said at I/O in May 2026 that it had passed 1 billion monthly users. Don't assume a citation in AI Overviews carries over: in Ahrefs' comparison of 540,000 query pairs the two features cite the same URLs only 13.7% of the time.
- AI OverviewsGoogle Search feature · formerly SGE
AI Overviews are the AI-written summaries Google Search shows above the regular results for some queries, with links to the pages the summary drew on. They replaced the Search Generative Experience and launched in the U.S. at I/O in May 2024.
The traffic question in one number: in Pew Research Center's March 2025 browsing data, people clicked a regular result on 8% of visits to a page with a summary and 15% without one. Links inside the summary itself got clicked on 1% of visits. At the same I/O keynote, Google put AI Overviews at more than 2.5 billion monthly users.
- AI SEOUmbrella term · also AI search optimization
AI SEO is a loose label for getting a brand into AI-powered search, whether that's answer engines like ChatGPT or AI features inside Google. Vendors also use it to mean doing ordinary SEO with AI tools, so ask which one they're selling.
If you need a precise term, use generative engine optimization, which at least has a paper behind it.
- AI visibilityMetric
AI visibility is how often, and how prominently, AI systems like ChatGPT, Gemini, Perplexity, Claude and Google's AI features name a brand or cite its pages for the questions its buyers ask.
In practice it's three numbers from a fixed prompt set: how often you're mentioned, how often you're cited, and your share of the brands named. Answers change from run to run, so a single check is an anecdote. For a quick first read, Arobis AI runs a free AI visibility checker; for weekly tracking, our roundup of the best AI visibility tools compares the paid platforms.
- Answer engineCategory
An answer engine is any search product that replies with a written answer and sources instead of ten blue links: Perplexity, ChatGPT search, Google AI Mode, Microsoft Copilot.
Each runs its own retrieval, which is why one page can be cited by Perplexity and missing from ChatGPT on the same day. More in how AI engines choose sources.
- Answer engine optimizationAEO
Answer engine optimization (AEO) is the practice of shaping content so an AI engine can lift it as a direct answer, using question-led headings, a one-sentence answer under each, clean definitions and lists.
It overlaps almost entirely with GEO. If there's a difference, AEO cares about the answer format and GEO about earning the citation. Our comparison of GEO vs AEO vs SEO argues the labels matter less than the work.
- Brand mentionMetric
A brand mention is your brand's name appearing in an AI answer, linked or not. It's the basic unit of AI visibility tracking, and it's not the same as a citation: ChatGPT can recommend you by name while linking only to a G2 review page.
Mentions on the wider web seem to matter more than links. In Ahrefs' study of 75,000 brands, branded web mentions correlated with AI Overview visibility at 0.664, against 0.218 for backlinks.
- ChunkingRetrieval
Chunking is splitting documents into smaller passages before they're embedded and indexed, so a retrieval system can match the one section that answers a question instead of a whole page.
No AI engine publishes its chunk sizes, and Google's guide to generative AI features says there's no requirement to break content into tiny pieces. The useful habit is older than AI: make each section make sense on its own, with the answer in its first sentence.
- CitationCore concept
A citation is the link or source reference an AI answer attaches to a claim. It tells the reader where the information came from and is the only part of an AI answer that can send you a visit.
How many slots there are depends on the engine. In an August 2026 test of 144 answers, Perplexity returned exactly 20 citations in 34 of 36 answers; Claude returned none in 16 of 36. The citability grader scores how easy a page is to cite.
- Citation rateMetric
Citation rate is the share of tracked AI answers that cite a given domain or URL, measured over a fixed set of prompts.
Example: you run 40 prompts through ChatGPT three times each, 120 answers. Your domain is cited in 18 of them. Citation rate is 15%. Run it again next week and get 21 of 120 (17.5%), and you can't call that a gain yet; answers vary that much between identical runs. Look for movement that holds over several weeks. The method is in how to measure AI visibility, and our test of the best free AI visibility checkers compares what each one actually measures.
- ClaudeBotCrawler · Anthropic
ClaudeBot is Anthropic's training crawler. It collects public web content that could be used to train Claude, and disallowing it tells Anthropic to leave your future content out.
Anthropic runs two more agents with their own tokens: Claude-SearchBot for search results and Claude-User for pages a person asks about (Anthropic help center). Blocking ClaudeBot alone doesn't affect search. The Guardian's robots.txt names all three, plus the older anthropic-ai token, in one
Disallow: /group.- CrawlerAlso bot, spider, user agent
A crawler is a program that requests web pages automatically to collect their content. Search engines use crawlers to build an index; AI companies use them to gather training data or fetch pages for live answers. Each one identifies itself with a user agent string.
Most AI crawling is still for training: Cloudflare found 80% of AI crawling in the year to July 2025 was for training, 18% for search and 2% for user actions. Your robots.txt isn't the only gate either. Cloudflare began blocking known AI crawlers by default for new domains on July 1, 2025, so check your CDN settings too. Full list in every AI crawler user agent.
- E-E-A-TGoogle quality framework
E-E-A-T stands for experience, expertise, authoritativeness and trustworthiness, the yardstick Google's search quality raters use to judge whether content comes from a credible source with first-hand knowledge.
Google added the first E, for experience, in December 2022. The guidelines evaluate Google's ranking systems; they aren't a ranking factor you can switch on. A product review that says "we ran it on 12 client accounts for six weeks" shows experience. One that restates the vendor's feature list doesn't, however many author bios sit next to it.
- EmbeddingsRetrieval · also vectors
Embeddings are lists of numbers that represent what a piece of text means, so passages with similar meaning end up close together. That's how a retrieval system finds "how much does it cost per user" when your page says "pricing per seat".
- EntitySemantic search
An entity is a specific, identifiable thing (a company, person, product, place or concept) that search engines and language models can recognize and connect to other things, however it's worded.
Example: "Example Co", "Example Co Inc." and "ExampleCo" should all resolve to one company. You make that easier by using one name everywhere, giving your Organization markup a fixed
@id, and listing your real profiles insameAs. Entities and their relationships live in a knowledge graph.- FAQ schemaStructured data · FAQPage
FAQ schema is FAQPage structured data: markup that labels each question and answer on a page so machines can read them as pairs.
It no longer earns anything visible in Google. FAQ rich results were limited to well-known government and health sites in August 2023, and Google's documentation records that they stopped appearing in Search on May 7, 2026. The markup is still valid and harmless, provided the visible FAQ and the markup say exactly the same thing.
- Generative engine optimizationGEO
Generative engine optimization (GEO) is the practice of making content more likely to be retrieved, quoted and cited in AI-generated answers from ChatGPT, Perplexity, Gemini, Claude, Google AI Mode and similar engines.
The term comes from a 2023 paper by researchers at Princeton and IIT Delhi (published at KDD 2024). They tested nine rewrites, and adding citations, quotations and statistics raised a source's visibility by up to 40%. Keyword stuffing did close to nothing. Start with what is generative engine optimization; if you're weighing outside help, our ranking of the best GEO agencies compares 12 of them on fit, published pricing and terms.
- Google-Extendedrobots.txt token · Google
Google-Extended is a robots.txt token, not a crawler. Disallowing it tells Google not to use content it crawls for Gemini training or to ground Gemini answers. Fetching is still done by Google's regular user agents.
The common mistake is expecting it to keep you out of AI Overviews. Google states it has no effect on inclusion in Google Search, and AI Overviews and AI Mode are part of Search. To limit what appears there, Google points to
nosnippet,max-snippetornoindex.- GPTBotCrawler · OpenAI
GPTBot is OpenAI's training crawler. Disallowing it in robots.txt tells OpenAI not to use your content to train its models.
It has nothing to do with whether ChatGPT search can cite you; that's OAI-SearchBot (OpenAI's crawler documentation). In the same Cloudflare data, GPTBot made up 28.1% of AI-only crawler traffic in July 2025, up from 11.9% a year earlier. Plenty of sites already use the split: in our 261-site crawler access study, 31 of the 59 sites that disallow GPTBot still allow OAI-SearchBot.
- GroundingAlso grounded generation
Grounding is basing an answer on information retrieved at the moment of the question, such as search results, instead of only on what the model learned in training. Grounded answers are the ones that carry citations.
Ask an assistant about a pricing change you made last week. If it searches, it can find the new price. If it answers from memory, it gives the old one, or a guess. Being citable means being reachable by the engine's search layer when the question is asked.
- HallucinationAlso confabulation
A hallucination is a confident statement from an AI model that is false or unsupported by its sources: an invented statistic, a feature your product doesn't have, a citation to a page that doesn't exist.
It's common enough to plan for. In a March 2025 Tow Center test of eight AI search tools, answers misidentified the source of news excerpts in more than 60% of 1,600 queries. If ChatGPT says your product lacks SSO and it doesn't, the fix is usually a clear, crawlable page stating the fact, plus third-party pages that agree.
- Knowledge cutoffAlso training cutoff
A knowledge cutoff is the date after which a model's training data stops. Anything newer, from a product launch to a price change, reaches its answers only if the engine searches the web. That's why grounding and crawlable pages matter most for recent facts.
- Knowledge graphSemantic search
A knowledge graph is a database of entities and the relationships between them. Search engines use it to work out who or what a query refers to and to show facts directly in results.
Google introduced its Knowledge Graph in May 2012, and knowledge panels draw on it. If two companies share your name, consistent entity details across your site and public profiles lower the odds of an engine mixing you up.
- Large language modelLLM
A large language model (LLM) is a neural network trained on very large amounts of text to predict the next token, which lets it write, summarize and answer in natural language. It's the engine inside ChatGPT, Claude and Gemini.
By itself an LLM answers from training data. AI search products bolt retrieval onto it so answers can cite current pages; see retrieval-augmented generation.
- LLM SEOAlso LLMO, LLM optimization
LLM SEO is mostly a synonym for generative engine optimization. Some people use it more narrowly for influencing what models learn in training, as opposed to what they retrieve when answering.
- llms.txtProposed file standard
llms.txt is a proposed Markdown file at a site's root that gives language models a short summary of the site and a curated list of its most useful pages.
Jeremy Howard proposed it on September 3, 2024, and it's still a proposal. Google's guide to generative AI features says Google Search ignores such files. It's most useful for developer docs. Read llms.txt: what it is and whether to bother.
- OAI-SearchBotCrawler · OpenAI
OAI-SearchBot is the crawler behind ChatGPT search. Sites that block it in robots.txt won't be shown in ChatGPT search answers.
Because it's separate from GPTBot, you can opt out of training and stay citable. The New York Times blocks both; The Guardian, which signed a content deal with OpenAI in February 2025, names neither. OpenAI says robots.txt changes take about 24 hours to apply.
- PerplexityBotCrawler · Perplexity
PerplexityBot is Perplexity's indexing crawler, used to surface and link pages in Perplexity answers. Perplexity says it isn't used to train foundation models.
Its sibling, Perplexity-User, fetches pages when someone asks a question and, per Perplexity's documentation, generally ignores robots.txt because a person requested the fetch.
- PromptCore concept
A prompt is the question or instruction someone types into an AI assistant. In AI visibility work it's the unit of measurement: the buyer question you test to see which brands an engine names.
Prompts are longer than searches. Semrush measured an average of 7.22 words per Google AI Mode query in 2025, against 4.0 for traditional Google searches. Compare "crm software" with "best CRM for a 20-person B2B sales team that uses HubSpot for marketing". Your prompt set should look like the second one.
- Prompt trackingMeasurement
Prompt tracking is running a fixed set of prompts through AI engines on a schedule and recording which brands, pages and sources appear, so changes in AI visibility can be measured.
Answers move between runs, so good tracking repeats each prompt several times and reports ranges. It's now standard practice: in SEOFOMO's 2026 survey of 171 search professionals, 92% said they track AI visibility and citations. Profound, which raised a $180M Series D in September 2026, sells it to enterprise brands, and our Profound alternatives guide covers when a smaller tool does the job.
- Query fan-outRetrieval technique
Query fan-out is the technique of splitting one question into many sub-queries, running them in parallel and combining the results into one answer.
Illustration: "best payroll software for a 50-person startup in Germany" might fan out into searches about German payroll compliance, payroll tools with EU data hosting, and pricing for small teams. A page that answers only one of those can still be cited. Google uses fan-out in AI Mode, and its Deep Search feature can issue hundreds of searches for one request.
- Retrieval-augmented generationRAG
Retrieval-augmented generation (RAG) is a method in which a system first retrieves relevant documents from an index, then has a language model write an answer grounded in that text.
The term comes from a 2020 paper by Patrick Lewis and colleagues at Facebook AI Research, University College London and New York University. AI search engines are RAG at web scale, which gives you the order of operations: a page that isn't retrieved can't be cited, however good it is.
- robots.txtProtocol · RFC 9309
robots.txt is a plain-text file at a site's root that tells crawlers which paths they may request, with rules addressed to specific user agents.
A three-line example that keeps ChatGPT search access while opting out of OpenAI training:
User-agent: GPTBot Disallow: / User-agent: OAI-SearchBot Allow: /The protocol was standardized as RFC 9309 in September 2022, which says its rules aren't access authorization: well-behaved bots follow them, others don't have to. Test your own file with the AI crawler robots.txt checker.
- Schema markupAlso JSON-LD, Schema.org
Schema markup is code, usually JSON-LD, that labels what's on a page (an article, a company, a product, a person) using the Schema.org vocabulary, so machines can read the facts without guessing.
It's the most common form of structured data. Google says no special markup is needed to appear in its AI features, so treat it as housekeeping. Generate Article and Organization markup with the schema generator.
- SentimentMetric
Sentiment is the tone an AI answer takes when it describes a brand: positive, neutral or negative. Tracked next to mentions, because "Acme is popular but has frequent outages" is a mention you'd rather not have.
About half of practitioners say they watch it: 87 of the 171 respondents to the same SEOFOMO survey picked brand sentiment in AI answers as one of their metrics.
- Source diversityCitation pattern
Source diversity is how many different domains an engine draws on for a topic. Low diversity means a few sites take most of the citations.
Engines differ a lot. Profound's analysis of 680 million citations found Wikipedia made up 47.9% of ChatGPT's top-10 sources, while Reddit made up 46.7% of Perplexity's. If one site dominates your topic in one engine, being present on that site may matter more than your own pages.
- Structured dataTechnical SEO
Structured data is information on a page in a standard machine-readable format, usually Schema.org vocabulary written as JSON-LD, describing what the page and its contents are.
Google uses it for rich results and to understand entities, and its policies are strict about one thing: don't mark up content readers can't see. Breaking that can cost a page its rich result eligibility. See schema markup.
- Training dataLLM concept
Training data is the text, code and other material a model learns from before release. It shapes what the model says when it answers without searching.
GPTBot and ClaudeBot collect public web pages for this. Blocking them keeps your future content out of training, but it doesn't affect AI search, which uses separate agents. For most B2B brands, being in training data is how a model learns your category includes you. See knowledge cutoff.
- Zero-click searchSearch behavior
A zero-click search is one that ends without a click to any website, because the answer was on the results page or in an AI summary.
AI Mode pushes this to the extreme: Semrush clickstream data put 92% to 94% of AI Mode sessions at zero clicks in mid-2025. More numbers are in our AI search statistics.
AI Ranked Editorial. "GEO glossary: 40 AI search and AEO terms defined." AI Ranked, October 11, 2026. https://airanked.ai/glossary/