- Start from 30 to 50 buyer prompts, not keywords. In a SparkToro and Gumshoe study, AI tools returned the same brand list for a prompt less than 1 time in 100, so you measure rates across runs.
- Pick engines by audience size and fit: ChatGPT reported 800 million weekly users in October 2025, and Google reported over 2 billion monthly users for AI Overviews in mid-2025.
- Off-site work usually carries the quarter. Branded web mentions correlated with AI Overview visibility at 0.664 in Ahrefs' 75,000-brand study, against 0.218 for backlinks.
- Technical access is the fastest win: OpenAI says robots.txt changes reach ChatGPT search in about 24 hours.
- Measure every four weeks with three runs per prompt per engine, and score each run on a 0 to 3 rubric so progress shows up before rankings do.
What is a GEO strategy?
A GEO strategy is the set of decisions that turns generative engine optimization from a list of tactics into a program: which questions you want to be the answer to, where AI engines get their information about your category, what you'll change on your site and off it, who owns each piece, and what number tells you it's working. (GEO here means generative engine optimization. If you were looking for geostrategy in the geopolitical sense, this isn't that page.)
Most published "GEO strategies" are tactic lists: add FAQs, use schema, write conversationally. Tactics are fine, but a list doesn't tell you which prompt to start with, which engine to ignore, or when to stop. The framework below does, in six workstreams over one quarter, with templates you can copy into a sheet.
It also helps to know what Google thinks. Its guidance says AI Overviews and AI Mode are "rooted in our core Search ranking and quality systems" and that there are no additional technical requirements to appear in them45. ChatGPT, Claude and Perplexity run their own retrieval, and Ahrefs found only about 12% of URLs cited by AI assistants ranked in Google's top 10 for the same prompt18 (Ahrefs). So your SEO program covers part of Google's AI surfaces and much less of the rest. The strategy is about the rest.
Why plan GEO in 90-day cycles?
Because the levers move at very different speeds, and a quarter is the shortest window in which all of them show up. OpenAI says a robots.txt change reaches ChatGPT search in about 24 hours1. Rewritten pages count once they're recrawled and retrieved, which takes weeks. Getting named on a third-party listicle or review page depends on when its editor updates, which takes one to three months. Answers built from training data change only when a new model ships, which you don't control at all.
There's also a measurement reason. AI answers vary so much from run to run that one month of data gives you a starting point and nothing more. Three readings, four weeks apart, is the minimum we'd use to call a trend. That's a quarter.
The 90-day GEO strategy framework at a glance
| Weeks | Workstream | Output | Done when |
|---|---|---|---|
| 1 to 2 | 1. Audit | Access, indexing, citability and entity findings | Every failed check has an owner and a date |
| 1 to 2 | 2. Prompt set and baseline | 40 weighted prompts, scored across engines | 3 runs per prompt per engine logged and scored |
| 2 | 3. Technical access | Crawler, CDN and snippet fixes live | All AI search bots get 200s on money pages |
| 3 to 4 | 4. Content gap map | Each low-scoring prompt tagged with a gap type and fix | A ranked backlog of page and outreach work |
| 4 to 12 | 5. Third-party presence | Placements on the most-cited URLs | Named on half of the top 20 cited URLs |
| 2, 6, 10, 13 | 6. Measurement | Monthly scorecard, quarter readout | Same prompts, same method, every cycle |
Workstream 1: run a GEO audit (weeks 1 to 2)
The audit answers one question: if an AI engine wanted to cite you, could it? Work through our GEO audit checklist, which turns this into 32 pass/fail checks across crawler access, indexing, page citability, structured data, off-site entity signals and measurement. For strategy purposes you need three outputs from it:
- A list of access failures (blocked bots, CDN challenges, pages not indexed in Google or Bing, snippet controls hiding content). These go straight to workstream 3.
- Your five to ten money pages, each graded for whether it states what you are, who it's for and what it costs in the first 100 words.
- The category label and core facts as they appear on your site, LinkedIn, Crunchbase, G2 or Capterra, and Wikidata if you have an entry. Mismatches go on the backlog.
Workstream 2: choose the prompt set (weeks 1 to 2)
Your prompt set is the strategy's scoreboard, so it's worth getting right once and then freezing it for the quarter. Write 40 prompts, give each a type, a funnel stage and a business weight, and track them across the engines your buyers use. Copy these columns:
id,prompt,type,stage,constraint,engines,weight,owner
C01,best scheduling software for multi-location clinics,category,consider,multi-location,"chatgpt,aio,perplexity",3,content
C04,scheduling tool for clinics that integrates with athenahealth,category,consider,integration,"chatgpt,aio",3,content
V02,Rostera vs Calendly for medical practices,comparison,decide,,"chatgpt,perplexity",2,product marketing
A01,alternatives to Acuity Scheduling for clinics,alternatives,consider,,"chatgpt,aio",2,content
P03,how to reduce no-shows at a physiotherapy clinic,problem,learn,,"aio,perplexity",1,content
B02,is Rostera HIPAA compliant,branded,decide,,"chatgpt,aio",1,product marketing
Rules that keep the set honest:
- No more than 20% branded. Branded prompts make the scorecard look good and tell you little about new buyers.
- At least a third with a constraint (team size, integration, industry, budget). Category leaders lock up generic prompts; constraints are where a challenger gets named.
- Weight by revenue, not volume. 3 for prompts that precede a demo or purchase, 2 for comparisons and alternatives, 1 for education and branded checks.
- Source the wording from buyers. Sales call notes, support tickets, the "best" and "vs" queries in Search Console, and the category names on G2.
Then pick engines. Our rule: include ChatGPT and Google AI Overviews for almost everyone, given their reach (OpenAI reported 800 million weekly ChatGPT users in October 202511; Google reported over 2 billion monthly AI Overviews users and over 100 million monthly AI Mode users in the US and India in its Q2 2025 results12). Add Perplexity if your buyers are technical or research-heavy, and Claude if they're developers or knowledge workers who already live in it. Two engines measured properly beat five measured once.
The scoring rubric
Mention rate alone hides the difference between "listed tenth" and "recommended first." Score each run instead, then average per prompt:
| Score | What the answer does |
|---|---|
| 0 | Doesn't mention you |
| 1 | Mentions you below fifth, or describes you wrongly (old pricing, wrong category) |
| 2 | Names you in the top five with an accurate description |
| 3 | Recommends you first, or as the best fit for the prompt's constraint, accurately |
Log whether your URL was cited as a separate yes/no column; a citation is worth tracking but it isn't the same as a recommendation. Then rank the work with one formula:
prompt_score = average of run scores (0 to 3), per engine
priority = weight x (3 - prompt_score)
quarter_target = move every weight-3 prompt to a score of 2 or higher on at least one engine
Three runs per prompt per engine is the floor. The reason is in the data: SparkToro and Gumshoe ran 12 recommendation prompts 60 to 100 times each and found the same brand list less than 1% of the time, and the same list in the same order less than 0.1%, while how often a brand appeared was far more stable6 (Search Engine Journal). For more on prompt tracking at scale, see the measurement section below.
Workstream 3: fix technical access (week 2)
This is the cheapest workstream and the one most often skipped. The checks, in order:
- Search crawlers allowed. OAI-SearchBot (ChatGPT search)1, Claude-SearchBot2, PerplexityBot3, plus Googlebot and Bingbot. Test robots.txt with our robots.txt AI crawler checker.
- The CDN agrees. Cloudflare announced on 1 July 2025 that it would block AI crawlers by default and ask every new domain whether to allow them17. Check the bot settings and your server logs, not only robots.txt.
- Training crawlers are a separate decision. GPTBot and ClaudeBot govern training, not search answers12. Block them only if your content is the product you sell. Our AI crawler access report shows how 261 major sites split that decision.
- Snippets aren't suppressed. For AI Overviews and AI Mode, Google requires only that the page be indexed and allowed to show a snippet4. A stray nosnippet or data-nosnippet on a template removes you quietly.
Workstream 4: map content gaps from the baseline (weeks 3 to 4)
Every prompt that scored under 2 has a reason, and the cited URLs usually tell you which one. Tag each low-scoring prompt with one gap type:
| Gap type | How you spot it | Fix | Typical time |
|---|---|---|---|
| No page | No page on your site answers the prompt; competitors' pages get cited | Build it: comparison, alternatives, integration or use-case page | 2 to 4 weeks |
| Page not cited | You have a page, but it's vague, buried or missing facts | Rewrite the first 100 words; add numbers, prices, named integrations | 1 to 2 weeks |
| Source gap | Cited URLs are third-party pages that don't name you | Outreach (workstream 5) | 1 to 3 months |
| Wrong facts | You're named but described wrongly | Find the cited page carrying the old fact; fix it or the profile it came from | Days to weeks |
| Training-only | Answer has no links and doesn't name you | Nothing direct; keep working source gaps for future models | Next model release |
For "page not cited" fixes, the GEO paper is the best evidence on what to change. On its 10,000-query benchmark, adding quotations, statistics and cited sources raised visibility, keyword stuffing lowered it, and citing sources helped pages ranked fifth by 115%19 (Aggarwal et al.). Grade rewrites with the citability grader before they ship. Google adds that you don't need to write in a special way for AI or chop content into tiny pieces5, which matches what we see: clear, specific pages win.
Freshness matters for pages you expect AI assistants to cite. Ahrefs found AI-cited URLs were 25.7% fresher than organic results on average across about 17 million citations, with Google's AI Overviews the exception10 (Ahrefs). Put the pages that carry your weight-3 prompts on a quarterly review where real facts get updated.
Workstream 5: build third-party presence (weeks 4 to 12)
For category and alternatives prompts, this is usually where the quarter is won. In Ahrefs' study of 75,000 brands, branded web mentions correlated with AI Overview visibility at 0.664, against 0.218 for backlink count7 (Ahrefs). The model repeats what the pages it reads agree on, so you need to be on those pages.
Don't guess which pages those are. Engine source mixes differ and swing: Profound found Wikipedia was 47.9% of ChatGPT's top 10 cited sources and Reddit 46.7% of Perplexity's8, and Semrush saw the share of ChatGPT responses citing Reddit drop from close to 60% to around 10% between early August and mid-September 20259. Build your own citation map from the baseline instead:
url | prompts_cited_on | engines | named_us (y/n) | page_type | owner_contact | ask | status | next_check
Sort by prompts_cited_on and work from the top. The asks that tend to land are specific and checkable: a trial login for a reviewer, a pricing page they can verify, a customer they can call. On review sites, fill out the category profiles the engines cited. In community threads, answer only where you have something concrete to add, and say who you are. Google cautions that inauthentic mentions aren't "as helpful as it might seem"5, and paid slots on thin listicles are the first thing a buyer, or a model, learns to discount. Track each brand mention you earn against the prompts its page is cited for.
Workstream 6: set the measurement cadence (weeks 2, 6, 10 and 13)
Run the full prompt set on a fixed schedule: baseline in week 2, then weeks 6 and 10, and a full readout in week 13. Same prompts, same engines, logged out or in a temporary chat, three runs each. Every scorecard shows four things next to last cycle's numbers: average score by prompt type, mention rate, share of voice against named competitors, and the top cited URLs.
Add first-party data where it exists:
- Google Search Console generative AI performance report: impressions in AI Overviews and AI Mode, rolled out to all sites on 31 August 202614.
- Bing Webmaster Tools AI Performance report (public preview since February 2026): citation counts for Copilot and Bing's AI summaries, plus sampled grounding queries15.
- GA4: ChatGPT adds utm_source=chatgpt.com to referral links16, so you can see which landing pages it sends people to.
Expect clicks to undercount influence. Pew found Google users clicked a link inside an AI summary on about 1% of visits where one appeared, and clicked any result on 8% of those visits against 15% without a summary13 (Pew Research Center). Our guide to measuring AI visibility covers sample sizes and a weekly report template.
Worked example: a 90-day plan for a clinic scheduling tool
This example is hypothetical. "Rostera" is an invented scheduling product for multi-location clinics, and the scores below are made up to show how the math drives decisions, not results from real runs.
Weeks 1 to 2. The audit finds Cloudflare challenging PerplexityBot and a nosnippet tag on the pricing page template. Rostera writes 40 prompts (14 category, 8 comparison, 6 alternatives, 6 problem, 6 branded), weights them, and runs them three times each in ChatGPT and Google AI Overviews: 240 scored runs.
Baseline, illustrative numbers.
| Prompt type | Prompts | Avg score ChatGPT | Avg score AIO | Main gap type |
|---|---|---|---|---|
| Category | 14 | 0.4 | 0.7 | Source gap |
| Comparison | 8 | 1.1 | 1.3 | No page (two pairs) |
| Alternatives | 6 | 0.2 | 0.5 | Source gap |
| Problem | 6 | 0.8 | 1.5 | Page not cited |
| Branded | 6 | 2.1 | 2.4 | Wrong facts (old pricing) |
Reading it. Category prompt C04 (athenahealth integration) has weight 3 and a ChatGPT score of 0.33, so its priority is 3 x 2.67 = 8.0, the highest in the set. Branded prompt B02 scores 2.67 at weight 1, priority 0.33, so it waits. The category and alternatives rows are mostly source gaps: the cited URLs are two review-site category pages, three listicles and a Reddit thread, and none names Rostera.
Weeks 3 to 12. Rostera fixes the access problems in week 2, builds the two missing comparison pages and an athenahealth integration page by week 6, corrects its pricing on G2 and Capterra, and pitches the three listicle authors with a trial account and a public pricing page. Weeks 6 and 10 re-run the set. At week 13 the readout compares average score by prompt type against the baseline, lists which cited URLs now name Rostera, and sets next quarter's prompt set, with the branded and problem prompts that hit a score of 2 swapped for new constrained category prompts.
Should you run GEO in-house, with tools, or with an agency?
It depends on two things you can count: how many prompts and engines you track, and whether you have someone who can do outreach. A rough rule:
- In-house with a spreadsheet works up to about 40 prompts on two engines. That's 240 runs a cycle, roughly a day of work by hand at two minutes a run.
- In-house with a tracker makes sense past that, or if you report weekly. Our list of the best AEO tools covers trackers and the content tools that help with workstream 4.
- An agency makes sense when workstream 5 is the bottleneck and nobody on the team owns outreach and review-site work. Agencies that specialize in AI search run this kind of program on retainer; our GEO agency comparison lists who does it. Our GEO agency pricing guide lays out published retainers and minimum terms so you can compare.
Where GEO strategies go wrong
- Tracking only branded prompts. "Is Rostera good" scores 2.4 and the dashboard looks great, while every category prompt names four competitors. Cap branded at 20% of the set.
- Changing the prompt set mid-quarter. New prompts make the trend line meaningless. Add them to next quarter's set.
- Reporting average position. With lists that almost never repeat, position is a noisy number. Lead with score and mention rate.
- Treating it as an SEO project. Google rankings help with AI Overviews, but with only about 12% of assistant citations coming from Google's top 1018, an SEO-only plan misses most of ChatGPT and Perplexity.
- Spending week 1 on llms.txt. Google says its Search doesn't use it5. Do access, pages and outreach first.
- Judging the quarter on training-only answers. Those move with model releases. Judge on prompts where the engine searched and cited sources.
Engine-specific playbooks
The framework stays the same across engines; what changes is where each one looks for sources. Each engine gets its own playbook: ChatGPT, Google AI Overviews, Perplexity, Gemini and Claude. If you want the underlying mechanics of training data versus retrieval first, start with our guide to LLM optimization.
Frequently asked questions
What is a GEO strategy?
A GEO strategy is a plan for getting your brand mentioned and cited in AI-generated answers. It defines the buyer prompts you want to win, the AI engines that matter for your audience, the pages and third-party sources you'll change, who owns each workstream and how you'll measure progress. It differs from a list of GEO tactics because it sets priorities and a timeline.
How long does a GEO strategy take to show results?
Plan on one quarter for a first read. Crawler fixes reach ChatGPT search in about a day, rewritten pages take weeks to be recrawled and cited, and third-party placements take one to three months. Because AI answers vary between runs, you need at least three monthly measurements to see a trend. Answers from training data only shift when new models are released.
How many prompts should a GEO prompt set have?
Thirty to fifty works for most brands, with 40 a good default. With fewer, one odd prompt skews the result; with more, manual tracking gets slow. Keep branded prompts under 20%, give at least a third a buyer constraint such as company size or an integration, and give every prompt three runs on each engine per cycle.
What's the difference between a GEO strategy and an SEO strategy?
An SEO strategy targets positions on a results page for keywords. A GEO strategy targets mentions and citations inside AI answers for prompts, and since those answers change every run, it measures rates instead of rank. It also leans harder on third-party pages: Ahrefs found branded web mentions tracked AI Overview visibility more closely than backlinks did. The technical base overlaps heavily.
Which AI engines should a GEO strategy cover?
Start with ChatGPT and Google AI Overviews for almost any business, because they reach the most people: 800 million weekly ChatGPT users in October 2025 and over 2 billion monthly AI Overviews users in mid-2025. Add Perplexity for technical or research-heavy buyers and Claude for developer audiences. Measuring two engines properly beats measuring five once.
Do you need an agency for GEO?
Not always. One person can run 40 prompts on two engines with a spreadsheet in about a day per cycle, and most teams can fix access and pages in-house. Agencies earn their fee on third-party work, meaning outreach to listicle authors, review-site profiles and community presence, which most marketing teams don't staff. Compare published retainers and minimum terms before signing.
Sources
- OpenAI. "Overview of OpenAI crawlers."
- Anthropic. "Does Anthropic crawl data from the web, and how can site owners block the crawler?"
- Perplexity. "Perplexity crawlers."
- Google Search Central. "AI features and your website."
- Google Search Central. "Google's guide to optimizing for generative AI features on Google Search."
- Search Engine Journal. "AI recommendations change with nearly every query: SparkToro" (SparkToro and Gumshoe study).
- Ahrefs. "An analysis of AI Overview brand visibility factors (75K brands studied)." 26 May 2025.
- Profound. "AI platform citation patterns" (680 million citations, August 2024 to June 2025).
- Semrush. "The most-cited domains in AI: a 3-month study" (230,000+ prompts, July to October 2025).
- Ahrefs. "AI assistants prefer to cite fresher content" (about 17 million citations). 28 July 2025.
- TechCrunch. "Sam Altman says ChatGPT has hit 800M weekly active users." 6 October 2025.
- Google. "Q2 2025 earnings call: CEO's remarks." July 2025.
- Pew Research Center. "Google users are less likely to click on links when an AI summary appears in the results." 22 July 2025.
- Google Search Central Blog. "Introducing Search generative AI performance reports in Search Console." 3 June 2026.
- Bing Webmaster Blog. "Introducing AI Performance in Bing Webmaster Tools public preview." 10 February 2026.
- OpenAI Help Center. "Publishers and developers FAQ."
- Cloudflare. "Cloudflare just changed how AI crawlers scrape the internet-at-large." 1 July 2025.
- Ahrefs. "Only 12% of AI cited URLs rank in Google's top 10 for the original prompt" (15,000 prompts). 11 August 2025.
- Aggarwal et al. "GEO: Generative Engine Optimization." KDD 2024.
AI Ranked Editorial. "GEO strategy: a 90-day framework for AI search visibility." AI Ranked, October 11, 2026. https://airanked.ai/guides/geo-strategy