- In AI Ranked's study of 261 major websites, 34% of those that gave a clear answer serve a valid llms.txt. B2B SaaS leads at 64%; news sits at 8%.
- Ahrefs found 28% of 137,210 domains in its analytics data publish an llms.txt, but 97% of those files received zero requests in May 20265.
- SE Ranking's study of about 300,000 domains found no relationship between having llms.txt and how often a domain is cited by LLMs6.
- Google says Google Search does not use llms.txt, and that creating one "will neither harm nor help" visibility in Search4.
- Stripe's llms.txt runs to 455 links across 26 sections and includes a block of instructions written for coding agents, which shows who the file actually serves.
What is llms.txt?
llms.txt is a plain Markdown file published at /llms.txt that gives language models a short, curated guide to a site. Jeremy Howard, co-founder of Answer.AI, proposed it on 3 September 20242. His case was that building context from a website is ambiguous, and that "site authors know best" which content a model should use2.
Think of it as a reading list. robots.txt says where crawlers may go and a sitemap lists every URL; llms.txt picks the 20 or 200 pages that explain you best, with a line on each, in a format a large language model can read without parsing HTML.
It's a community proposal, and the spec at llmstxt.org calls itself one1. No standards body has adopted it. robots.txt, for comparison, is an IETF standard, RFC 93099.
What does an llms.txt file look like?
It's Markdown with a fixed order, and the spec at llmstxt.org requires only one element: an H1 with the site or project name1. After that H1 comes a blockquote with a short summary, then any number of paragraphs or lists with no headings, then H2 sections of Markdown links, each with an optional colon and note. A final H2 called "Optional" holds links a tool can skip when it needs a shorter context.
Here's a complete, valid file for a fictional software company. Copy it and swap in your own pages:
# Northwind Analytics
> Northwind Analytics is a product analytics platform for B2B SaaS teams. This file lists the pages that best explain the product, pricing and API.
Northwind is used by product and growth teams to track activation and retention. Prices are in USD.
## Product
- [Product overview](https://northwind.example/product): What the platform does and who it is for
- [Pricing](https://northwind.example/pricing): Plans, limits and annual discounts
## Docs
- [Quickstart](https://northwind.example/docs/quickstart.md): Install the SDK and send a first event
- [API reference](https://northwind.example/docs/api.md): REST endpoints and authentication
## Optional
- [Changelog](https://northwind.example/changelog)
- [Company](https://northwind.example/about)
The spec also suggests serving a clean Markdown copy of each page at the same URL with .md appended, and lets a file sit at a subpath to cover only the URLs below it1.
What real llms.txt files look like: Stripe and Vercel
We fetched several live files with curl while writing this guide. Two show the range.
Vercel's llms.txt is about 5 KB and follows the spec closely, H1 and blockquote first:
# Vercel
> Vercel is a cloud platform for building, deploying, and scaling web applications and AI workloads.
Use this index to find machine-readable documentation and platform resources. Follow the linked indexes when you need individual pages.
Stripe's is a different animal. It's about 92 KB, served as text/markdown, with 455 links across 26 H2 sections, running from "Docs" and "Payment Methods" through "Terminal" to a closing "Optional". It skips the blockquote and opens with instructions aimed at coding agents. One section is titled outright for them:
## Instructions for Large Language Model Agents: Best Practices for integrating Stripe
As an LLM, you should always default to the latest version of the API and SDK unless the user specifies otherwise.
Further down, the same section tells agents to "never recommend the Charges API" and steers them to Checkout Sessions. Every link points at a .md URL, and those return Markdown too (we checked docs.stripe.com/testing.md).
That's the real use case in one file. Stripe wrote it for coding agents, and the instructions exist to keep those agents from building integrations on APIs Stripe wants retired. If your site has nothing a developer's agent would load into context, you're writing a file for an audience that isn't coming.
What is llms-full.txt?
llms-full.txt puts the full text of a site's key pages in one Markdown document instead of linking to them. It isn't in the llmstxt.org spec1. Mintlify, a documentation platform, says it first built llms-full.txt with Anthropic, which wanted a cleaner way to feed its docs to models, then rolled it out to all customers3.
Howard's original post described something related: a tool that expands llms.txt into context files called llms-ctx.txt and llms-ctx-full.txt2. Full-text files get big fast. Anthropic's platform.claude.com/llms-full.txt was about 43 MB when we fetched it, and Cloudflare's developer docs version about 68 MB, far more than fits in a single model context. At that size they're corpora for tools to search, not something a model reads front to back.
Who has adopted llms.txt?
Developer-facing companies, mostly. When we checked, Markdown llms.txt files were live at Anthropic's developer docs, OpenAI's developer site, Perplexity's API docs, Stripe's docs, Cloudflare's developer docs and Vercel. Even the proposal's own site publishes one: llmstxt.org/llms.txt is three links and well under 1 KB.
Across the wider web, the numbers depend on the sample. SE Ranking found the file on 10.13% of roughly 300,000 domains in its November 2025 study6. Ahrefs found it on 28% of 137,210 domains in its Web Analytics data in May 2026, and notes its customers skew technical, so treat 28% as an upper bound5.
Among big brands the split by industry is stark. Our study of 261 major websites' robots.txt and llms.txt files found a valid llms.txt on 81 of the 241 sites that gave a clear answer (34%). B2B SaaS led at 59 of 92 (64%), including Stripe, HubSpot, Atlassian and Shopify. Only 3 of 40 news sites had one, and no government, search or social site did. Six more SaaS sites, Salesforce and Notion among them, serve a /llms.txt that doesn't open with an H1, which fails the spec's one requirement.
Note what's missing from the adopters' statements. OpenAI, Anthropic and Perplexity publish llms.txt for their own developer docs. None of them has said its search product reads other sites' llms.txt files to decide what to cite.
Who actually reads llms.txt?
Coding agents and SEO tools, mostly. AI search crawlers rarely fetch the file, and when they do there's no sign it changes what gets cited. Google's statements, Ahrefs' server logs and SE Ranking's citation data point the same way.
Start with Google. In April 2025 John Mueller wrote on Reddit that, as far as he knew, none of the AI services had said they use llms.txt, and compared it to the keywords meta tag: a site owner's claim about what the site is about7. Search Engine Land reported that Gary Illyes said at Google's Search Central Deep Dive in Asia Pacific that Google doesn't support llms.txt and isn't planning to8. Google's AI optimization guide now says you don't need machine-readable files, AI text files or Markdown to appear in Search, that "Google Search itself doesn't use them," and that creating llms.txt "will neither harm nor help" Search visibility4.
The logs say the same. Ahrefs analyzed requests to llms.txt across 137,210 domains in May 2026. Of the roughly 38,000 sites with a file, 97% got no requests at all. Named AI tools made up 19.5% of fetches, with GPTBot the most frequent at 4.51%, and AI retrieval bots (the kind that fetch pages for live answers) made up just 1.1%5. About 12% of requests came from GEO tools, checkers and researchers studying the file itself5. So the file gets roughly ten times more visits from tools auditing it than from the bots that fetch pages for AI answers.
And the citation data. SE Ranking tested about 300,000 domains with correlation tests and an XGBoost model and found no relationship between llms.txt and AI citation frequency. The model got slightly more accurate when llms.txt was removed as a feature6. Its conclusion: llms.txt "doesn't seem to directly impact AI citation frequency. At least not yet"6.
Some systems do read it. Coding assistants and agents that a developer points at a docs site can use it, which is why Stripe writes instructions into it. The engines that produce AI search answers pick sources from their own indexes, as covered in how AI engines choose which sources to cite.
Should you add an llms.txt file?
Yes if developers or their agents use your docs; otherwise it's optional housekeeping. Three questions settle it:
- Do developers paste your docs URLs into ChatGPT, Claude, Cursor or a similar tool? If yes, add one and keep it maintained. Include an instructions section if agents keep getting your API wrong.
- Have you confirmed OAI-SearchBot, Claude-SearchBot and PerplexityBot can reach your site? If not, fix that first using our AI crawler robots.txt list. A blocked crawler costs you citations today; a missing llms.txt costs you nothing measurable.
- Do you have ten minutes and a stable set of key pages? Then adding one is harmless, per Google4. Just don't count it as AI visibility work. The things that do count are covered in our guide to generative engine optimization.
Keep it honest whatever you decide. A file that describes pages differently from what they say is exactly the kind of self-reported claim Mueller compared to the keywords meta tag. If you want to know whether any of this changes how AI engines describe you, test the prompts your buyers ask before and after, for example with the free Arobis AI visibility checker.
A 10-minute llms.txt setup, with pass/fail checks
This works for a typical marketing or product site. Big docs sites should generate the file from their docs platform instead (Mintlify, for one, generates it for its customers3), or start from our llms.txt generator.
- Check what's already there (1 minute). Run
curl -sI https://yourdomain.com/llms.txt. Pass: a 404, or a file you meant to publish. Fail: a 200 that returns your HTML homepage or a soft-404 page; fix that routing first, since tools will read the HTML as your llms.txt. - Pick the pages (3 minutes). Choose 10 to 20 URLs: product overview, pricing, the three to five pages that answer your buyers' biggest questions, docs quickstart, comparison pages. Pass: every URL returns 200 without a redirect. Fail: any URL redirects or is noindexed.
- Write the header (1 minute). H1 with your exact brand name, then a one-sentence blockquote saying what you sell and to whom. Pass: a stranger could repeat it accurately. Fail: it contains words like "leading" or "best-in-class."
- Write one line per link (3 minutes). Format
- [Title](URL): what the page answers. Pass: each note states a fact from the page ("Plans from $49 a month, annual billing"). Fail: notes repeat the title or make claims the page doesn't support. - Move the extras to Optional (30 seconds). Changelog, careers, press, legal. Pass: the main sections fit on one screen.
- Publish at the root (1 minute). Save as UTF-8 plain text at
/llms.txton the canonical host. Pass:curl -sIshows 200 and atext/plainortext/markdowncontent type, and the www and non-www versions resolve to the same file. Fail:text/html, or a redirect chain. - Check robots.txt (30 seconds). Pass: no rule disallows
/llms.txtfor the agents you want reading it. - Set a review trigger. Pass: someone owns the file and updates it when pricing or core pages change. A stale price in llms.txt is worse than no file.
Thirty days later, grep your access logs for /llms.txt. If Ahrefs' numbers hold for you, you'll most likely see a handful of SEO tools and maybe GPTBot, which is useful to know before anyone asks for budget to "optimize" it.
Frequently asked questions
What is llms.txt?
llms.txt is a proposed Markdown file published at a site's root, at /llms.txt, that gives large language models a short summary of the site and a curated list of its most useful pages with descriptions. Jeremy Howard of Answer.AI proposed it on 3 September 2024. The only required element is an H1 with the site or project name.
Does llms.txt improve AI search visibility?
There is no evidence that it does. SE Ranking's study of about 300,000 domains found no relationship between having llms.txt and AI citation frequency. Ahrefs found 97% of llms.txt files received no requests in May 2026. Google says Google Search does not use the file, and that creating one will neither help nor harm Search visibility.
Does ChatGPT or Google read llms.txt?
Neither has said it uses llms.txt to choose sources. Google's AI optimization guide says Google Search does not use files like it, and OpenAI has announced no support. In Ahrefs' May 2026 log data, GPTBot was the most frequent named AI fetcher of llms.txt at 4.51% of requests, but a crawler fetching a file shows nothing about whether it shapes answers.
What is the difference between llms.txt and llms-full.txt?
llms.txt is a short index: a summary plus links to key pages. llms-full.txt puts the full text of those pages in one Markdown file. The llmstxt.org spec defines llms.txt but not llms-full.txt; Mintlify says it developed llms-full.txt with Anthropic. On big documentation sites the full file can run to tens of megabytes, far more than a model can read at once.
Is llms.txt the same as robots.txt?
No. robots.txt is an IETF standard, RFC 9309, that tells crawlers which URLs they may fetch, and AI companies document how their bots follow it. llms.txt is an unofficial proposal that recommends what to read and grants or blocks nothing. If you want to allow or block AI crawlers, that still happens in robots.txt.
How long does it take to set up llms.txt?
About ten minutes for a typical marketing or product site: pick 10 to 20 key URLs, write a one-sentence summary and one line per link, save it as a plain text file at your root, and check that it returns a 200 status with a text content type. Large documentation sites usually generate it from their docs platform instead of writing it by hand.
Sources
- llmstxt.org, The /llms.txt file (specification)
- Jeremy Howard, Answer.AI, /llms.txt: a proposal to provide information to help LLMs use websites (3 September 2024)
- Mintlify, The value of llms.txt: hype or real? (9 May 2025)
- Google Search Central, AI optimization guide
- Ahrefs, We analyzed 137K sites: 97% of llms.txt files never get read (15 June 2026)
- SE Ranking, llms.txt study (covered by Search Engine Journal, 20 November 2025)
- Search Engine Journal, Google says LLMs.txt comparable to keywords meta tag (17 April 2025)
- Search Engine Land, Google says normal SEO works for ranking in AI Overviews and LLMS.txt won't be used
- IETF, RFC 9309: Robots Exclusion Protocol
AI Ranked Editorial. "llms.txt: what it is and does it help AI visibility?" AI Ranked, October 11, 2026. https://airanked.ai/guides/llms-txt