SEO and AI visibility 6 min
llms.txt: what it is really for
The llms.txt file is a 2024 proposal: a table of contents at the root of a site that tells an AI agent which pages to read. No major engine says it uses it to decide whom to cite, and Google writes that no special file is needed for its AI features. It helps agents reading a site on demand.
What the proposal says
The llms.txt file is described on llmstxt.org, the proposal’s own site, published by Jeremy Howard on 3 September 2024 and revised on 10 August 2026. The text calls itself “a proposal to standardise on using an /llms.txt file”. The word matters: this is not a standard from a standards body, it is a convention anyone is free to adopt.
The format is deliberately simple. A Markdown file at the root of the site, containing:
- a level 1 heading, the only required element, with the name of the site or project;
- a short summary in a blockquote;
- optionally a few paragraphs of explanation;
- titled sections listing links, each with an optional note. A section named “Optional” groups, by convention, what an agent can skip when it is short on space.
The proposal also recommends serving a clean Markdown version of useful pages, at the same address with .md appended. The underlying idea is sound: a modern HTML page is full of menus, scripts and banners, and a model that must pull out the essentials within a limited context window benefits from receiving text that is already clean.
The site is careful to separate the file from its neighbours. robots.txt says who may read; sitemap.xml lists every page for indexing; llms.txt picks what an agent should read to understand the site, on demand.
What can be verified about real usage
This is where enthusiastic articles get carried away. Here is what we could verify ourselves, with the sources opened.
On the publishing side, adoption is real. Version 2 of the proposal states that thousands of sites publish an llms.txt and that documentation platforms generate one automatically. We did not count those thousands of sites, but one example is easy to check: OpenAI’s documentation on its crawlers itself points to its /llms.txt as the index of its documentation, and offers Markdown versions of its pages by appending .md to the URL.
On the engine side, nothing verifiable. What a company publishes for its own documentation says nothing about what its crawlers read elsewhere. The official pages we opened all point the same way:
- The same OpenAI documentation describes its crawlers (OAI-SearchBot for ChatGPT search, GPTBot for its models) in terms of robots.txt, never with llms.txt as a criterion.
- Anthropic’s help page on its crawlers does not mention llms.txt.
- Google is the most explicit: its documentation on AI features states that you do not need to create new machine readable files or AI text files to appear in AI Overviews or AI Mode.
The proposal itself stays careful: it says the file was expected to be useful mainly when a model answers (inference) rather than for training, and it never claims to influence ranking or the choice of cited sources.
Our conclusion, based on what can be verified as of 2 October 2026: no primary source allows anyone to claim that a major engine uses llms.txt to decide whom to cite. Those who promise it are extrapolating. If a vendor ever announces it in writing, this article will be updated.
Where it actually helps
The file does have a concrete use, more modest than the one it is credited with: agents that read a site on demand. A coding assistant that needs to learn a library, an agent given a site’s address and asked to answer from its content, an internal tool querying a documentation set. In those cases someone has already decided to read your site; llms.txt saves them from guessing where to start.
What we did
Challenge Us: a table of contents and a corpus
Challenge Us is our group challenge app. Its site publishes two complementary files, both written at build time from the same lists as the pages:
/llms.txt, the table of contents: a summary of the app, then help, blog, challenge pages and legal pages, each with its link./aide/index.json, the knowledge base corpus: for each of the 45 pages, the question, the short answer, the plain text, the section, the canonical URL, the date it was checked and the app version in which the answer was verified.
The first says where to go, the second says what to read. Because they are produced from the same list as the published pages, they cannot drift away from the site: that is the only condition under which we accept one more file. A hand-maintained llms.txt always ends up lying about what exists.
We also decided what we would not do: no vector database, no chunking for a semantic search engine. At a few dozen pages, the whole corpus fits in a model’s context window. Building the infrastructure before having the content leaves you with the infrastructure and no content.
Palmora: a gap found in the audit, then filled
On Palmora, our six-language real estate agency in Phuket, the SEO and GEO audit of 22 July 2026 noted the absence of an llms.txt (the request returned a 404). We ranked it low priority, estimated at two or three hours together with a blog RSS feed, far behind the real problem of the day: a robots.txt that blocked AI crawlers without our knowledge, which we describe in the article on how to get cited by ChatGPT. The file exists today, generated on each request: agency overview, main areas, guides by language and nationality, property portals.
And this site?
Honestly: brainplusai.com does not have one yet. That is not a pose, it is the order of priorities we recommend: structure first, table of contents second. The blog and offers are still being structured, and a table of contents published before the structure is stable would have to be redone.
Should you publish one?
Yes, if the following three conditions are met, and without expecting anything in terms of citations:
- Everything else is in order. AI crawlers allowed to read the site, pages that answer a question at the top, sourced facts, structured data faithful to what is displayed. An llms.txt makes up for none of these gaps. We go through these foundations in what an SEO agency should deliver in the age of AI.
- It is generated, not handwritten. Produced at build time from your content collections or catalogue, it stays accurate without effort. Written by hand, it goes stale as soon as the second page is added.
- It points to pages that answer. A table of contents of hollow pages is useless. On a multilingual site, it points to pages that exist in the announced language: hreflang mistakes call for the same care.
To find out whether it is read, look at your server logs: who requests /llms.txt, and how often. That is the only honest measurement available today.
Checking these foundations is part of our SEO audit, and putting them in place over time is the job of our SEO and GEO engine.
Frequently asked questions
Does llms.txt replace robots.txt? No. robots.txt says who may read your site, and the major vendors’ crawlers honour it. llms.txt neither blocks nor allows anything: it suggests what to read to someone who has already decided to read. If your robots.txt blocks AI crawlers, no llms.txt will change that.
Do I also need an llms-full.txt? Some sites publish a version containing all the text. That is useful when the corpus is small and structured, like documentation or a help base. For Challenge Us we preferred structured JSON, easier for a program to use than one long text file.
How long does it take to create one? For a site whose content lives in collections or a database, a few hours are enough to generate it automatically; that was our estimate for Palmora. The real cost is keeping it accurate, hence the point of generating it every time the site is built rather than writing it once.
Jim · founder of Brain Plus AI
Also in seo and ai visibility
- 7min
Practical guide
How to get cited by ChatGPT and AI search engines
What actually gets a site cited by ChatGPT, Claude or Perplexity: let the bots read it, answer straight away, source every fact, mark up only what is shown. And how to measure it without fooling yourself.
- 5min
Choose and priceThe topic guide
AI SEO agency: what it should actually deliver
"AI" has become every SEO agency's sales pitch. Here is what an AI SEO agency should concretely deliver, the questions that expose the others, and why human review remains non-negotiable.
- 6min
Practical guide
Hreflang and multilingual SEO: the costly mistakes
Hreflang pointing to pages that do not exist, a poorly chosen x-default, 404s in the wrong language, untranslated titles, a language routed but forgotten elsewhere: the mistakes we found on our own multilingual sites, and how to avoid them.