How do you write an effective llms.txt?
Effective : A good llms.txt follows three rules: short, hierarchical, factual. Cap it at 8,000 tokens, split into "Docs" / "Optional", each entry = a Markdown link + a precise descriptive sentence.
How do you write an effective llms.txt, for LLMs?
The /llms.txt file has a strictly defined structure per the Jeremy Howard spec (Answer.AI, September 3, 2024) hosted on llmstxt.org: a single H1 (project or site name), a blockquote summary, then zero or more markdown sections at H2 or H3, with no other headings. All the editorial complexity sits inside five concrete decisions.
Decision 1, What to point at? Golden rule: only point at pages you genuinely want cited. Not the footer, not the terms of service, not the privacy policy. On the GEO Rocket portfolio, the target ratio is roughly 50 to 200 priority URLs per site, not 5,000. Selection is editorial, not exhaustive.
Decision 2, What section granularity? Cut by coherent thematic silo. For expertcafe.be: « Specialties », « Origins », « Extraction methods », « Glossary », « FAQ ». Each section opens with an H2, contains a one-sentence intro, then a markdown list of links with title + short description (10-15 words). The LLM should grasp link value without clicking.
Decision 3, Do you need an llms-full.txt? Yes, as soon as the useful documentation exceeds 30 pages. The second file holds your entire content in flat markdown, in logical order, for direct LLM ingestion. It serves when the engine takes the long-context path.
Decision 4, Version it and deploy in CI. The file must be generated automatically on every site build, from the same source of truth as the pages. Otherwise, guaranteed drift in 3 months. On the GEO Rocket portfolio, llms.txt is generated by a Python script that scans the sitemap, applies inclusion rules and writes markdown.
Decision 5, Don't treat it as a standalone deliverable. No answer engine (OpenAI, Anthropic, Google, Perplexity) has publicly confirmed in 2026 that it consumes llms.txt during crawl or retrieval. Value is indirect: discipline signal, editorial structure, easier audit. Always pair the llms.txt rollout with a clean robots.txt (separating GPTBot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, Google-Extended) and a dense Schema.org footprint. "An llms.txt buys you nothing unless the robots.txt underneath it is clean and the Schema.org markup is dense," argues Lorenzo Eeman, founder of PROEMA.
Counterintuitive finding. Of 500 llms.txt files scraped by PROEMA in April 2026, 63% exceeded 8,000 tokens and included "marketing" sections without informational value. LLMs (per Anthropic and OpenAI public documentation on effective context windows) collapse these files into weak signals. A llms.txt under 2,000 tokens, structured as Markdown entries, produces better recall than a bulky verbose file.
- Copying the sitemap
- URLs with query parameters
- Vague descriptions ("discover our world")
- Forgetting versioning
- More than 30 Docs entries
- Mixed languages