Large language models increasingly read websites to answer questions, but a typical page is full of navigation menus, scripts, cookie banners, and footers. An llms.txt file is a short Markdown file that gives those models a clean map of your site: what it is, and which pages matter most. This guide covers what the file is, where it goes, how to format it, what to avoid, and how it differs from robots.txt.
What llms.txt is
llms.txt is a proposed convention, described at llmstxt.org since 2024. It is not an official standard from a standards body, and no AI provider has promised to read it in every case. Treat it as low-cost documentation for tools that choose to use it. Its main appeal is that it is easy to write and easy to maintain.
Where it goes
Place the file at the root of your domain, so it is available at https://example.com/llms.txt. A file in a subdirectory such as /docs/llms.txt will not be found by convention. Serve it as plain text with a 200 status code, and make sure it is not behind a login or a redirect chain.
The Markdown format
The format is deliberately simple. It has four parts, in this order:
- An H1 with the name of the site or project. This is the only required element.
- A blockquote summary. One or two sentences that say what the site is and who it is for.
- Optional paragraphs of context. Plain text that explains anything the links cannot, such as scope or how the documentation is organized.
- H2 sections containing link lists. Each entry is a Markdown link, a colon, and a short description.
The skeleton looks like this:
# Site Name
> One or two sentences that say what the site is and who it is for.
Optional context paragraph.
## Section name
- [Page title](https://example.com/page): Short description of what the page covers.
A complete example
Here is a file for a fictional note-taking app called Quillpad:
# Quillpad
> Quillpad is a note-taking app for teams that stores notes as plain Markdown and syncs them across devices. Setup guides, API docs, and pricing are linked below.
Quillpad saves every note as a standard .md file. The API is versioned, and the current version is v2.
## Documentation
- [Getting started](https://example.com/docs/getting-started): Install Quillpad and create your first workspace.
- [Markdown syntax](https://example.com/docs/markdown): Supported formatting, including tables and task lists.
- [API reference](https://example.com/docs/api): Endpoints, authentication, and rate limits for v2.
## Product
- [Pricing](https://example.com/pricing): Plans, seat limits, and what each tier includes.
- [Security](https://example.com/security): Encryption, data residency, and audit reports.
## Optional
- [Changelog](https://example.com/changelog): Release notes by date.
- [Careers](https://example.com/careers): Open roles.
The Optional section has a specific meaning in the convention. It marks secondary content that a tool can skip when it has limited context space.
Common mistakes
- A missing or wrong H1. The H1 anchors the file. Without it, the file does not follow the format.
- Links with no descriptions. A bare list of URLs tells a model very little. Each link needs a sentence about what is there.
- Serving HTML instead of Markdown. A catch-all route can return your HTML error page at /llms.txt. Check the response with curl or your browser's network tab.
- Dead or redirecting links. Use final URLs, and check them whenever you restructure the site.
- Listing everything. A 2,000-line file is less useful than a focused 30-line one. Prioritize the pages you most want people and tools to find.
- Marketing copy in the summary. The blockquote should describe the site, not sell it. Phrases like "the best" or "revolutionary" add nothing.
- Letting it go stale. A file that points to removed pages is worse than no file at all.
- Blocking it. Do not disallow /llms.txt in robots.txt, and do not require authentication to read it.
How it differs from robots.txt
These two files do different jobs, and most sites benefit from having both.
robots.txt tells crawlers what they may request. It holds allow and disallow rules for each user agent and follows the Robots Exclusion Protocol (RFC 9309). It controls crawling, and compliance is voluntary, though major crawlers generally honor it.
llms.txt does not restrict access. It is a curated map that points a model to the content you consider most useful and briefly explains each part. If you want to keep a crawler away from pages, use robots.txt or real access controls. If you want to point a model toward your best documentation, write an llms.txt.
A short checklist
- Put the file at /llms.txt on your root domain.
- Start with an H1 and a one- or two-sentence blockquote.
- Use H2 sections with full URLs and one-line descriptions.
- Move secondary material under "Optional."
- Confirm that every link returns a 200 status.
- Review the file whenever you make major changes.
If you would rather not write the file by hand, the free llms.txt generator can produce a starting draft. Edit it, then check the result against the rules above before you publish.
Top comments (0)