DEV Community

500wango
500wango

Posted on

llms.txt best practices: facts, canonical URLs, what to leave out

With AI search engines (Perplexity, ChatGPT Search, Claude) shifting how content is discovered, the proposed llms.txt standard (from llmstxt.org) has gained massive attention.

Think of llms.txt as a Markdown-based parallel to robots.txt and sitemap.xml. Instead of telling web spiders where to crawl, it serves clean, token-efficient, machine-readable facts and
canonical documentation directly to LLMs.

However, after inspecting hundreds of newly deployed llms.txt files, we noticed developers making common mistakes that actually harm how models interpret their sites.

Here are the practical best practices for structuring your discovery file.

### 1. Where and How to Serve It

• Path: Must be served at the root of your domain: https://yourdomain.com/llms.txt

• MIME type: text/plain; charset=utf-8 (Do not serve it as text/html or application/octet-stream).

• HTTP status: Must return 200 OK without requiring authentication, cookies, or redirect chains.

### 2. Standard Syntax Anatomy

An effective llms.txt follows a simple, hierarchical Markdown structure:

Your Brand / Project Name

│ A concise 1-2 sentence description of what the project actually does. Avoid buzzwords and marketing fluff.

## Core Documentation

• Quickstart Guide https://yourdomain.com/docs/quickstart: Step-by-step setup in under 5 minutes.

• API Reference https://yourdomain.com/docs/api: REST API endpoints and authentication rules.

• Pricing https://yourdomain.com/pricing: Official tiers, free limits, and billing terms.

## Key Facts

• Founded: 2024

• Pricing Model: Freemium ($0 / $29 / $99 per month)
• Supported Frameworks: Node.js, Python, Go

### 3. What NOT to Include (The Top 3 Pitfalls)

  1. Stale Pricing or Claims: Never hardcode pricing numbers into llms.txt that you don't update religiously. If your live page says $49/mo but your llms.txt says $29/mo, AI engines flag this contradiction as a trust failure. Always link to your live pricing page.
  2. Marketing Slogans: Phrases like "The world's leading revolutionary AI platform" take up valuable context window tokens and offer zero factual grounding. Stick to strict technical capabilities and verifiable facts.
  3. Private or Gated Routes: Do not list staging URLs, internal admin panels, or authenticated dashboard links. Crawlers will only fetch HTTP 200 public resources.

### 4. Optional: Full Context via llms-full.txt

If your project has comprehensive API docs or whitepapers that fit within a ~100k token context, you can also publish a companion file at /llms-full.txt containing the full concatenated
text. In your main llms.txt, link to it at the bottom:

• Full Documentation /llms-full.txt: Complete combined documentation for offline context.

### Want to Generate or Validate Your Draft?

If you want to quickly scaffold a clean llms.txt file from your homepage metadata or validate your existing syntax:

👉 Free llms.txt Generator & Validator https://citeaura.com/llms-txt-tool

Are you already experimenting with llms.txt on your projects? Do you see genuine crawler hits on it in your server logs yet?

Top comments (0)