llms.txt Explained: Does Your Site Actually Need One
Another File That's Supposed to Save Your Site?
You already know robots.txt — the plain text file in a site's root that tells crawlers where they're allowed to go and where they're not. You know sitemap.xml too — the list of every page you want indexed, so nothing gets missed. Both have been sitting quietly in site roots for decades, and nobody seriously argues about whether they're useful. They are, backed by years of practice and official documentation from Google, Bing, and everyone else.
Sometime around 2024–2025, a third file joined the conversation: llms.txt. The pitch spread fast through SEO blogs and niche newsletters — another plain text file, except this one is supposedly for language models and AI agents instead of search crawlers: ChatGPT, Perplexity, and the rest. The logic sounds reasonable on the surface: if there's a file for crawlers, why not one for AI? Plenty of people read a single blog post, put a file together in an evening from a template, dropped it at the site root, and started watching their ChatGPT and Perplexity mentions, waiting for them to climb.
A week goes by, then two. You check the server logs and the file did get fetched — a handful of times, mostly by people double-checking it didn't 404. None of those hits come from a recognizable AI-system user agent. Brand mentions in ChatGPT and Perplexity haven't budged either way in that window. Which raises the obvious question: is the file simply not doing anything, or is two weeks just too soon to tell?
A few things are feeding the hype at once: AI as a topic sells better than almost anything else right now, the format itself is dead simple — you get a visible result immediately, unlike the slow grind of fixing content structure — and there's a real appeal to being able to say you did it early. None of that is wrong exactly. It just has little to do with what the file actually delivers.
If you're in that spot — already shipped the file, or still deciding whether it's worth the hour — here's the honest short version: llms.txt by itself is unlikely to move the needle, because the major AI systems don't reliably read it yet. That doesn't make the idea worthless, though. Let's go through what it actually is, why it got this much attention, and when it's genuinely worth your time.
What llms.txt Actually Is
llms.txt is a plain markdown file you put at the site root — /llms.txt, right next to robots.txt and sitemap.xml. The idea was formalized in 2024 on llmstxt.org as an open community proposal — not a W3C standard, not part of the HTTP spec, not a requirement documented by Google or OpenAI. Just a convention: if enough sites put a file at this address in this format, language models have an easier time understanding them.
Markdown wasn't an arbitrary choice: it's dramatically cheaper in tokens than the same text wrapped in HTML tags, navigation chrome, and scripts, which matters for a model that effectively pays in tokens for every character of context it has to read.
The proposal came from Jeremy Howard, the fast.ai co-founder who later started Answer.AI. His team published the spec on llmstxt.org in September 2024 and pitched it to the industry as a voluntary convention, the way sitemap.xml eventually became one. Several documentation-hosting platforms picked it up almost immediately, auto-generating an llms.txt for every site built on top of them — so the file quietly showed up on hundreds of products without any team actually deciding to add it. Showing up automatically and actually getting used when a model answers a question, though, are two different stories.
- An H1 title — the site or product name
- A short one- or two-sentence blockquote describing what the site does
- A handful of H2 sections with bullet-point links to the pages that matter most — docs, API reference, pricing, blog
- A one-line, plain-language note on each link, not a page's SEO title
Picture the same request going to an agent two different ways: "open this SaaS product's site and find out what the 10-seat team plan costs." Without llms.txt, the agent lands on the homepage, tries to make sense of the nav, and guesses at the likely section — Pricing, For Teams, Plans — and can easily miss if the menu is unconventional or the price sits behind a contact-sales form. With llms.txt, in theory it's simpler: the agent sees a Product section pointing straight at /pricing with a one-line description. In practice, the gap comes down to whether that particular agent reads the file at all — and today the honest answer is still "not consistently."
The file isn't written for a human visitor or a classic search crawler — it's written for an AI agent or model parsing a site in the moment, say, when a user asks an assistant to 'open this product's docs and find out how to configure X.' Instead of the agent crawling the nav menu and parsing rendered HTML, it gets a compact map up front: here's the documentation, here's the API, here's the blog, here's what to skip. It's the same logic as a course syllabus saving you from reading the whole textbook — just aimed at a model instead of a student.
Is It a Standard? No — and That Matters
Here's where honesty matters, because there's a lot of marketing noise around llms.txt. robots.txt and sitemap.xml have something llms.txt doesn't yet: broad, documented support from the systems they're written for. Googlebot has official documentation on how it reads robots.txt. Google Search Console officially ingests sitemap.xml. The major AI crawlers and agents — from OpenAI, Anthropic, Perplexity, Google — haven't made an equivalent public commitment to read and act on llms.txt when answering users.
That's not to say AI companies ignore root-level text files altogether — it's just that robots.txt is usually the one actually in play, not llms.txt. Every major provider runs its own named crawler: GPTBot for OpenAI, PerplexityBot for Perplexity, Google-Extended for Google, ClaudeBot for Anthropic. All of them publicly commit to honoring Disallow directives addressed to them in robots.txt, which means a site can already restrict what these crawlers index or use for model training today. That mechanism works precisely because it's built on a protocol with decades of track record behind it, not something invented from scratch.
In practice, that means most answers from ChatGPT, Perplexity, Copilot, Gemini, and Google AI today still lean on far more conventional signals: whether a crawler could technically reach and parse the page, structured data present on it, the clarity and structure of the actual text, domain authority — broadly the same set of signals behind ordinary AEO/GEO work (see what AEO is and GEO explained). Whether llms.txt exists is, at best, a minor factor in that mix right now.
None of that makes llms.txt worthless — it just means your expectations need calibrating. Even if the big providers aren't reading the file directly, curating a shortlist of your most important pages is a useful exercise in its own right: it forces you to actually decide what matters on the site, and that same clarity pays off in structured data work and in ordinary site navigation too.
| Aspect | robots.txt | sitemap.xml | llms.txt |
|---|---|---|---|
| Purpose | Allows or blocks crawler access to pages | Lists URLs for search engine indexing | Summarizes a site and links to its key pages for AI systems |
| Audience | Search and other web crawlers | Search engines (Google, Bing, etc.) | Language models and AI agents — in theory |
| Format | Plain text, User-agent/Disallow directives | XML with URLs and metadata | Markdown: title, short description, link lists |
| Who honors it | Officially supported by major search crawlers | Officially ingested by Search Console and equivalents | Not formally committed to by any major AI provider |
| Status | De facto industry standard since 1994 | Officially supported protocol, adopted by Google | Community-proposed convention (llmstxt.org, 2024) |
We covered this in more depth when llms.txt's second version shipped — see our note on llms.txt v2 adoption: no sign yet of major providers switching over en masse. That could change; conventions do catch on over a year or two, the way sitemap.xml eventually did. But building a strategy around a file that isn't consistently read today is a bad bet.
What an llms.txt File Looks Like in Practice
Easier to show than describe in paragraphs. Here's a minimal but realistic llms.txt for a hypothetical SaaS product (the example uses a fictional Example Brand on mysite.com — don't confuse it with a real site):
# Example Brand > Example Brand is a SaaS platform for monitoring SEO and brand visibility in AI system answers. It tracks rankings, mentions, and citations across ChatGPT, Perplexity, Gemini, and other systems. ## Docs - [Quickstart](https://mysite.com/docs/quickstart): connect your first project in 5 minutes - [API reference](https://mysite.com/docs/api): endpoints, auth, rate limits ## Product - [Pricing](https://mysite.com/pricing): current plans and limits - [Blog](https://mysite.com/articles): articles on AEO, GEO, and AI search visibility ## Optional - [Status page](https://mysite.com/status): uptime and incident history
Notice three things. First, it doesn't try to list every page on the site — it's not a sitemap.xml replacement, it's a curated shortlist, realistically 5 to 15 links. Second, every link gets a short plain-language note — not an SEO title, an explanation of why you'd click it. Third, there's no markup, no scripts, no inline styling — just clean markdown a model can consume without parsing a page's HTML.
The format has an optional companion: llms-full.txt. Same list of sections, except instead of links it bundles the full text of every page into one file, so an agent can grab everything it needs in a single request instead of following links one by one. In practice, llms-full.txt balloons into hundreds of kilobytes for any site with real depth, and nobody's obligated to read all of it either — models still work within a bounded context per request, so an oversized file just gets truncated somewhere in the middle.
Where llms.txt Sits in the Bigger SEO Picture
To avoid overrating the file, it helps to see where it actually sits. llms.txt is a thin, optional layer on top of what actually determines whether a site shows up in an AI answer.
The lower a layer sits on this stack, the more it actually affects whether ChatGPT or Perplexity cites you. llms.txt sits near the top, next to the technical scaffolding, not the substance. If you want a real read on how readable and well-structured your pages are for AI, there's a dedicated audit for that — AI Readiness parses a page roughly the way a model extracting information would, and points out concrete problems instead of guessing.
Here's the analogy worth keeping in mind: fixing structural content problems with one optional root-level file is like polishing the front door handle on a house with a leaking roof. It looks tidy. It doesn't fix what actually matters.
Do You Actually Need One
The answer splits into two scenarios, and it matters which one you're actually in before you spend time on the file.
If your technical foundation is already solid — pages crawl cleanly, render human-readable HTML without a hard JavaScript dependency, structured data is in place (FAQPage, Article, Product — whatever applies), and the content itself is written to clearly answer real questions — then adding llms.txt is fine. It's cheap: one markdown file, half an hour of work, essentially zero risk of breaking anything. It won't hurt, and it might marginally help with the few agents that do read it today. But it's a low-priority nice-to-have, not a lever that moves AI visibility. That checklist is quick to run in practice: structured data can be checked with Google's Rich Results Test, and JavaScript dependency by just disabling JavaScript in your browser's devtools and seeing what's left on the page.
If you've got real problems instead — half your content only renders after JavaScript executes, there's no structured data anywhere, sitemap.xml hasn't been touched in a year, and your pages are written for keywords rather than for answering a question — llms.txt won't compensate for any of that. It doesn't hack rankings and it doesn't substitute for doing the content work. It's like putting up nice directional signs outside a store with half-empty shelves: the signs won't make a shopper find a product that isn't there. Fix that first, using something like our AEO checklist and the piece on how AI engines choose sources, then come back to llms.txt — if it's even still worth it by then.
More often it's a mixed bag: part of the site is fine, part isn't. Marketing landing pages get server-rendered and read cleanly, while the docs or the blog run on an SPA framework that serves an empty div until JavaScript executes. llms.txt doesn't save you here either — it can point to the docs page, but if the actual content behind that link stays invisible without JS, the agent hits the exact same wall it would have hit without the file.
Figuring out which scenario you're in doesn't take much. Pull up a key page with curl -A "Mozilla/5.0 (compatible; GPTBot/1.0)" https://your-site/page, or use your browser's "view page source" without letting JavaScript run. If what comes back is empty, or just an application shell with no actual text, crawlers and AI agents are almost certainly hitting the same wall you just did. That's the first thing worth checking, well before llms.txt even enters the conversation.
You don't have to guess. Track whether ChatGPT, Perplexity, Copilot, Gemini, and Google AI actually mention and cite you on real prompts, and compare the trend before and after changes. AI Control keeps a history of visibility, Share of Voice, and cited links across your tracking profiles for exactly this.
Questions People Actually Ask About llms.txt
A few quick answers to what people usually ask.
Is llms.txt a replacement for sitemap.xml?
No. sitemap.xml is an exhaustive list of URLs for search engine indexing, while llms.txt is a short, curated list of a handful of key pages aimed at language models. Use them together, not as substitutes for each other.
Can llms.txt hurt a site?
Just having the file causes no harm by itself. The real risks show up if it accidentally links to private or internal pages, or if it goes stale for years and starts pointing at removed sections — at that point it's just useless, not harmful.
Do ChatGPT, Perplexity, and the rest actually read it?
There's no public commitment from any major provider confirming consistent support, and observed behavior is uneven and inconsistent. It shouldn't be treated as a primary channel for AI search visibility.
Where should I start if I have neither llms.txt nor solid technical SEO?
Start with the technical basics: crawlability, structured data, clearly structured content that answers real questions. Add llms.txt at the end, once everything else is in order — it's a half-hour task.
Does llms.txt need updating as often as the site's content?
Yes, if you decide to keep one. A stale file pointing at removed pages or old sections is worse than having none at all — it hands out a wrong map instead of a useful one. Updating it by hand whenever the site's structure changes meaningfully is enough for most projects.
Is llms.txt worth doing for a small brochure site?
Probably not. The format is built for sites with a meaningful amount of structured content — documentation, a blog, an API reference — where an agent actually has something worth curating. For a 5-to-10-page site, llms.txt barely changes anything over what an AI agent already gets from normal navigation.
Checklist: Should You Bother With llms.txt
If you want the short version, here's what to check before spending time on the file — or before deciding you don't need it. None of it needs special tooling, just an honest read on where the site stands right now.
- Pages crawl cleanly and serve content without a hard dependency on JavaScript execution — if not, start there, not with llms.txt
- Key pages carry structured data (FAQPage, Article, and similar) — a signal both search engines and some AI systems actually use
- sitemap.xml is current and free of broken or stale links
- Page content is written to directly answer the reader's question, not just to hit keywords
- If all of that is in order, add llms.txt as a cheap bonus: 5 to 15 links to docs, pricing, the blog, and other key sections, each with a short description
- Don't expect a surge in citations from the file alone — it's hygiene, not a growth lever
- Track actual brand mentions and citations in AI answers so you know what's really moving the needle, instead of betting on one file
Want to check this in your market?
AI Control regularly collects AI responses, brand positions, competitors and cited sources for your prompt library.
Is your own content set up for stories like these?
Use 50 welcome credits for an AI Readiness check — retrieval, extractability, schema.org signals, and a prioritized rewrite brief, scored the way an AI assistant actually reads your page.