← Back to Blog

llms.txt Explained: Should Your Business Have One?

Key Takeaways

  • llms.txt is a proposed markdown file at your site's root, put forward by Answer.AI's Jeremy Howard in September 2024 — it is not an official web standard.
  • Google has stated on the record that it doesn't use llms.txt for Search, including AI Overviews, and updated its own documentation to say so on 15 June 2026.
  • Ahrefs studied 137,210 domains and found that of the ~38,000 with a valid llms.txt file, 97% received zero fetch requests in May 2026.
  • No major AI company — OpenAI, Anthropic, Perplexity, or Microsoft — has confirmed its crawlers actually read llms.txt for retrieval or citation.
  • For most UK small businesses, llms.txt is optional and low-priority. The fundamentals — answer-first content, named sources, off-site consistency — do far more for AI visibility.

If you've seen "llms.txt" mentioned in an SEO newsletter or a developer's blog and wondered whether your business is missing out, the short answer is: probably not. llms.txt is a plain-text file some websites publish to describe themselves to AI systems — but the AI systems it's meant for mostly aren't reading it, and Google has confirmed it plays no part in Search or AI Overviews.

That doesn't mean it's a scam or entirely pointless. It means the file has been widely misunderstood as an SEO or "AI visibility" tactic when the evidence doesn't support that framing. This guide explains what llms.txt actually is, what the data says about whether anyone uses it, and what a UK small business should genuinely prioritise instead. For the broader picture of how AI search visibility actually works, see our guide to generative engine optimisation.


What is llms.txt, exactly?

llms.txt is a markdown file published at a website's root (yoursite.co.uk/llms.txt) that summarises the site's purpose and lists links to its most important pages, written specifically to be easy for a large language model to read. It follows a fixed structure: an H1 with the site or project name, a short blockquote summary, optional descriptive paragraphs, and then H2-headed sections of markdown links — with an "Optional" section at the end for lower-priority pages that can be dropped if an AI system needs a shorter version.

The idea is that a human-readable webpage — full of navigation menus, cookie banners, and marketing copy — is inefficient for a language model to parse. llms.txt strips all of that away and hands over a clean, curated list of what matters. Some sites go further and publish a .md (markdown) version of individual pages alongside the normal HTML, for the same reason.

It sits in the same conceptual family as robots.txt (which tells crawlers what they may access) and sitemap.xml (which lists every URL on a site) — but it isn't enforced or read by the same infrastructure, and that distinction matters more than it might first appear.


Where did llms.txt come from?

The proposal came from Jeremy Howard, co-founder of Answer.AI, published on 3 September 2024 at answer.ai and mirrored at llmstxt.org. Howard's pitch was straightforward: as AI assistants increasingly answer questions by reading websites directly, site owners should have a way to present their content in a format built for that purpose, rather than leaving it to the model to scrape and guess.

It's worth being precise about what it was and wasn't. It was a proposal — a specification anyone could implement voluntarily. It was never submitted to, or ratified by, a standards body such as the IETF or W3C, and no major AI company committed to supporting it at launch. That's a meaningfully different status from robots.txt, which every major search and AI crawler recognises and (mostly) respects, or sitemap.xml, which Google Search Console explicitly asks site owners to submit.

Two years on, that gap between "proposed" and "adopted" hasn't closed — if anything, the evidence below shows it's widened.


Do AI tools like ChatGPT and Google actually use llms.txt?

No AI company has confirmed that its crawlers read llms.txt to retrieve or cite content, and Google has explicitly said it doesn't use it for Search. Gary Illyes, speaking at Google's Search Central Deep Dive event in Asia-Pacific on 23 July 2025, said Google "doesn't support llms.txt and isn't planning to," and that ordinary SEO — not a separate llms.txt file — is what drives visibility in AI Overviews (reported by Search Engine Land, July 2025).

Google made that position official on 15 June 2026, updating its Search Central documentation to state that llms.txt files aren't needed for Search — including its generative features — and that having one "won't harm (nor help)" a site's visibility, though maintaining one for other purposes is "fine" (Google Search Central, updated 15 June 2026). Search Engine Journal covered the update and noted an internal inconsistency worth knowing about: Google's own Lighthouse tool (version 13.3, 2026) added an "agentic browsing" audit that checks for an llms.txt file — a signal aimed at browser AI agents navigating a page on a user's behalf, which is a genuinely different use case from search ranking (Search Engine Journal, 2026).

For OpenAI, Anthropic, Perplexity, and Microsoft/Bing, we found no equivalent official statement either way. That absence is itself informative: if any of these companies' crawlers gave llms.txt meaningful weight in retrieval or citation, it would be a straightforward, low-cost thing for them to say publicly — and none have. Anthropic does publish llms.txt files on its own documentation sites, but that reflects Anthropic acting as a publisher of the file, not a confirmation that its models fetch and prioritise other sites' versions.


What does the data actually say about llms.txt usage?

Adoption has grown steadily since 2024, but adoption is not the same as usage — and the traffic data is the part most coverage of llms.txt leaves out.

Ahrefs analysed 137,210 domains in May 2026 and found that 28% (roughly 38,000 domains) publish a valid llms.txt file. Of those, 97% received zero fetch requests during the month studied — nothing, human or bot, ever read them. Of the small remainder that did see traffic, only about 19.5% of those requests came from named AI tools; the rest came from unrelated bots (Ahrefs, "llms.txt Study," 15 June 2026; also covered by Search Engine Journal, 2026). Separately, Search Engine Land tracked ten sites with llms.txt files over a 180-day window and found eight showed no measurable change in AI-driven traffic; the two that did gain traffic attributed it to PR coverage and content restructuring, not the file itself (Search Engine Land, 2026).

The chart below shows just how lopsided that is.

Does anyone actually read llms.txt files? ~38,000 domains with a valid file, May 2026 0% 50% 100% Received zero requests 97% Received any request 3%
Source: Ahrefs, "llms.txt Study," 15 June 2026 (137,210 domains analysed; 38,000 with a valid file).

Reported adoption rates also vary a lot depending on who's sampling, which is worth knowing before you treat any single number as "the" adoption rate: SE Ranking found 10.13% adoption across a broader domain sample (SE Ranking, November 2025), while Rankability found 8.7% adoption among the global top 1,000 sites (Rankability, June 2026). Ahrefs' higher 28% figure likely reflects its user base skewing towards technically-minded website owners, docs platforms, and SEO practitioners who are more likely to have heard of llms.txt in the first place.

Reported llms.txt adoption varies by sample 0% 15% 30% Ahrefs (137,210 domains) 28% SE Ranking (broad sample) 10.13% Rankability (global top 1,000) 8.7%
Sources: Ahrefs (15 Jun 2026), SE Ranking (Nov 2025), Rankability (Jun 2026). Different samples, different results — treat any single adoption figure with caution.

Platforms such as Mintlify, GitBook, and Wix now auto-generate llms.txt for hosted documentation sites, which explains a chunk of the adoption growth without implying any corresponding uplift in AI traffic (Mintlify, 2025–2026).


How is llms.txt different from robots.txt and sitemap.xml?

The names sound similar, but the three files do genuinely different jobs — and only two of them are read by the crawlers your business actually depends on.

FilePurposeRead by Google Search?Read by AI crawlers?
robots.txtTells crawlers which parts of a site they may or may not accessYes — foundational, always respected by compliant crawlersMostly — though compliance is inconsistent (see below)
sitemap.xmlLists every indexable URL on a site, to help discovery and crawlingYes — Google Search Console explicitly requests itIndirectly, if a crawler follows normal links
llms.txtA curated, human-authored summary of a site written for LLMsNo — Google confirmed this on 15 June 2026Unconfirmed for any major AI company

That "compliance is inconsistent" caveat on robots.txt is worth flagging, because it undercuts the idea that any text file is a reliable control mechanism. Cloudflare found in August 2025 that Perplexity was using undeclared, browser-spoofing crawlers to access tens of thousands of domains that had explicitly blocked it via robots.txt — leading Cloudflare to strip Perplexity of its verified-bot status (Cloudflare, via Search Engine Journal, 4 August 2025). If a well-established standard like robots.txt isn't reliably honoured by every AI crawler, it's a reasonable inference that a two-year-old, voluntary proposal with no enforcement mechanism fares no better.

For the crawlers that do respect it, robots.txt directives (allowing or blocking GPTBot, ClaudeBot, Google-Extended, PerplexityBot, and similar) remain the actual lever a business has over AI crawling — not llms.txt.


Should your small business have an llms.txt file?

For the vast majority of UK small businesses — local service providers, e-commerce shops, agencies, trades — no, not as a priority, and it's fine to have never heard of it until now. The evidence is fairly clear-cut: Google has said on the record that it doesn't affect Search or AI Overviews, no other major AI company has confirmed using it, and the traffic data shows almost nothing is fetching it in practice. Time spent building and maintaining an llms.txt file is time not spent on things with demonstrated impact.

There are two situations where it's more reasonable to consider:

You run a developer-facing product, API, or documentation site. AI coding assistants and agentic tools increasingly navigate technical documentation directly, and Google's own Lighthouse "agentic browsing" audit checks for llms.txt on exactly this basis. If your audience is developers using AI tools to integrate with your product, an llms.txt file is low-cost and plausibly useful — separate from any Search or GEO benefit.

You already have a documentation platform that generates one for free. If you're on Mintlify, GitBook, or a similar platform that creates llms.txt automatically, there's no reason to remove it. It's genuinely harmless — Google's own wording — and costs you nothing to leave in place.

Outside those two cases, adding llms.txt to a typical small-business website is a low-value use of time, not because it's harmful, but because nothing is currently reading it.


What should you prioritise instead of llms.txt?

If you want your business to show up in AI-generated answers — ChatGPT, Google AI Overviews, Perplexity — the evidence points firmly at content structure and off-site presence, not a technical file. We cover this in detail in what generative engine optimisation actually is and how GEO and SEO compare in 2026, but the short version is:

  • Answer-first content structure. Open each section with a direct answer to its heading's implicit question, in the first sentence. AI systems extract from the start of a passage, not the middle. Our guide to writing answer-first content for AI search covers the mechanics in detail.
  • Named, dated sources for every claim. "Research shows" is invisible to an AI system checking credibility. "Ahrefs, June 2026" is not.
  • Off-site consistency. Your Google Business Profile, directory listings, and social presence carry real weight in AI citation decisions — considerably more than any file you can add to your own server.
  • Ordinary technical SEO. Fast pages, clean crawlability, a working sitemap.xml, and sensible robots.txt rules for AI crawlers you do or don't want visiting. These are the fundamentals every AI system's guidance — including Google's own — keeps pointing back to.

For a full walkthrough of how to structure a site for AI visibility rather than chasing individual files, see our complete guide to optimising a website for AI search. And if you'd rather have this assessed and implemented properly than DIY it from blog posts, that's exactly the kind of practical, ship-it work covered by our AI consultancy for small businesses.


Frequently Asked Questions

Does Google use llms.txt to rank my site or show it in AI Overviews?

No. Google updated its Search Central documentation on 15 June 2026 to state explicitly that llms.txt isn't needed for Search, including its generative features, and that having one won't help or harm your visibility. Gary Illyes had already said as much at a Google event in July 2025. Ordinary SEO and content quality are what drive AI Overviews visibility, not this file.

Will ChatGPT, Perplexity, or Microsoft Copilot read my llms.txt file?

There's no official confirmation either way from OpenAI, Perplexity, or Microsoft. What we do have is real traffic data: Ahrefs found that 97% of published llms.txt files received zero fetch requests in May 2026, and of the small remainder, only around a fifth of that traffic came from named AI tools. In practice, treat it as unread until you see evidence otherwise for your own site.

Is llms.txt the same as robots.txt or sitemap.xml?

No, though the names invite confusion. robots.txt controls crawler access and is respected by all major compliant crawlers. sitemap.xml lists your indexable pages and is actively used by Google Search. llms.txt is a voluntary, unenforced proposal that no major AI company has confirmed reading. Only the first two have a demonstrated effect on how your site is crawled and indexed.

Should I remove my llms.txt file if I already have one?

No need. If a platform like Mintlify or GitBook generated it automatically, or you added it out of caution, it's harmless to leave in place — that's Google's own characterisation. Just don't spend further time maintaining or expanding it under the assumption it's driving AI visibility, because the evidence doesn't support that.

What should I do instead to improve my AI search visibility?

Focus on answer-first content structure, named and dated sources for every statistic, and consistent business information across your Google Business Profile, directories, and social channels. These have measurable links to AI citation rates. Our guides to generative engine optimisation and answer-first content walk through exactly how to apply them.


Sources

#ClaimSourceYear
1Origin of the llms.txt proposalJeremy Howard / Answer.AI3 Sep 2024
2Google Search doesn't use or need llms.txt (official documentation)Google Search CentralUpdated 15 Jun 2026
3Illyes: Google "doesn't support llms.txt and isn't planning to"Search Engine Land23 Jul 2025
4Google's Search vs. Lighthouse guidance differs by productSearch Engine Journal2026
597% of ~38,000 llms.txt files received zero requests in May 2026 (137,210-domain study)Ahrefs15 Jun 2026
6Coverage of Ahrefs' zero-request findingSearch Engine Journal2026
710-site tracked study found no attributable AI traffic gainSearch Engine Land2026
88.7% adoption among the global top 1,000 sitesRankabilityJun 2026
910.13% broader adoption rateSE RankingNov 2025
10Mintlify/GitBook/Wix auto-generate llms.txtMintlify2025–2026
11Cloudflare found Perplexity using undeclared crawlers, bypassing robots.txtCloudflare, via Search Engine Journal4 Aug 2025