---
title: Does llms.txt Get You Cited by AI? I Checked the Data.
canonical: "https://www.rankinghacks.com/does-llms-txt-get-you-cited/"
pubDate: "2026-08-11T09:00:00.000Z"
updatedDate: "2026-08-11T00:00:00.000Z"
author: Andreas De Rosi
description: "Do AI engines cite sites with llms.txt more? 42% of domains ChatGPT and Perplexity cited had one vs 4% of indie blogs, a 10x correlation. But the causation claim breaks down. The full data."
categories: [ai-search]
---

A few weeks ago I [checked 319 independent blogs for llms.txt and found only 12 had one](/llms-txt-in-the-wild/). That answered the adoption question: almost nobody has shipped the file. But it left the bigger question open, the one people actually argue about. Does having an llms.txt get you *cited* by AI answer engines? If the format does nothing, the 4% base rate is a curiosity. If it moves citations, that empty field is the cheapest opportunity in SEO right now.

I am in an unusual position to check, because I already run [my own AI-citation tracker](/track-ai-citations-chatgpt-perplexity/). It logs which URLs ChatGPT and Perplexity actually cite across a fixed set of GEO and SEO questions. So instead of guessing, I took the domains those two engines really cited and probed every one of them for an llms.txt file. Here is what the data says, including the part that stops it from meaning what you want it to mean.

## The method

The cited set is not a hand-picked list. My tracker ran 25 fixed GEO/SEO queries through Perplexity and OpenAI's web-search models, producing 50 answers. Those answers carried **479 citation URLs, which resolve to 258 unique third-party domains** once you strip the self-citations to RankingHacks. That is the population: the domains AI reached for when answering questions in my niche.

I then probed each of those 258 domains for a `/llms.txt` file using the exact same content-based classifier from the [first study](/llms-txt-in-the-wild/). That detail matters. Roughly a fifth of sites return `200 OK` with an HTML page for any path you ask for, so a status-code check would count a homepage as an llms.txt hit. I required the response to actually look like the file: Markdown, an H1, real sections. The comparison base rate is the first study's **4.2%** adoption among neutral indie blogs.

## The correlation is real and large

Of the 258 cited domains, 254 were reachable, and **107 of them had a well-formed llms.txt. That is 42.1%.**

<figure>
  <img src="/images/posts/does-llms-txt-get-you-cited/01-cited-vs-baseline.png" alt="Bar chart comparing llms.txt adoption: 4.2% of independent blogs versus 42.1% of the domains ChatGPT and Perplexity cited" loading="lazy" />
  <figcaption>Domains that AI actually cites are ten times more likely to carry a well-formed llms.txt than a neutral sample of independent blogs. <strong>42.1%</strong> versus <strong>4.2%</strong>. A real, large correlation.</figcaption>
</figure>

Ten times the base rate. If you stopped reading here you would conclude that llms.txt is a citation magnet and everyone should ship one tomorrow. Plenty of posts have stopped exactly there. It is a clean, quotable, and misleading number, and the honest thing to do with it is to try to break it.

## The finding that kills the causation claim

The moment you look at *which* domains get cited most, the story changes.

<figure>
  <img src="/images/posts/does-llms-txt-get-you-cited/02-top-cited-llmstxt.png" alt="Bar chart of the most-cited domains, colored by whether they have an llms.txt. The seven most-cited (youtube, linkedin, developer.chrome.com, searchengineland, reddit, searchenginejournal, wikipedia) are all grey, meaning no llms.txt." loading="lazy" />
  <figcaption>The most-cited domains, colored by llms.txt presence. The <strong>seven most-cited sources have no llms.txt at all</strong>. The file only starts appearing further down, among mid-tier SEO and marketing sites.</figcaption>
</figure>

The seven domains AI cited most in my data are youtube.com, linkedin.com, developer.chrome.com, searchengineland.com, reddit.com, searchenginejournal.com, and en.wikipedia.org. Between them they account for nearly 30% of every third-party citation. **Not one of them has an llms.txt.** I re-checked all seven the day I published this: four return a genuine 404, and three serve an HTML soft-404. AI plainly does not need the file to cite you. It cites the biggest, most authoritative sources it can find, file or no file.

That flips the weighting. Count *domains* and 42% have an llms.txt. Count *citations*, so that a domain cited 45 times counts 45 times, and the share drops to **29.9%**, because the heavy hitters at the top have none.

<figure>
  <img src="/images/posts/does-llms-txt-get-you-cited/03-domains-vs-weighted.png" alt="Bar chart showing 42.1% when counting domains falls to 29.9% when weighting by citation volume" loading="lazy" />
  <figcaption>The same data, weighted two ways. Counting sites overstates the effect. Weighting by how often each domain is actually cited corrects it downward.</figcaption>
</figure>

And the domains that *are* cited with an llms.txt tell you what the correlation is really made of. They are semrush, conductor, hubspot, github, yoast, brightlocal, vendasta, and a long tail of GEO and AI-visibility SaaS. These are exactly the domains that have two things at once: the authority to get cited, and the marketing awareness to ship a new file the week it trended. That is a textbook confound. The llms.txt is not causing the citation. Both are downstream of the same thing, an established, marketing-savvy, well-resourced site.

## The honest answer

Correlated strongly, about tenfold. Not causally proven. And the sources AI cites most do not have the file at all.

A well-formed llms.txt is a **marker** that travels with the kind of site AI tends to cite, not a demonstrated **cause** of the citation. It is a signal of a site that is on top of its technical marketing, in the same way that a tidy sitemap or fast Core Web Vitals correlates with sites that rank without being the reason they rank. An observational snapshot like this one cannot separate the file's own effect from the authority of the sites that happen to ship it. Anyone selling you "add llms.txt and get cited 10x more" is reading the first chart and ignoring the next two.

This does not contradict the case I made in the [first study](/llms-txt-in-the-wild/). There the point was that a well-formed llms.txt is cheap and uncrowded, an afternoon of work in a field where 96% of your peers have skipped it. That is still true, and it is still worth doing. What this second study adds is the ceiling: do it because it is cheap table-stakes among sites that take AI visibility seriously, not because the file itself buys you citations. If you are already authoritative, it is a tidy, low-cost way to be legible to the machines. If you are not, shipping the file will not vault you above wikipedia in a Perplexity answer.

## What to actually do with this

Ship the llms.txt, because it is nearly free and the downside is zero. But spend your real effort on the thing the data says actually earns the citation: being a source worth citing. In my own tracker the citations that matter follow original data, first-hand experience, and specific numbers, the same things that earned the citations in the first place. The file is the wrapper. The reason to cite you is the content.

If you want to measure this on your own site rather than trust my sample, that is the entire point of running a [citation tracker](/track-ai-citations-chatgpt-perplexity/) and a [GEO audit](/geo-audit-own-site/). Watch which of your pages ChatGPT and Perplexity pull, and optimize for that, not for a checklist item.

## Method, caveats, and how to reproduce it

The cited set is a 2026-07-27 snapshot of my tracker: 25 queries, two engines, 479 citation URLs, 258 domains. The llms.txt probe uses the same open, content-based classifier as the first study, so it dodges the soft-404 trap that inflates status-only counts. [Download the full dataset (CSV)](/downloads/does-llms-txt-get-you-cited.csv): one row per cited domain with its citation count, which engines cited it, and its llms.txt classification, size, and shape. The honest caveats: this is one publisher's query set in one niche, it is a snapshot rather than a trend, and it is observational, so it can show correlation but cannot prove the file's independent effect. That last point is not a hedge, it is the finding. The correlation is real. The causation is not established, and the most-cited sources in the entire dataset carry no llms.txt at all.

Which is a more useful thing to know than a clean 10x would have been.

Related reading: [I checked 319 indie blogs for llms.txt](/llms-txt-in-the-wild/), [does llms.txt actually work](/llms-txt-does-it-work/), [how I track AI citations](/track-ai-citations-chatgpt-perplexity/), [how I GEO-audited my own site](/geo-audit-own-site/), and [what happened when I GEO-optimized RankingHacks](/we-geo-optimized-our-own-site/).
