Brands with an llms.txt file scored 9.8 points lower on presence than brands without one. We found no evidence of benefit, and we deleted our own recommendation.
Brands with an llms.txt file on their site averaged 29.3% presence in AI answers. Brands without one averaged 39.1%. That is a gap of 9.8 points in the wrong direction, and we found no positive association anywhere in the data.
We are not claiming llms.txt is harmful. We think that reading is almost certainly wrong. The finding we are prepared to defend is narrower and duller: in our corpus there is no evidence of benefit. That is enough to stop recommending it, and it was enough for us to stop, ten weeks before we had this measurement.
Alongside presence, we run technical and entity readiness checks on tracked brands: whether the site renders server-side, whether it carries schema.org markup, whether the brand has a Wikipedia page, a Trustpilot profile, an llms.txt file. Comparing average presence for brands that have each signal against brands that do not gives a crude but informative ranking of what travels with visibility.
| Signal | Presence, brands with it | Presence, brands without | Gap |
|---|---|---|---|
| Wikipedia page | 58.3% | 25.0% | +33 points |
| Trustpilot profile | 45.8% | 22.4% | +23 points |
| schema.org markup | 37.3% | 30.8% | +6.5 points |
| Server-side rendering | 34.2% | 28.9% | +5.3 points |
| llms.txt file | 29.3% | 39.1% | −9.8 points |
These sit on 18 to 23 brands, with some cells as small as three. They are directional and nothing more. But the ordering is consistent across every cut we have taken of the data, and the ordering is the interesting part: being an established entity in the world correlates three to five times more strongly with AI visibility than any on-site technical signal. The technical layer is real and small. The entity layer is where the variance lives.
llms.txt is a remedy. Nobody adds one because things are going well. It is adopted by brands that already suspect they have an AI visibility problem and are looking for something cheap to try — which means a remedial signal will correlate with the condition it is meant to remedy, in the same way that carrying painkillers correlates with headaches. That is the reading we hold, and it fits the data better than any causal story.
It is also worth stating what is not in dispute: no major model provider has confirmed that its crawlers read the file. That is not proof of nothing happening. It is an absence of the evidence that would be needed to justify recommending it, and after two years of the format existing, the absence has become informative.
Our audit used to tell customers to create an llms.txt file. The more useful part of this report is what happened to that recommendation.
In June 2026 we commissioned an internal review of the evidence behind every claim our audit methodology made. It was not a marketing exercise and it was not kind. It reported on 12 June 2026, and on the same day the recommendation was pulled: llms.txt was removed from the playbook and demoted to informational-only, so that a brand's not having one no longer affects any score we produce. Three content rules went with it, deleted outright rather than softened:
Each traced back to a single unverifiable vendor page, or had no evidence behind it that any engine consumed the thing being recommended. The review also corrected two figures that had been misread rather than invented: a widely circulated statistic that Wikipedia accounts for 47.9% of ChatGPT citations turned out to be its share within the top ten cited sources, not overall — the overall figure is 7.8% — and the same misreading applied to a claim about Reddit and Perplexity, where the overall share is 6.6%. A third claim, that 92% of AI Overviews citations come from top-ten organic results, was contradicted by per-citation measurement closer to 38%.
The detail that decides whether a correction is genuine is what happens to the weight the deleted claim was carrying. Ours moved: multi-modal coverage fell from 13.5% to 5% of the audit's composite score, and authority and brand signals rose from 18% to 24%. Deleting a statistic and quietly keeping the weight it justified would have been cosmetic. A new rule now requires every statistic in the methodology to carry a metric and a date, with a scheduled re-verification cadence, because claims in this category decay in weeks.
The −9.8 point measurement arrived in August, ten weeks after the deletion. It confirmed a decision already taken on other grounds. We think the order matters, and we are reporting it in that order deliberately: the recommendation was removed because we could not source it, not because we had a number that made it look bad.
Do not buy llms.txt as a visibility lever, and be wary of anyone selling it as one. The file takes a few minutes to write and does no measurable damage; if you want one for its own sake, have one. What it should not do is occupy a place on a prioritised action list, or displace work on the things that our data associates with much larger differences.
Those things, in order of the size of the association: presence in the reference layer, a Wikipedia-class entry being the clearest case, at +33 points; depth on review platforms, at +23; and above both, being retrieved at all. Where a brand's own domain is cited in an answer, presence runs between 76% and 100% depending on the engine, against 41% to 45% where it is not — a lift of 35 to 56 points, and by a wide margin the strongest relationship in our corpus. Technical hygiene sits an order of magnitude below that. It is worth doing and not worth selling.
There is a broader point here about how this category argues. A recommendation that is cheap, plausible and unfalsifiable will survive indefinitely if nobody measures it, because nobody who followed it can prove it did not help. The only defence is to publish the measurement and the method together, so that the claim can be checked and, when necessary, withdrawn.
A model provider documenting that its crawlers read llms.txt and act on it would change the question immediately, and we would say so. That is the cleanest possible refutation of our position and it requires nothing from us.
A controlled before-and-after test would be better evidence than anything in this report. Take a set of brands, add the file, change nothing else, and measure presence across a full sampling cycle either side. Our finding is a cross-sectional comparison of brands that differ in many ways at once; an intervention study on the same brands would separate selection from effect. We have not run it. Anyone could.
A materially larger sample could also overturn the direction. Twenty-two brands with cells as small as three is a thin basis for anything, and the honest description of our result is that it fails to find a benefit rather than that it establishes a penalty. If a corpus ten times the size returned a positive association, we would publish that and reinstate the recommendation with the evidence attached.
What would not change our mind is a vendor page, an agency case study without a control, or a plausible mechanism unaccompanied by a measurement. That is precisely the class of evidence we deleted from our own methodology, and we cannot consistently reject it in our work and accept it in someone else's.
22,012 completed AI answer runs between 15 April and 21 August 2026 — 11,554 production and 10,458 staging — across five engines (ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews), approximately 110 brands, 65 cold scans and 5,440 captured citation links. The readiness comparison in this report sits on a subset of roughly 22 brands for which both presence and site-readiness checks were available.
Presence is the share of completed prompt runs in which the brand was mentioned, computed with read-only SQL against the live databases under the pipeline's own scoping rules. Mentions are resolved by entity resolution with a three-state verdict, adopted 18 August 2026, which replaced a pure word-boundary string matcher that counted any appearance of the brand's name as a mention regardless of whether the answer was about that company. The verdict is asymmetric by design: it prefers to record uncertainty over asserting a match, and it lowered reported scores when it shipped.
For each signal, brands are split into those where the check returned present and those where it returned absent, and mean presence is compared between the two groups. This is a cross-sectional association between brands, not a before-and-after test on the same brand. It cannot distinguish a signal that causes visibility from a signal adopted by brands that already lack it — which is exactly the ambiguity we flag on llms.txt.
Production and staging were kept separate throughout and never pooled. Staging, a rehearsal environment with a different brand set, is used strictly as a replication check on findings computed in production. The readiness correlates were not independently replicated on staging; they rest on the production subset only, which is one reason this report is graded directional.
The June 2026 evidence review was an internal audit of every empirical claim in our audit methodology, checking each against its original source. Its findings were applied on 12 June 2026, the day it reported, and are recorded in the framework's own dated version history, which names each refuted claim and what replaced it. The current methodology carries a standing rule requiring a metric and a date on every statistic, with scheduled re-verification.