Owned sources account for 10.2% of the citations behind AI answers in our corpus. Roughly nine in ten links point somewhere a brand does not control.
Owned sources — a brand's own domains — account for 10.2% of the 5,440 citation links we have captured. Roughly nine in ten links behind an AI answer point somewhere the brand does not control. You cannot write your way in from your own site alone.
That is the headline, and on its own it is misleading, so the rest of this report is the correction to it. Owned content is a small share of citations and a decisive one. It behaves differently on every engine. And the thing that actually predicts whether a brand is named is not what it published but whether anything of its was retrieved at all.
Where an engine exposes the sources behind an answer, we capture and classify them. Across roughly 515 runs we have 5,440 links, classified against a shared dictionary into owned domains, editorial press, reference works, user-generated content, and a large residual of retail, directories and institutions.
| Source type | All links | ChatGPT | Claude | Gemini | Perplexity | AI Overviews |
|---|---|---|---|---|---|---|
| Other (retail, directories, institutions) | 53.3% | 61.5% | 55.8% | 58.1% | 47.6% | 51.0% |
| Unclassified | 14.0% | 12.2% | 9.8% | 14.7% | 17.1% | 13.3% |
| Editorial (press and media) | 11.4% | 0.6% | 8.7% | 7.7% | 19.1% | 6.6% |
| Owned (brand's own domains) | 10.2% | 22.7% | 11.3% | 10.5% | 4.5% | 16.8% |
| User-generated (social, forums, reviews) | 7.8% | 1.3% | 5.1% | 7.4% | 10.9% | 10.2% |
| Reference (Wikipedia-class) | 3.4% | 1.7% | 9.4% | 1.5% | 0.8% | 2.2% |
| Links cited per run | — | 4.3 | 13.7 | 9.6 | 17.3 | 6.9 |
The engines are not variations on one behaviour. They have different diets. ChatGPT builds answers substantially from brands' own sites — 22.7% owned — and barely reads the press at 0.6%. Perplexity is close to the exact inverse: 4.5% owned against 19.1% editorial. Claude is the reference-layer reader, citing Wikipedia-class sources at 9.4%, several times any other engine. Gemini and AI Overviews sit in between with a real appetite for user-generated content.
The sharpest single consequence: you cannot get into Perplexity through your own website. The channel that works on ChatGPT is close to closed there, and the channel that works on Perplexity is close to invisible on ChatGPT. One content strategy cannot serve five engines, and a brand with strong owned content and no earned coverage will show exactly the engine-shaped hole this table predicts.
The strongest relationship in the entire corpus is not about source type. It is about whether the brand's own domain appeared among the sources at all.
| Engine | Own domain cited in | Presence when cited | Presence when not | Difference |
|---|---|---|---|---|
| Claude | 23.5% of runs | 100.0% | 43.6% | +56 points |
| Gemini | 28.4% | 96.8% | 44.9% | +52 points |
| ChatGPT | 23.6% | 96.2% | 45.2% | +51 points |
| Perplexity | 20.2% | 78.3% | 40.7% | +38 points |
| AI Overviews | 21.3% | 76.5% | 41.3% | +35 points |
On Claude, every single run in which the brand's own domain was retrieved produced a mention of the brand. Read from the other direction, the same fact is starker: in answers that named the brand, owned sources made up 19.9% of the links; in answers that did not, 0.8%. Losing answers are assembled from long-tail and unclassified domains — 27.5% of the links in answers that named nobody relevant were unclassified — while winning answers are assembled from the brand's own pages plus real press and reference entries.
So owned content is 10.2% of the volume and close to all of the swing. Both statements are true and most commentary in this category carries only one of them.
Two further measurements complete the picture. Brands with a Wikipedia page average 58.3% presence against 25.0% for brands without one, a gap of 33 points. Brands with schema.org markup average 37.3% against 30.8%, a gap of 6.5 points. Both are correlational and both rest on a small sample of 18 to 23 brands, but the ordering is stable across every cut: the entity layer correlates three to five times more strongly than any on-site technical signal. Schema markup is cheap hygiene worth doing, not a lever worth selling.
And the prize is binary. Between 64% and 78% of all brand mentions in our corpus sit at position one, and 79% to 93% are in the top three; almost nothing appears past position five. There is no page two of an AI answer. The realistic outcomes are that you are the answer, you are one of two or three alternatives, or you do not exist — which is why moving from position six to position four is not a game worth playing.
The first conclusion is the practical one. Earned placement dominates. If nine in ten citations are somebody else's material, then the work that changes a brand's standing looks more like digital public relations, review-platform depth and reference-layer presence than like publishing more pages on your own domain. The budget shape that follows is closer to a communications programme than a content programme.
The second is that owned content still matters, and matters most where it is retrieved. The instruction is not to stop publishing. It is to stop treating publication as the goal. The question worth asking of any page is not whether it exists or how it is optimised but whether it gets fetched for the questions that matter — a question that can be checked page by page rather than asserted as a general rule.
The third is that the advice must be per engine. Earn coverage is right for Perplexity and Gemini and close to useless for ChatGPT, where owned domains are the dominant citable material. A single ranked action list applied to five engines will be wrong on at least three of them.
The fourth is about where the winnable ground actually is. Of the cited domains appearing around the brands we track, roughly one in eight is realistically earnable inventory — editorial, review and reference placements a brand could plausibly obtain — and roughly one in four is a competitor's own website, which is not available at any price. That ratio is a better guide to what a visibility programme can achieve than any score. It also comes with a large caveat: 62% of those source records are not yet assessed for whether they can be influenced at all, so the ratio describes the assessed minority.
Citation capture is the weakest part of this report and we would rather say so than bury it. It runs on about 4.5% of runs, so every figure here describes that subset. Widening capture is the top data gap we have identified in our own corpus, and it would do two things: test whether the ordering survives, and make the retrieval relationship computable per brand rather than only in aggregate. If a wider sample moved the engine ordering — Perplexity's owned share rising materially, say — the playbook above would change with it.
The classification dictionary is the second weakness. 14% of links are unclassified and 53% fall into a residual bucket of retail, directories and institutions that is doing far too much work. A better dictionary could move the owned share in either direction.
The retrieval relationship is correlational and the direction runs both ways: an engine that has already decided to discuss a brand will also cite it, so we cannot separate retrieval causing the mention from the decision to mention causing the citation. What survives that objection is the negative form, which is the operationally useful one — we observe no path to being named without being findable. Absence from the retrieval set is structurally disqualifying even if presence in it is not strictly causal.
One engine's figures are not comparable with the rest. Claude's captured links include sources consulted but not ultimately cited. That is deliberate, because consultation is what retrieval means, but it inflates Claude's link counts relative to engines that expose only final citations.
Finally, model transitions reset the board. Following an engine update on 7 July 2026, same-prompt visibility across our corpus moved by between 15% and 74% while the number of brands named per answer roughly doubled. A generation of models that began citing brand websites far more heavily would overturn the central figure in this report within weeks, which is why we date every statistic and re-verify on a schedule rather than publishing a number as though it were a constant.
22,012 completed AI answer runs between 15 April and 21 August 2026 — 11,554 production and 10,458 staging — across five engines (ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews), approximately 110 brands and 65 cold scans. Citation analysis in this report rests on 5,440 links captured across roughly 515 production runs, which is about 4.5% of production runs; every citation figure describes that subset rather than the whole corpus.
Presence is the share of completed prompt runs in which the brand was mentioned, computed with read-only SQL against the live databases under the pipeline's own scoping rules. Mentions are resolved by entity resolution with a three-state verdict, adopted on 18 August 2026. It replaced a pure word-boundary string matcher which counted any occurrence of a brand's name as a mention, whether or not the answer concerned that company. The three-state verdict is asymmetric by design — it prefers recording uncertainty to asserting a match — and it lowered reported scores when it shipped. In this corpus, 40 of 97 checked verdicts returned a different entity.
Captured links are classified against a shared source dictionary into owned, editorial, reference, user-generated and a residual other category, with an unclassified bucket for domains the dictionary does not cover. A link counts as owned where the citation domain matches the brand's domain by equality, subdomain or suffix. Position is model-extracted from the answer text and carries extraction noise. Claude's captured links include sources consulted but not cited, which is deliberate — consultation is the retrieval signal — but makes its diet non-comparable with engines that expose final citations only.
Production and staging were kept separate throughout and never pooled. Staging is a rehearsal environment with a different brand set, used strictly as a replication check; approximately 4,300 further citation links were captured there. Findings are reported as replicated only where they reproduce independently on that separate corpus.
Figures are quoted as measured rather than recomputed for presentation. The join behind every presence figure is published internally in full, caveats are stated before findings, correlates are labelled correlational by the analysts who produced them, and data gaps found in our own records — citation capture coverage among them — are listed rather than omitted.