Brand presence in AI answers falls from 90% on branded prompts to 12% on upper-funnel ones, measured across 2,860 tagged evaluations and replicated on a second corpus.
Presence in AI answers falls from 90.2% to 11.6% as prompts move away from a brand's own name. Branded prompts return 100% presence on every conversational engine we measure. The prompts where buying decisions are actually made return a third of that, or a tenth.
The sequence is 90 → 47 → 33 → 24 → 12, measured across 2,860 funnel-tagged evaluations inside a corpus of 22,012 AI answer runs, and replicated on an independent staging corpus built from a different set of brands. We call it the funnel cliff. It is the most consequential thing we know about how AI visibility gets measured, because it means a brand can hold a strong-looking presence figure and still be absent from every question a prospective customer would actually ask.
Every prompt in the corpus carrying a funnel tag was grouped into one of five stages, from prompts that name the brand through to educational prompts with no purchase intent at all. Presence is the share of completed runs in which the brand was mentioned. Position is the brand's average rank within the answers that mentioned it.
| Prompt stage | Evaluations | Presence | Average position |
|---|---|---|---|
| Branded bottom-funnel | 748 | 90.2% | 1.00 |
| Question prompts | 997 | 46.6% | 1.75 |
| Unbranded bottom-funnel | 218 | 32.6% | 3.01 |
| Head terms | 630 | 24.4% | 2.06 |
| Upper funnel | 267 | 11.6% | 3.16 |
Replication on staging returned 89.4 → 38.3 → 32.7 → 34.3 → 15.1. The shape holds: saturated at the top, near the floor at the bottom, a wide gap between. The ordering of the two middle steps does not hold — on staging, head terms sit fractionally above unbranded bottom-funnel prompts. We report that rather than smoothing it. The honest reading is that steps two, three and four form a band rather than a strict ranking, and that the cliff itself — the drop from branded to everything else, and the second drop into upper funnel — is what replicates.
Branded bottom-funnel — 90.2%, average position 1.00. Prompts containing the brand name: what is this company like, is it any good, what do people say about it. The buyer already knows who you are. Every conversational engine returns the brand essentially every time, first in the answer. Google AI Overviews manages 19.8% on the same prompts, and that is not a visibility failure — Google frequently serves no overview at all for navigational queries, so there is no answer to be present in. Scoring that gap as a deficiency penalises a brand for a Google product decision.
Question prompts — 46.6%, position 1.75. Specific, natural-language questions that describe a need rather than a name. Roughly half of these answers name the brand. This is the widest realistic opportunity in the dataset: specific enough that category giants thin out, commercial enough that the person asking is close to a decision.
Unbranded bottom-funnel — 32.6%, position 3.01. Purchase-intent phrasing without a brand name. Note what happens to position: brands that survive this drop arrive third rather than first. Given that between 64% and 78% of all mentions across our corpus sit at position one, arriving third is a materially different outcome from being the answer.
Head terms — 24.4%, position 2.06. Short category terms. This is incumbent territory. The shortlist slots go to the largest names in the category, and a mid-sized brand competing here is competing for a seat that is usually already taken before the question is asked.
Upper funnel — 11.6%, position 3.16. Educational and explanatory prompts, where a brand hopes to be named as an authority. The platform split here is the most surprising figure in the set: ChatGPT names brands in 27.9% of these answers and AI Overviews in 25.0%, while Claude manages 3.2%, Gemini 6.6% and Perplexity 4.8%. Three of five engines essentially refuse to name commercial brands in educational answers.
The first consequence is that a single presence number is close to uninformative. A brand reporting 30% presence might be at 100% on branded prompts and 5% on everything else, or evenly at 30% across the funnel. Those are completely different business situations and the aggregate cannot tell them apart. Any visibility figure quoted without a stage breakdown should be treated as an unresolved question rather than an answer.
The second is that prompt-set design drives perceived performance. In our tagged pool, 748 of 2,860 evaluations — 26% — are branded prompts held near maximum by the shape of the question. A brand that fills its tracking set with prompts naming itself will look healthy without doing anything, and will keep looking healthy while losing every unbranded answer in its category. This is not a subtle measurement artefact. It is the single easiest way to produce a flattering AI visibility report, and it requires no dishonesty from anyone.
The third is that branded prompts are the wrong instrument for presence and the right instrument for something else. There is no variance left in them to measure, so they cannot distinguish a strong brand from a weak one. What they can do is show whether the engine describes the brand accurately — which is a different question, and one a customer is far more likely to notice going wrong. Our position is that branded prompts should be reported separately from a presence score, or excluded from it, and used to monitor how the brand is characterised.
The fourth is a targeting judgement. If presence roughly doubles between head terms and specific questions, the winnable ground for most brands is question prompts. Head terms are a later and harder campaign, and upper-funnel visibility — the thing usually sold as thought leadership — currently pays on ChatGPT and AI Overviews and almost nowhere else. Anyone selling it as a five-platform play is describing a market that our data does not show.
Funnel stage in this corpus is derived from free-text tags that cover about a quarter of evaluations. The remaining 8,693 untagged evaluations return 26.3% presence, which is consistent with the question-and-head-term blend we assume they are, but that is an assumption rather than a measurement. A structured funnel field applied across the whole corpus could move the middle of the cliff in either direction, and we would publish the revised figures.
Two of the five cells are small. Unbranded bottom-funnel rests on 218 evaluations and upper funnel on 267, and the platform splits inside them are thinner still. A larger sample showing unbranded bottom-funnel prompts at parity with question prompts would collapse two steps into one and simplify the picture.
The top of the cliff is the part most exposed to engine behaviour rather than brand behaviour. If a future model generation became more willing to say it does not recognise a company, branded presence would fall below saturation and branded prompts would start carrying information again. We would treat that as a change in the instrument, not a change in the brands, and report it as a separate series.
Finally, 86% of these runs are Ireland and Great Britain. A comparable measurement on a United States-weighted corpus that produced a materially flatter gradient would tell us the cliff is a feature of this market rather than of AI answering generally. That test has not been run.
22,012 completed AI answer runs between 15 April and 21 August 2026: 11,554 in production and 10,458 in staging. Five engines — ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews — across approximately 110 brands, 895 production prompts, and 65 anonymous cold scans. Country mix is Ireland and Great Britain heavy, at 86% of runs, with Hong Kong, South Africa and United States tails.
Presence is the share of completed prompt runs in which the brand was mentioned, computed with read-only SQL against the live databases, using the pipeline's own scoping rules — a brand is evaluated only on its own prompts. Position is model-extracted from the answer text and carries extraction noise.
Mentions are resolved by entity resolution with a three-state verdict, adopted on 18 August 2026. It replaced a pure word-boundary string matcher that counted any appearance of the brand's name as a mention, whether or not the answer was about that company. The three-state verdict is asymmetric by design: it is more willing to record uncertainty than to claim a match, which lowers reported scores rather than raising them. Of 97 checked verdicts in this corpus, 40 returned a different entity — mentions demoted because the engine was discussing another company with a similar name.
Funnel stage is derived by text matching against free-text prompt tags. Approximately 25% of production evaluations — 2,860 rows — carry such tags; there is no structured funnel field in the schema. The tagged pool is what the table above describes.
Production and staging were kept separate throughout and never pooled. Staging is a rehearsal environment with 21 organisations and 38 brands, used strictly as a replication check: a finding counts as replicated when it reproduces there on a different brand set. The funnel cliff replicated; the ordering of its two middle steps did not.
The join that produced every presence figure in this report is published internally in full, and the figures are quoted as measured rather than recomputed for presentation.