← All research
90% → 12%

The funnel cliff: from 90% presence to 12%

Brand presence in AI answers falls from 90% on branded prompts to 12% on upper-funnel ones, measured across 2,860 tagged evaluations and replicated on a second corpus.

Report period
15 April – 21 August 2026
Sample
2,860 funnel-tagged evaluations within 22,012 runs across five engines
Confidence
High

Presence in AI answers falls from 90.2% to 11.6% as prompts move away from a brand's own name. Branded prompts return 100% presence on every conversational engine we measure. The prompts where buying decisions are actually made return a third of that, or a tenth.

The sequence is 90 → 47 → 33 → 24 → 12, measured across 2,860 funnel-tagged evaluations inside a corpus of 22,012 AI answer runs, and replicated on an independent staging corpus built from a different set of brands. We call it the funnel cliff. It is the most consequential thing we know about how AI visibility gets measured, because it means a brand can hold a strong-looking presence figure and still be absent from every question a prospective customer would actually ask.

What we measured

Every prompt in the corpus carrying a funnel tag was grouped into one of five stages, from prompts that name the brand through to educational prompts with no purchase intent at all. Presence is the share of completed runs in which the brand was mentioned. Position is the brand's average rank within the answers that mentioned it.

Prompt stageEvaluationsPresenceAverage position
Branded bottom-funnel74890.2%1.00
Question prompts99746.6%1.75
Unbranded bottom-funnel21832.6%3.01
Head terms63024.4%2.06
Upper funnel26711.6%3.16

Replication on staging returned 89.4 → 38.3 → 32.7 → 34.3 → 15.1. The shape holds: saturated at the top, near the floor at the bottom, a wide gap between. The ordering of the two middle steps does not hold — on staging, head terms sit fractionally above unbranded bottom-funnel prompts. We report that rather than smoothing it. The honest reading is that steps two, three and four form a band rather than a strict ranking, and that the cliff itself — the drop from branded to everything else, and the second drop into upper funnel — is what replicates.

What each step represents

Branded bottom-funnel — 90.2%, average position 1.00. Prompts containing the brand name: what is this company like, is it any good, what do people say about it. The buyer already knows who you are. Every conversational engine returns the brand essentially every time, first in the answer. Google AI Overviews manages 19.8% on the same prompts, and that is not a visibility failure — Google frequently serves no overview at all for navigational queries, so there is no answer to be present in. Scoring that gap as a deficiency penalises a brand for a Google product decision.

Question prompts — 46.6%, position 1.75. Specific, natural-language questions that describe a need rather than a name. Roughly half of these answers name the brand. This is the widest realistic opportunity in the dataset: specific enough that category giants thin out, commercial enough that the person asking is close to a decision.

Unbranded bottom-funnel — 32.6%, position 3.01. Purchase-intent phrasing without a brand name. Note what happens to position: brands that survive this drop arrive third rather than first. Given that between 64% and 78% of all mentions across our corpus sit at position one, arriving third is a materially different outcome from being the answer.

Head terms — 24.4%, position 2.06. Short category terms. This is incumbent territory. The shortlist slots go to the largest names in the category, and a mid-sized brand competing here is competing for a seat that is usually already taken before the question is asked.

Upper funnel — 11.6%, position 3.16. Educational and explanatory prompts, where a brand hopes to be named as an authority. The platform split here is the most surprising figure in the set: ChatGPT names brands in 27.9% of these answers and AI Overviews in 25.0%, while Claude manages 3.2%, Gemini 6.6% and Perplexity 4.8%. Three of five engines essentially refuse to name commercial brands in educational answers.

What it means for a brand

The first consequence is that a single presence number is close to uninformative. A brand reporting 30% presence might be at 100% on branded prompts and 5% on everything else, or evenly at 30% across the funnel. Those are completely different business situations and the aggregate cannot tell them apart. Any visibility figure quoted without a stage breakdown should be treated as an unresolved question rather than an answer.

The second is that prompt-set design drives perceived performance. In our tagged pool, 748 of 2,860 evaluations — 26% — are branded prompts held near maximum by the shape of the question. A brand that fills its tracking set with prompts naming itself will look healthy without doing anything, and will keep looking healthy while losing every unbranded answer in its category. This is not a subtle measurement artefact. It is the single easiest way to produce a flattering AI visibility report, and it requires no dishonesty from anyone.

The third is that branded prompts are the wrong instrument for presence and the right instrument for something else. There is no variance left in them to measure, so they cannot distinguish a strong brand from a weak one. What they can do is show whether the engine describes the brand accurately — which is a different question, and one a customer is far more likely to notice going wrong. Our position is that branded prompts should be reported separately from a presence score, or excluded from it, and used to monitor how the brand is characterised.

The fourth is a targeting judgement. If presence roughly doubles between head terms and specific questions, the winnable ground for most brands is question prompts. Head terms are a later and harder campaign, and upper-funnel visibility — the thing usually sold as thought leadership — currently pays on ChatGPT and AI Overviews and almost nowhere else. Anyone selling it as a five-platform play is describing a market that our data does not show.

What would change our mind

Funnel stage in this corpus is derived from free-text tags that cover about a quarter of evaluations. The remaining 8,693 untagged evaluations return 26.3% presence, which is consistent with the question-and-head-term blend we assume they are, but that is an assumption rather than a measurement. A structured funnel field applied across the whole corpus could move the middle of the cliff in either direction, and we would publish the revised figures.

Two of the five cells are small. Unbranded bottom-funnel rests on 218 evaluations and upper funnel on 267, and the platform splits inside them are thinner still. A larger sample showing unbranded bottom-funnel prompts at parity with question prompts would collapse two steps into one and simplify the picture.

The top of the cliff is the part most exposed to engine behaviour rather than brand behaviour. If a future model generation became more willing to say it does not recognise a company, branded presence would fall below saturation and branded prompts would start carrying information again. We would treat that as a change in the instrument, not a change in the brands, and report it as a separate series.

Finally, 86% of these runs are Ireland and Great Britain. A comparable measurement on a United States-weighted corpus that produced a materially flatter gradient would tell us the cliff is a feature of this market rather than of AI answering generally. That test has not been run.

How we measured this

Corpus

22,012 completed AI answer runs between 15 April and 21 August 2026: 11,554 in production and 10,458 in staging. Five engines — ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews — across approximately 110 brands, 895 production prompts, and 65 anonymous cold scans. Country mix is Ireland and Great Britain heavy, at 86% of runs, with Hong Kong, South Africa and United States tails.

How presence was determined

Presence is the share of completed prompt runs in which the brand was mentioned, computed with read-only SQL against the live databases, using the pipeline's own scoping rules — a brand is evaluated only on its own prompts. Position is model-extracted from the answer text and carries extraction noise.

Mentions are resolved by entity resolution with a three-state verdict, adopted on 18 August 2026. It replaced a pure word-boundary string matcher that counted any appearance of the brand's name as a mention, whether or not the answer was about that company. The three-state verdict is asymmetric by design: it is more willing to record uncertainty than to claim a match, which lowers reported scores rather than raising them. Of 97 checked verdicts in this corpus, 40 returned a different entity — mentions demoted because the engine was discussing another company with a similar name.

Funnel staging

Funnel stage is derived by text matching against free-text prompt tags. Approximately 25% of production evaluations — 2,860 rows — carry such tags; there is no structured funnel field in the schema. The tagged pool is what the table above describes.

Replication

Production and staging were kept separate throughout and never pooled. Staging is a rehearsal environment with 21 organisations and 38 brands, used strictly as a replication check: a finding counts as replicated when it reproduces there on a different brand set. The funnel cliff replicated; the ordering of its two middle steps did not.

Disclosure

The join that produced every presence figure in this report is published internally in full, and the figures are quoted as measured rather than recomputed for presentation.

What this does not show

  • 52.5% of production own-brand telemetry sits under an internal pilot account. The AI answers behind it are real and the platform behaviour it records is a valid observation of the outside world, but row counts must not be read as customer traction.
  • Funnel tags cover about a quarter of evaluations. The cliff is measured on 2,860 tagged rows; the 8,693 untagged rows are assumed to behave like a question and head-term blend, which is an assumption and not a measurement.
  • Two cells are small. Unbranded bottom-funnel rests on 218 evaluations and upper funnel on 267. Platform splits within those rows are thin and should be read as directional.
  • Client prompt sets are curated. Presence rates measured on tracked brands are not market rates. Our cold-scan corpus shows market rates running three to five times lower.
  • A methodology discontinuity falls inside the period. On 7–8 July 2026 the scan models were upgraded and the scoring panel moved from v1 to v2. Raw figures either side of that boundary are not comparable without a cohort control.
  • Lever findings elsewhere in this research programme are correlational, not causal, and are labelled as such.
  • Geographic concentration. 86% of runs are Ireland and Great Britain. United States dynamics are under-sampled.
  • Position figures are model-extracted and carry extraction noise.

AI visibility monitoring across the major AI engines.