← All research
46.2% vs 17%

Why B2B brands are losing the answer

E-commerce brands appear in 46.2% of AI answers and B2B brands in roughly 17%, and B2B is the only segment out-mentioned by its own competitors.

Report period
15 April – 21 August 2026
Sample
22,012 runs across five engines
Confidence
High

E-commerce brands appear in 46.2% of the AI answers we measured. Business-to-business brands appear in roughly 17%. B2B is the only segment whose brands are out-mentioned by their own tracked competitors, 17.2% against 21.4%, and shut out entirely of 16.2% of answers.

What we measured

Every own brand in the corpus carries a business model classification. We grouped those into e-commerce, B2B services, B2B product, and local or hospitality, then computed presence per segment: the share of completed prompt runs in which the brand was mentioned, evaluated only on its own prompt set.

SegmentEvaluationsPresenceAverage positionPerplexityAI Overviews
E-commerce4,88546.2%1.8336.9%24.0%
B2B services2,58017.5%2.5314.4%14.6%
B2B product88516.4%2.461.7%26.3%
Local and hospitality2987.4%2.093.6%6.2%

The gap replicated on our staging corpus, which is a separate environment with a different brand set and independent runs: 48.5% for e-commerce against 16.2% for B2B. That matters, because a gap this large invites the obvious objection that it is an artifact of one client mix. It appears in both environments, with different brands.

We also ran the harder control. Branded prompts, which almost every brand wins, exist only in the e-commerce portfolios, and they inflate the gap. Restricted to unbranded prompts alone, e-commerce leads 38.2% against 17.5% and 16.4%. The honest structural gap is therefore about 2.2 to 2.3 times rather than 2.6. It is smaller than the headline and still very large.

B2B brands lose their own matchups

Presence tells you whether a brand appeared. It does not tell you who appeared instead. For that we compared each brand against the competitors it named itself, on its own prompts, run by run.

SegmentOwn presenceCompetitor presenceShut-out rateExclusive-win rate
E-commerce46.2%18.3%7.9%35.7%
B2B17.2%21.4%16.2%12.0%

Shut-out rate is the share of answers where the brand is absent and at least one of its tracked competitors is present. Exclusive-win rate is the share where the brand is present and no tracked competitor is.

This is the most commercially useful comparison in our corpus. B2B is the only segment where the competitors a brand chose to track are named more often than the brand itself. B2B brands are shut out roughly twice as often as e-commerce brands and win exclusively about a third as often. The e-commerce brands we track are mostly ahead of their chosen rivals. The B2B brands are mostly behind theirs.

Why the gap exists

Incumbent gravity is a different force in each segment

AI answers default to category giants and institutions. That is true in both segments, but the giants are not the same kind of entity, and that asymmetry is the mechanism behind the gap.

In B2B the incumbents are advice brands. The recurring entities in B2B answers are the Big Four accountancy firms, the top-tier strategy consultancies, and professional bodies. They occupy the exact recommendation slot a mid-market challenger needs. A consultancy answering the question of who should help with a given problem is competing against the largest advisory brands in the world for the same sentence, and the model has no reason to prefer the challenger.

In e-commerce the incumbents are distribution brands. The recurring entities are national supermarkets and marketplaces. A direct-to-consumer brand appears beside them in a list of where to buy something, rather than being replaced by them. They provide context, not substitution. That single structural difference does a great deal of the work in the presence gap.

The same logic explains why local and hospitality is the worst segment we measure at 7.4%. Answers to local-intent questions name directories, maps and review platforms rather than individual businesses. The institutional layer is the answer.

Professional questions are asked more generically

Consumer buying questions tend to be specific: a product, a use, a constraint, a price. Professional buying questions tend to arrive as broad requests for a category of help. Incumbent gravity is strongest exactly where the question is generic, because a generic question is best answered with a safe, large, well-known name. The more specific the question, the thinner the giants become.

The earned layer is missing

Citations in B2B brand runs contain 3.2% editorial coverage, against 9.9% in e-commerce runs. Editorial coverage is precisely the input that the retrieval-heavy engines consume, and our B2B brands do not have much of it. Their answers lean instead on directories, institutional domains and long-tail sources.

This produces one of the more counter-intuitive findings in the corpus. B2B product brands run at 1.7% presence on Perplexity, their worst engine by a wide margin, and 26.3% on Google AI Overviews, their best. E-commerce is the mirror image. A sensible B2B visibility strategy is AI Overviews and Gemini first, and Perplexity last, which is the opposite of the large-language-model-first instinct almost everyone brings to this problem.

What it means for a brand

If you sell to businesses, the first useful thing to know is that a 20% presence figure is a poor e-commerce result and a respectable B2B one. Graded on a single curve, the same number misleads in both directions. Segment-aware benchmarks are not a refinement; without them the score is close to uninterpretable.

The second is that the competitor comparison is more urgent than the score. Being out-mentioned by the rivals you named yourself, and absent from one in six of the answers they appear in, is a concrete commercial fact in a way that a composite visibility figure is not.

The third is where the winnable ground actually is. Contesting the head term against the largest firms in your category is not a campaign; it is a budget disposal method. The challenger's route is to be the precise answer to a precise question, which means building coverage on the narrow questions your buyers actually ask before their category is settled. Combine that with the engine ordering above and the near-term programme is unusually specific: narrow questions, earned editorial coverage to feed retrieval, and AI Overviews and Gemini as the realistic battleground.

What would change our mind

The segment classification is enriched by a language model and 18 of 72 own brands are unclassified, which is a quarter of the portfolio sitting outside these cells. A manual classification pass could move the boundaries, though it is unlikely to close a gap of this size.

The local and hospitality segment rests on four brands. We state the number because the gap is enormous and worth knowing about. We would not price anything on it.

The most likely challenge to the finding is prompt-set curation. Prompt sets are built per brand and they differ by segment, so some of this gap could reflect harder prompts rather than harder conditions. The unbranded-only control reduces the gap honestly and does not eliminate the concern. A shared, standardised prompt set applied across segments would settle it, and that is the test we would want to run before treating the ratio as a fixed property of the market rather than of our current portfolio.

How we measured this

Corpus. 11,554 completed production prompt runs and 10,458 staging runs, between 15 April and 21 August 2026, across five answer engines: ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews. Approximately 110 own brands, 153 tracked competitor brands, 895 prompts, 65 anonymous cold scans. 22,012 runs in total, from a database snapshot taken on 22 August 2026 and queried read-only.

Presence. The share of completed prompt runs in which the brand was mentioned, with an own brand evaluated only against its own prompt set. Competitor figures use the same join restricted to tracked competitor brands on the same prompts.

Shut-out and exclusive-win. Computed per run rather than per mention. For each completed run we take whether any own-brand row was mentioned and whether any tracked-competitor row was mentioned. A shut-out is a run where the own brand is absent and at least one tracked competitor is present. An exclusive win is a run where the own brand is present and no tracked competitor is.

Segments. Derived from a business-model field enriched by a language model at onboarding. E-commerce combines business-to-consumer and direct-to-consumer e-commerce, 30 active own brands. B2B combines software, niche product, professional services and agency, 19 own brands. Local and hospitality combines local service and hospitality, 4 own brands. 18 own brands are unclassified and excluded from the segment cells.

Controls. Branded prompts exist only in the e-commerce portfolios and inflate the headline gap, so the comparison was re-run on unbranded prompts alone, where e-commerce leads 38.2% against 17.5% and 16.4%.

Entity resolution. Brand mentions were originally detected by a word-boundary string matcher, which counted any answer containing the brand's name as a mention whether or not the answer concerned that company. On 18 August 2026 we replaced it with a three-state verdict that can return uncertainty rather than forcing a match, and which is deliberately more willing to abstain than to claim. Generic and colliding names are more common in B2B than in e-commerce, so this correction matters more to the segment reported here.

Replication. Staging is a separate environment with an overlapping but distinct brand set and independent runs. It reproduced the segment gap at 48.5% against 16.2%. It is used as a replication check, never pooled with production numbers.

What this does not show

  • 52.5% of production own-brand telemetry sits under an internal pilot account tracking roughly eleven real brands. The AI answers observed are entirely real, but the row counts are not customer traction, and this qualifies every commercial reading of the corpus.
  • Segments come from a business-model field enriched by a language model. 18 of 72 own brands are unclassified, so a quarter of the portfolio sits outside these cells.
  • The local and hospitality segment rests on four brands. The gap is large enough to state and too small to price.
  • Prompt sets are curated per brand and differ by segment, so part of the gap may reflect prompt difficulty rather than market conditions. The unbranded-only control reduces the gap honestly and does not eliminate the concern.
  • Client prompt sets are curated, so these presence rates are not market rates. Our anonymous cold-scan corpus indicates market rates are three to five times lower.
  • All lever findings in this work, including the source-diet and editorial-coverage comparisons, are correlational rather than causal. Strong brands are both cited more and mentioned more, and nothing here separates the two.
  • The tracked-competitor comparison uses competitors the brand named itself. Those lists tend to name aspirational rivals rather than the entities actually crowding the answers, so the true competitive field is harder than the tracked lens shows.
  • Scan models and our scoring panel both changed on 7 and 8 July 2026. Figures spanning that boundary require a same-prompt cohort control.
  • Country mix is Ireland and Great Britain heavy at 86% of runs. United States dynamics are under-sampled.
  • AI Overviews mixes no-overview-shown with overview-without-brand in the client corpus, so its presence figures should be read as a floor.

AI visibility monitoring across the major AI engines.