← All research
−15% to −74%

The July model event

Identical prompts lost 15% to 74% of their presence across five engines after a July 2026 model change, while the uncontrolled trend line rose.

Report period
15 April – 21 August 2026
Sample
22,012 runs across five engines
Confidence
Moderate

Identical prompts, run before and after a model change on 7 July 2026, lost between 15% and 74% of their presence across five answer engines. Over the same four months the naive trend line rose. Both readings come from the same database, and only one of them is true.

What we measured

Two model changes are recorded in our own logs on 7 July 2026: the ChatGPT scans moved from GPT-4o to GPT-5.5, and the Claude scans moved from Haiku 4.5 to Sonnet 5. Gemini, Perplexity and Google AI Overviews had no recorded change on our side.

To test what the change did, we took only the prompts that ran in both windows, 1 June to 6 July and 8 July to 21 August, and compared presence on that fixed cohort. Same prompts, same brands, before and after. Anything that changed is a property of the answers, not of who we were tracking.

EnginePresence beforePresence afterRelative change
Google AI Overviews22.0%5.8%−74%
Perplexity33.9%16.8%−50%
Claude44.9%28.4%−37%
Gemini59.8%40.9%−32%
ChatGPT49.6%42.3%−15%

Every engine got harder to win on the same questions. The two engines with a recorded model change moved, which is expected. The three without a recorded change moved the same way, which is not, and which makes an ecosystem-wide shift in answer style the most economical explanation for a five-engine coincidence in one week.

The answers grew while the brands shrank

On the same fixed cohort, the number of distinct brands named per answer rose sharply.

EngineBrands per answer beforeBrands per answer afterMultiple
Claude3.08.02.7x
Perplexity2.95.92.0x
Gemini5.410.31.9x
ChatGPT4.07.21.8x
Google AI Overviews3.24.61.4x

The share of answers naming any brand at all rose too, on four of the five engines. So answers did not get shyer about naming companies. They got longer and denser, and the new slots went to the incumbent field rather than to the brands we track. Longer lists, fewer of our brands in them.

The composition illusion

Here is the part that matters more than the event itself.

Across the same four months, our uncontrolled weekly presence figure rose from around 21% to somewhere between 41% and 59%, depending on the engine. A dashboard reading that number would have reported a strong, sustained improvement. The like-for-like measurement over exactly the same period fell on every engine.

Both numbers are correct. The uncontrolled figure rose because the mix changed: new brands were onboarded over the period, with prompt sets that fitted them better. The average moved because the population moved. This is Simpson's paradox arriving in a production analytics product, and it is the single most dangerous failure mode in this category.

The general principle is the composition illusion. When the number of brands named in an answer grows, a brand's share of that answer can fall without anything about the brand changing at all. A visibility drop may be a change in the composition of the answer rather than a change in the brand's standing. The reverse holds too: a visibility rise may be a change in the composition of the measurement rather than a change in the brand's standing.

There is a second-order consequence for the metric itself. If answers routinely name ten brands where they once named five, being named is a weaker outcome than it was. Readers still act on the first two or three names. Binary presence therefore becomes a less honest metric over time, and position and exclusivity have to carry more of the weight. That is a metric-design problem arriving on a schedule.

What it means for a brand

A point-in-time audit cannot tell you what you want to know. If you commissioned an AI visibility audit in June 2026 and acted on it in August, you were acting on a picture of a different world, and nothing in the audit would have told you so. The ground moved under it. This is not a criticism of any particular audit; it is a property of measuring a system whose behaviour is reset by releases you are not told about in advance.

What separates a real change from a model change is not more data. It is the right comparison. Three disciplines follow from this event, and we apply all three.

  1. Cohort-control every trend claim. Any statement that a brand's visibility improved, made on a prompt set that changed underneath the series, is not credible. It is also the kind of claim that survives right up until a customer checks it.
  2. Treat model releases as measurement events. A confirmed release, with a similar magnitude of change visible across other brands in the category, reads as a platform-driven shift rather than a brand-specific performance signal. It should be reported that way, and the two regimes should be reported as separate series rather than stitched into one continuous line.
  3. Report rolling multi-week bands, not day-over-day deltas. Weekly presence in our corpus swings by 15 to 25 points. Stability only emerges at three to four weeks of aggregation. A daily delta is noise with a chart around it.

The commercial conclusion follows directly. Repeated measurement with delta detection is not a nicer version of a one-off audit. It is the only configuration that can distinguish a change you caused from a change that was done to you, and the difference between those two determines whether your next quarter of work is aimed at anything real.

What would change our mind

We want to be careful about what this is. It is one event. A single dated model transition, well instrumented and cohort-controlled, on five engines. That is a good deal more than the anecdotes circulating in this category, and it is still one observation. Two observations of the pattern, upgrade followed by tighter like-for-like visibility, would make it a general claim. One makes it a well-measured incident.

There are specific ways it could be less than it looks. The new model generation simply writes longer answers, so the brands-per-answer figure partly measures verbosity rather than a market shift. The AI Overviews collapse from 22.0% to 5.8% may partly reflect Google serving overviews differently over the summer rather than any model effect. The two cohort windows carry unequal run counts. And three of the five engines had no recorded model change, which is either evidence of something broader or evidence that upstream changes happened and nobody logged them.

The next major model release is the test. We will re-run the same cohort analysis against it and publish the result whichever way it points.

How we measured this

Corpus. 11,554 completed production prompt runs and 10,458 staging runs, between 15 April and 21 August 2026, across five answer engines: ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews. Approximately 110 own brands, 895 prompts, 65 anonymous cold scans. 22,012 runs in total, from a database snapshot taken on 22 August 2026 and queried read-only.

Cohort construction. The before-and-after comparison uses only prompts with at least one completed run in both windows, 1 June to 6 July 2026 and 8 July to 21 August 2026. Presence was then computed per engine within each window on that fixed prompt set. This removes the effect of brands and prompts being added or removed over the period, which is the mechanism behind the uncontrolled trend pointing the opposite way.

Recorded events. Two scan-model changes are logged on 7 July 2026: ChatGPT GPT-4o to GPT-5.5, and Claude Haiku 4.5 to Sonnet 5. Our own scoring panel moved from version 1 to version 2 on 8 July 2026, which is a second discontinuity in the same week. Score-level series are therefore treated as two regimes rather than one trend. The run-level presence figures reported here are computed from raw completed runs and are not affected by the scoring panel change. The mention-extraction model was unchanged across the boundary.

Entity resolution. Brand mentions were originally detected by a word-boundary string matcher, which counted any answer containing the brand's name as a mention whether or not the answer concerned that company. On 18 August 2026 we replaced it with a three-state verdict that can return uncertainty rather than forcing a match, and which is deliberately more willing to abstain than to claim.

Replication. Staging is a separate environment with an overlapping but distinct brand set and independent runs. It is used as a replication check on the shape of findings, never pooled with production numbers.

What this does not show

  • 52.5% of production own-brand telemetry sits under an internal pilot account tracking roughly eleven real brands. The AI answers observed are entirely real, but the row counts are not customer traction, and this qualifies every commercial reading of the corpus.
  • This is one model event, not a general law. A second observation of the same pattern would make it a general claim about model releases. One observation makes it a well-instrumented incident.
  • The two cohort windows carry unequal run counts.
  • The new model generation writes longer answers, so brands per answer partly measures verbosity rather than a market shift. Whether denser answers are a durable ecosystem change or an artifact of one model generation is unresolved.
  • Three of the five engines had no recorded model change and moved the same way, which argues either for a broader shift or for unlogged upstream changes.
  • The AI Overviews decline from 22.0% to 5.8% may partly reflect Google serving overviews differently over the summer rather than a model effect. AI Overviews also mixes no-overview-shown with overview-without-brand in this corpus.
  • Our scoring panel changed on 8 July 2026, one day after the model changes, so score-level series straddle two methodology regimes. The run-level presence figures here are unaffected, but no score-level trend should be read across that boundary.
  • All lever findings in this work are correlational, not causal.
  • Prompt sets are curated per brand, so these presence rates are not market rates.
  • Country mix is Ireland and Great Britain heavy at 86% of runs. United States dynamics are under-sampled.

AI visibility monitoring across the major AI engines.