Back 10 minute read

I Studied When Google And LLM Results Disagreed. Here is What I Found And Shared At Google Summit 2025.

I Studied When Google And LLM Results Disagreed. Here is What I Found And Shared At Google Summit 2025. 10
minute
read

After doing a 6 month long study into AEO and SEO results, I recently shared these findings with live audiences at Google Summit 2025 and First Page Digital Summit 2025. In particular, i was interested in Brands that were consistently ranking first on Google were not being recommended by AI, while other brands that barely appeared on page one were showing up confidently in ChatGPT and Google AI Overviews.

Fabian sharing on a research study done by him on AEO at First Page Summit 2025

Fabian Seow speaking at First Page Digital Summit 2025, sharing his findings from the study done comparing Google Search and LLM results.

Fabian speaking on seo and aeo correlation at google summit 2025

Fabian Seow speaking at Google Summit 2025, sharing his findings from the study done comparing Google Search and LLM results.

At first, this felt like noise. However, once the same mismatches repeated across industries, prompts, and clients, it became clear that discovery itself was fragmenting. Google Search was no longer the only system deciding visibility.

This article documents what we studied across Google Search, Google AI Overviews, and ChatGPT, how we measured agreement and disagreement, and why the idea of a single AEO strategy breaks down the moment you look closely.

What We Studied and Why Granularity Matters

  • 44 commercial topics
  • 560 unique websites
  • 20 industry niches
  • 4 business types
  • Google Search results
  • Google AI Overviews (AIO)
  • ChatGPT answers and citations
  • E-commerce and B2B ecommerce
  • Services and home services
  • Location-based businesses
  • YMYL sectors such as law and preschool

We deliberately designed the study to span multiple industries and prompt types. Early testing showed that high-level averages hide more than they reveal. Once we segmented by niche, business model, and platform, behaviours that initially looked inconsistent started to form clear patterns.

Granularity was not a methodological detail. It was the difference between insight and confusion.

How We Measured Agreement and Divergence

  • Jaccard similarity
  • Spearman rank correlation
  • Top-1 recommendation agreement

Before interpreting outcomes, it was important to define how similarity and disagreement were measured. No single metric captures how recommendation systems behave, so we used multiple measures to reflect different dimensions of agreement.

Jaccard Similarity: Do the Same Brands Even Appear?

  • Overall AI-to-AI overlap: ~0.206
  • B2C ecommerce: ~0.206
  • B2B ecommerce: ~0.356
  • Services: ~0.285
  • YMYL: ~0.101

Jaccard similarity measures whether the same brands appear at all, regardless of order. The low values across most segments showed that ChatGPT and Google AIO often do not even evaluate the same pool of brands. This is especially pronounced in YMYL niches, where overlap is close to zero, indicating that each system relies on different authority ecosystems.

Spearman Rank Correlation: When Brands Overlap, Do They Rank Similarly?

  • Overall rank correlation: ~0.26

Spearman correlation measures whether platforms rank overlapping brands in a similar order. The low correlation indicates that even when systems agree on which brands matter, they often disagree on how strongly they matter.

Top-1 Agreement: Who Actually Wins?

  • ChatGPT #1 = Google #1: ~20–25%
  • Google AIO #1 = Google #1: ~8–11%
  • ChatGPT #1 = AIO #1: 18%

Top-1 agreement reflects real-world impact. Users tend to remember the first recommendation, and the consistently low agreement rates confirm that being the best is now platform-specific.

What Our Study Found: Where AI and Google Actually Align

  • Average AI #1 = Google #1: ~15.6–18%
  • ChatGPT #1 = Google #1: ~20–25% depending on segment
  • Google AIO #1 = Google #1: ~8–11%
  • ChatGPT and AIO agree on #1 brand: 18% of topics
  • Overall rank alignment (Spearman): ~0.26

At a surface level, these numbers suggest weak alignment, but the more important takeaway is that disagreement is the default state rather than an edge case. In most scenarios, Google Search, Google AI Overviews, and ChatGPT surfaced different brands as their primary recommendation, even when evaluating the same prompt.

This tells us that rankings alone are no longer a reliable proxy for visibility. AI systems are not ranking pages in the traditional sense. They are selecting answers based on defensibility, pattern repetition, and evidence availability, which produces fundamentally different outcomes.

Services vs E-commerce: Two Different Systems of Trust

  • Cross-inclusion for services: ~31–32%
  • Cross-inclusion for e-commerce: ~17–18%
  • Higher ChatGPT–AIO agreement in services than e-commerce

When segmented by business model, the divergence became much clearer. Services categories showed higher overlap between AI platforms, suggesting shared logic around reputation, experience, and proof. E-commerce categories, however, showed much weaker overlap and frequent reversals, where brands recommended by AI were not the ones ranking well on Google.

This difference reflects how trust is evaluated. Services rely on outcomes, credibility, and perceived expertise, while e-commerce relies on comparison, validation, and third-party judgement. AI mirrors these trust models closely, which is why applying the same optimisation strategy across both models consistently underperforms.

What Our Study Found: Citations Did Not Depend On Rankings

  • Consumer electronics listicle citation rate: 68–72%
  • Home services listicle reliance (ChatGPT): ~67%
  • Home services brand page reliance (AIO): ~67%
  • Videography commercial page citations: AIO 84.4%, ChatGPT 73.8%
  • Social and GBP citations in lifestyle (AIO): ~23%

Once we shifted focus from answers to citations, platform behaviour became much easier to explain. Google AIO tended to rely more on brand-owned commercial pages and traditional Google-style signals, while ChatGPT leaned more heavily on editorial content, listicles, and aggregators, particularly in services and lifestyle niches.

Even within the same business model, citation behaviour changed by industry. This is why citation targeting must be segmented not only by platform, but also by niche and prompt type. Without this segmentation, brands often invest heavily in sources that the AI platform does not meaningfully value.

High-Trust (YMYL) Niches: Authority as a Gatekeeper

  • Law authority citations: AIO 33.3%, ChatGPT 25.0%
  • Preschool authority citations: AIO 13.6%, ChatGPT 25.0%
  • Dominant sources: government directories, regulators, ranking platforms

YMYL niches behaved differently from almost every other category. In these sectors, authority listings acted as gatekeepers rather than enhancers. If a brand or firm was not present in recognised directories or official registries, AI visibility was inconsistent regardless of content quality.

This reframes AEO in YMYL categories as an ecosystem participation problem rather than a traditional optimisation problem. Content supports authority, but it rarely substitutes for it.

Why AI Rewards the “Average” Brand

  • Next-word prediction as the core decision engine
  • Concept proximity over qualitative judgment
  • Preference for statistically safe, repeatable answers

When people hear that AI uses next-word prediction, it often sounds abstract and disconnected from real-world outcomes. However, this mechanism directly explains why certain brands are repeatedly recommended while others are consistently ignored.

AI does not evaluate brands the way a human reviewer would. It does not assess originality, innovation, or differentiation in a qualitative sense. Instead, it predicts which sequence of words is most likely to appear next based on everything it has seen before, and for recommendation prompts this means selecting brands that most commonly appear alongside a topic with familiar attributes in trusted contexts.

This is where concept proximity becomes critical. Brands that sit closest to the centre of a topic’s conceptual space are easier for AI to predict because they use familiar language, repeat widely accepted attributes, and appear consistently across multiple sources, which reduces uncertainty.

By contrast, brands that differentiate aggressively often fragment their signals. Unique positioning, novel terminology, or niche claims may convert users, but if they appear inconsistently or only in brand-controlled content, they introduce risk that AI systems avoid.

This is why AI tends to reward what looks like the average option. Average does not mean mediocre. It means statistically central and defensible, supported by overlapping evidence that makes the recommendation easy to justify.

The Supernode Effect and Why Distribution Wins

  • Large editorial and lifestyle publications
  • Aggregator and “best of” listicles
  • Competitor-written roundups and rankings
  • Authority and institutional directories

As we analysed citations across prompts and industries, a relatively small group of websites kept appearing repeatedly in AI answers. These sites were not always the strongest in a traditional SEO sense, but they were consistently referenced across many different contexts.

We started referring to these sites as supernodes.

Supernodes act as stable reference hubs for AI systems. Because they appear frequently across different prompts, they become statistically reliable sources that AI can reuse to justify recommendations with low risk.

Supernode websites observed in the study

  • TechRadar
  • RTINGS
  • Legal500
  • Chambers
  • Government and regulatory directories
  • Straits Times
  • Third-party and competitor “best of” roundups

What makes supernodes powerful is repetition rather than authority in isolation. Each additional appearance increases conceptual certainty, making brands easier to recommend again in future answers.

This explains why some brands consistently outperform their apparent authority. Their advantage lies in distribution across trusted hubs rather than superior on-site optimisation, which is why distribution eventually beats optimisation once baseline quality is met.

Industry and Platform Playbooks

Playbook: E-commerce

  • AI rewards expert and niche listicles
  • Retailer and category page presence matters
  • Attribute parity with reviewed competitors is critical

For e-commerce, success comes from aligning product pages with the attributes AI repeatedly sees in expert reviews, while simultaneously securing placements on authoritative listicles that AI uses as evidence. Rankings alone are insufficient without third-party validation.

Playbook: Services

  • AI rewards proof of outcomes and experience
  • Commercial pages matter more than blogs
  • Third-party validation varies by niche

For services, strong commercial pages with portfolios, case studies, and client proof form the foundation, but off-site validation must be selectively layered based on how each platform evaluates the specific service category.

Playbook: Home Services

  • ChatGPT favours third-party and competitor listicles
  • Google AIO favours brand-owned commercial pages

Home services require a balanced strategy. Over-investing in one platform’s signals creates blind spots on the other. Brands that perform best reinforce both listicle presence and on-site proof.

Playbook: YMYL Niches

  • Authority listings are non-negotiable
  • Official registries and rankings dominate citations

In YMYL categories, the first priority is securing authority placements. Content optimisation supports these signals but rarely compensates for missing institutional validation.

Playbook: AEO Analysis by Prompt

  • Extract entities from AI answers
  • Identify overlapping entities in citations
  • Segment by platform and industry
  • Replicate entities across trusted sources
  • Track AI visibility rather than rankings

This prompt-level approach reflects how AI systems actually reason and avoids the false comfort of generic optimisation tactics.

Final Notes From the Field

SEO still matters, but it no longer explains discovery on its own. AI systems reward consistency, distribution, and contextual fit, and they do so differently depending on industry, platform, and prompt type.

There is no one-size-fits-all AEO approach. The brands that win stop asking how to optimise for AI in general and start asking how AI reasons about their specific market, which is where predictable outcomes begin.

If you are interested to get a deep dive into your industry’s prompts’ citations, engage our AI SEO services. First Page Digital has been offering AEO in Singapore for a year now and have built up a proven process for getting you recommended. Contact us today.

Suggested Articles