The AI Source Preference Matrix: What We Actually Know About How AI Chooses Sources

Reading Time: 8 minutes

Last updated July 2026

Nobody actually knows how AI chooses sources. That’s exactly why I built this framework. Every week someone publishes another post claiming they have cracked AI search. One says schema markup is the answer. Another says it is backlinks. Someone else insists llms.txt is the future. Most of those claims share one thing in common: they are presented with far more certainty than the evidence supports.

The reality is messier. Google AI Overviews, ChatGPT Search, Perplexity, Claude, Copilot, and Gemini do not publish complete ranking algorithms. Much of what the SEO industry shares today comes from a mix of official documentation, patents, independent testing, observation, and educated inference. Those are not the same thing, and treating them as if they were is where most AI visibility advice goes wrong.

So instead of asking “what is the AI ranking algorithm,” I started asking a different question: how much do we actually know? That question became the research behind this article.

The Biggest Problem in AI Visibility Isn’t Missing Data

It is mixing different kinds of evidence together as if they carry equal weight.

Google officially documents some aspects of crawling and structured data.1 Perplexity openly shows its citations as part of the product experience.4 Independent researchers and practitioners have run prompt tests across platforms, though methodologies vary widely. But nobody outside these companies knows the complete weighting of every signal that goes into a citation decision.

That is why every finding below is classified as one of four evidence types:

  • Documented: officially stated by the platform
  • Observed: consistently seen across independent testing
  • Inferred: our current interpretation based on available evidence
  • Unknown: conflicting evidence or insufficient data

That distinction matters more than any individual platform recommendation you will read this year. It is a philosophy, not just a framework, and it shapes every research question below.

Most AI visibility advice tries to reduce uncertainty. Good research should expose it. Those are not the same goal, and confusing them is how so much confident-sounding advice in this space ends up being wrong.

AI Visibility Is Becoming an Evidence Problem

Five years ago, most SEO debates were about tactics.

Should you build more backlinks? Publish longer articles? Improve Core Web Vitals?

Today, AI visibility debates increasingly revolve around something else: evidence.

Two experienced practitioners can recommend completely opposite strategies because they’re relying on different types of evidence. One may cite official platform documentation. Another points to patent filings. Someone else shares results from their own testing, while another references client campaigns or anecdotal observations.

None of those sources are worthless.

But they aren’t equally reliable either.

That’s why this article doesn’t ask, “Who’s right?” It asks a different question:

What kind of evidence supports this claim?

That single question changed how I evaluate almost every new AI visibility recommendation I come across.

Before you study the matrix, remember what it is—and what it isn’t. This is a synthesis of current evidence, not a reverse-engineered ranking algorithm.

The AI Source Preference Matrix explains how AI engines select sources, highlighting evidence categories, confidence levels, and working hypotheses for source trustworthiness.

Rather than walking through six platforms individually, I wanted to organize the evidence around four questions marketers actually care about.

Research Question 1: Where Do AI Systems Actually Find Information?

This is the strongest area of evidence in our current research, meaning it sits closest to officially stated fact rather than inference.

Google AI Overviews draws from the Google Index, Knowledge Graph, trusted publishers, and official sources.1 ChatGPT Search uses the Bing index along with partner sources, trusted publishers, and real-time web access.2 Claude’s web search pulls from the live web when enabled, plus academic sources and any files the user provides directly, which is a meaningfully different retrieval path than the others.3 Perplexity uses the live web through its own crawler, with a visible lean toward academic sources, forums, news, and documentation.4 Microsoft Copilot blends the Bing index with Microsoft 365 content, trusted partners, and official documentation.5 Gemini uses the Google Index, Knowledge Graph, the open web, and Google’s own product ecosystem.6

Worth a quick note here: Google AI Overviews and Gemini get separate treatment in this research even though they share underlying infrastructure. They serve different use cases and retrieval contexts, which leads to different citation behavior and source weighting, so collapsing them into one column would have hidden more than it revealed.

The practical answer to Research Question 1: if you want visibility across all six systems, you cannot optimize for one index and assume the rest follow. Bing-dependent platforms and Google-dependent platforms are pulling from genuinely different pools, and a strategy built entirely on one will leave real gaps in the other.

Research Question 2: What Content Do They Appear to Cite?

This is classified as observed evidence rather than documented, meaning it comes from consistent patterns across independent testing rather than a platform’s own statement.

A theme shows up across every platform: clear, well-structured, original content wins. Google AI Overviews favors top-ranking, authoritative, fresh, well-structured pages. ChatGPT favors clear answers with original insight, direct comparisons, and current news or documentation. Claude favors in-depth explanations with nuanced perspective, especially when it can weigh multiple points of view. Perplexity is the most citation-forward of the group, consistently favoring forums, studies, and data-backed content. Copilot leans toward Microsoft-ecosystem sources like enterprise sites, whitepapers, and its own documentation. Gemini favors Help Center content, community forums, and generally high-quality sites.

Interestingly, the platforms do not simply reward “high quality content.” They appear to reward different expressions of quality. Perplexity frequently favors cited research. Claude often surfaces nuanced explanations. Google AI Overviews frequently intersects with traditional search authority. ChatGPT often blends authoritative publishers with clear, answer-oriented content. That suggests there may not be a single universal “AI optimized page.” There may never be a single AI-optimized page because the retrieval goals of these systems are fundamentally different.

Layered on top of that observed behavior is what we call a strategic hypothesis, our current best interpretation of what to prioritize. For AI Overviews, that means SEO fundamentals plus topical authority. For ChatGPT, original content plus clear answers plus brand mentions. For Claude, depth and clarity paired with real-world examples. For Perplexity, research data and citations with topical depth. For Copilot, authoritative sourcing plus structured documentation. For Gemini, topical authority plus technical clarity plus freshness.

I want to be direct about what that paragraph is and is not. It is reasoning, not proven fact. It is the part of this research most likely to change as new evidence comes in, and we labeled it that way on purpose instead of dressing up a working theory as a conclusion.

Research Question 3: How Important Is Freshness?

Perplexity rates very high on freshness, which tracks with its heavy reliance on live crawling and its citation-first design. Google AI Overviews, ChatGPT, and Copilot all rate high. Claude and Gemini sit at medium.

If your content strategy leans on evergreen pages that rarely get touched, Perplexity and the other high-freshness platforms are the ones most likely to de-prioritize you over time. This research question is a reasonable guide to where update effort pays off fastest, and where it matters less.

Research Question 4: How Well Do We Actually Understand These Systems?

This is the most important question in the entire analysis, and the honest answer is: not as well as most content in this space implies.

Behavioral understanding sits at medium for AI Overviews, ChatGPT, Perplexity, and Gemini. It drops to low for Claude and Copilot. In plain terms, even after sustained testing and documentation review, our grasp of the actual ranking and selection logic behind these platforms is partial at best, and thin in a couple of cases.

Every platform still carries a real list of open questions. For AI Overviews: exact ranking factors, signal weighting, and the role of the Knowledge Graph versus standard web results. For ChatGPT: retrieval scoring details and how source weighting varies by domain type. For Claude: retrieval scoring, source weighting, and how memory interacts with search influence. For Perplexity: crawler depth and coverage plus personalization impact. For Copilot: how the Microsoft Graph influences results and the weighting of Microsoft 365 sources. For Gemini: how Google blends its various systems and the exact role of real-time data versus the index.

Anyone selling a definitive playbook for exactly how Claude or Copilot rank sources is overstating what is currently knowable. That is not a criticism of those platforms. It is just where the evidence currently stands.

How We Classified the Evidence

Not every claim deserves the same level of confidence.

To reduce our own bias, we classified each finding using a simple decision process we are calling The AI Visibility Evidence Model:

Documented: The platform has explicitly stated or documented the behavior.

Observed: Multiple independent researchers or repeated testing consistently report similar behavior.

Inferred: The conclusion is our interpretation based on the available evidence, but the platform has not confirmed it.

Unknown: Evidence is conflicting, insufficient, or simply doesn’t exist yet.

If a claim couldn’t clearly fit one of those categories, we intentionally treated it as lower confidence rather than higher confidence.

Visual diagram of the AI Visibility Evidence Model showing evidence types, classification, and decision flow for AI source validation.

Not Everyone Agrees About Schema

Schema markup comes up constantly in discussions of AI optimization. Some practitioners believe it directly improves AI citations. Others argue its benefit is largely indirect, working through better search visibility and cleaner structured data rather than any direct citation boost.

There is not enough public evidence right now to say either position is definitively correct. That is exactly why schema shows up as an inferred tactic in our thinking rather than a documented fact. The same caution applies to several other popular claims floating around right now, including specific llms.txt benefits and precise backlink-to-citation correlations. Confident advice on any of these should be read as someone’s current interpretation, not settled science, until the platforms themselves say otherwise or independent testing produces a consistent pattern.

Limitations

This analysis has important limitations. Platform behavior changes frequently. Some systems personalize answers for individual users. Some retrieve live information while others rely more heavily on pretrained knowledge. Many ranking signals remain proprietary, and none of the six companies publish a complete account of how those signals are weighted.

Accordingly, this research should be read as a synthesis of current evidence rather than a definitive explanation of any platform’s internal algorithm.

Five Ways This Research May Change

Here is where I expect the next real shifts to come from:

  • ChatGPT expands its retrieval sources or partnerships again
  • Google changes how AI Overviews sources or weights citations
  • Claude publishes more documentation on search and retrieval behavior
  • Independent testing contradicts one of the current strategic hypotheses
  • Platform weighting shifts as usage patterns and competitive pressure change

Good research should evolve. If our conclusions never change, we are probably not paying attention.

Patterns Worth Watching

Google is the most documented platform, not the easiest to understand. There is a lot of public material on Search Central, patents, and AIO guidance, but documentation availability and behavioral understanding turned out to be two separate things entirely. High transparency did not translate into high clarity on the exact ranking logic.

Perplexity is the most transparent by design, not by disclosure. Its citation-first product structure means you can watch sourcing behavior in real time simply by using the product, which is a very different kind of transparency than a company publishing a whitepaper.

Claude may be the hardest platform to reverse engineer, despite producing excellent answers. Strong output quality and clear behavioral understanding are not the same thing, and this project was a good reminder that you can trust a tool’s answers without being able to explain exactly how it got there.

What I Would Not Over-Optimize For Yet

Based on the current evidence, I would be cautious about treating schema markup, llms.txt, AI keyword density, or any proprietary GEO (Generative Engine Optimization) score as a guaranteed path to citations. Some may help indirectly. Some may prove useful over time. But none should be treated as settled ranking factors across AI answer engines.

What This Means for Your Strategy

A few working rules came out of this research. Understand which sources each platform actually relies on before building a strategy around any single one. Focus effort on what is documented and consistently observed rather than chasing the inferred findings as if they were settled fact. Treat every strategic hypothesis as a hypothesis, not a guarantee, which is why we intentionally left those without confidence scores instead of faking precision we do not have. Build authority in the specific places where each platform is actually looking for trust signals. And retest regularly, because this space moves fast enough that research from six months ago is already stale in places.

The Real Lesson

The biggest lesson from this research was not learning how AI chooses sources. It was realizing how little anyone outside these companies truly knows.

That is not a weakness. That is the current state of the field. The opportunity is not pretending certainty exists where it does not. It is being disciplined enough to separate documented facts from repeated observations, reasonable inferences, and genuine unknowns, and being honest about which category any given claim actually belongs to.

The future of AI visibility probably will not belong to the people making the boldest claims. It will belong to the people willing to update those claims when better evidence arrives. That is the standard we are trying to hold ourselves to.

Uncertainty isn’t a bug in AI visibility research. Pretending uncertainty doesn’t exist is.


1 Google Search Central, “AI Features and Your Website”
2 OpenAI Help Center, “ChatGPT Search”
3 Anthropic, “Web search tool”
4 Perplexity, “Perplexity Search API Quickstart”
5 Microsoft Support, “What information does Copilot use to answer my prompt?”
6 Google AI for Developers, “Grounding with Google Search”

Leave a Comment

Your email address will not be published. Required fields are marked *