Skip to content

The neutral index of AI-visibility & GEO tools

Research report · geo

How AI answer engines cite sources: what the studies show

What Profound, Otterly and Ahrefs citation studies actually measure, where their samples and denominators differ, and what they do not prove about AI source selection.

Jul 16, 2026updated Sep 5, 20268 min readSource-linked research

Marketers asking how to get cited by AI answer engines face two different questions: what sources appeared in a particular sample, and why an engine selected them. Published citation studies help with the first. They do not, on their own, reveal a complete source-selection algorithm or prove that copying a pattern will cause a brand to be cited.

This article brings together selected findings from Profound, Otterly and Ahrefs. Source dates, populations and denominators matter. These are separate observational studies, not a controlled experiment or a single comparable time series.

The short version

In Profound’s analysis of 680 million citations from August 2024 through June 2025, Wikipedia accounted for 47.9% of citations among ChatGPT’s ten leading source domains, but 7.8% of all its citations. Reddit accounted for 46.7% of Perplexity’s top-ten-source citations, but 6.6% of all citations. Those different denominators must stay attached to the numbers.

Otterly reports different brand-owned-domain shares across three engines, but its overall brand/community split is internally inconsistent. Ahrefs reports about 76% top-ten overlap in its July 2025 study and 37.1% top-ten organic overlap in its March 2026 study. The samples and parsing changed, so that contrast is not proof that Google’s behavior alone halved the overlap. Measure the engines and query panel relevant to you, with clear definitions, rather than assume one source mix or headline percentage applies everywhere.

Different source concentrations in Profound’s sample

Profound’s AI Platform Citation Patterns describes 680 million citations collected from August 2024 to June 2025. The page was published in June 2025 and updated in August 2025. Its figures describe the sampled citation distribution, not the probability that an engine will cite a source for an arbitrary question.

Engine Leading source: share among its top-ten source domains That source across all its citations Additional reported context
ChatGPT Wikipedia: 47.9% Wikipedia: 7.8% .com domains: 80.4%; .org: 11.3% of citations
Perplexity Reddit: 46.7% Reddit: 6.6% YouTube: 2.0% of all citations
Google AI Overviews Reddit: 21.0% Reddit: 2.2% Wikipedia: 0.6% of all citations

“47.9%” does not mean that nearly half of all ChatGPT citations went to Wikipedia. It is Wikipedia’s share within the top-ten-source subset. The 7.8% figure uses the broader denominator. Both can be true while the long tail remains large and commercial domains account for most citations.

The table also does not establish that an engine reaches for Wikipedia or Reddit first. Citation concentration is an observed output pattern, not a trace of retrieval order. Topic mix, query selection, language, location, time and classification choices can affect aggregate shares. A high domain share may justify examining that domain’s actual coverage of your subject; it does not establish a universal optimization priority or a causal benefit from editing it.

Brand versus community: useful engine detail, inconsistent overall totals

Otterly’s AI Citations Report 2026 says in its summary that it analyzed over one million citations in January–February 2026 across ChatGPT, Perplexity and Google AI Overviews. Elsewhere the report uses a broader “2025–2026” description. This article preserves the stated summary window as the vendor’s description, not an independently reconstructed collection log.

The overall split needs a correction: the report’s summary says community 52.5%, brand 47.5%, while its body assigns 52.5% to brands and 47.5% to other sources. Those statements conflict. The previous version of this article silently selected the first one. Neither overall split is treated here as a settled finding, and the engine rows below should not be reverse-engineered into a total without the engine populations and weighting.

Engine Brand-owned-domain share reported by Otterly
Google AI Overviews 59.8%
ChatGPT 44.7%
Perplexity 28.9%

These are the report’s per-engine figures, not the chance that your brand will be cited, evidence that the same page performs differently across engines, or a matched-query experiment. They support keeping engine-level reporting separate rather than interpreting one pooled share as a universal source mix. Whether a brand should focus on its own site or third-party coverage still requires topic-level evidence.

Otterly also reports AI Overviews on roughly 33% of the queries it tracked. That denominator is its tracked panel. It does not establish the incidence of AI Overviews across all Google searches, so the previous inference about “most searches” globally is withdrawn.

Organic-rank overlap: distinguish a newer sample from a controlled trend

Ahrefs’ July 21, 2025 study analyzed 1.9 million citations from one million AI Overviews, considering the three most visible citations in each response. It reported 76.10% of cited pages ranking in the top ten, 9.50% in positions 11–100 and 14.40% outside the top 100. The median rank across those three cited URLs was 3; the medians for the first, second and third citation positions were 2, 4 and 5 respectively. The previous article’s blanket “median cited position was 4” did not accurately describe that study.

Ahrefs’ March 2, 2026 study analyzed 863,000 keyword SERPs and four million cited URLs. It distinguishes 37.9% in the top ten SERP blocks, which can include features as well as blue links, from 37.1% in the top ten organic blue links. In the organic-ranking breakdown, another 26.2% ranked 11–100 and 36.7% were outside the top 100.

Published study Stated population and metric Reported top-ten overlap
Ahrefs, July 21, 2025 1.9M citations from 1M AI Overviews; three most visible citations per response 76.10%
Ahrefs, March 2, 2026 863K keyword SERPs, 4M cited URLs; organic blue-link positions 37.1%
Same March 2026 study, different ranking definition Top-ten SERP blocks, including features 37.9%

Ahrefs explicitly says it improved citation parsing between the studies. The newer sample therefore differs in measurement as well as date. The large difference is worth investigating, but these tables do not isolate how much came from engine changes, different queries, more complete citation capture or another factor. Calling it a demonstrated six-month collapse in Google’s reliance on organic ranking overstates the evidence.

The March sample shows that many cited URLs were not top-ten organic results for the measured keyword. It does not show that conventional SEO is irrelevant, that those URLs were invisible for related queries, or that a lower organic rank caused citation inclusion. A previously included BrightEdge “17%” comparison is withdrawn because a supporting primary study and comparable metric were not established in this correction; it is not used to corroborate a trend.

What engine documentation adds

Google’s AI-features guidance says AI Overviews and AI Mode may use query fan-out, issuing multiple related searches across subtopics and data sources. It also says the features use different models and techniques, so responses and supporting links can differ. For a supporting link, a page must be indexed and eligible to appear with a snippet; Google says there are no additional technical requirements for these AI features.

That documentation gives a plausible reason why the original keyword’s blue-link ranking need not enumerate every supporting page. It does not prove that fan-out explains a specified share of the Ahrefs result, and it is not a specification for ChatGPT or Perplexity. Nor do these source-distribution studies establish that a particular schema type, word count, Wikipedia edit or forum post guarantees selection.

What this means for a measurement plan

  1. Keep engine and answer surface visible. Report ChatGPT, Perplexity and AI Overviews separately before presenting a combined metric. Document queries, locations, language, dates and failed/missing answers behind each denominator.
  2. Inspect cited pages instead of optimizing to a domain league table. A domain share can point to material worth reviewing, but the actual page and claim determine relevance. An emitted citation is not automatic proof that its destination supports the answer. Follow platform rules and disclosure requirements for any third-party contribution; the studies are not a reason to manufacture endorsements.
  3. Repeat a defined panel and record method changes. Preserve raw answers and citation URLs where permitted. Keep prompt selection, engine settings, repetition and classification consistent enough to interpret a change. If a parser or sampling rule changes, mark the break rather than presenting it as an engine effect.
  4. Separate outcomes. Brand mentions, links to your domain, visits and conversions answer different questions. More citations in a sampled panel are not by themselves proof of increased customer demand or revenue caused by a content change.

For practical next steps, see how to rank in ChatGPT, getting cited by Perplexity and measuring AI share of voice. Our State of GEO tools 2026 and Peec vs Profound vs Scrunch vs Otterly discuss measurement products. Those pages have their own scope and dates; this correction does not re-verify them. For another product category, our sister directory The Agents Index is a browsing resource, not evidence that these citation distributions generalize to coding-agent recommendations.

Methodology & September 5, 2026 correction note

This is a synthesis of vendor-published observational research, not original measurement or independent replication. Primary study pages were reviewed for the specifically quoted figures and methodological limits. Corrected top-ten versus all-citation denominators, Ahrefs’ publication/sample/median descriptions, the Otterly overall-split contradiction and overbroad inference from its tracked query panel. Withdrew unsupported BrightEdge corroboration and causal claims that changing source shares reveal an engine algorithm or prove an optimization tactic. Historical sample sizes and dates remain attached to their results; no new citation dataset was collected and no raw vendor dataset was audited. Original publication and stored verification dates are preserved. The modification date records this scoped editorial correction, not fresh execution of the studies or full re-verification of linked guides.

Get the next report

New tools rankings and fresh data reports. One short email, one-click unsubscribe.