Study 001

What sites do ChatGPT, Perplexity and Google cite?

A combined, privacy-filtered index of the domains and URLs exposed as sources across monitored AI answers. It shows source patterns across the panel, not provider market share or provider-by-provider preferences.

45.90M
citation appearances
502K
unique cited domains
1,617
tracked brands
2.47M
unique URLs
Last updated · Aug 18, 2026
[01]

Executive findings

Provenance

Evidence behind the findings

The index is broad but not dominated by one source. Its biggest analytical limit is classification coverage: 86.75% of citation appearances sit outside the current named source taxonomy.

Verified snapshot2026-08-18

48.4M

Citation appearances

Sum of citation appearance counts in the latest valid snapshot for each tracked brand.

Verified snapshot2026-08-18

529.2K

Unique cited domains

Distinct normalized domains present in the current brand citation snapshots.

Verified snapshot2026-08-18

7%

Top three domain share

Share of citation appearances attributed to the three most-cited domains in the current snapshot.

Verified snapshot2026-08-18

86.75%

Unclassified source share

Citation share assigned to “other” because the source domain is not in the hand-maintained category map.

Ten most-cited domains
youtube.com3.47%

1,680,655 appearances

en.wikipedia.org1.83%

886,607 appearances

reddit.com1.74%

842,349 appearances

linkedin.com0.64%

309,013 appearances

google.com0.62%

301,892 appearances

facebook.com0.58%

280,486 appearances

pmc.ncbi.nlm.nih.gov0.42%

201,135 appearances

instagram.com0.35%

171,552 appearances

forbes.com0.34%

165,230 appearances

edmunds.com0.31%

150,833 appearances

How to read this

Count: one source can appear repeatedly across responses and tracked brands, so citation appearances exceed unique URLs.

Concentration: the top three domains hold 7%. This describes the tracked panel, not the full web.

Selection: the index reflects prompts tracked for participating brands. It is not a random sample of all questions asked of AI systems.

Attribution: model and geography fields are not retained in this public aggregate. They are not inferred.

[02]

Press summary and practical use

Citation appearances

48,422,016

Unique cited URLs

2,602,557

Top-three share

7%

Unclassified share

86.75%

The short version for journalists and analysts

The current index combines citation appearances observed across monitored ChatGPT Search, Perplexity, and Google AI surfaces. YouTube, English Wikipedia, and Reddit lead the domain table, but together the top three account for only 7% of appearances.

The long tail is the larger story: the snapshot covers 529,165 unique cited domains and 2,602,557 URLs across 1,643 tracked brands.

This public aggregate does not retain provider attribution. It cannot say that ChatGPT prefers one source while Perplexity or Google prefers another, and it cannot establish that a page trait caused a citation.

Use the data without overclaiming

  1. Search the domain table to see whether owned, competitor, publisher, community, or reference sites appear.
  2. Use the separate page-type study to inspect formats. Raw page URLs are not published in this privacy-filtered aggregate.
  3. Track the same prompt set over time before treating a source gap as an action target.
[03]

Domain index

The explorer covers the top 4,301 domains in the public slim export. Rows with fewer than 3 contributing brands are excluded here and from the download.

4,301 privacy-qualified domains in the slim exportMinimum 3 brands
AI Citation Index domain ranking
RankDomainCitationsShareURLsBrandsType
1youtube.com1,680,6553.47%99,2551,137Social / UGC
2en.wikipedia.org886,6071.83%23,5331,549Reference
3reddit.com842,3491.74%48,5681,321Social / UGC
4linkedin.com309,0130.64%22,3591,119Social / UGC
5google.com301,8920.62%26,410565Other
6facebook.com280,4860.58%18,449910Social / UGC
7pmc.ncbi.nlm.nih.gov201,1350.42%10,187423Academic
8instagram.com171,5520.35%13,759804Social / UGC
9forbes.com165,2300.34%6,245928News & Media
10edmunds.com150,8330.31%5,813101Other
11gartner.com149,8860.31%6,800499Other
12g2.com147,1280.30%8,225609Review Sites
13cars.com115,9920.24%4,917109Other
14techradar.com101,7980.21%2,242687Tech Publications
15clutch.co101,5020.21%3,596276Other
16autotrader.com87,9560.18%3,85589Other
17carfax.com86,0260.18%2,67793Other
18yelp.com84,3120.17%3,312279Review Sites
19kbb.com82,1060.17%2,552113Other
20cargurus.com79,6740.16%3,61996Other
21nerdwallet.com69,8130.14%2,376257Other
22tripadvisor.com68,2840.14%3,055116Review Sites
23medium.com63,3430.13%5,291657Social / UGC
24truecar.com57,4920.12%1,81787Other
25f6s.com56,0320.12%2,999556Other
26dealerrater.com55,8650.12%1,55083Other
27play.google.com55,7110.12%3,800361App Stores
28sciencedirect.com55,5960.11%4,014395Academic
29wifitalents.com53,9550.11%3,498443Other
30apps.apple.com53,3170.11%3,531426App Stores
31caranddriver.com53,1630.11%1,456121Other
32consumerreports.org50,8310.10%1,492222Review Sites
33dqsglobal.com45,4510.09%1,49516Other
34gitnux.org44,5640.09%2,985426Other
35salesforce.com42,4340.09%1,783326Other
36capterra.com41,0130.08%2,634376Review Sites
37worldmetrics.org39,3580.08%2,546402Other
38ibm.com38,1660.08%2,040269Other
39caredge.com36,3020.07%72787Other
40quora.com35,9260.07%2,720539Social / UGC
41mandg.com35,2620.07%1,01110Other
42m.yelp.com35,2130.07%2,173202Review Sites
43trmlabs.com34,4930.07%35915Other
44designrush.com34,3040.07%1,420217Other
45zapier.com34,0650.07%926303Other
46healthline.com33,6060.07%1,224151Other
47learn.microsoft.com32,9160.07%2,141206Official Docs
48elliptic.co32,8960.07%34512Other
49tiktok.com31,9750.07%2,318356Social / UGC
50learn.g2.com31,4260.06%660375Review Sites
[04]

Source types

Named types cover a minority of the index

The classifier uses a hand-maintained domain map. Social sources account for 7.2% and reference sources for 1.98% in the current snapshot, but “other” still holds 86.75%. That makes named category comparisons directional, not comprehensive.

[05]

Prompt intent

Prompts are assigned to rule-based intent groups. “Best of” represents 33% of classified prompt records, while 51.9% remain “other”. The index publishes aggregate prompt counts, never prompt text.

Based on 42,200 unique prompts
Best Of33.0% of queries

"Best of" and "top X" queries often surface review sites and curated lists. Reddit and user-generated content sites perform well here.

Top sources
youtube.comen.wikipedia.orgreddit.comforbes.comgartner.comwirecutter.comtechradar.comcnet.comtomsguide.compcmag.comengadget.comtheverge.comgizmodo.comzdnet.comtrustpilot.comg2.comcapterra.comsoftwareadvice.comgetapp.comproducthunt.comalternativeto.netslant.cotrustradius.comforrester.comconsumerreports.orggoodhousekeeping.combuzzfeed.combusinessinsider.comwired.comlifehacker.commakeuseof.comhowtogeek.comdigitaltrends.comlaptopmag.comandroidauthority.com9to5mac.commacrumors.comimore.comwindowscentral.comandroidcentral.comxda-developers.comslickdeals.netnytimes.comwsj.comtheguardian.combbc.combloomberg.comtechcrunch.comarstechnica.comanandtech.comtomshardware.comnotebookcheck.netrtings.comsoundguys.comwhatifi.comreviewed.combestreviews.comthespruce.combobvila.comfamilyhandyman.compopularmechanics.comwired.co.uktechspot.comguru3d.comoverclock.netlinus.techgamersNexus.neteurogamer.netpolygon.comkotaku.com
[06]

Index history

Daily points are plotted as observed. Missing snapshots remain gaps. Movement can reflect new citation behavior, a changed tracked panel, or both, so the chart is descriptive rather than causal.

Loading observed historical snapshots

The verified current snapshot and aggregate CSV remain available above. Historical points are loaded separately, and missing dates are never interpolated.

[07]

Coverage and limits

What this dataset can answer

Cited domains and source shareAvailable

Available for the 5,000 domains in the public slim export. The download applies a minimum of three contributing brands.

Source categoriesLimited

Available through a hand-maintained domain map, but most citations remain in “other”. Category shares are partial.

Prompt intentLimited

Available as aggregate prompt counts across six rule-based intent groups. Raw prompts are not published.

Historical movementLimited

Daily aggregate snapshots show index movement. Changes can reflect both citation behavior and panel composition.

Model differencesNot available

Provider attribution is not retained in this public aggregate, so the index does not claim model-by-model source differences.

Response citation rateNot available

The public index has citation appearances but no denominator for all responses, including responses with no citations.

Page typeNot available

This export holds domain and URL counts, not page classifications. A separate reviewed page-type study is linked below.

Geography and languageNot available

Prompt geography and language are not preserved in the public aggregate.

Causal ranking factorsNot available

This is observational data. It cannot establish that a source trait caused a citation.

[08]

Method and evidence

Methodology

Method and dataset coverage

AI Citation Index

Verified snapshot

Verified 2026-08-18

Records
48,422,016 citation appearances, 2,602,557 unique URLs, 529,165 domains, and 1,643 tracked brands
Dates
2025-10-03 to 2026-08-18
Sampling
Latest valid citation snapshot per tracked brand. Counts are appearance counts within those snapshots, not a census of all AI answers.
Model scope
The public aggregate does not retain provider-level attribution, so model comparisons and response-level citation rates are unavailable.
Geography
Prompt and response geography is not retained in the public aggregate.
Privacy
Downloads include aggregate domain rows seen across at least 3 contributing brands. Raw URLs and customer data are excluded.
Not included
Raw prompts, Raw responses, Customer identities, Account identifiers, Model-level denominators, Geography

Metric ledger

Open any metric to inspect its calculation, sample, scope, and limits.

Definition
Sum of citation appearance counts in the latest valid snapshot for each tracked brand.
Numerator
Not applicable
Denominator
Not applicable
Sample
1,643 tracked brands
Dates
2025-10-03 to 2026-08-18
Model scope
The public aggregate does not retain provider-level attribution, so model comparisons and response-level citation rates are unavailable.
Geography
Prompt and response geography is not retained in the public aggregate.
Limits and confidence
Not unique citations and not a census of all AI answers. A URL can appear more than once. Direct aggregate count from the verified public snapshot.
Review
Verified 2026-08-18. Reviewed by Mack Grenfell.

Download the aggregate data

Aggregate, privacy-filtered files. No customer identities, prompts, account IDs, raw URLs, or absolute traffic volumes.

Definitions

Citation appearance
One observed instance of a source URL appearing in a tracked AI response snapshot. It is not necessarily a unique URL or a unique answer.
Unique cited domain
A normalized hostname with at least one citation appearance in the current brand snapshots.
Verified snapshot
A point-in-time value regenerated from an attributable public aggregate and checked against its numerator, denominator, and coverage metadata.

Update history

2026-08-18

Flagship evidence refresh

Rebuilt the three public research surfaces from 48,422,016 citation appearances and 4,928 qualifying GA4 properties.

2026-05-06

Markdown experiment added

Added the randomized 9,033-page crawler experiment as a reviewed State of AI Search dataset.

2026-03-30

Longitudinal citation study added

Added citation persistence and volatility findings with an explicit 177-day citation observation window.

Corrections

2026-08-18

Citation headline and Wikipedia share

Removed the stale 1.3M citation and 17% Wikipedia claims. The page now resolves current values from the verified Citation Index and records the snapshot date.

2026-08-18

Traffic loading values

Removed invented placeholder growth and source-share values. Loading and failed requests now render as unavailable.

2026-08-18

Live-data description

Replaced “every number is live” with dataset-level status because several reviewed studies are point-in-time snapshots.

Author and reviewer: Mack Grenfell. Generated 18 August 2026 at 09:02 UTC. To refresh: run npm run research:build in the frontend directory, then npm run research:validate and npm run privacy:public-research.