
TL;DR: We analyzed 24,035 citations produced by AI assistants across 2,314 answers between June 4 and August 10, 2026. 79.7% of everything AI models cite is a page the cited company owns and can edit today — 62.5% ordinary company websites plus 17.2% vendor blogs and resource pages. Reddit accounts for 3.5% of citations and YouTube for 3.4%. The standard "get on Reddit" advice optimizes a channel roughly twenty times smaller than the one you already control.
The dataset
Every number below comes from our own tracking pipeline: we ask AI assistants the kinds of questions buyers ask ("best X for Y", "A vs B", "what is Z"), collect the grounded answers, and record every source the model cites.
- 24,035 citations across 2,314 analyzed answers, June 4 – August 10, 2026
- 1,626 distinct prompts over 197 websites spanning SaaS, e-commerce, agencies, developer tools and media — tracked through our own onboarding pipeline for QA, not customer data
- 5,770 distinct cited hosts
- Answers carry a mean of 10.4 citations; 14.3% of answers cite nothing at all
- Proportions below carry 95% Wilson confidence intervals where we make a comparative claim
This is not a random sample of real user queries — the prompts were generated by our pipeline for a fixed set of brands. We state that up front because most studies in this space do not.
Finding 1 — AI answers are built on pages vendors control
The share of citations by source kind:
| Source kind | Share of 24,035 citations |
|---|---|
| Company websites (product, pricing, feature pages) | 62.5% |
| Vendor blogs and resource pages | 17.2% |
| The tracked brand's own site | 3.8% |
| Forums and communities (Reddit is 96% of this) | 3.5% |
| Video (entirely YouTube) | 3.4% |
| Blog platforms (Medium, Dev.to, Substack) | 2.9% |
| Directories and review sites (G2, Capterra, Clutch) | 2.7% |
| News media | 2.4% |
| Search engines, code repositories, documentation | 1.5% combined |
Add the first two rows and the picture inverts most GEO advice: ~80% of what AI models cite is vendor-controlled surface. Not Reddit threads, not review aggregators, not press — ordinary pages that some company published and can rewrite this afternoon.
The practical reading: before investing in community seeding or digital PR, make the pages you already own citable — answer-shaped openings, comparison tables, visible dates, structured data. You can measure how citable a given page is with our free Citation Probability Checker.
Finding 2 — Reddit and YouTube are different instruments
The two loudest names in GEO advice sit at the top of the host table, but they behave differently:
| Host | Citations | Answers containing it |
|---|---|---|
| reddit.com | 813 | 487 (21% of answers) |
| youtube.com | 811 | 767 (33% of answers) |
Nearly identical citation counts — very different reach. Reddit gets cited several times within an answer; YouTube gets cited once across more answers. Any "top sources" chart ranked by raw citation count silently understates YouTube by a third.
Finding 3 — brand confusion concentrates in comparison questions
We flag an answer as a brand mismatch when the model names the brand but its citations point at a different entity — a look-alike domain or a namesake. Rates by prompt type:
| Prompt type | Mismatch rate | 95% CI | n |
|---|---|---|---|
| Organic ("best tools for…") | 1.2% | 0.7–1.9% | 1,446 |
| Comparison ("A vs B") | 10.9% | 8.4–13.9% | 496 |
| Brand-specific ("is X good?") | 12.6% | 9.6–16.4% | 364 |
Comparison and brand-specific questions confuse models roughly ten times more often than organic ones, and the intervals against organic do not overlap. (Comparison vs brand-specific is statistically inseparable in our data, so we make no claim about which of the two is worse.)
The mechanism is intuitive: an organic question lets the model talk about whoever it wants; a question naming your brand forces it to resolve your identity — and that is where look-alike domains and namesakes intercept the citation. If your brand shares its name with anything, this is your risk surface. Our free Entity Confusion Checker runs this exact test for any brand.
What we refused to publish
Three cuts of this data did not survive our own bar, and we think saying so matters more than a longer findings list:
- Platform-vs-platform comparisons. On a matched set of prompts answered by four or more platforms, all ten pairwise mismatch-rate comparisons have overlapping confidence intervals. Claims like "model X confuses brands twice as often as model Y" do not survive the data we have — our non-Gemini samples are ~37 answers each.
- Per-website rates. 184 of 197 sites fall under our 30-answer minimum. A mismatch rate computed on 12 answers is a coin flip, not a finding.
- Month-over-month trends. The mix of tracked brands changed sharply between months, so an apparent trend is a sampling artifact until measured on a stable panel.
What this means for your GEO strategy
- Your own pages are the primary citation surface. ~80% of citations point at vendor-controlled pages. Comparison pages, honest feature tables, dated updates and schema markup on your domain are the highest-leverage work — the same pages Google leans on, which you can verify one query at a time by checking whether a search returns an AI Overview and whose domains it cites.
- Community is real but small. Reddit's 3.5% is worth having — it is not worth being your strategy.
- If your name is ambiguous, comparison queries are where you lose. A ten-fold concentration of confusion in brand-naming questions means disambiguation (consistent entity data, an about page, llms.txt) protects exactly the queries where buyers compare you.
- Expect ~10 citations per answer. You are competing for one of ten slots, not one of one — being a secondary source in many answers beats being absent from all of them.
Methodology notes
Answers were collected through grounded AI assistants configured to search the live web; citations are the URLs the models themselves attached to their answers. Source kinds were assigned by a rule-based classifier over the registrable domain and path (vendor blog paths counted separately from product pages). "The tracked brand's own site" means citations pointing at the brand each prompt was about. All brands in the corpus were tracked through our own service accounts; no customer data is included. Proportions are shares of total citations unless stated otherwise; comparative claims carry 95% Wilson intervals computed on answer-level denominators.
We will re-run this study on a stable daily panel; if the numbers move, the update will say so.


