Last updated August 2026.
To find out which pages ChatGPT and Perplexity actually cite, you need to run your target queries, parse every source link and in-text attribution in the response, and tabulate results by URL, domain, and page type. That audit tells you two things: which of your own pages earn citations, and which third-party domains are winning credit for queries where your brand should appear.
Most brands skip the second half. That is the expensive mistake.
Why this matters before you touch a single page
According to Previsible’s 2025 AI Discovery Report (covering 19 GA4 properties), AI-referred sessions grew 527% comparing January to May 2025 against the same period in 2024. That growth is concentrated in sources the engines already trust, and those sources are not always brand-owned.
When you audit your citation footprint, you discover the real competitive landscape: not which sites outrank you on Google, but which URLs ChatGPT and Perplexity actually pull from when someone asks about your category. That list includes your own pages, competitor pages, and third-party editorial sources. In that order, roughly, with the third-party sources often dominating.
Before you invest another hour in on-page optimization, you need to know what the current citation map actually looks like.
Step 1: Separate mentions from cited sources
The first distinction to lock in: a mention and a cited source are not the same thing.
A mention is when an AI engine writes your brand name into its response without pointing to a specific URL. Useful for brand awareness, but not auditable at the page level.
A cited source is a URL the engine surfaces as a reference. ChatGPT with Browse shows these as numbered footnotes. Perplexity shows them as source cards at the top of each response. Google AI Overviews shows them inline. Each of these is traceable: you know the exact URL, the domain, and the page type.
For the purposes of citation auditing, focus on cited sources. Mentions matter for brand visibility, but source citations are what you can diagnose and improve.
Step 2: Run the queries and parse the output
Manual audit (small prompt sets)
For fewer than 20 queries, a manual audit is fast enough to be worth starting today.
- Open ChatGPT with Browse enabled (or Perplexity, which always cites sources).
- Enter each target query exactly as a user would ask it.
- Record every source URL in the response. In ChatGPT, expand the footnote list. In Perplexity, copy the source cards.
- Note whether each URL is your domain, a competitor, or a third-party editorial source.
- Run each query at least three times on different days. Citation sets vary between runs.
Tabulate results in a simple spreadsheet: query, cited URL, domain, page type (blog post, comparison page, product page, community thread, etc.), and whether the URL belongs to your brand or a third party.
Automated audit (larger prompt sets)
For 50 or more queries across multiple engines, manual parsing becomes impractical. Several platforms automate this:
| Tool | What it shows | Entry price |
|---|---|---|
| Profound | Citation maps by URL, domain share, prompt volumes from real queries | $99/mo (ChatGPT only); $399/mo full coverage |
| Peec AI | Page-level citation attribution across ChatGPT, Perplexity, Gemini, and AI Overviews | Contact for pricing |
| Semrush AI Toolkit | Citation source breakdown alongside rank tracking | Included in Guru/Business plans |
| Otterly.AI | GEO Audit across 20+ citation-readiness factors per URL | $29/mo (15 prompts); $189/mo Standard |
| Temso | Citation monitoring showing source URLs per engine, plus gap diagnosis and fix workflow | From $89/mo, all 8 engines |
Profound gives you the deepest citation source attribution: visual maps that show exactly which URLs appear in AI answers for your query set, with domain-level and page-level breakdowns. Peec AI goes further with granular attribution across five platforms. For teams already in Semrush, the AI Toolkit surfaces citation sources without requiring a new tool. Otterly.AI audits individual URLs for citation-readiness. Temso closes the full loop from monitoring to gap diagnosis to execution on a single flat subscription.
None of these replace running the queries yourself. Use the automated output to scale what the manual audit reveals.
Step 3: Tabulate by URL, domain, and page type
Once you have the raw citation data, sort it three ways.
By URL
Which specific pages on your domain earn citations? Which specific competitor pages appear consistently? A URL-level view tells you what content is already working (protect and model it) and where a specific page needs improvement (or needs to exist at all).
By domain
What is the share of citations going to your domain versus competitors versus third-party editorial sources? This is the closest thing to a citation share number for your query cluster. See the glossary for the full definition of citation share.
A common pattern: your domain captures a small fraction of citations, competitor domains capture another slice, and the majority goes to third-party sources such as community platforms, review sites, industry publications, and editorial comparison pages.
By page type
Group cited URLs by content format: product pages, blog posts, comparison pages, community threads, review profiles, and so on. This tells you what content types the engine weights most for your query cluster.
For example, a product-comparison query for your category may produce citations that are:
- 40% community platform threads (Reddit, Quora, specialist forums)
- 30% editorial comparison posts on industry blogs
- 20% review platform profiles (G2, Capterra)
- 10% brand-owned content
That breakdown tells you where to invest: not necessarily more product pages, but more presence on the platforms the engine already trusts.
Step 4: Identify the citation gap
The citation gap is the distance between where you appear and where you should appear, given your category relevance.
After tabulating your citation data, ask:
- Which third-party domains appear for queries where your brand is the most relevant answer?
- Are your own pages being cited at all, or is all the credit going to editorial sources?
- For queries where you do appear, does the citation link to your best page for that query, or to a secondary or outdated page?
A critical insight from Profound’s analysis of 100,000 prompts across ChatGPT and Perplexity: only about 11% of cited domains overlap between the two engines. Building citation presence on ChatGPT does not transfer automatically to Perplexity. Each engine has a largely separate source pool.
This means your citation gap is not a single number. It is a per-engine map.
Step 5: Understand why third-party domains win
When a third-party domain consistently earns citations that should belong to your brand, there is almost always a structural reason. The most common patterns:
Community platforms dominate product-comparison queries. According to Profound’s Q2 2025 analysis, Reddit accounts for 46.7% of Perplexity’s top-10 most-cited sources for commercial queries. These sources aggregate real user opinions at scale. Your product page cannot replicate that social proof signal. The lever is earning mentions and links within those community threads.
Editorial comparison pages rank above brand pages for category queries. When someone asks “best tools for X,” AI engines tend to cite editorial rankings rather than vendor pages. An external editorial source that compares your tool against alternatives will often earn the citation ahead of your own site. This is one of the reasons earning mentions on trusted third-party domains carries more weight than optimizing your own pages.
Your pages are not structured for direct-answer retrieval. According to Kevin Indig’s 2026 analysis of ChatGPT citations (reported by Search Engine Land), 44.2% of citations came from the first 30% of a page’s content. Pages that bury the answer get paraphrased out of the citation. Pages that front-load a concise, self-contained answer to the query get quoted verbatim and credited.
Your pages may not be crawlable by AI systems. Some technical configurations block AI crawlers at the CDN or robots.txt level. See /methodology for how we evaluate crawlability as part of citation readiness.
Step 6: Prioritize what to fix
Not every citation gap is equally worth closing. Prioritize based on two factors:
- Query commercial value. Focus on queries that drive buyer decisions, not informational queries where citation volume is high but intent is low.
- Closeable gap. Some third-party domain dominance cannot be displaced by better on-page content. If Reddit threads own the top citations for your query, the path to citations runs through those threads, not a better blog post. Prioritize gaps where your own content has a realistic path to citation.
For each priority query, document:
- The current top-cited URL
- Why that URL earns the citation (content format, depth, directness, third-party authority)
- What you need to change or create to compete for that citation slot
The full GEO tool ranking at /rankings/geo-tools evaluates each platform on how well it supports this diagnosis step, not just the monitoring step.
The citation audit in one table
| Step | What you do | Tool options |
|---|---|---|
| 1. Separate mentions from citations | Define which outputs you are tracking | Manual or any citation tracker |
| 2. Run queries and parse output | Run queries, record every source URL | Manual (small sets); Profound, Peec AI, Semrush, Otterly.AI, Temso (large sets) |
| 3. Tabulate by URL, domain, page type | Group and sort the data | Spreadsheet or platform dashboard |
| 4. Identify the citation gap | Find which third-party domains win your queries | Profound citation maps; Peec AI attribution; Temso gap analysis |
| 5. Understand why third parties win | Diagnose the structural reason | Review content format, crawlability, third-party authority |
| 6. Prioritize fixes | Rank gaps by commercial value and closability | Internal prioritization; Temso fix workflow |
What to track going forward
A one-time citation audit tells you where you stand. A weekly citation pulse tells you whether you are moving.
Set up tracking for your 20 to 30 highest-value queries across at least ChatGPT and Perplexity. Record citation share (what percentage of runs return your domain as a cited source) for each query, each week. Look for two signals:
- Your own citation rate trending up as you publish and earn mentions.
- Third-party domains losing citation share for queries you are now competing on.
Citation share movement at the query level is the most direct signal available that your GEO work is having an effect. Rank-position changes in Google will not tell you this.
Start here
Run your three most important category queries in Perplexity right now. Copy every source URL from the response. Check how many of those URLs belong to your domain. If the answer is zero, or close to it, you have your citation gap in plain view.
The tools above will help you scale that audit and track movement over time. The full GEO platform ranking at /rankings/geo-tools evaluates each option on how well it covers citation source attribution specifically, not just share-of-voice counts.