What a Citation Audit is
A Citation Audit is the GEO equivalent of a baseline keyword report. It's the snapshot — across all five engines — of where a brand appears in buyer answers today, scored against Citation Share.
Without one, every GEO program is flying blind. With one, every move has a measurable starting point. The audit produces a single defensible number — Citation Share — that a CMO can put on a slide, a CFO can attach to a budget line, and an agency can be measured against quarter over quarter.
The Citation Audit is increasingly a media-buying line item, not a PR one. As Ronn Torossian argued in MediaPost: more than a third of US consumers now begin product research inside ChatGPT, Claude, Gemini, Perplexity, or Google AI Overviews — and the 2026 media plan still doesn't have a line for that. Ranking #1 on Google used to correlate with AI visibility at roughly 70%. That correlation has collapsed to under 20%.
The media plan and the visibility plan are now two different documents, and the planner who optimizes against the old correlation is allocating capital against a chart that no longer holds.
Context on the shift in buyer behavior: ChatGPT crossed 800M weekly active users in 2025. Google confirmed AI Overviews reach 1.5B users monthly across 100+ countries. Perplexity crossed 22M monthly active users by mid-2025 and closed a funding round at an $18B valuation. Gartner projects a 25% decline in traditional search engine volume by 2026 as generative AI chatbots absorb the query. The buyer moved. The measurement stack hasn't.
The CMO question for 2026 board meetings is the only one that matters: What is our share of the answer? The Citation Audit produces that number.
The six steps
Step 1 — Build the prompt set
50 to 100 buyer questions in the brand's category. Not generic queries. The actual questions a buyer types into ChatGPT before making a decision. Under 50 and the sample is noise; over 100 and marginal signal drops off. The tight range is deliberate.
Sources for the prompt set:
- Sales team — what prospects ask on first calls. Pull from Gong, Chorus, or raw call recordings.
- Customer success — what new customers needed to know before they bought. Onboarding-call transcripts and first-30-day support tickets are the richest seam.
- Google Search Console — high-volume queries in the category. Export the Performance report and filter by question modifiers (what, how, why, best, vs).
- SparkToro and AnswerThePublic — buyer intent data. AlsoAsked and Ahrefs' "Questions" report are strong supplements.
- Reddit, Quora — long-tail comparative questions. Reddit especially: it's in the training data for every major model and it's where comparative queries originate.
- Competitor review sites — G2, Capterra, Trustpilot, Sephora, Amazon Q&A. Every "does X work for Y" question is a prompt candidate.
Cover all four query types: definitional ("what is X"), comparative ("X vs Y"), procedural ("how do I X"), recommendation ("best X for Y"). Weight the mix toward recommendation queries — that's where purchase intent concentrates.
Worked example — B2B SaaS category (project management software):
Distribution across four query types for a 60-prompt audit:
- Definitional (10 prompts, 17%): "What is a Kanban board?" · "What is agile project management?" · "What is a Gantt chart?" · "Difference between task and project management?"
- Comparative (18 prompts, 30%): "Asana vs Monday.com" · "ClickUp vs Notion for teams" · "Best Jira alternatives" · "Trello vs Asana for small business"
- Procedural (12 prompts, 20%): "How to set up a project timeline in Monday" · "How to migrate from Trello to Asana" · "How to build a sprint board"
- Recommendation (20 prompts, 33%): "Best project management tool for remote teams" · "Best PM software for creative agencies under $15/user" · "Top project management tools 2026" · "Project management tool for engineering leaders"
The recommendation weighting is deliberate — that's where the purchase decision compresses. Definitional prompts anchor entity understanding; comparative and procedural surface competitor positioning; recommendation captures shortlist inclusion.
Step 2 — Run the prompts
Each prompt, each engine. Five engines: ChatGPT, Claude, Perplexity, Gemini, Google AI Overviews. That's 250–500 calls per audit. Run each prompt at least twice per engine — LLM outputs vary run-to-run, and a single call is not a signal.
Automate via API where possible: OpenAI Responses API, Anthropic Messages API, Perplexity Sonar API, Gemini API, SerpApi for Google AI Overviews.
For engines with browsing enabled (ChatGPT with search on, Perplexity, Gemini with grounding, Claude with the web search tool), enable web access — that is where citations get emitted. Log the full response object: model version, timestamp, tool calls, source URLs. Record every answer verbatim. Don't summarize. The verbatim record is the audit trail.
Budget note: at retail API rates, a 500-call audit across the five engines runs roughly $40–$120 in raw inference costs. The labor cost — scoring, diagnosis, planning — is where the hours go.
Step 2b — Run-to-run drift by engine
The five engines drift differently across identical prompts, and a serious audit accounts for each drift profile in its scoring. This is why single-run scoring produces a screenshot, not a measurement.
- ChatGPT — highest drift with browsing enabled; source URL sets can differ 40–60% between two calls of the same prompt in the same hour. Averaging across ≥3 runs is the floor for a defensible score. With browsing off, drift compresses to ~10–20% (paraphrase variation only, no new sources).
- Claude — lowest drift on entity-list prompts; roughly 10–15% variance in cited-source set. Prose framing varies more than the source set does. Two runs is usually sufficient for a stable score.
- Perplexity — moderate drift with a specific signature: the top 3 citations tend to be stable, positions 4–10 rotate substantially. If you're scoring lead-position citations, two runs is fine. If you're scoring full source-set inclusion, three or more.
- Gemini — moderate drift with a Google-grounding overlay; drift is higher when the query has fresh news attached (retrieval is time-sensitive) and lower on evergreen definitional prompts.
- Google AI Overviews — drift is a function of AIO's own eligibility model, not just model temperature. On some prompts AIO doesn't fire at all in a re-run; on others the source set is nearly identical. Track AIO trigger rate separately from cited source stability.
The practical rule: three runs per prompt per engine is the audit-grade floor. Two runs is defensible on tight budgets. One run is a demo, not a measurement.
Step 3 — Score the results
For each answer, mark:
- Brand cited (yes/no).
- Position of citation (lead, mid, footnote). Lead-position citations carry roughly 3× the click-through weight of footnote mentions in Perplexity source panels.
- Source quoted (which page, which property). Track the exact URL — this tells you which of your pages the model is pulling from.
- Competitors cited. Log every competitor in every answer, not just the top one.
- Citation sentiment (neutral, positive, negative). A negative citation is worse than no citation.
Aggregate into Citation Share — weighted across the five components: Citation Frequency (40%), Cross-Engine Breadth (20%), Query-Type Breadth (20%), Extractability (15%), Crawl Access (5%). This is the locked scoring formula, now on the record in MediaPost.
Component definitions. Citation Frequency: share of prompts in which the brand is mentioned at all. Cross-Engine Breadth: coverage across all five engines, not concentration in one. Query-Type Breadth: coverage across definitional / comparative / procedural / recommendation. Extractability: whether the page structure — headings, schema, clean HTML — lets the model quote you cleanly. Crawl Access: whether GPTBot, ClaudeBot, PerplexityBot, Google-Extended, and Applebot can reach the page at all — checked against robots.txt and llms.txt.
Step 4 — Compare against competitors
A brand's own Citation Share is half the story. The other half is who else is in the answers.
For every prompt where the brand isn't cited, log which competitor is. That's the competitive intelligence — where the brand is losing share, to whom, and on what query type. Build the competitor matrix as a heat map: rows are competitors, columns are query types, cells are citation frequency. The hot spots are where the market is being lost.
Include category-adjacent citations too — Wikipedia, Wikidata, Reddit threads, trade publications, Statista, industry association pages. These are the retrieval anchors the models trust, and they are frequently the actual source behind a citation attributed to a brand.
Step 5 — Diagnose the gaps
Five diagnostic questions:
- Is Citation Frequency low because of low overall mentions, or low extractability? Test by pulling the top-cited pages in the category and diffing schema, heading structure, and word count against yours.
- Is Cross-Engine Breadth low because of crawl access or signal mismatch? Check robots.txt for GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Applebot line by line.
- Is Query-Type Breadth low because of missing anchor types? Definitional gap = no /what-is/ page. Comparative gap = no "X vs Y" table. Procedural gap = no how-to. Recommendation gap = no "best of" list you're featured on.
- Is Extractability low because of schema or page structure? Run the page through Google's Rich Results Test and the Schema.org validator. Missing FAQPage, Article, and Organization schema is the most common failure.
- Are competitors winning on assets the brand could replicate? Usually the answer is yes, and usually the asset is one of three things: a Wikipedia page, a well-cited industry-report PDF, or a comparison page ranking in Google that Perplexity is scraping.
Step 6 — Build the 90-day plan
The audit ends in a plan, not a report. Three priorities, sequenced by impact:
- Month 1 — fix the floor. Crawl access, llms.txt, schema, the biggest extractability gaps. Whitelist GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Applebot, Amazonbot, Meta-ExternalAgent, Bytespider. Ship JSON-LD for Article, FAQPage, Organization, and Product where relevant.
- Month 2 — build the anchors. Definition pages, comparison tables, FAQ blocks, methodology pieces. One retrieval anchor per query type. Anchor pages are the extraction targets — write them for citation, not for scroll depth.
- Month 3 — push entity authority. Earn media, file the Wikipedia draft, populate Wikidata, reinforce structured data. Get on third-party lists. LLMs disproportionately trust entities that appear in multiple independent authoritative sources.
Re-measure at day 90. Compare to baseline. Iterate. Well-run programs show 2–4× Citation Share lift by day 90 when the floor was broken at baseline; 20–40% lift when the floor was already clean.
What a Citation Audit costs in time
Done manually: 40–60 hours for a tight 50-prompt audit. Done with API automation and a scoring framework: 8–12 hours. A 100-prompt audit with full competitor mapping and a written 90-day plan: 60–80 hours end to end, even automated.
Most agencies don't run them at all. That's the gap.
Common audit mistakes
- Prompt set too generic. "What is the best beauty brand" tells you nothing. "Which sunscreen brands does a dermatologist recommend for sensitive skin under $30" tells you everything.
- Single-engine measurement. Only running ChatGPT misses the rising Perplexity user base, Gemini's Google Workspace footprint, and Google AI Overviews' 1.5B monthly reach entirely.
- No competitor mapping. Citation Share without competitive context is half a metric.
- One-shot, never re-run. GEO compounds — you have to measure the compounding. Quarterly minimum, monthly for competitive categories.
- Audit without a plan attached. The report isn't the deliverable. The 90-day plan is.
- Single-run scoring. Model outputs drift. Run each prompt at least twice — three times is the audit-grade floor — a single call is a screenshot, not a measurement.
- No brand-safety layer. Track sentiment. A brand cited negatively 40% of the time has a bigger problem than a brand not cited at all.
- Confusing prompt count with prompt quality. A 200-prompt set that skews 80% definitional is worse than a 50-prompt set with balanced coverage across all four query types. Prompt-mix diversity matters more than volume past 50 prompts.
- Scoring without competitor cell-mapping. Recording "brand not cited" without recording which competitor won that cell forfeits the strategic half of the output. Every miss should log the winner.
- Ignoring engine-specific drift signatures. Averaging ChatGPT scores from browsing-on and browsing-off runs blends two different retrieval systems into one meaningless number. Segment by engine mode, not just engine name.
The bottom line
A Citation Audit is the single highest-leverage move in a GEO program. It tells you exactly where the engine knows you, where it doesn't, and what to do about it.
Run it now. Re-run it in 90 days. Re-run it every quarter after that.
The Citation Audit Family
This piece is the methodology anchor. The applied audits run the methodology against specific source layers:
- Wire Service Citation Audit 2026 — PR Newswire, Business Wire, GlobeNewswire, Newsfile, ACCESS Newswire. How AI engines actually cite the wires.
- IR Page Citation Audit 2026 — Corporate investor relations pages. Financial services ~100%, biotech ~92%, mega-cap consumer tech and junior mining ~0%.
- Vertical Citation Audit — Seven verticals: luxury hotels, universities, fashion, hospitals, restaurants, law firms, financial advisors.
- LLM Citation Audits: 100 Consumer Brands — Legacy brands underperform digital challengers by 32% on Citation Share.
- Podcast Citation Index 2026 — Which podcasts the five engines actually cite, ranked.
- Paywall Visibility Index 2026 — The visibility penalty paid by paywalled publishers inside AI answers.
In the Press
- "The Media Plan Is Now A Citation Plan" — Ronn Torossian, MediaPost, June 25, 2026. The Citation Audit scoring formula and the three-line brief that should appear on every 2026 media plan: citation share by engine, retrieval anchors, prompt-level reporting.
The conceptual hub: What Is AI Communications? The Definitive Guide for 2026
The complete research catalog: Everything-PR Research Index