Why measuring AI citations is no longer optional
Without measurement, GEO is faith. With measurement, it’s an acquisition channel as steerable as any other. Three concrete reasons: (1) prioritise actions, you’ll know which engine cites you least and why; (2) defend the GEO budget to leadership with hard KPIs; (3) detect regressions, a model swap (GPT-4 to GPT-5) can wipe your citations overnight.
The stakes are documented: the founding Princeton study (Aggarwal et al., "GEO: Generative Engine Optimization", KDD 2024) shows that targeted optimisations can increase a content’s visibility by up to 40% in generative engine answers. You still need to measure that visibility to see the gain. NEXUS GEO recommends putting 20 to 30% of the GEO budget into measurement, with the rest going to editorial and technical production: that split is what makes steering possible.
Free tools, where to start
Google Search Console: AI Overviews and AI Mode data
Contrary to a common belief, Google does not offer a separate report for AI Overviews: impressions and clicks from AI Overviews and AI Mode are counted in the Performance report (search type "Web"), as the official Google Search Central documentation explains. Free, official, the first place to look. Limits: you cannot precisely isolate the share that comes from AI answers, and it tells you nothing about ChatGPT, Claude or Perplexity.
Bing Webmaster Tools: the Copilot and ChatGPT proxy
Microsoft Copilot relies on the Bing index, and ChatGPT also uses Bing for its searches in browsing mode. Monitoring your performance in Bing Webmaster Tools (free) therefore gives a useful proxy of your visibility in the Microsoft/OpenAI ecosystem; Microsoft has said that Copilot traffic is included in the performance data. Turn it on today if you haven’t already.
Manual prompt tests
The most basic approach but ruthlessly effective: list 20 to 30 prompts your prospects would type ("best agency X in Europe", "solution Y for SMBs"). Test them every month in ChatGPT, Claude, Gemini and Perplexity. Note: (1) are you cited? (2) at what position? (3) which competitors are cited instead? Cost: €0, time: about 2 hours per month.
Paid tools, when to level up
Profound (tryprofound.com), the US reference
Profound is the US pioneer of GEO measurement platforms. Very wide coverage, checked in June 2026: ChatGPT, Claude, Gemini, Perplexity, Microsoft Copilot, Grok, Meta AI, DeepSeek and Google AI Overviews, with analysis of prompt volumes and AI bot crawling. Enterprise pricing on demo (an entry price of around $999/month was reported in 2025). Suited to brands with large prompt volumes to track and a substantial GEO budget.
Goodie AI
Direct competitor to Profound, positioned on enterprise (higoodie.com). Goodie AI tracks more than 11 models (including ChatGPT, Claude, Perplexity, Gemini, Copilot, Grok and Amazon Rufus) and focuses on brand mention analysis: who is cited, from what angle, with what sentiment. Published pricing: $399/month (Core, 7-day trial) and $999/month (Pro), enterprise on quote. A good alternative if you want analytical depth more than volume.
AthenaHQ
AthenaHQ (athenahq.ai) focuses on AI share of voice, with dashboards that marketing leaders can read at a glance. Coverage: ChatGPT, Perplexity, Google AI Overviews and AI Mode, Gemini, Claude, Copilot, Grok. It offers a free plan limited in credits, a Starter plan at $295/month and an enterprise plan on quote (prices checked on 29 September 2026). Interesting for teams that want a visual monthly deliverable without an enterprise budget.
The DIY approach: an in-house tracker is doable
The principle fits in one sentence: query the AI engines regularly through their APIs on a panel of industry prompts, detect whether your brand is mentioned, and store the results to follow the trend. API costs stay affordable for tracking at a reasonable scale. Advantages over the US tools: full control of the prompt scope, and the ability to cover European models such as Mistral’s Le Chat.
The real difficulty lies less in the architecture than in long-term reliability. Plan for an initial development effort of several weeks, then regular upkeep: detecting brand-name variants, constant changes in APIs and models, ongoing maintenance. That industrialisation is what separates a weekend script from a steering instrument.
The 3 metrics that actually matter
1. Citation rate
Percentage of prompts in the panel where your brand is cited at least once. The simplest and most telling metric. Orders of magnitude vary widely with sector, competition and starting notoriety: set a progression path measured month after month rather than an absolute threshold.
2. Average position in the answer
When you’re cited, in what position? First brand mentioned, or third out of five? Positions 1-2 capture most of the attention. Track the distribution of positions, not just the average: moving from "cited at the end of the list" to "cited at the top" changes the business value of the citation.
3. AI share of voice
Across all brands cited in your prompt panel, what share do you represent? If you get 40 citations and your 3 direct competitors total 120, your AI share of voice is 25%. The reference metric for steering against competition.
Monthly methodology: the principles of reliable measurement
Whatever tooling you choose, a usable monthly measurement rests on five principles:
- A stable prompt panel that reflects the real buying journey, large enough to smooth out the models’ natural variance, repeated identically from one month to the next.
- Engine-by-engine measurement: the same prompt produces very different answers in ChatGPT, Claude, Gemini or Perplexity; aggregating without distinguishing them loses the actionable information.
- A constant protocol: same period, same wording, same test conditions, otherwise the variations you observe mean nothing.
- Robust mention detection: brand-name variants, indirect citations, linked sources; under-counting is the most common mistake of amateur setups.
- A causal reading of the gaps: every significant variation should be matched to an event (an action shipped, a competitor’s publication, a model update).
The detailed construction of the prompt panel, the alert thresholds and the weighting of engines are part of the proprietary methodology NEXUS GEO applies, in line with the GEO-47 framework (8 pillars, 47 criteria). The €1,750 GEO audit includes the initial benchmark against 3 named competitors, in 5 AI engines: ChatGPT, Gemini, Perplexity, Mistral and Google AI Overviews.
FAQ
Sources
- Google Search Central: "AI features and your website", official documentation (developers.google.com/search/docs/appearance/ai-features), consulted in June 2026.
- Microsoft Bing Webmaster Tools: performance data including Copilot (bing.com/webmasters).
- Profound: tryprofound.com, engine coverage checked in June 2026.
- Goodie AI: higoodie.com, engine coverage checked in June 2026, pricing on 29 September 2026.
- AthenaHQ: athenahq.ai/pricing, prices checked on 29 September 2026.
- Aggarwal P. et al. ("GEO: Generative Engine Optimization") KDD 2024 (arxiv.org/abs/2311.09735).
- NEXUS GEO AI-citation tracker: the agency’s in-house tool, not sold separately.
Would you rather hand measurement and execution to a specialist agency? The NEXUS GEO audit sets your citation baseline and the action plan that goes with it. Plan details are on the pricing page.
GEO audit
Want your AI citation baseline?
The €1,750 NEXUS GEO audit analyses your visibility in 5 AI engines (ChatGPT, Gemini, Perplexity, Mistral and Google AI Overviews) against 47 criteria, with a benchmark against 3 named competitors. 30-page PDF report, score per pillar, 6-month action plan and 60-minute debrief, delivered within 10 business days.
