Methodology · v2.3 · last revised July 2026
Same prompts. Same week. Every tool.
Rankings you can't audit are marketing. Here is exactly how a tool earns its score — and how to catch us being wrong.
The score, decomposed
Share of our observed panel citations the tool surfaced: per-prompt detail, competitor visibility, sentiment, and alerting latency.
First-class tracking across ChatGPT, Perplexity, AI Overviews, Gemini, Claude, and Copilot — not just one engine with extrapolation.
Does the tool tell you what to do — and does doing it work? Measured by the 8-week lift test below.
Capability per dollar at the entry and mid tiers, including prompt caps and seat limits.
Setup time, export quality, integrations, and whether a non-specialist can read the dashboard.
Each criterion is scored 0–10 by two reviewers independently; disagreements over 1.5 points trigger a joint re-test.
The 400-prompt panel
Every Monday we run the same panel of 400 buyer-intent prompts — 80 each across SaaS, e-commerce, finance, healthcare, and local services — through ChatGPT, Perplexity, Google AI Overviews, Gemini, Claude, and Copilot. Each tool is then judged on how much of that observed reality it surfaced, and how fast.
400
Prompts, fixed panel
6
Answer engines queried
52
Runs per year, every Monday
5
Industries covered
The 8-week lift test
Tracking scores measure seeing; the lift test measures acting. For every tool that ships recommendations, we apply its top suggestions to a section of our own test site and measure citation-share change over eight weeks against an untouched control section. It's a small sample and we say so — but it's the only apples-to-apples action test in the category, and it's why optimization-capable tools like CiteCue score ahead of pure monitors.
Rules we hold ourselves to
- Weights and criteria change only at quarterly reviews, announced in advance — never mid-cycle.
- No vendor sees its score before publication, and no vendor can pay to be re-tested sooner than anyone else.
- Every score links to the week it was measured; historical scores are never silently edited.
- When our test site's lift data is too noisy to call, we publish 'inconclusive' rather than a number.
Disclosure
The AEO Report is reader-supported. Some outbound links, including to CiteCue, are partner links that may earn us a commission. Partnerships never touch scores: the prompt panel is automated, criteria and weights are published above, and we'll re-test any score on request — desk@theaeoreport.org .