Every AI search conversation eventually reaches the same problem: nobody agrees what to count. One vendor reports mentions, another reports citations, a third reports a proprietary visibility score with no published formula. The numbers cannot be compared, and the trend lines quietly change meaning when someone tweaks the method.
This page publishes the metric definitions we use, so you can hold any vendor, including us, to the same wording. Six metrics, what each one measures, how to instrument it, and the failure mode attached to each.
Metric 1: citation share
Citation share is the percentage of your tracked prompt set where a page on your own domain appears in the answer's cited source list, measured per engine and per prompt.
Instrument it by running a fixed prompt set on a schedule, archiving the verbatim answer and the full source list, then counting prompts where your domain appears at least once. Count domains, not URLs, at the headline level, and keep the URL detail underneath so you can see which page earned it.
Failure mode: counting a brand mention inside the prose as a citation. Those are different metrics with different fixes. A mention without a citation usually means the model knows you from third-party sources and found nothing citable on your site, which is an on-site content gap.
Metric 2: share of voice
Share of voice is the percentage of your tracked prompt set where your brand is named anywhere in the answer, cited or not, measured against a fixed competitor set of no more than five names.
Instrument it with string matching on the archived answers, including known misspellings and legacy product names, then compute the same figure for each competitor so the number has a denominator that means something. Freeze the competitor set for at least two quarters. Swapping competitors is the easiest way to manufacture improvement.
Failure mode: treating share of voice as the primary KPI when your goal is qualified traffic. Share of voice is a brand-presence metric. It tells you whether you are in the consideration set the model generates, which is upstream of clicks. This is the distinction between winning the citation and winning the representation, covered in GEO versus AEO.
Metric 3: prompt coverage
Prompt coverage is the percentage of your tracked prompt set that has a genuinely citable answer published on your own domain: a page whose structure answers that specific question in its own words, near the top, with the entity clarity a model needs.
Instrument it by scoring each prompt as covered, partially covered, or uncovered, and recording the URL that covers it. This is the only metric in the list entirely inside your control, which makes it the best weekly production target. Citation share follows coverage with a lag of weeks to months, depending on crawl and index behavior.
Failure mode: marking a prompt covered because a page mentions the topic. Coverage means a reader or a model could lift a correct answer from the page without reading the rest of it. The mechanics of that structure are in the SEO, AEO, and GEO guide.
Metric 4: description accuracy
Description accuracy is whether the sentence a model generates about your category position, product capabilities, pricing, and stage is factually correct.
Instrument it by capturing verbatim output for a small set of brand prompts ("what is X", "how much does X cost", "who is X for") on every engine, then logging each inaccuracy with its likely source. Do not reduce this to a score. The fix depends on the exact wrong wording and on which third-party page is feeding it.
Failure mode: assuming inaccuracies come from your site. They usually come from an outdated review profile, an old funding article, or a comparison page written before your last release. Correcting them is off-site work, which is why GEO is a production and outreach discipline rather than a reporting one.
Metric 5: engine splits
Engine splits are the same four metrics above, reported separately for each engine you track, with month-over-month change.
Instrument it by never aggregating first. Each engine has a different retrieval mix, different freshness behavior, and different tolerance for thin pages. A blended number hides the fact that you gained on Perplexity, held on ChatGPT, and lost on AI Overviews, which are three different next actions.
Failure mode: a single composite visibility score. Composites are convenient for slides and useless for decisions, and when the formula is proprietary they cannot be audited at all.
Metric 6: the pipeline endpoint
The pipeline endpoint is sessions, conversions, and qualified conversations attributable to AI search referrers and to the pages built for the prompt set, reconciled to your analytics and CRM with the attribution limits stated.
Instrument it with referrer segmentation for the known AI surfaces, a landing-page cohort for pages built against tracked prompts, and branded search plus direct traffic as companion series. Expect under-attribution: the conversation happened off your site, and many buyers arrive later through a branded search. Report the referral number as a floor, not a total.
Failure mode: demanding clean attribution before investing. That standard would have ruled out most brand and content investment for two decades. Use cohorts and directional lift, and write the limits into the report.
Metrics that look useful and are not
- Total mentions across all engines. No denominator, so growth can come from adding prompts.
- Proprietary visibility scores. Unauditable formula, incomparable across vendors, changes silently.
- Sentiment as a single number. Useful as a flag, useless as a target. Read the verbatim quotes instead.
- Crawler hits from AI bots. Interesting for infrastructure, weakly related to citation.
- Prompt count itself. Tracking more prompts is not progress, and metered pricing rewards inflating it.
A reporting template that holds up
The monthly artifact we send has one page per metric family and looks like this.
| Section | Contents | Decision it supports |
|---|---|---|
| Coverage | Covered, partial, uncovered counts with URLs | What to publish next month |
| Citation share | Per engine, with the prompt-level win and loss list | Which existing pages to restructure |
| Share of voice | Against a frozen five-competitor set | Whether off-site work is needed |
| Accuracy | Verbatim quotes, each inaccuracy plus likely source | Which third-party sources to correct |
| Pipeline | AI referrers, page cohort, branded and direct series | Whether to hold or expand the investment |
Ours is described metric by metric on the AEO service page, including the tracked prompt counts at each tier and the cadence. If a vendor cannot show you this level of definition before you sign, the reporting will be decided after the fact.
Setting a baseline you can defend
- Write the prompt set with the sales team, using language buyers actually type, and freeze it.
- Run every prompt on every engine twice in the same week to see how volatile each answer is.
- Archive verbatim answers and source lists, because retroactive recomputation is otherwise impossible.
- Score prompt coverage on your own domain, honestly, before publishing anything new.
- Publish the definitions above in the engagement document, then do not change them for two quarters.
For a sense of where most companies start, our AI Search Readiness Benchmark crawled 160 live software sites and scored ten readiness signals. The median site had not implemented the structural basics, which is usually the reason a citation-share number looks flat regardless of publishing volume.
What to do with all of this
Pick the six metrics, write the definitions, freeze the prompt set and the competitor set, and make prompt coverage the weekly production target. Review the prompt set quarterly and log what you retire. If you are choosing between buying instrumentation and buying execution, the trade-offs are laid out in platform versus agency, and the tooling landscape is covered in AEO tools.
Want a read on your current baseline and which of these six is actually your constraint? Bring your site and your prompt list to a 30-minute growth call and we will score coverage with you on the spot.