AEO · 11 min read

AI Visibility Metrics That Actually Mean Something

Definitions for the metrics that matter in AI search: citation share, share of voice, prompt coverage, description accuracy, engine splits, and the pipeline endpoint, plus how to instrument each one.

The short answer

Six metrics carry almost all the signal in AI search: citation share (how often a page on your domain is cited), share of voice (how often your brand is named at all), prompt coverage (how much of your prompt set has a citable answer on your domain), description accuracy (whether the generated sentence about you is true), engine splits (the same metrics per engine), and the pipeline endpoint (sessions and qualified conversations from AI referrers). Everything else is either a leading indicator of those or vanity. The definition matters more than the dashboard, because undefined metrics drift exactly when results flatten.

  • Write metric definitions down before the first report, not after results flatten
  • Citation share and share of voice measure different things and move independently
  • Prompt coverage is the only metric fully inside your control, so it is the best weekly target
  • Description accuracy needs verbatim capture, not a score, because the fix depends on the wording
  • Always report per engine: gains rarely land on all of them in the same month
  • AI referral traffic is under-attributed by design, so treat direct and branded lift as part of the signal

Free original research

AI Search Readiness Benchmark 2026

We crawled 160 live software sites across 10 AI search readiness signals. Median score is 7/10, 52% serve a real llms.txt, and 38% publish no JSON-LD at all. Read the findings, or drop your email and take the raw dataset.

Read the benchmark

Fielded 2026-08-09. Free to cite and republish under CC BY 4.0 with a link.

Every AI search conversation eventually reaches the same problem: nobody agrees what to count. One vendor reports mentions, another reports citations, a third reports a proprietary visibility score with no published formula. The numbers cannot be compared, and the trend lines quietly change meaning when someone tweaks the method.

This page publishes the metric definitions we use, so you can hold any vendor, including us, to the same wording. Six metrics, what each one measures, how to instrument it, and the failure mode attached to each.

Metric 1: citation share

Citation share is the percentage of your tracked prompt set where a page on your own domain appears in the answer's cited source list, measured per engine and per prompt.

Instrument it by running a fixed prompt set on a schedule, archiving the verbatim answer and the full source list, then counting prompts where your domain appears at least once. Count domains, not URLs, at the headline level, and keep the URL detail underneath so you can see which page earned it.

Failure mode: counting a brand mention inside the prose as a citation. Those are different metrics with different fixes. A mention without a citation usually means the model knows you from third-party sources and found nothing citable on your site, which is an on-site content gap.

Metric 2: share of voice

Share of voice is the percentage of your tracked prompt set where your brand is named anywhere in the answer, cited or not, measured against a fixed competitor set of no more than five names.

Instrument it with string matching on the archived answers, including known misspellings and legacy product names, then compute the same figure for each competitor so the number has a denominator that means something. Freeze the competitor set for at least two quarters. Swapping competitors is the easiest way to manufacture improvement.

Failure mode: treating share of voice as the primary KPI when your goal is qualified traffic. Share of voice is a brand-presence metric. It tells you whether you are in the consideration set the model generates, which is upstream of clicks. This is the distinction between winning the citation and winning the representation, covered in GEO versus AEO.

Metric 3: prompt coverage

Prompt coverage is the percentage of your tracked prompt set that has a genuinely citable answer published on your own domain: a page whose structure answers that specific question in its own words, near the top, with the entity clarity a model needs.

Instrument it by scoring each prompt as covered, partially covered, or uncovered, and recording the URL that covers it. This is the only metric in the list entirely inside your control, which makes it the best weekly production target. Citation share follows coverage with a lag of weeks to months, depending on crawl and index behavior.

Failure mode: marking a prompt covered because a page mentions the topic. Coverage means a reader or a model could lift a correct answer from the page without reading the rest of it. The mechanics of that structure are in the SEO, AEO, and GEO guide.

Metric 4: description accuracy

Description accuracy is whether the sentence a model generates about your category position, product capabilities, pricing, and stage is factually correct.

Instrument it by capturing verbatim output for a small set of brand prompts ("what is X", "how much does X cost", "who is X for") on every engine, then logging each inaccuracy with its likely source. Do not reduce this to a score. The fix depends on the exact wrong wording and on which third-party page is feeding it.

Failure mode: assuming inaccuracies come from your site. They usually come from an outdated review profile, an old funding article, or a comparison page written before your last release. Correcting them is off-site work, which is why GEO is a production and outreach discipline rather than a reporting one.

Metric 5: engine splits

Engine splits are the same four metrics above, reported separately for each engine you track, with month-over-month change.

Instrument it by never aggregating first. Each engine has a different retrieval mix, different freshness behavior, and different tolerance for thin pages. A blended number hides the fact that you gained on Perplexity, held on ChatGPT, and lost on AI Overviews, which are three different next actions.

Failure mode: a single composite visibility score. Composites are convenient for slides and useless for decisions, and when the formula is proprietary they cannot be audited at all.

Metric 6: the pipeline endpoint

The pipeline endpoint is sessions, conversions, and qualified conversations attributable to AI search referrers and to the pages built for the prompt set, reconciled to your analytics and CRM with the attribution limits stated.

Instrument it with referrer segmentation for the known AI surfaces, a landing-page cohort for pages built against tracked prompts, and branded search plus direct traffic as companion series. Expect under-attribution: the conversation happened off your site, and many buyers arrive later through a branded search. Report the referral number as a floor, not a total.

Failure mode: demanding clean attribution before investing. That standard would have ruled out most brand and content investment for two decades. Use cohorts and directional lift, and write the limits into the report.

Metrics that look useful and are not

  • Total mentions across all engines. No denominator, so growth can come from adding prompts.
  • Proprietary visibility scores. Unauditable formula, incomparable across vendors, changes silently.
  • Sentiment as a single number. Useful as a flag, useless as a target. Read the verbatim quotes instead.
  • Crawler hits from AI bots. Interesting for infrastructure, weakly related to citation.
  • Prompt count itself. Tracking more prompts is not progress, and metered pricing rewards inflating it.

A reporting template that holds up

The monthly artifact we send has one page per metric family and looks like this.

SectionContentsDecision it supports
CoverageCovered, partial, uncovered counts with URLsWhat to publish next month
Citation sharePer engine, with the prompt-level win and loss listWhich existing pages to restructure
Share of voiceAgainst a frozen five-competitor setWhether off-site work is needed
AccuracyVerbatim quotes, each inaccuracy plus likely sourceWhich third-party sources to correct
PipelineAI referrers, page cohort, branded and direct seriesWhether to hold or expand the investment

Ours is described metric by metric on the AEO service page, including the tracked prompt counts at each tier and the cadence. If a vendor cannot show you this level of definition before you sign, the reporting will be decided after the fact.

Setting a baseline you can defend

  1. Write the prompt set with the sales team, using language buyers actually type, and freeze it.
  2. Run every prompt on every engine twice in the same week to see how volatile each answer is.
  3. Archive verbatim answers and source lists, because retroactive recomputation is otherwise impossible.
  4. Score prompt coverage on your own domain, honestly, before publishing anything new.
  5. Publish the definitions above in the engagement document, then do not change them for two quarters.

For a sense of where most companies start, our AI Search Readiness Benchmark crawled 160 live software sites and scored ten readiness signals. The median site had not implemented the structural basics, which is usually the reason a citation-share number looks flat regardless of publishing volume.

What to do with all of this

Pick the six metrics, write the definitions, freeze the prompt set and the competitor set, and make prompt coverage the weekly production target. Review the prompt set quarterly and log what you retire. If you are choosing between buying instrumentation and buying execution, the trade-offs are laid out in platform versus agency, and the tooling landscape is covered in AEO tools.

Want a read on your current baseline and which of these six is actually your constraint? Bring your site and your prompt list to a 30-minute growth call and we will score coverage with you on the spot.

Ready to run this playbook?

Momentence runs the full engine, one team, one dashboard, six-month minimum. 30-minute call, no pitch deck.

Free original research

AI Search Readiness Benchmark 2026

We crawled 160 live software sites across 10 AI search readiness signals. Median score is 7/10, 52% serve a real llms.txt, and 38% publish no JSON-LD at all. Read the findings, or drop your email and take the raw dataset.

Read the benchmark

Fielded 2026-08-09. Free to cite and republish under CC BY 4.0 with a link.

FAQ

Common questions.

Citation share is the percentage of your tracked prompt set where a page on your own domain appears as a cited source in the generated answer, measured per engine. It is the closest AI search analogue to a ranking, and unlike a ranking it is binary per prompt: you are either in the source list or you are not.

Share of voice counts how often your brand is named anywhere in the answer, cited or not, against a fixed competitor set. Citation share counts only source-list appearances from your domain. A brand can have high share of voice and near-zero citation share when models describe it from third-party sources, which is a signal to invest in on-site answer coverage.

Enough to be representative, few enough to act on. Five to ten prompts is a baseline for an early-stage company, twenty covers most mid-market categories across buyer stages, and fifty or more suits multi-product or multi-market portfolios. Beyond that you get noise unless you have someone whose job is analysis.

Partly. Referrals from ChatGPT, Perplexity, and similar surfaces do appear in analytics, but a large share of AI-influenced sessions arrive as direct or branded search later, because the answer happened off your site. Treat AI referral sessions as a floor, and watch branded search volume and direct traffic alongside it rather than expecting one clean number.

There is no universal benchmark, and anyone quoting one is guessing. What matters is direction against your own baseline on a stable prompt set with a fixed definition. In our own 2026 crawl of 160 software sites, the median site had not implemented the basics that make citation possible at all, so most companies have room to move before benchmarks are the interesting question.

Keep reading

More from Learn.

Search · 18 min read

The Definitive SEO, AEO, and GEO Guide for Software

One operating system for SEO, AEO, and GEO, with actionable checklists for technical foundations, answer coverage, entity data, off-page corpus work, and measurement.

AEO · 12 min read

AEO Tools in 2026, an Honest Landscape Review

What AEO and GEO tools actually do, the five categories worth knowing, how to evaluate one, what no tool can do for you, and a minimum viable stack you can run for under $500 a month.

AEO · 12 min read

AI Visibility Platform vs Agency, How to Decide

Buy an AI visibility platform, hire an agency, or both. A decision framework covering scope, total cost of ownership with the operator counted, break-even points, and the failure modes of each model.

Paid · 11 min read

B2B PPC Agency Pricing: Fee Models and Real Cost

How B2B PPC and paid media agencies charge, what each fee model rewards, worked total cost math at $15K, $40K, and $100K monthly spend, and the break-even point against an in-house buyer.

Growth · 11 min read

What Makes the Best B2B Marketing Agency for You

Best is situational, not a ranking. The five decision axes that separate B2B marketing agencies, a scoring rubric, disqualifiers, and how to read the shortlist you already have.

Growth · 12 min read

B2B Marketing Firm vs In-House Team: How to Decide

A build-versus-buy framework for B2B software companies: real total cost of ownership, a scoring model, the hybrid setups that work, and how to hand execution back in-house later.

Growth · 11 min read

How to Evaluate B2B Marketing Companies

A vendor evaluation scorecard for B2B marketing companies: the five types, how to read a proposal, the red flags that predict failure, and the questions to ask on the first call.

Growth · 12 min read

How to Choose a B2B Marketing Agency

A diligence framework for choosing a B2B marketing agency: the four agency archetypes, what each fixes, the questions that expose weak shops, and the real cost of fragmentation.

Buying Guide · 12 min read

SaaS Marketing Agency Pricing: What Each Retainer Buys

What SaaS marketing agency pricing actually covers at every retainer band, how agencies build a price, and how to audit a proposal before you sign it.

Paid Media · 12 min read

How to Choose a B2B SaaS Paid Advertising Agency

How to evaluate a B2B SaaS paid advertising agency: fee models, minimum viable spend, measurement prerequisites, and the diligence questions that matter.

AEO · 10 min read

GEO vs AEO: Two Different Jobs Inside AI Search

GEO and AEO are related but not the same. AEO wins the citation on a specific answer, GEO shapes how models describe your brand. How to run both together.

Demand Gen · 8 min read

Inbound vs Outbound for B2B SaaS: How to Split the Budget

A practical framework for splitting growth budget between inbound demand capture and outbound demand creation, with the failure modes of each.

SEO · 10 min read

Best SaaS SEO Agencies in 2026

A criteria-driven comparison of the SaaS SEO agencies worth shortlisting in 2026: where each is strongest, where it is not, and how to pick one.

AEO · 9 min read

AEO vs SEO: The Difference and Why You Need Both

AEO is the discipline of getting cited by ChatGPT, Perplexity, Gemini, and AI Overviews. How it differs from SEO and how to run both together.

Website · 10 min read

How to Build an AI-Native Marketing Website

Ship a modern software marketing website in weeks: AI in the production loop, SSR performance from day one, and a design system marketing can extend.

Demand Gen · 11 min read

Software Demand Generation Playbook for 2026

How B2B and B2C software companies combine site, SEO, AEO, content, and paid into one demand gen system that compounds instead of resetting.