Metrics / 10 September 2026 / 5 min read

Why one AI visibility score hides the thing you need to know

Averaging three engines into one number destroys the only signal that tells you what to do next.

Citarra field note02An average can be mathematically correct and operationally useless.

Most AI visibility tools lead with one number. It is reassuring, it fits on a slide, and it is close to useless for deciding what to do on Monday.

The engines are not measuring the same thing

Each assistant has different training data, different retrieval behaviour, and a different relationship with the live web. One leans heavily on current search. Another answers mostly from what it already knows. A third may weight community discussion more than editorial content.

So the same brand, measured in the same week with the same questions, can look strong in one place and absent in another. That is not measurement error. It is the finding.

What the average destroys

Consider a brand at 46% on one engine and 71% on another. The average is around 58%, which describes neither. It hides the actionable fact: there is a specific gap on a specific engine, and the sources that engine cites are a different list.

Track the average and a line moves for reasons you cannot explain. Track each engine and a drop tells you where to look.Per-engine evidence

The second number people conflate

There are two distinct things worth measuring and they are frequently merged.

  • Visibility is the share of measured answers that name you at all. It asks whether you are in the conversation.
  • Share of voice is your slice of all brand mentions against a defined competitor set. It asks how much of the conversation is yours.

They move independently. A brand can be named in most answers while holding a small share because several competitors appear alongside it. Collapse them into one score and you cannot tell whether the problem is absence or crowding.

What to ask a vendor

If a tool shows a single number, ask which engines were averaged, how many samples produced it, and what happens when one engine fails during a run.

If an engine returns errors for half a run and the tool silently averages the rest, the score can move for reasons that have nothing to do with your brand.

What to carry forward.

Per-engine differences are information, not noise.
Visibility and share of voice answer different questions.
Every score needs its sample size and failure handling.
Next field noteYou cannot prove AI visibility drove revenue