Chapter 2 of 6 · updated 2026-10-02

Measuring AI visibility

How to build a prompt set, which metrics to record, how often to run them, and how to read a baseline without fooling yourself.

Build the prompt set

Start from the questions buyers ask before they shortlist you, not from your keywords. Four types cover most categories: category prompts ("best X for Y"), comparison prompts ("X vs Y"), use-case prompts ("how do I do Z") and problem prompts ("why does W happen"). Twenty to fifty prompts grouped into topics is enough to start; branded prompts ("what is Acme") teach you little because the brand is almost always mentioned.

Record the right things

For every answer, store the full text and the citations, then derive the metrics: whether the brand is mentioned, where in the answer it appears, how it is described, and which pages are cited. Do the same for named competitors from the same answers. Visibility (share of answers mentioning you), share of voice (your mentions over all tracked brands' mentions), and citation counts by domain all come from that stored record.

Keeping the text matters more than it seems. A score you cannot open cannot be explained to a stakeholder, and the sentence that mentions you is where you will find the stale review or the competitor comparison that shaped it.

Run daily, read weekly

Because answers vary, one run per prompt per engine per day is the practical floor. After two weeks you have a baseline with enough runs to tell a real change from noise; after that, read the series weekly and treat day-to-day movement as what it mostly is.

Expect zero

A new or small brand usually starts at zero on unbranded prompts across every engine. That is the honest baseline, not a tracking failure. The useful output of the first two weeks is the list of prompts where competitors appear and you do not, and the sources those answers cite.

Key takeaways

  • Twenty to fifty unbranded prompts in four types, grouped into topics.
  • Store full answers and citations; derive mentions, position, sentiment, share of voice from them.
  • One run per prompt per engine per day; two weeks for a baseline.
  • Zero is a normal starting point; the prompt-and-source gap list is the real output.

Sources and further reading

Get early access to LeapScope

Join the waitlist and we will invite you in batches, with early-access pricing for everyone on the list.