Free tool · no email required

Track your AI visibility honestly

Every tool in this category reports a share-of-voice percentage. Almost none of them report how many questions it came from. This does the same arithmetic with the sample size, the error bar and the significance test attached — so you can tell a result from a fluctuation.

What AI visibility measurement is

AI visibility is how often a model names your brand when someone asks a question in your category. It is measured by running a fixed set of buying questions, recording which answers name you, and dividing. Two numbers come out: a visibility rate — answers naming you, over answers collected — and a share of voice, your mentions over every brand mentioned.

What this does today. It computes and compares runs you collected yourself. It does not query the models for you — that needs a service behind the page and it does not exist yet. Generate the question set with the ChatGPT product checker, run it logged out, then bring the counts back here. Checked 2026-09-09.

Work out where you stand.

Prefilled with an illustrative run — not a client and not a benchmark. Replace the counts with your own; everything stays in this browser tab.

One row per answer collected, not per question in your list — if you ran the set across three models, that is three times the questions.

Competitor mentions in the same answers

Log every brand that appeared, not only the ones you already worry about. The new name in the list is usually the finding.

Compare against an earlier run

Visibility rate

25%

6 of 24 answers · 95% interval 8%–42%

Share of voice your mentions ÷ all brand mentions
12.0%
Brand mentions counted yours + every competitor logged
50
Error bar on the rate ± points, at this sample size
±17.3

Who owns the answer

  • AG1 21
  • Huel 14
  • Bloom 9
  • Example Brand 6

Against the earlier run

+12.5 points

Not distinguishable from noise at this sample size.

p = 0.267 · 150 questions per run needed to call a change this size

What to record, and what goes wrong

Six columns. The traps beside them are the reasons most in-house tracking produces a chart nobody trusts by the third month.

The measurement sheet, and the mistake that ruins each column
Record Why The trap
Named The primary measure. Everything else is a breakdown of it. Counting a passing mention the same as a recommendation. Decide once, write it down, keep it consistent.
Position in the answer First of three and last of nine are different outcomes with the same yes. Models reorder between runs. Track it, but do not react to a single move.
Competitors named Turns a rate into a share, and shows who owns the category answer. Only logging the ones you already worry about, which guarantees you never discover a new one.
Sources cited The most actionable column: it names the third parties the model trusts in your category. Ignoring it because it is not a number. It is the working target list for outreach.
Accuracy of what was said Being named with a wrong price is worse than not being named. Recording only presence, so a growing accuracy problem stays invisible.
Model, date and session state Answers vary by model version, region and account history. Without these the series is not comparable. Running logged in. Your own history will flatter every result you collect.

How often to run it

When What Why
Baseline The full prompt set, twice in one week, logged out. Two runs establish how much this set moves on its own. Without that you cannot read the next one.
Monthly The same set, same conditions, same day of month. Frequent enough to catch a change, rare enough that you are not measuring noise repeatedly.
After a data fix Two runs, two weeks apart, starting two weeks after the change ships. Models do not re-read a page the moment you edit it, and one run after a fix tells you nothing.
Never Daily. Day-to-day variance is larger than almost any real movement. Daily tracking produces a chart that looks like signal and is not.

The arithmetic, in full

Four formulas. Published because a visibility percentage with no method behind it is a number you cannot argue with or act on.

  1. 1

    Visibility rate

    answers naming you ÷ answers collected. The primary measure. Everything else is a breakdown of it.

  2. 2

    Share of voice

    your mentions ÷ all brand mentions. Answers the different question of who owns the category answer, and moves when a competitor gains even if you have not lost.

  3. 3

    The error bar

    A 95% Wald interval on the rate. On 24 questions it is roughly ±17 points, which is the whole reason a small run cannot support a confident claim. It is an approximation and gets rough near 0% and 100% — the sample size is printed next to it so you can judge it yourself.

  4. 4

    Is the change real

    A two-proportion z-test between the two runs, two-tailed, at p < 0.05, plus the sample size per run that would be needed to detect a change of the size you observed at 80% power. Usually the answer is a bigger set run less often.

No benchmark appears anywhere on this page. We have not measured a cross-brand distribution of AI visibility and will not publish an invented one, so there is no “good score” here to compare yourself against. Your own series under constant conditions is the only comparison this data supports. metric to confirm.

What this cannot tell you

It is not a rank. Rank implies a stable ordered list. Ask a model the same question twice and the order moves, so presence and share across a set are the strongest claims the data supports. Any product reporting a precise daily rank inside a model answer is reporting noise with a decimal point on it.

It is not traffic. Being named in an answer is not a visit, and nobody publishes how often the questions in your category are actually asked. A visibility rate cannot be converted into sessions, and any tool offering you an “AI search volume” figure has invented it.

It cannot attribute a cause. If the rate moves after you fixed your product data, the fix is a plausible explanation, not a demonstrated one — the model changed too, and so did your competitors. Hold the prompt set and the conditions constant and you at least remove the variables you control.

It measures one half of the work. Visibility follows from data a model can read and from third parties it trusts. This measures the outcome; the readiness checker measures the half you control outright.

Definitions

Visibility rate
Answers naming your brand, over answers collected, for a fixed question set.
Share of voice
Your mentions over all brand mentions in the same answers. A relative measure — it falls when a competitor rises, even if you have not moved.
Answer engine optimisation (AEO)
Getting cited inside an AI-generated answer rather than ranked in a list of links.
Generative engine optimisation (GEO)
The same work under a different name. The terminology is unsettled; the mechanics are not.
Statistical significance
Here: the observed change is larger than the run-to-run variation of the same prompt set would comfortably produce, at p < 0.05.

Questions about AI visibility tracking

What does an AI visibility tool actually measure?

How often a brand is named in the answers a model gives to a set of questions. That is it. The percentage on any dashboard in this category is a count of mentions over a count of questions asked — which is why the number is meaningless without the sample size beside it, and why almost none of them show you one.

Does this tool run the queries for me?

No, and it says so above the calculator rather than hiding it. Everything here runs in your browser, and a browser cannot query a model without an API key we would have to ship to every visitor. So you run the prompt set — the ChatGPT product checker generates one — and this does the arithmetic that turns the results into something you can defend. When a service sits behind it, the page will say that it has changed.

What is a good share of voice?

We do not publish one, because we have not measured a cross-brand distribution and an invented benchmark would be worse than no benchmark. What is defensible is your own series: the same prompt set, the same conditions, measured over time. A number you can compare to yourself last month beats a number somebody else made up about your industry.

How is this different from an AI rank tracker?

Rank implies a stable ordered list, which is not what a model produces — ask the same question twice and the order moves. This measures presence and share across a set of answers, with an error bar, which is the strongest claim the underlying data supports. Anything reporting a precise daily rank for a model answer is reporting noise.

Is this answer engine optimisation or generative engine optimisation?

Both names describe the same work: getting cited inside an AI-generated answer instead of ranked in a list of links. The terminology is unsettled and the acronym does not matter. What matters is that the work splits into two halves — data a model can read, which you control outright, and third-party sources it trusts, which you influence slowly.

Why is my 12-point improvement “not significant”?

Because on a 24-question set, moving from 3 mentions to 6 is well inside the range the same set produces on a quiet week. The calculator tells you how many questions per run you would need for a change that size to be distinguishable from variance. Usually the honest answer is to run a bigger set less often rather than a small set every week.

How often should we run this?

Baseline twice in a week to learn how much the set moves on its own, then monthly under identical conditions, plus a pair of runs a fortnight apart after any significant data fix. Not daily — day-to-day variance is larger than almost any real movement, and daily tracking produces a chart that looks like signal and is not.

Find out what you’re losing.

Before you commit to anything, we tell you exactly what you’re losing and what it costs to stop it. Two weeks. Fixed fee. Credited in full against any build you go ahead with.

Fee
$1,500–$3,000, fixed
Duration
Two weeks
Credited
In full, against any build
You supply
Read access + one 45-minute call