An instrument, not a dashboard
Anything can tell you your number went up. This tells you whether your work is what moved it, then hands the next job to your agent.

Four decisions that make the number worth having
Search is never forced
The web search tool is there, but the model decides whether to use it and how often. If you force it, you’ve destroyed the signal, because sometimes a real answer just comes from training data. We record whether search actually fired on every single probe.
A rate, not a yes or no
Each measurement fires a batch of probes and reports the share that recommended you, with a Wilson confidence interval on it. One probe is basically a coin flip. The variance is the thing you’re trying to see.
Control prompts
Mark a prompt as a control and it never gets optimisation work. So when the tracked rate climbs and the control stays flat, you can actually attribute the difference. Control-origin fanouts are locked too, so nobody can quietly work them.
Grounded and training, split
Run the same prompt with retrieval and without it. One shows you what the model finds today, the other shows what it already believes about you. They move for different reasons, and they need different work.
Track what you want, as often as you want
Add as many prompts as you need and set the cadence on each one. You bring your own API keys, so you’re paying the API directly and you can see what every run costs.
- No per-prompt pricing, so the prompt set can match the business
- Cadence from daily to monthly, set on each prompt
- Bring your own keys, and see exactly what each run costs
- Any prompt can be a control

Prompt, to fanout, to source
That middle column is the part mention-rate tools leave out, and it’s the only part that tells you which page to go and change.

Graded, counted, and locked when they’re controls
Every query the model fired, how often it fired, and whether it was discovery or verification. Start working one and it gets flagged, so what you changed and when sits right next to the measurement.
- Discovery and verification queries separated
- Fire counts per query, tracked over time
- Control-origin rows locked against optimisation

The pages that decide the answer
Every source is attributed to the fanout that surfaced it, and ranked by how often the model actually read it. Your own pages get marked, so the gap between what you own and what decides the answer becomes your work list.
- Ranked by how often the model read them
- Attributed to the query that surfaced them
- Your own pages marked against everyone else's
Then hand it to your agent and get it done
On its own, measurement is just a report. So you pick the fanout you want to win, pass it to your own agent over MCP, and the skills to act on it come with it.
Update your site
Rewrite the pages that should already be winning the query, against the ones that actually are.
Target existing sources
Go after the third-party pages the model keeps reading, and get your brand into them.
Create competing sources
Where there’s no good source yet, build the one the model cites next.
The work lands back on the same chart
Every action the agent takes gets logged with the page, the date and the target it was aimed at, and the marker lands on the same timeline as the mention rate. Control prompts stay untouched, so a lift stays attributable.

Rank is not citation
We ran the experiment ourselves. Across 91 days on a real site, Google position and AI citation had basically no relationship at all.