The reporting problem in AI search is that the easy metrics are the weak ones. Citations counted without context go up when you publish more and tell you nothing about whether buyers hear your name.
Metrics worth reporting
Recommendation rate on buying prompts is the headline. Underneath it: share of voice against a named competitor set, and the direction of both over a period long enough to mean something.
Report every rate with its sample size and a confidence band. A rate of 40% from five runs and 40% from two hundred are different claims, and a stakeholder who later discovers they were the same chart will stop trusting all of it.
Use a control
Keep a few prompts you are deliberately not working on, tracked alongside the rest. When everything moves at once, the control tells you it was the platform and not you. Without one, every model update looks like a win or a crisis.
Table stakes
- Recommendation rate, not just mentions.
- Sample size and confidence band shown with every rate.
- A control group of untouched prompts.
- A fixed cadence, changed rarely, because changing it breaks the trend.
- A log of what you changed and when, so a move can be explained.
The last point is the one that turns reporting into learning.