Docs navigation
Benchmark questions
Where benchmark questions come from, how the budget is split across topics by real buyer demand, and what a given library size can honestly claim.
Updated
Your question library is the set of buyer-style questions every benchmark asks. It’s the measuring stick, so it is built to be realistic, stable, and yours to curate.
Grounded in real demand
Questions aren’t synthetic templates. Generation starts from real Google search queries in your category, with monthly search volume attached, and turns them into the questions a buyer would ask an AI assistant: “best X for Y”, “compare A vs B”, “does A support Z”. Each question records why it was generated : the keyword, capability, or buyer role behind it : so you can always trace a question back to the demand it represents.
The mix leans deliberately competitive: most questions put you in a buying context against named competitors, because that is where deals are won and lost. The rest are open landscape questions that measure organic visibility: whether AI brings you up without being told to.
How the budget is split: topics first, by demand
Your keywords are grouped into topics : the handful of things buyers in your category actually search for. A topic is the unit of demand, because search volume is measured per topic, not per individual question. Your library size is then split across those topics in proportion to that demand.
The split is deliberately compressed, not strictly proportional. A topic with 750 searches a month gets roughly three times the questions of one with 10 : not seventy-five times. Head topics should lead; they shouldn’t swallow the whole library and leave you blind everywhere else.
Within a topic, the allocation is spent on depth rather than breadth. Competitors and buyer roles are how a topic varies its questions, not separate budgets of their own. A topic big enough to fund four competitors properly carries four; a smaller one carries two at full depth instead of one shallow question about each.
Every measured slice gets at least three questions
A topic either earns at least three questions or it isn’t measured : and when it isn’t, we say so rather than report on it anyway.
Three is the floor because one question is one question put to three assistants. If those answers come back badly, you’ve learned that one phrasing of one question went badly : not that AI is silent on the topic. Publishing that as a finding would be publishing noise as fact.
The same floor applies at reporting time. A finding is never drawn from a slice measured by fewer than three questions, so if a topic falls under the bar in a given week : never allocated, or an assistant failed mid-run : it reads as not measured this period. Absence of measurement is not absence from the answer, and we won’t print one as the other.
The consequence is that a low-demand topic can drop out of a small library entirely. That’s intentional: nine topics measured well beats thirteen measured unreliably.
Your own capabilities are the exception. Keyword tools have no search volume for plenty of things buyers genuinely evaluate you on. A topic that represents one of your actual product capabilities is measured anyway, on estimated demand, and Benchmarks → Questions → Topics marks it estimated. A keyword topic nobody searches for is weak evidence and still drops; a real capability is not.
What your library size can claim
Questions buy confidence, and different cuts of your data cost different amounts. One question counts toward its topic, its competitor, and its buyer role simultaneously : which is why all three can be reported at once. Crossing two of them doesn’t overlap, so it costs far more.
| What you want to read | Typical cells | Questions needed | 75 | 150 |
|---|---|---|---|---|
| Performance by topic | 9 | 27 | ✅ | ✅ |
| Performance by competitor | 4 | 12 | ✅ | ✅ |
| Performance by buyer role | 4 | 12 | ✅ | ✅ |
| Topic × competitor (“how do we do against Acme on pricing”) | 36 | 108 | ❌ | ✅ |
| Topic × competitor × buyer role | 144 | 432 | ❌ | ❌ |
A 75-question library answers “how are we doing on preventive maintenance” and “how are we doing against ServiceTitan” : separately, and well. Crossing the two becomes reportable at 150. That is the concrete thing a larger library buys.
Persistent, not regenerated
The library persists across runs. Every benchmark snapshots the currently active questions and asks the same questions as the last one, which is what makes scores comparable over time : movement in your numbers reflects changes in AI’s answers, not changes in the questionnaire.
We also generate more questions than your plan measures at once. The extras sit on the bench and keep their history. When demand shifts, or you raise your plan, or you add a competitor, the library rebalances by promoting from that bench : no new generation, no cost : and questions that already carry run history are kept first, so your trend lines survive the change.
Questions you wrote yourself are never switched off by a rebalance. The demand split governs generated questions; a question you typed is a deliberate choice, so it stays active until you archive it. Same for anything you’ve starred.
You can add, edit, archive, and re-generate questions between runs. Archived questions stop being benchmarked but keep their history; new questions start contributing from the next run.
Per-subject libraries
Each subject : your brand and each product : has its own library, generated from its own category and keywords. A benchmark for a product runs that product’s questions against that product’s competitors.
Curating well
- Volume beats cleverness. A question tied to a query real buyers search monthly is worth ten hypothetical ones.
- Keep the set stable. Resist rewording questions between runs; every edit resets that question’s comparability.
- Prune what you’d never act on. If a question’s outcome wouldn’t change what you write or fix, archive it : it’s diluting the questions that matter.
- Depth beats coverage. Three questions on one comparison tell you something. One question on each of three comparisons tells you nothing, three times.