KnitKnot
You are reading the agent-optimized layer of this page: the literal markdown we serve to AI crawlers and assistants, shipped in the page source of every visit. Making sure AI reads the right facts about a company is literally what KnitKnot does.

# Benchmark questions

Where benchmark questions come from, how the budget is split across topics by real buyer demand, and what a given library size can honestly claim.


Your question library is the set of buyer-style questions every benchmark asks. It's the measuring stick, so it is built to be realistic, stable, and yours to curate.

## Grounded in real demand

Questions aren't synthetic templates. Generation starts from real Google search queries in your category, with monthly search volume attached, and turns them into the questions a buyer would ask an AI assistant: "best X for Y", "compare A vs B", "does A support Z". Each question records why it was generated : the keyword, capability, or buyer role behind it : so you can always trace a question back to the demand it represents.

The mix leans deliberately competitive: most questions put you in a buying context against named competitors, because that is where deals are won and lost. The rest are open landscape questions that measure organic visibility: whether AI brings you up without being told to.

## How the budget is split: topics first, by demand

Your keywords are grouped into **topics** : the handful of things buyers in your category actually search for. A topic is the unit of demand, because search volume is measured per topic, not per individual question. Your library size is then split across those topics in proportion to that demand.

The split is deliberately **compressed**, not strictly proportional. A topic with 750 searches a month gets roughly three times the questions of one with 10 : not seventy-five times. Head topics should lead; they shouldn't swallow the whole library and leave you blind everywhere else.

Within a topic, the allocation is spent on **depth** rather than breadth. Competitors and buyer roles are how a topic varies its questions, not separate budgets of their own. A topic big enough to fund four competitors properly carries four; a smaller one carries two at full depth instead of one shallow question about each.

## Every measured slice gets at least three questions

**A topic either earns at least three questions or it isn't measured** : and when it isn't, we say so rather than report on it anyway.

Three is the floor because one question is one question put to three assistants. If those answers come back badly, you've learned that one phrasing of one question went badly : not that AI is silent on the topic. Publishing that as a finding would be publishing noise as fact.

The same floor applies at reporting time. A finding is never drawn from a slice measured by fewer than three questions, so if a topic falls under the bar in a given week : never allocated, or an assistant failed mid-run : it reads as **not measured this period**. Absence of measurement is not absence from the answer, and we won't print one as the other.

The consequence is that a low-demand topic can drop out of a small library entirely. That's intentional: nine topics measured well beats thirteen measured unreliably.

**Your own capabilities are the exception.** Keyword tools have no search volume for plenty of things buyers genuinely evaluate you on. A topic that represents one of your actual product capabilities is measured anyway, on estimated demand, and **Benchmarks → Questions → Topics** marks it *estimated*. A keyword topic nobody searches for is weak evidence and still drops; a real capability is not.

## What your library size can claim

Questions buy confidence, and different cuts of your data cost different amounts. One question counts toward its topic, its competitor, and its buyer role simultaneously : which is why all three can be reported at once. Crossing two of them doesn't overlap, so it costs far more.

| What you want to read | Typical cells | Questions needed | 75 | 150 | |---|---|---|---|---| | Performance by topic | 9 | 27 | ✅ | ✅ | | Performance by competitor | 4 | 12 | ✅ | ✅ | | Performance by buyer role | 4 | 12 | ✅ | ✅ | | Topic × competitor ("how do we do against Acme *on pricing*") | 36 | 108 | ❌ | ✅ | | Topic × competitor × buyer role | 144 | 432 | ❌ | ❌ |

A 75-question library answers "how are we doing on preventive maintenance" and "how are we doing against ServiceTitan" : separately, and well. Crossing the two becomes reportable at 150. That is the concrete thing a larger library buys.

## Persistent, not regenerated

The library persists across runs. Every benchmark snapshots the currently active questions and asks the same questions as the last one, which is what makes scores comparable over time : movement in your numbers reflects changes in AI's answers, not changes in the questionnaire.

We also generate more questions than your plan measures at once. The extras sit on the bench and keep their history. When demand shifts, or you raise your plan, or you add a competitor, the library rebalances by promoting from that bench : no new generation, no cost : and questions that already carry run history are kept first, so your trend lines survive the change.

**Questions you wrote yourself are never switched off by a rebalance.** The demand split governs generated questions; a question you typed is a deliberate choice, so it stays active until you archive it. Same for anything you've starred.

You can add, edit, archive, and re-generate questions between runs. Archived questions stop being benchmarked but keep their history; new questions start contributing from the next run.

## Per-subject libraries

Each [subject](/docs/subjects/) : your brand and each product : has its own library, generated from its own category and keywords. A benchmark for a product runs that product's questions against that product's competitors.

## Curating well

  • - **Volume beats cleverness.** A question tied to a query real buyers search monthly is worth ten hypothetical ones.
  • - **Keep the set stable.** Resist rewording questions between runs; every edit resets that question's comparability.
  • - **Prune what you'd never act on.** If a question's outcome wouldn't change what you write or fix, archive it : it's diluting the questions that matter.
  • - **Depth beats coverage.** Three questions on one comparison tell you something. One question on each of three comparisons tells you nothing, three times.

Raw mirror of this content: https://knitknot.ai/docs/benchmark-questions.md. Site-wide summary: /llms.txt · full content: /llms-full.txt

Docs navigation
Docs Core concepts

Benchmark questions

Where benchmark questions come from, how the budget is split across topics by real buyer demand, and what a given library size can honestly claim.

Updated

Your question library is the set of buyer-style questions every benchmark asks. It’s the measuring stick, so it is built to be realistic, stable, and yours to curate.

Grounded in real demand

Questions aren’t synthetic templates. Generation starts from real Google search queries in your category, with monthly search volume attached, and turns them into the questions a buyer would ask an AI assistant: “best X for Y”, “compare A vs B”, “does A support Z”. Each question records why it was generated : the keyword, capability, or buyer role behind it : so you can always trace a question back to the demand it represents.

The mix leans deliberately competitive: most questions put you in a buying context against named competitors, because that is where deals are won and lost. The rest are open landscape questions that measure organic visibility: whether AI brings you up without being told to.

How the budget is split: topics first, by demand

Your keywords are grouped into topics : the handful of things buyers in your category actually search for. A topic is the unit of demand, because search volume is measured per topic, not per individual question. Your library size is then split across those topics in proportion to that demand.

The split is deliberately compressed, not strictly proportional. A topic with 750 searches a month gets roughly three times the questions of one with 10 : not seventy-five times. Head topics should lead; they shouldn’t swallow the whole library and leave you blind everywhere else.

Within a topic, the allocation is spent on depth rather than breadth. Competitors and buyer roles are how a topic varies its questions, not separate budgets of their own. A topic big enough to fund four competitors properly carries four; a smaller one carries two at full depth instead of one shallow question about each.

Every measured slice gets at least three questions

A topic either earns at least three questions or it isn’t measured : and when it isn’t, we say so rather than report on it anyway.

Three is the floor because one question is one question put to three assistants. If those answers come back badly, you’ve learned that one phrasing of one question went badly : not that AI is silent on the topic. Publishing that as a finding would be publishing noise as fact.

The same floor applies at reporting time. A finding is never drawn from a slice measured by fewer than three questions, so if a topic falls under the bar in a given week : never allocated, or an assistant failed mid-run : it reads as not measured this period. Absence of measurement is not absence from the answer, and we won’t print one as the other.

The consequence is that a low-demand topic can drop out of a small library entirely. That’s intentional: nine topics measured well beats thirteen measured unreliably.

Your own capabilities are the exception. Keyword tools have no search volume for plenty of things buyers genuinely evaluate you on. A topic that represents one of your actual product capabilities is measured anyway, on estimated demand, and Benchmarks → Questions → Topics marks it estimated. A keyword topic nobody searches for is weak evidence and still drops; a real capability is not.

What your library size can claim

Questions buy confidence, and different cuts of your data cost different amounts. One question counts toward its topic, its competitor, and its buyer role simultaneously : which is why all three can be reported at once. Crossing two of them doesn’t overlap, so it costs far more.

What you want to readTypical cellsQuestions needed75150
Performance by topic927
Performance by competitor412
Performance by buyer role412
Topic × competitor (“how do we do against Acme on pricing”)36108
Topic × competitor × buyer role144432

A 75-question library answers “how are we doing on preventive maintenance” and “how are we doing against ServiceTitan” : separately, and well. Crossing the two becomes reportable at 150. That is the concrete thing a larger library buys.

Persistent, not regenerated

The library persists across runs. Every benchmark snapshots the currently active questions and asks the same questions as the last one, which is what makes scores comparable over time : movement in your numbers reflects changes in AI’s answers, not changes in the questionnaire.

We also generate more questions than your plan measures at once. The extras sit on the bench and keep their history. When demand shifts, or you raise your plan, or you add a competitor, the library rebalances by promoting from that bench : no new generation, no cost : and questions that already carry run history are kept first, so your trend lines survive the change.

Questions you wrote yourself are never switched off by a rebalance. The demand split governs generated questions; a question you typed is a deliberate choice, so it stays active until you archive it. Same for anything you’ve starred.

You can add, edit, archive, and re-generate questions between runs. Archived questions stop being benchmarked but keep their history; new questions start contributing from the next run.

Per-subject libraries

Each subject : your brand and each product : has its own library, generated from its own category and keywords. A benchmark for a product runs that product’s questions against that product’s competitors.

Curating well

  • Volume beats cleverness. A question tied to a query real buyers search monthly is worth ten hypothetical ones.
  • Keep the set stable. Resist rewording questions between runs; every edit resets that question’s comparability.
  • Prune what you’d never act on. If a question’s outcome wouldn’t change what you write or fix, archive it : it’s diluting the questions that matter.
  • Depth beats coverage. Three questions on one comparison tell you something. One question on each of three comparisons tells you nothing, three times.