KnitKnot
You are reading the agent-optimized layer of this page: the literal markdown we serve to AI crawlers and assistants, shipped in the page source of every visit. KnitKnot finds the claims AI repeats about a company, traces them to their sources, and helps the company correct the record.

# Run your first benchmark

Check your company profile, competitors, and question library, then start a run and follow it live.


A benchmark run takes your active question library and executes every question across ChatGPT, Claude, Perplexity, and Gemini, then scores each response. Before your first run, spend a few minutes checking the inputs : they determine what the benchmark measures.

## 1. Check your company profile

Open **Benchmarks → Benchmark model** in the console to review verified facts, products, capabilities, requirements, buyer roles, competitors, and research evidence. This model gives scoring its company context, so correct missing or inaccurate inputs before running.

If your company has multiple products, each product gets its own profile, its own competitor set, and its own question library : scores are tracked per product.

## 2. Check your competitors

Your competitor set should be the vendors AI actually shortlists against you : not every company in your market. Fewer, credible competitors beat a long speculative list: every competitor you add expands the question library and the run.

Competitors are managed per product under **Benchmark model → Competitors**. Adding one triggers the same research pass used for your own company, so its profile is benchmark-ready.

## 3. Review the question library

Open **Benchmarks → Questions** to review what buyers ask AI: head-to-head comparisons ("X vs Y for mid-market compliance"), open category questions ("best automated compliance platforms"), and capability-specific questions. KnitKnot generates them from real Google search queries with monthly volume, layered with your capabilities and buyer roles.

The library is persistent. Archive questions that don't fit, add your own, and star the ones that matter most : but keep the core set stable so run-over-run trends stay meaningful.

## 4. Start the run

Start a benchmark from **Benchmarks → Results**. Each question is sent to each engine as a fresh conversation : no history, no account context : the same way a first-time buyer would ask.

You can follow progress live: each answer streams its status, and failed cells (an engine timing out, for example) can be retried individually without re-running everything else.

## 5. What happens during scoring

Every response is scored by a semantic pipeline, not keyword matching:

  • - Each claim the AI made about you or a competitor is extracted, with the verbatim quote and the source the AI cited for it.
  • - Claims are verified against your researched profile : this is where factual errors and outdated information get caught.
  • - Head-to-head questions get a win/loss/tie outcome based on which vendor the AI actually recommended.
  • - Coverage, sentiment, and positioning are scored per response.

Scoring is deterministic: the same response always produces the same score, so movement between runs reflects the AI's answers changing : not scoring noise.

When the run finishes, your report and score update automatically. Next: [read your report](/docs/read-your-report/).

Raw mirror of this content: https://knitknot.ai/docs/run-your-first-benchmark.md. Site-wide summary: /llms.txt · full content: /llms-full.txt

Docs navigation
Docs Getting started

Run your first benchmark

Check your company profile, competitors, and question library, then start a run and follow it live.

Updated

A benchmark run takes your active question library and executes every question across ChatGPT, Claude, Perplexity, and Gemini, then scores each response. Before your first run, spend a few minutes checking the inputs : they determine what the benchmark measures.

1. Check your company profile

Open Benchmarks → Benchmark model in the console to review verified facts, products, capabilities, requirements, buyer roles, competitors, and research evidence. This model gives scoring its company context, so correct missing or inaccurate inputs before running.

If your company has multiple products, each product gets its own profile, its own competitor set, and its own question library : scores are tracked per product.

2. Check your competitors

Your competitor set should be the vendors AI actually shortlists against you : not every company in your market. Fewer, credible competitors beat a long speculative list: every competitor you add expands the question library and the run.

Competitors are managed per product under Benchmark model → Competitors. Adding one triggers the same research pass used for your own company, so its profile is benchmark-ready.

3. Review the question library

Open Benchmarks → Questions to review what buyers ask AI: head-to-head comparisons (“X vs Y for mid-market compliance”), open category questions (“best automated compliance platforms”), and capability-specific questions. KnitKnot generates them from real Google search queries with monthly volume, layered with your capabilities and buyer roles.

The library is persistent. Archive questions that don’t fit, add your own, and star the ones that matter most : but keep the core set stable so run-over-run trends stay meaningful.

4. Start the run

Start a benchmark from Benchmarks → Results. Each question is sent to each engine as a fresh conversation : no history, no account context : the same way a first-time buyer would ask.

You can follow progress live: each answer streams its status, and failed cells (an engine timing out, for example) can be retried individually without re-running everything else.

5. What happens during scoring

Every response is scored by a semantic pipeline, not keyword matching:

  • Each claim the AI made about you or a competitor is extracted, with the verbatim quote and the source the AI cited for it.
  • Claims are verified against your researched profile : this is where factual errors and outdated information get caught.
  • Head-to-head questions get a win/loss/tie outcome based on which vendor the AI actually recommended.
  • Coverage, sentiment, and positioning are scored per response.

Scoring is deterministic: the same response always produces the same score, so movement between runs reflects the AI’s answers changing : not scoring noise.

When the run finishes, your report and score update automatically. Next: read your report.