KnitKnot
You are reading the agent-optimized layer of this page: the literal markdown we serve to AI crawlers and assistants, shipped in the page source of every visit. KnitKnot finds the claims AI repeats about a company, traces them to their sources, and helps the company correct the record.

# Introducing KnitKnot

KnitKnot runs the questions software buyers ask across four AI engines, saves every answer, checks the claims, and turns repeated problems into work a company can ship.


The old Drata answer from our [first post](/blog/why-we-pivoted/) has a KnitKnot score of 28.

That number is useful. It is also the least interesting part of the record.

What matters is everything underneath it. I can open the answer ChatGPT gave, read the exact line that favored Vanta, see which pages it cited, and compare the same question across Claude, Perplexity, and Gemini. The number tells me where to look. The saved answer tells me what happened.

This distinction shaped most of the product. We wanted a record that could survive after the answer changed, rather than another dashboard that tells a company it has an AI visibility problem and leaves the marketing team to guess what caused it.

## Start with the questions

A benchmark begins with a company profile. KnitKnot researches the products, features, positioning, competitors, and buyer personas associated with the company. Each fact has a source, and the company can correct anything we got wrong before the run starts.

Those inputs become a prompt library. Some prompts ask open category questions, such as which compliance tools are best for a mid-market company. Others compare two vendors or drill into a feature that matters to a particular buyer.

We ground the library in search queries with measured demand, then layer in the company's actual products and competitors. This keeps us from inventing a tidy set of questions that nobody asks. It also gives the benchmark a stable shape. A team can archive a bad prompt or add one we missed, but the useful prompts persist from one run to the next.

That persistence matters. If the questions change every week, the score can move even when the answers do not. You end up measuring your own prompt generator.

## Run each question as a fresh conversation

KnitKnot sends the active prompt library to ChatGPT, Claude, Perplexity, and Gemini. Each prompt starts in a fresh conversation with no account history or prior messages. We save the returned answer along with the engine, model version, timestamp, and any citations the engine exposed.

The same question often produces four different buying stories. One engine may recommend you, another may leave you out, and a third may repeat a claim that disappeared from your website months ago. An average hides most of that. We keep the engines separate so the company can read each answer.

This is a controlled benchmark, not a recording of every buyer's screen. Location, account history, personalization, and a model update can all change what an individual sees. A fresh conversation gives us a repeatable starting point. AI still varies.

## Keep the answer before scoring it

The complete response is the first artifact we store. We do this before trying to summarize or score it.

That sounds obvious, but it changes what the product can explain later. If a benchmark only stores a score, there is no way to tell whether a movement came from the recommendation or from a change in the facts and sources behind it. The original answer is gone.

Once the response is saved, KnitKnot extracts the parts a company can inspect: the recommendation, feature comparisons, factual claims, coverage, and cited domains.

We check material claims against the researched company profile and its dated sources. A claim we can verify gets a factual verdict. A claim without enough evidence stays unverifiable. We would rather leave the field unresolved than invent certainty because a report looks cleaner with every row filled in.

## Separate showing up from being chosen

We made one early mistake by treating every mention as visibility.

If the prompt says, "Compare Drata and Vanta," ChatGPT is almost guaranteed to mention Drata. Counting that as organic visibility gives the company credit for appearing in a question that forced its name into the answer.

KnitKnot separates open discovery prompts from competitive prompts. Visibility measures whether an engine surfaces the company when the prompt does not name it. Win rate covers the head-to-head answers where the engine made a choice. A company can do well at one and badly at the other.

This is why the report has more than a single percentage. Coverage shows how central the company was to an answer. Competitive outcomes record wins, losses, and ties. Accuracy tracks the claims we could check. The AI Presence Score combines these signals into a top-line number, but every part of it links back to the evaluations that produced it.

The dashboard does not ask a new judge to reinterpret the answer every time somebody opens a chart. It reads the evaluation facts saved when that response was scored. The report and the drill-down start from the same rows.

## Let problems keep their history

One strange answer is evidence, but it may also be noise. The same problem appearing across several questions or later benchmarks is different.

KnitKnot turns material gaps into persistent issues. An issue might track an outdated claim, a missing feature, a losing comparison, or a category where competitors appear and you do not. It keeps the captured answers and supporting evidence attached as the reach changes over time.

This prevents the work from resetting after every run. If an issue disappears for a while and later returns, it reconnects to the existing record instead of arriving as an unrelated alert.

The issue then becomes a playbook. For a page that needs work, the playbook gathers the buyer language and relevant evidence into a brief for what to create or revise. A technical problem, such as a page an engine cannot read, gets a repair plan instead of an article draft.

For Drata, publishing ten posts about questionnaire automation would miss the problem. We would inspect the sources that taught ChatGPT to associate the feature with Vanta, check how clearly Drata's current pages state the capability, and decide which page needs to carry the correction.

<aside class="inline-cta"> <p>Want to see the answers behind your own score?</p> <a href="#" data-cta="waitlist">Run a free benchmark &rarr;</a> </aside>

## Measure after something ships

When a team marks a playbook shipped, KnitKnot records the before-state from the completed benchmark. A later full benchmark shows whether the answers, citations, score, or linked issues changed.

We are careful with the word caused. AI answers vary, and a model can change without a company doing anything. When a later answer cites the new page and corrects the targeted claim, KnitKnot preserves that receipt. Other movement stays descriptive. We publish our measurement rules and noise tests in the [open lab notebook](/experiments/) because that boundary belongs in the product.

I still expect the scoring and issue system to change as we collect more history. The part I feel certain about is the record underneath it. A company should be able to move from any score to the answer, and from the answer to the evidence that explains it.

When I open that old Drata evaluation now, the bad answer has not vanished. We know what ChatGPT told the buyer that day. We can see why the score was low and where the comparison went wrong.

We built KnitKnot so companies can keep that record and decide what to change before the next buyer asks.

Raw mirror of this content: https://knitknot.ai/blog/introducing-knitknot.md. Site-wide summary: /llms.txt · full content: /llms-full.txt

← Field notes
Published
Apr 14, 2026
Reading time
6 minutes
Filed under
product · benchmarks

KnitKnot field note

Introducing KnitKnot

KnitKnot runs the questions software buyers ask across four AI engines, saves every answer, checks the claims, and turns repeated problems into work a company can ship.

Kevin Kho

Co-founder, KnitKnot

A governed page revision bounded by approved facts, review marks, and release recordsEvidence plate / productObserved · sourced · reviewed

The old Drata answer from our first post has a KnitKnot score of 28.

That number is useful. It is also the least interesting part of the record.

What matters is everything underneath it. I can open the answer ChatGPT gave, read the exact line that favored Vanta, see which pages it cited, and compare the same question across Claude, Perplexity, and Gemini. The number tells me where to look. The saved answer tells me what happened.

This distinction shaped most of the product. We wanted a record that could survive after the answer changed, rather than another dashboard that tells a company it has an AI visibility problem and leaves the marketing team to guess what caused it.

Start with the questions

A benchmark begins with a company profile. KnitKnot researches the products, features, positioning, competitors, and buyer personas associated with the company. Each fact has a source, and the company can correct anything we got wrong before the run starts.

Those inputs become a prompt library. Some prompts ask open category questions, such as which compliance tools are best for a mid-market company. Others compare two vendors or drill into a feature that matters to a particular buyer.

We ground the library in search queries with measured demand, then layer in the company’s actual products and competitors. This keeps us from inventing a tidy set of questions that nobody asks. It also gives the benchmark a stable shape. A team can archive a bad prompt or add one we missed, but the useful prompts persist from one run to the next.

That persistence matters. If the questions change every week, the score can move even when the answers do not. You end up measuring your own prompt generator.

Run each question as a fresh conversation

KnitKnot sends the active prompt library to ChatGPT, Claude, Perplexity, and Gemini. Each prompt starts in a fresh conversation with no account history or prior messages. We save the returned answer along with the engine, model version, timestamp, and any citations the engine exposed.

The same question often produces four different buying stories. One engine may recommend you, another may leave you out, and a third may repeat a claim that disappeared from your website months ago. An average hides most of that. We keep the engines separate so the company can read each answer.

This is a controlled benchmark, not a recording of every buyer’s screen. Location, account history, personalization, and a model update can all change what an individual sees. A fresh conversation gives us a repeatable starting point. AI still varies.

Keep the answer before scoring it

The complete response is the first artifact we store. We do this before trying to summarize or score it.

That sounds obvious, but it changes what the product can explain later. If a benchmark only stores a score, there is no way to tell whether a movement came from the recommendation or from a change in the facts and sources behind it. The original answer is gone.

Once the response is saved, KnitKnot extracts the parts a company can inspect: the recommendation, feature comparisons, factual claims, coverage, and cited domains.

We check material claims against the researched company profile and its dated sources. A claim we can verify gets a factual verdict. A claim without enough evidence stays unverifiable. We would rather leave the field unresolved than invent certainty because a report looks cleaner with every row filled in.

Separate showing up from being chosen

We made one early mistake by treating every mention as visibility.

If the prompt says, “Compare Drata and Vanta,” ChatGPT is almost guaranteed to mention Drata. Counting that as organic visibility gives the company credit for appearing in a question that forced its name into the answer.

KnitKnot separates open discovery prompts from competitive prompts. Visibility measures whether an engine surfaces the company when the prompt does not name it. Win rate covers the head-to-head answers where the engine made a choice. A company can do well at one and badly at the other.

This is why the report has more than a single percentage. Coverage shows how central the company was to an answer. Competitive outcomes record wins, losses, and ties. Accuracy tracks the claims we could check. The AI Presence Score combines these signals into a top-line number, but every part of it links back to the evaluations that produced it.

The dashboard does not ask a new judge to reinterpret the answer every time somebody opens a chart. It reads the evaluation facts saved when that response was scored. The report and the drill-down start from the same rows.

Let problems keep their history

One strange answer is evidence, but it may also be noise. The same problem appearing across several questions or later benchmarks is different.

KnitKnot turns material gaps into persistent issues. An issue might track an outdated claim, a missing feature, a losing comparison, or a category where competitors appear and you do not. It keeps the captured answers and supporting evidence attached as the reach changes over time.

This prevents the work from resetting after every run. If an issue disappears for a while and later returns, it reconnects to the existing record instead of arriving as an unrelated alert.

The issue then becomes a playbook. For a page that needs work, the playbook gathers the buyer language and relevant evidence into a brief for what to create or revise. A technical problem, such as a page an engine cannot read, gets a repair plan instead of an article draft.

For Drata, publishing ten posts about questionnaire automation would miss the problem. We would inspect the sources that taught ChatGPT to associate the feature with Vanta, check how clearly Drata’s current pages state the capability, and decide which page needs to carry the correction.

Measure after something ships

When a team marks a playbook shipped, KnitKnot records the before-state from the completed benchmark. A later full benchmark shows whether the answers, citations, score, or linked issues changed.

We are careful with the word caused. AI answers vary, and a model can change without a company doing anything. When a later answer cites the new page and corrects the targeted claim, KnitKnot preserves that receipt. Other movement stays descriptive. We publish our measurement rules and noise tests in the open lab notebook because that boundary belongs in the product.

I still expect the scoring and issue system to change as we collect more history. The part I feel certain about is the record underneath it. A company should be able to move from any score to the answer, and from the answer to the evidence that explains it.

When I open that old Drata evaluation now, the bad answer has not vanished. We know what ChatGPT told the buyer that day. We can see why the score was low and where the comparison went wrong.

We built KnitKnot so companies can keep that record and decide what to change before the next buyer asks.