KnitKnot intelligence
Field notes.
What we are building, testing, measuring, and learning about the answers AI agents assemble.
Featured evidence plate Aug 14, 2026 · 5 min read
Why AI recommends only 4-7 vendors in your category
Google returns 10 results. AI usually returns 4-7 names. There's no page two. If you're not in the set, you're not in the evaluation.
The archive
17 notes
Plate 02sourcesAug 03, 2026 · 16 min
Most of your citations do nothing
The average AI answer cites 11.6 sources. Across 76,301 measured claims, fewer than 3 of them actually carry anything the buyer reads. We score every owned page on what its citations earn and trace every false or damaging claim back to the page that supplied it.
Plate 03sourcesJul 25, 2026 · 9 min
AI is citing pages that don't exist
We probed every URL that ChatGPT, Claude, Perplexity, and Gemini cited across nearly 10,000 buyer-question answers. One citation in eighteen points at a page that is permanently gone, and nearly a third of answers lean on at least one.
Plate 04AI presenceJul 15, 2026 · 6 min
The hidden cost of AI misinformation
When AI gets a fact wrong about your company, it doesn't show up in your CRM as a lost deal. It shows up as a deal that never existed. We tried to quantify what that costs.
Plate 05researchJun 28, 2026 · 10 min
72% of brands have factual errors in AI responses
We analyzed 33,000 AI evaluations across ChatGPT, Claude, Perplexity, and Gemini for 47 B2B companies. 72% had at least one verifiably wrong factual claim. The errors cluster into five predictable, fixable patterns.
Plate 06AI presenceJun 18, 2026 · 6 min
The 10 questions AI buyers ask that your website can't answer
We generate benchmark prompts grounded in real Google search data, with search volume attached to each one. The questions buyers ask ChatGPT, Claude, Perplexity, and Gemini are more adversarial, more specific, and more comparative than anything your website was designed to handle. Here are the ten patterns that show up most.
Plate 07researchJun 12, 2026 · 9 min
Why AI recommends your competitor instead of you
We analyzed 33,000 AI evaluations across four models. The most surprising finding: models disagree with each other on who to recommend 48.6% of the time. Which model the buyer opens matters more than most companies realize.
Plate 08benchmarksJun 10, 2026 · 8 min
What ChatGPT says when a buyer asks to compare you
We ran the same comparison prompt across ChatGPT, Claude, Perplexity, and Gemini for a B2B company. Four models gave four different answers. Two got the pricing wrong. One recommended the competitor based entirely on the competitor's own blog post.
Plate 09sourcesJun 06, 2026 · 6 min
Not all citations are equal
A source that shaped the AI's recommendation carries more weight than one that provided a background fact. We model which sources had the most influence over what the buyer heard.
Plate 10AI presenceJun 03, 2026 · 8 min
AI is lying about your company
We pulled every factual claim from our first 2,000 benchmark evaluations and checked them against reality. The error rate was higher than we expected, and the errors weren't random.
Plate 11methodologyMay 28, 2026 · 7 min
A customer told us our benchmark was rigged
We designed adversarial prompts to show companies where AI was misrepresenting them. Customers kept getting defensive about the prompts themselves. So we rebuilt the whole thing around real buyer behavior.
Plate 12benchmarksMay 21, 2026 · 5 min
Prompt libraries are coverage optimization problems
A bigger prompt library doesn't mean a better benchmark. We had hundreds of prompts and still missed the buyer situations that mattered most.
Plate 13methodologyMay 14, 2026 · 5 min
Approximating the Claude Engine
ChatGPT, Perplexity, and Gemini all have incognito search. Claude doesn't. To benchmark how Claude represents companies, we had to find a way that respects Anthropic's terms instead of working around them. Here's what we built.
Plate 14measurementMay 07, 2026 · 6 min
Confident lies are worse than hedged ones
Accuracy and conviction are independent axes. Most AI benchmarks only measure the first one. We model the interaction between what AI knows and how sure it sounds.
Plate 15productApr 29, 2026 · 4 min
What a candidate asks AI about your company
A senior leader at a mid-size company asked us to track how AI describes them to candidates weighing offers. It wasn't our use case. With barely any changes, it worked.
Plate 16scoringApr 22, 2026 · 8 min
We stopped asking AI who wins
Most LLM-as-judge systems ask one question: who's better? We decompose into structured signals and derive the outcome deterministically. Here's why.
Plate 17productApr 14, 2026 · 6 min
Introducing KnitKnot
KnitKnot runs the questions software buyers ask across four AI engines, saves every answer, checks the claims, and turns repeated problems into work a company can ship.
Plate 18launchApr 07, 2026 · 7 min
Why we pivoted KnitKnot
We started KnitKnot as a digital sales room. Buyers liked it, but nobody needed it. Then we began asking what a sales room should look like when the buyer is an AI agent.