KnitKnot
You are reading the agent-optimized layer of this page: the literal markdown we serve to AI crawlers and assistants, shipped in the page source of every visit. Making sure AI reads the right facts about a company is literally what KnitKnot does.

# Most of your citations do nothing

The average AI answer cites 11.6 sources. Across 76,301 measured claims, fewer than 3 of them actually carry anything the buyer reads. We score every owned page on what its citations earn and trace every false or damaging claim back to the page that supplied it.


## Eleven receipts, three that matter

Getting cited is not the win. Getting cited is the ticket to the room where the win happens.

We can now put a number on that. Our benchmark corpus contains 76,301 individual claims extracted from 8,345 AI answers to real B2B buyer questions. **The average answer cites 11.6 sources, but only 2.8 of them supply anything the answer actually says about a company.** Every other citation sits in the footer doing nothing. Seventeen percent of citations shape a claim. The other 83% are decoration.

This is the gap that makes most AI visibility tooling misleading. Citation counts improve, dashboards look healthy, and nothing about the answer changes. So we stopped counting citations and started scoring them. We call the metric **effectiveness**: when the engine cited this page, what happened in the answer?

## What effectiveness actually measures

Effectiveness is scored per page, not per domain, over every scored answer that cited it. Each citing answer contributes one outcome, and the outcome depends on which kind of buyer question it was.

For a visibility question such as *"what are the best tools for X,"* the outcome is how much of the answer was actually about you. We already score that as a coverage enum on every evaluation, and it converts directly to a weight:

<div style="margin: 2em 0; border-block: 1px solid hsl(var(--border)); overflow-x: auto;"> <table style="display: table; width: 100%; border-collapse: collapse; font-size: 14px;"> <thead> <tr style="background: hsl(var(--muted) / 0.5);"> <th style="text-align: left; padding: 10px 16px; font-weight: 600; font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; color: hsl(var(--muted-foreground)); border-bottom: 1px solid hsl(var(--border));">Answer outcome</th> <th style="text-align: center; padding: 10px 16px; font-weight: 600; font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; color: hsl(var(--muted-foreground)); border-bottom: 1px solid hsl(var(--border)); white-space: nowrap;">Weight</th> <th style="text-align: left; padding: 10px 16px; font-weight: 600; font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; color: hsl(var(--muted-foreground)); border-bottom: 1px solid hsl(var(--border));">What the citation earned</th> </tr> </thead> <tbody> <tr> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); font-weight: 500;">Primary / substantial</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: center; font-family: var(--font-mono, monospace); color: hsl(var(--color-brand-teal));">1.0</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); color: hsl(var(--muted-foreground));">The answer is about you</td> </tr> <tr> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); font-weight: 500;">Peripheral</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: center; font-family: var(--font-mono, monospace); color: hsl(var(--color-brand-primary));">0.4</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); color: hsl(var(--muted-foreground));">A sentence, in passing</td> </tr> <tr> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); font-weight: 500;">Incidental</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: center; font-family: var(--font-mono, monospace); color: hsl(var(--color-brand-primary));">0.1</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); color: hsl(var(--muted-foreground));">Your name in a list</td> </tr> <tr> <td style="padding: 10px 16px; font-weight: 500;">Absent</td> <td style="padding: 10px 16px; text-align: center; font-family: var(--font-mono, monospace); color: hsl(var(--outcome-loss));">0.0</td> <td style="padding: 10px 16px; color: hsl(var(--muted-foreground));">Your page was read. You were not mentioned.</td> </tr> </tbody> </table> </div>

For a head-to-head question, the outcome is the decisive result we already derive from [structured comparison signals](/blog/we-stopped-asking-ai-who-wins): win 1.0, tie 0.5, loss 0.0. We discard "not compared" rather than score it as a zero.

Then two guards, because a metric that lies on thin data is worse than no metric:

**A sample floor.** Below eight scorable citing answers, a page gets no effectiveness score at all. It gets counts and an explicit *"not enough citations yet."* A page cited twice cannot earn a confident 100.

**Shrinkage toward the mean.** Even above the floor, the raw average is pulled toward the workspace-wide mean with a pseudo-count of eight. A page that went 3-for-3 doesn't outrank a page that went 40-for-46. This is empirical-Bayes, not editorial judgment, and it is the difference between a leaderboard and a measurement.

The result is a 0–100 number per page. Across the owned pages in our corpus that clear the floor, effectiveness runs from **46 to 95, with a median of 69**. The pages do not cluster around one middling score.

And we deliberately do not band it green/yellow/red. There is no threshold at which a page becomes "effective." The number is a rate, shown as a rate, next to the counts it came from.

## Cited constantly. Invisible anyway.

The most useful thing effectiveness surfaces is a failure mode that citation counting cannot see, because by citation counting it looks like a triumph.

One page in our corpus was cited in 1,127 answers to buyer questions. It was a buyer's-guide post on a customer's own blog, exactly the kind of asset a GEO consultant would call a home run. **In 662 of those answers, the company was never mentioned at all.**

Fifty-nine percent silent. The engines read the page, used it to frame the category, listed the vendors it discussed, and left out the vendor that wrote it. That page is doing real work in the market. It's doing it for everyone else.

On the same domain, a product feature page cited 45 times was silent in 4. Same brand, same authority, same buyers, same crawler. One page was silent 9% of the time; the other, 59%. Across every owned page in the corpus, **28.3% of citing answers never name the company whose page was cited.**

We file that as its own tracked issue, `cited_but_silent`, when a page clears the sample floor and is silent in the majority of the answers that cite it. It ranks by raw silent *count* before share, on purpose: a page silent in 662 of 1,127 answers is a bigger problem than one silent in 18 of 18. A ranking that says otherwise is optimizing for tidy percentages instead of buyers.

The fix is almost never "get cited more." The page is structured so the engine can extract the category without extracting you. You're one row among ten in the comparison table, item six in your own listicle, or a single mention in the closing paragraph that the summarizer drops.

## The page is the unit. Not the domain.

Domain-level authority scores are the other thing effectiveness kills.

Take the domain with the widest measured spread in our corpus. Nine of its pages clear the sample floor. Their pull-through, the share of citing answers where the company was actually present, runs from **8 to 91 out of 100**. One domain. One brand. An 11× gap between its worst-performing cited page and its best.

<div style="margin: 2em 0; border-block: 1px solid hsl(var(--border)); overflow-x: auto;"> <table style="display: table; width: 100%; border-collapse: collapse; font-size: 14px;"> <thead> <tr style="background: hsl(var(--muted) / 0.5);"> <th style="text-align: left; padding: 10px 16px; font-weight: 600; font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; color: hsl(var(--muted-foreground)); border-bottom: 1px solid hsl(var(--border));">Domain (owned pages over the floor)</th> <th style="text-align: right; padding: 10px 16px; font-weight: 600; font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; color: hsl(var(--muted-foreground)); border-bottom: 1px solid hsl(var(--border));">Weakest page</th> <th style="text-align: right; padding: 10px 16px; font-weight: 600; font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; color: hsl(var(--muted-foreground)); border-bottom: 1px solid hsl(var(--border));">Strongest page</th> </tr> </thead> <tbody> <tr> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); font-weight: 500;">B2B SaaS vendor (9 pages)</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; font-variant-numeric: tabular-nums;"><span style="color: hsl(var(--outcome-loss)); font-weight: 600;">8</span></td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; font-variant-numeric: tabular-nums;"><span style="color: hsl(var(--color-brand-teal)); font-weight: 600;">91</span></td> </tr> <tr> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); font-weight: 500;">OSS infrastructure project (5 pages)</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; font-variant-numeric: tabular-nums;"><span style="color: hsl(var(--color-brand-primary)); font-weight: 600;">51</span></td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; font-variant-numeric: tabular-nums;"><span style="color: hsl(var(--color-brand-teal)); font-weight: 600;">100</span></td> </tr> <tr> <td style="padding: 10px 16px; font-weight: 500;">Developer database (12 pages)</td> <td style="padding: 10px 16px; text-align: right; font-variant-numeric: tabular-nums;"><span style="color: hsl(var(--color-brand-teal)); font-weight: 600;">82</span></td> <td style="padding: 10px 16px; text-align: right; font-variant-numeric: tabular-nums;"><span style="color: hsl(var(--color-brand-teal)); font-weight: 600;">100</span></td> </tr> </tbody> </table> </div>

That last row matters as much as the first. Every cited page on a tightly scoped technical site can pull the company into the answer. A site with a large content-marketing surface is more likely to vary from page to page. If you only have a domain-level number, you cannot tell those two situations apart, much less tell which of your nine pages to fix.

## Where damaging claims actually come from

Effectiveness answers *did this page help*. The harder question is *did this page hurt*, and answering it requires knowing which cited source supplied which sentence.

Most systems guess. We used to. In a 2026 audit we found that asking a judge model *"which source backs this claim?"* collapsed every claim onto a single comparison-titled URL on a third of evaluations, and contradicted the response's own inline citations on 69% of the claims we hand-checked. Semantic attribution is confident and wrong.

So we replaced it with a deterministic ladder. No model involved, runs at zero cost, and it reads the citation signals the engine itself emitted:

<div style="margin: 2em 0; border-block: 1px solid hsl(var(--border)); overflow-x: auto;"> <table style="display: table; width: 100%; border-collapse: collapse; font-size: 14px;"> <thead> <tr style="background: hsl(var(--muted) / 0.5);"> <th style="text-align: left; padding: 10px 16px; font-weight: 600; font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; color: hsl(var(--muted-foreground)); border-bottom: 1px solid hsl(var(--border));">Tier</th> <th style="text-align: left; padding: 10px 16px; font-weight: 600; font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; color: hsl(var(--muted-foreground)); border-bottom: 1px solid hsl(var(--border));">Signal</th> <th style="text-align: right; padding: 10px 16px; font-weight: 600; font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; color: hsl(var(--muted-foreground)); border-bottom: 1px solid hsl(var(--border));">Share of attributions</th> </tr> </thead> <tbody> <tr> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); font-weight: 500;">1. Character offset</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); color: hsl(var(--muted-foreground));">Engine stamps the exact text position of the citing sentence</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; font-variant-numeric: tabular-nums;">19.4%</td> </tr> <tr> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); font-weight: 500;">2. Footnote marker</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); color: hsl(var(--muted-foreground));">Literal <code>[3]</code> in the prose, mapped to the third source</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; font-variant-numeric: tabular-nums;">21.0%</td> </tr> <tr> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); font-weight: 500;">3. Inline link</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); color: hsl(var(--muted-foreground));">The claim's own sentence carries a markdown link</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; font-variant-numeric: tabular-nums;">59.5%</td> </tr> <tr> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); font-weight: 500;">4. Sole source</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); color: hsl(var(--muted-foreground));">Exactly one distinct URL cited — it can only be that one</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; font-variant-numeric: tabular-nums;">0.1%</td> </tr> <tr> <td style="padding: 10px 16px; font-weight: 500;">5. Null</td> <td style="padding: 10px 16px; color: hsl(var(--muted-foreground));">No signal. We record nothing rather than guess.</td> <td style="padding: 10px 16px; text-align: right; color: hsl(var(--muted-foreground)); font-variant-numeric: tabular-nums;">—</td> </tr> </tbody> </table> </div>

The result is that **only 44.3% of claims get a source.** The rest have no defensible link, so they carry none. A null is more useful than a plausible URL because a downstream finding such as *this competitor's page is feeding the AI false claims about you* has to survive a click.

Traceability varies enormously by engine, and it's purely a function of how each one formats its answers. Claude links inline, so three-quarters of its claims are attributable. Perplexity mostly doesn't, so most of its aren't.

<div style="margin: 2em 0; border-block: 1px solid hsl(var(--border)); overflow-x: auto;"> <table style="display: table; width: 100%; border-collapse: collapse; font-size: 14px;"> <thead> <tr style="background: hsl(var(--muted) / 0.5);"> <th style="text-align: left; padding: 10px 16px; font-weight: 600; font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; color: hsl(var(--muted-foreground)); border-bottom: 1px solid hsl(var(--border));">Engine</th> <th style="text-align: right; padding: 10px 16px; font-weight: 600; font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; color: hsl(var(--muted-foreground)); border-bottom: 1px solid hsl(var(--border));">Claims measured</th> <th style="text-align: right; padding: 10px 16px; font-weight: 600; font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; color: hsl(var(--muted-foreground)); border-bottom: 1px solid hsl(var(--border));">Traceable to a source</th> </tr> </thead> <tbody> <tr> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); font-weight: 600; font-size: 13px;">Claude</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; color: hsl(var(--muted-foreground)); font-variant-numeric: tabular-nums;">23,011</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; font-variant-numeric: tabular-nums;"><span style="color: hsl(var(--color-brand-teal)); font-weight: 600;">74.4%</span></td> </tr> <tr> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); font-weight: 600; font-size: 13px;">ChatGPT</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; color: hsl(var(--muted-foreground)); font-variant-numeric: tabular-nums;">18,840</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; font-variant-numeric: tabular-nums;">38.3%</td> </tr> <tr> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); font-weight: 600; font-size: 13px;">Gemini</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; color: hsl(var(--muted-foreground)); font-variant-numeric: tabular-nums;">17,459</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; font-variant-numeric: tabular-nums;">37.2%</td> </tr> <tr> <td style="padding: 10px 16px; font-weight: 600; font-size: 13px;">Perplexity</td> <td style="padding: 10px 16px; text-align: right; color: hsl(var(--muted-foreground)); font-variant-numeric: tabular-nums;">16,991</td> <td style="padding: 10px 16px; text-align: right; font-variant-numeric: tabular-nums;"><span style="color: hsl(var(--outcome-loss)); font-weight: 600;">17.5%</span></td> </tr> </tbody> </table> </div>

## Who owns the page decides what it says about you

Here is the finding that reframed how we think about the whole problem.

Once every claim is bound to the page that supplied it, and every page is classified as yours, a tracked competitor's, or a third party's, you can ask a question nobody could answer before: **does it matter whose page the engine was reading?**

It matters enormously.

<div style="margin: 2em 0; border-block: 1px solid hsl(var(--border)); overflow-x: auto;"> <table style="display: table; width: 100%; border-collapse: collapse; font-size: 14px;"> <thead> <tr style="background: hsl(var(--muted) / 0.5);"> <th style="text-align: left; padding: 10px 16px; font-weight: 600; font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; color: hsl(var(--muted-foreground)); border-bottom: 1px solid hsl(var(--border));">Source of the claim</th> <th style="text-align: right; padding: 10px 16px; font-weight: 600; font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; color: hsl(var(--muted-foreground)); border-bottom: 1px solid hsl(var(--border));">Claims</th> <th style="text-align: right; padding: 10px 16px; font-weight: 600; font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; color: hsl(var(--muted-foreground)); border-bottom: 1px solid hsl(var(--border));">Damaging</th> <th style="text-align: right; padding: 10px 16px; font-weight: 600; font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; color: hsl(var(--muted-foreground)); border-bottom: 1px solid hsl(var(--border));">Provably false</th> </tr> </thead> <tbody> <tr> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); font-weight: 600; font-size: 13px;">Your own page</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; color: hsl(var(--muted-foreground)); font-variant-numeric: tabular-nums;">14,512</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; font-variant-numeric: tabular-nums;"><span style="color: hsl(var(--color-brand-teal)); font-weight: 600;">2.7%</span></td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; font-variant-numeric: tabular-nums;"><span style="color: hsl(var(--color-brand-teal)); font-weight: 600;">0.01%</span></td> </tr> <tr> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); font-weight: 600; font-size: 13px;">A competitor's page</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; color: hsl(var(--muted-foreground)); font-variant-numeric: tabular-nums;">5,815</td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; font-variant-numeric: tabular-nums;"><span style="color: hsl(var(--outcome-loss)); font-weight: 600;">11.6%</span></td> <td style="padding: 10px 16px; border-bottom: 1px solid hsl(var(--border) / 0.5); text-align: right; font-variant-numeric: tabular-nums;"><span style="color: hsl(var(--outcome-loss)); font-weight: 600;">1.03%</span></td> </tr> <tr> <td style="padding: 10px 16px; font-weight: 600; font-size: 13px;">Everything else</td> <td style="padding: 10px 16px; text-align: right; color: hsl(var(--muted-foreground)); font-variant-numeric: tabular-nums;">13,476</td> <td style="padding: 10px 16px; text-align: right; font-variant-numeric: tabular-nums;"><span style="color: hsl(var(--color-brand-primary)); font-weight: 600;">11.5%</span></td> <td style="padding: 10px 16px; text-align: right; font-variant-numeric: tabular-nums;">0.78%</td> </tr> </tbody> </table> </div>

**A claim the engine drew from a competitor's page is 4.3× more likely to damage you than one it drew from yours.** That is the entire competitive positioning thesis, measured, in one row.

One caveat: the "provably false" gap is wider than it looks because our judge presumes a claim sourced from your own site is true marketing and refuses to accuse it. The damaging column has no such rule. Polarity is assigned before any truth judgment, so 2.7% versus 11.6% is the useful comparison.

Our [confidence calibration rule](/blog/confident-lies-are-worse-than-hedged-ones) matters for the second caveat: of the 5,599 damaging claims in the corpus, only **396 are provably false**. Ninety-three percent are accurate but unflattering. The AI is framing a documented limitation, gap, or tradeoff in a way that costs you the deal. We keep those categories separate. A **misrepresentation** is a specific checkable fact that a dated line in your own record directly negates, and each one includes the quoted contradiction, its URL, and the date it was true as of. **Negative framing** covers the rest. It gets a "counter this," never an "AI is lying about you." A tool that calls a true criticism a lie will be caught once and then ignored.

There's a third pattern worth naming. Some of the sources feeding damaging claims aren't competitor marketing at all. They're the surfaces you might assume were neutral. Grouped by source type, the claims engines pulled from social posts (15.4% damaging), academic writing (14.1%), independent blogs (12.3%) and community threads (11.5%) were roughly twice as damaging as the ones they pulled from vendor sites (6.8%), registries (6.2%) or news (6.6%). Someone's Reddit comment from 2023 is a primary source in your category, and it is aging.

## What the measurement is for

None of this is worth much as a report card. It's worth something as a work queue.

Every finding above resolves to a specific, tracked issue with a stable identity, so it can be watched across runs rather than rediscovered every month. A page cited constantly but silent in most of its answers opens a `cited_but_silent` issue naming that URL and that count. A false claim opens a `misrepresentation` issue keyed on the *content* of the claim, not its wording. It is content-addressed and semantically deduped, so the same lie reworded three ways stays one issue instead of splitting into three. Both are gated: an issue has to survive two consecutive runs before it opens, because a single run's noise is not a finding.

From there they become different kinds of work. A misrepresentation is a *fix*: plant a dated, sourced, quotable correction of that specific false fact on a page the engine already reads. A cited-but-silent page is a *reinforce*: the citation is already yours, but the extraction is broken. A competitor page grounding damaging claims about you is a *create*: the ground is theirs and it needs contesting. Same evidence, three different verbs. Choose the wrong one and a content team can spend a quarter without moving the answer.

Then the next benchmark re-runs the same buyer questions and tells you whether the page's effectiveness moved. Not whether it got more citations. Whether the citations started earning something.


Citation counts are becoming the default metric for this category. The number that matters is what happens in the answer after the engine opens your page. We can measure it for each page and show the receipts.

If you want to see which of your pages the engines actually read, what each one earns, and which competitor page is grounding the claims costing you deals, [a benchmark shows you in a day](/examples/). The first one is free.

Raw mirror of this content: https://knitknot.ai/blog/most-of-your-citations-do-nothing.md. Site-wide summary: /llms.txt · full content: /llms-full.txt

← Field notes
Published
Aug 03, 2026
Reading time
16 minutes
Filed under
sources · methodology · research

KnitKnot field note

Most of your citations do nothing

The average AI answer cites 11.6 sources. Across 76,301 measured claims, fewer than 3 of them actually carry anything the buyer reads. We score every owned page on what its citations earn and trace every false or damaging claim back to the page that supplied it.

Max Wiesner

Co-founder, KnitKnot

A source trail linking claim fragments to citations and evidence receiptsEvidence plate / sourcesObserved · sourced · reviewed

Eleven receipts, three that matter

Getting cited is not the win. Getting cited is the ticket to the room where the win happens.

We can now put a number on that. Our benchmark corpus contains 76,301 individual claims extracted from 8,345 AI answers to real B2B buyer questions. The average answer cites 11.6 sources, but only 2.8 of them supply anything the answer actually says about a company. Every other citation sits in the footer doing nothing. Seventeen percent of citations shape a claim. The other 83% are decoration.

This is the gap that makes most AI visibility tooling misleading. Citation counts improve, dashboards look healthy, and nothing about the answer changes. So we stopped counting citations and started scoring them. We call the metric effectiveness: when the engine cited this page, what happened in the answer?

What effectiveness actually measures

Effectiveness is scored per page, not per domain, over every scored answer that cited it. Each citing answer contributes one outcome, and the outcome depends on which kind of buyer question it was.

For a visibility question such as “what are the best tools for X,” the outcome is how much of the answer was actually about you. We already score that as a coverage enum on every evaluation, and it converts directly to a weight:

Answer outcome Weight What the citation earned
Primary / substantial 1.0 The answer is about you
Peripheral 0.4 A sentence, in passing
Incidental 0.1 Your name in a list
Absent 0.0 Your page was read. You were not mentioned.

For a head-to-head question, the outcome is the decisive result we already derive from structured comparison signals: win 1.0, tie 0.5, loss 0.0. We discard “not compared” rather than score it as a zero.

Then two guards, because a metric that lies on thin data is worse than no metric:

A sample floor. Below eight scorable citing answers, a page gets no effectiveness score at all. It gets counts and an explicit “not enough citations yet.” A page cited twice cannot earn a confident 100.

Shrinkage toward the mean. Even above the floor, the raw average is pulled toward the workspace-wide mean with a pseudo-count of eight. A page that went 3-for-3 doesn’t outrank a page that went 40-for-46. This is empirical-Bayes, not editorial judgment, and it is the difference between a leaderboard and a measurement.

The result is a 0–100 number per page. Across the owned pages in our corpus that clear the floor, effectiveness runs from 46 to 95, with a median of 69. The pages do not cluster around one middling score.

And we deliberately do not band it green/yellow/red. There is no threshold at which a page becomes “effective.” The number is a rate, shown as a rate, next to the counts it came from.

Cited constantly. Invisible anyway.

The most useful thing effectiveness surfaces is a failure mode that citation counting cannot see, because by citation counting it looks like a triumph.

One page in our corpus was cited in 1,127 answers to buyer questions. It was a buyer’s-guide post on a customer’s own blog, exactly the kind of asset a GEO consultant would call a home run. In 662 of those answers, the company was never mentioned at all.

Fifty-nine percent silent. The engines read the page, used it to frame the category, listed the vendors it discussed, and left out the vendor that wrote it. That page is doing real work in the market. It’s doing it for everyone else.

On the same domain, a product feature page cited 45 times was silent in 4. Same brand, same authority, same buyers, same crawler. One page was silent 9% of the time; the other, 59%. Across every owned page in the corpus, 28.3% of citing answers never name the company whose page was cited.

We file that as its own tracked issue, cited_but_silent, when a page clears the sample floor and is silent in the majority of the answers that cite it. It ranks by raw silent count before share, on purpose: a page silent in 662 of 1,127 answers is a bigger problem than one silent in 18 of 18. A ranking that says otherwise is optimizing for tidy percentages instead of buyers.

The fix is almost never “get cited more.” The page is structured so the engine can extract the category without extracting you. You’re one row among ten in the comparison table, item six in your own listicle, or a single mention in the closing paragraph that the summarizer drops.

The page is the unit. Not the domain.

Domain-level authority scores are the other thing effectiveness kills.

Take the domain with the widest measured spread in our corpus. Nine of its pages clear the sample floor. Their pull-through, the share of citing answers where the company was actually present, runs from 8 to 91 out of 100. One domain. One brand. An 11× gap between its worst-performing cited page and its best.

Domain (owned pages over the floor) Weakest page Strongest page
B2B SaaS vendor (9 pages) 8 91
OSS infrastructure project (5 pages) 51 100
Developer database (12 pages) 82 100

That last row matters as much as the first. Every cited page on a tightly scoped technical site can pull the company into the answer. A site with a large content-marketing surface is more likely to vary from page to page. If you only have a domain-level number, you cannot tell those two situations apart, much less tell which of your nine pages to fix.

Where damaging claims actually come from

Effectiveness answers did this page help. The harder question is did this page hurt, and answering it requires knowing which cited source supplied which sentence.

Most systems guess. We used to. In a 2026 audit we found that asking a judge model “which source backs this claim?” collapsed every claim onto a single comparison-titled URL on a third of evaluations, and contradicted the response’s own inline citations on 69% of the claims we hand-checked. Semantic attribution is confident and wrong.

So we replaced it with a deterministic ladder. No model involved, runs at zero cost, and it reads the citation signals the engine itself emitted:

Tier Signal Share of attributions
1. Character offset Engine stamps the exact text position of the citing sentence 19.4%
2. Footnote marker Literal [3] in the prose, mapped to the third source 21.0%
3. Inline link The claim's own sentence carries a markdown link 59.5%
4. Sole source Exactly one distinct URL cited — it can only be that one 0.1%
5. Null No signal. We record nothing rather than guess.

The result is that only 44.3% of claims get a source. The rest have no defensible link, so they carry none. A null is more useful than a plausible URL because a downstream finding such as this competitor’s page is feeding the AI false claims about you has to survive a click.

Traceability varies enormously by engine, and it’s purely a function of how each one formats its answers. Claude links inline, so three-quarters of its claims are attributable. Perplexity mostly doesn’t, so most of its aren’t.

Engine Claims measured Traceable to a source
Claude 23,011 74.4%
ChatGPT 18,840 38.3%
Gemini 17,459 37.2%
Perplexity 16,991 17.5%

Who owns the page decides what it says about you

Here is the finding that reframed how we think about the whole problem.

Once every claim is bound to the page that supplied it, and every page is classified as yours, a tracked competitor’s, or a third party’s, you can ask a question nobody could answer before: does it matter whose page the engine was reading?

It matters enormously.

Source of the claim Claims Damaging Provably false
Your own page 14,512 2.7% 0.01%
A competitor's page 5,815 11.6% 1.03%
Everything else 13,476 11.5% 0.78%

A claim the engine drew from a competitor’s page is 4.3× more likely to damage you than one it drew from yours. That is the entire competitive positioning thesis, measured, in one row.

One caveat: the “provably false” gap is wider than it looks because our judge presumes a claim sourced from your own site is true marketing and refuses to accuse it. The damaging column has no such rule. Polarity is assigned before any truth judgment, so 2.7% versus 11.6% is the useful comparison.

Our confidence calibration rule matters for the second caveat: of the 5,599 damaging claims in the corpus, only 396 are provably false. Ninety-three percent are accurate but unflattering. The AI is framing a documented limitation, gap, or tradeoff in a way that costs you the deal. We keep those categories separate. A misrepresentation is a specific checkable fact that a dated line in your own record directly negates, and each one includes the quoted contradiction, its URL, and the date it was true as of. Negative framing covers the rest. It gets a “counter this,” never an “AI is lying about you.” A tool that calls a true criticism a lie will be caught once and then ignored.

There’s a third pattern worth naming. Some of the sources feeding damaging claims aren’t competitor marketing at all. They’re the surfaces you might assume were neutral. Grouped by source type, the claims engines pulled from social posts (15.4% damaging), academic writing (14.1%), independent blogs (12.3%) and community threads (11.5%) were roughly twice as damaging as the ones they pulled from vendor sites (6.8%), registries (6.2%) or news (6.6%). Someone’s Reddit comment from 2023 is a primary source in your category, and it is aging.

What the measurement is for

None of this is worth much as a report card. It’s worth something as a work queue.

Every finding above resolves to a specific, tracked issue with a stable identity, so it can be watched across runs rather than rediscovered every month. A page cited constantly but silent in most of its answers opens a cited_but_silent issue naming that URL and that count. A false claim opens a misrepresentation issue keyed on the content of the claim, not its wording. It is content-addressed and semantically deduped, so the same lie reworded three ways stays one issue instead of splitting into three. Both are gated: an issue has to survive two consecutive runs before it opens, because a single run’s noise is not a finding.

From there they become different kinds of work. A misrepresentation is a fix: plant a dated, sourced, quotable correction of that specific false fact on a page the engine already reads. A cited-but-silent page is a reinforce: the citation is already yours, but the extraction is broken. A competitor page grounding damaging claims about you is a create: the ground is theirs and it needs contesting. Same evidence, three different verbs. Choose the wrong one and a content team can spend a quarter without moving the answer.

Then the next benchmark re-runs the same buyer questions and tells you whether the page’s effectiveness moved. Not whether it got more citations. Whether the citations started earning something.


Citation counts are becoming the default metric for this category. The number that matters is what happens in the answer after the engine opens your page. We can measure it for each page and show the receipts.

If you want to see which of your pages the engines actually read, what each one earns, and which competitor page is grounding the claims costing you deals, a benchmark shows you in a day. The first one is free.