KnitKnot
You are reading the agent-optimized layer of this page: the literal markdown we serve to AI crawlers and assistants, shipped in the page source of every visit. Making sure AI reads the right facts about a company is literally what KnitKnot does.

# 002 — The placebo backtest: what does our estimator report when nothing happened?

Take a time window in which a customer shipped no changes, pretend a change shipped, and run our own impact estimator. What "effects" does it report?


**Question.** Take a time window in which a customer shipped *no* changes. Pretend a change shipped, and run our own impact estimator — the exact production code path — over random cohorts of buyer questions. What "effects" does it report?

**Why it matters.** Whatever the estimator reports on a null window is pure noise. That distribution is the *noise floor*: no per-fix impact number smaller than it deserves to be published. Most tools in this category report before/after deltas with no such floor; the independent literature (bootstrap confidence intervals of 3–6pp on citation share; run-to-run answer overlap sometimes below 1-in-100) says the floor is large.

## Method

  • - A production workspace (anonymized) with a **stable question library** across two benchmark
    runs ten days apart, and verifiably zero shipped changes in between.
  • - 200 random cohorts of 15 questions each ("pretend this cohort was treated"), each run through
    the production difference-in-differences path: treated-cohort delta minus untouched-holdout
    delta.
  • - Gate: on a null window, the mean effect must be statistically indistinguishable from zero.

## Results

| Metric | Mean null effect | 95% null interval | |---|---|---| | Win rate | −1.1pp | **[−24.9, +22.2]pp** | | Visibility | −0.1pp | [−19.8, +19.2]pp |

The gate passes (no *systematic* bias) — but the interval is the story: **a 15-question cohort can swing ±23 points of win rate between two runs with zero real change.** Roughly one in five null cohorts shows an apparent effect of 10 points or more.

## What we changed because of this

  • - **We do not publish per-fix effect sizes at this cohort size.** Every one of our own nine
    historical per-fix impact numbers fell inside the null band; they were retired.
  • - The placebo backtest is now a **standing build gate**: any change to the estimator must pass
    it on a null window before shipping.
  • - Proof of impact is reported as **receipts** (the AI answer that changed, citing the shipped
    page, with fetch dates) and **sustained multi-run trend**, with the noise floor printed next
    to any aggregate movement.

## Honest caveats

  • - The null window is assumed treatment-free because no tracked change shipped in it; customer
    changes made outside the platform aren't logged and could contaminate it. If they did, the
    true floor is *lower* than measured — which would not change the conclusion's direction.
  • - One workspace, one run pair. The floor will be re-estimated per customer as clean run pairs
    accumulate; nothing about the method is workspace-specific.

Raw mirror of this content: https://knitknot.ai/experiments/002-placebo-noise-floor.md. Site-wide summary: /llms.txt · full content: /llms-full.txt

Measurement lab Experiment 002

Experiment 002

002 — The placebo backtest: what does our estimator report when nothing happened?

Take a time window in which a customer shipped no changes, pretend a change shipped, and run our own impact estimator. What "effects" does it report?

Headline result: 95% null interval ≈ ±23pp win-rate on 15-prompt cohorts — the published noise floor.

Question. Take a time window in which a customer shipped no changes. Pretend a change shipped, and run our own impact estimator — the exact production code path — over random cohorts of buyer questions. What “effects” does it report?

Why it matters. Whatever the estimator reports on a null window is pure noise. That distribution is the noise floor: no per-fix impact number smaller than it deserves to be published. Most tools in this category report before/after deltas with no such floor; the independent literature (bootstrap confidence intervals of 3–6pp on citation share; run-to-run answer overlap sometimes below 1-in-100) says the floor is large.

Method

  • A production workspace (anonymized) with a stable question library across two benchmark runs ten days apart, and verifiably zero shipped changes in between.
  • 200 random cohorts of 15 questions each (“pretend this cohort was treated”), each run through the production difference-in-differences path: treated-cohort delta minus untouched-holdout delta.
  • Gate: on a null window, the mean effect must be statistically indistinguishable from zero.

Results

MetricMean null effect95% null interval
Win rate−1.1pp[−24.9, +22.2]pp
Visibility−0.1pp[−19.8, +19.2]pp

The gate passes (no systematic bias) — but the interval is the story: a 15-question cohort can swing ±23 points of win rate between two runs with zero real change. Roughly one in five null cohorts shows an apparent effect of 10 points or more.

What we changed because of this

  • We do not publish per-fix effect sizes at this cohort size. Every one of our own nine historical per-fix impact numbers fell inside the null band; they were retired.
  • The placebo backtest is now a standing build gate: any change to the estimator must pass it on a null window before shipping.
  • Proof of impact is reported as receipts (the AI answer that changed, citing the shipped page, with fetch dates) and sustained multi-run trend, with the noise floor printed next to any aggregate movement.

Honest caveats

  • The null window is assumed treatment-free because no tracked change shipped in it; customer changes made outside the platform aren’t logged and could contaminate it. If they did, the true floor is lower than measured — which would not change the conclusion’s direction.
  • One workspace, one run pair. The floor will be re-estimated per customer as clean run pairs accumulate; nothing about the method is workspace-specific.