← All research

Why we report raw counts, not percentages

At N=5, "0%" and "20%" differ by a single answer yet read completely differently. Here's why we insist on X/Y raw counts with a Wilson 95% confidence interval.

Open almost any GEO report and you’ll find the sentence: “your visibility on Doubao is 20%.” It sounds rigorous. The problem: behind that number there are often just 5 measurements.

A percentage over 5 runs is noise

Suppose we run the same set of questions on one engine 5 times, and the brand is mentioned once. The report says “20%.” But if that one mention becomes zero, the number drops from 20% to 0% — a single answer’s jitter manufacturing a percentage that looks wildly different.

The same in reverse: 3/5 becomes “60%,” 4/5 becomes “80%,” separated by a single answer. When the sample is in the single digits, a percentage dresses up measurement error as a firm conclusion. A legal team will see through it, and so will a competitor.

“Percentages on small samples are noise that won’t survive scrutiny — so we report raw counts with a Wilson 95% confidence interval.”

So we write X/Y, with an interval

Public reports show only raw counts: “0 out of 5,” “15 out of 40.” It doesn’t hide the sample size; a reader immediately knows how much weight the conclusion carries.

On top of the raw count we add a Wilson 95% confidence interval — a closed-form, deterministic interval (E.B. Wilson, 1927) whose small-sample coverage beats the common Wald interval and needs no randomized procedure like bootstrapping. In other words: the same evidence yields the same interval, no matter who computes it.

The discipline has support in the literature:

  • Don’t Measure Once (arXiv 2604.07585): one measurement is untrustworthy; binary outcomes must carry uncertainty.
  • Quantifying Uncertainty in Answer Engine Visibility (arXiv 2603.08924): quantifying uncertainty is a prerequisite for visibility research.

This isn’t just honesty — it’s self-protection

Writing noise as a percentage looks good short term and is a time bomb long term: a client waves “20%→35%” and asks what you did, when the truth may be jitter. We’d rather write “1 of 5 → 2 of 5 (overlapping intervals)” and have the conclusion hold up.

The value of measurement is precisely that it doesn’t exaggerate on your behalf. A report that won’t collapse under scrutiny is the only kind worth deciding on.


Read our full methodology, or see what we measure.