Deterministic, not AI scoring
Four of the five dimensions (mention, rank, cited-as-source, source platforms) are pure code — exact string matching, position tertiles, source-domain matching, zero LLM, zero randomness. The only subjective dimension, “sentiment / frame,” is locked under a confidence discipline: auto-drafts are always recorded as L3, and only human review promotes them to L1 before they enter a report. The result: a different person on a different day, given the same batch, reaches a byte-for-byte identical conclusion.
- Mention: exact match of all brand aliases in the answer body, yes/no.
- Rank: ordinal in numbered lists + whether it falls in the first 600 characters.
- Cited as source: URL / title / platform name matched against brand aliases.
- Source platforms: inherited from the cited source domains of that answer.
Raw counts, no public percentages
“0 mentions out of 5” is far more honest than “0%.” A percentage at N=5 is noise that won’t survive your legal team or a competitor. So public reports show only raw counts (X/Y) with a Wilson 95% confidence interval — a closed-form, deterministic interval whose small-sample coverage beats Wald (E.B. Wilson, 1927). We cite "Don’t Measure Once" for the principle: one measurement is untrustworthy and must carry uncertainty.
“Listed as a source” ≠ “entered the answer body”
These are two independent signals. Share of Voice defaults to absorption (entering the body); selection (listed as a source but not entering the body) is reported separately and never mixed in. This distinction comes straight from the Citation Selection → Absorption research.
Site audit: three fixed-weight tiers, 45 / 30 / 25
The site audit splits a domain’s “AI-readability” into three tiers with fixed weights that are identical for every client: basic access (robots / llms.txt / HTTPS / server-readability) is 45, structured data and discoverability is 30, content and Chinese entities is 25. Basic access carries the most weight because C-SEO Bench (NeurIPS 2025) found infrastructure dominates content tweaks. The band cut-lines 80 / 60 / 40 are anchored to the effect-size distributions of two public benchmarks (KDD 2024 and C-SEO Bench), not pulled from thin air.
| 80–100 | Excellent |
| 60–79 | Good |
| 40–59 | Foundation |
| 0–39 | Critical |
Agent reachability: a parallel L0–L3 ladder
AI is moving from reading-for-you to buying-for-you. The site audit runs a parallel agent-reachability dimension (kept out of the total score) judged on endpoint (30) / product (30) / commerce (40) signals into L0–L3: can an agent discover you, read structured products and prices, and follow the path to checkout. It is reported separately because it measures a different future.
Traceability is a discipline, not a feature
Evidence is stored append-only by date, resumable, never overwritten; re-measuring is a diff between two batches. Every measured number in a report is wrapped and anchored back to the screenshot of that exact answer. A quality gate automatically blocks bare numbers, dangling references and public percentages — and it cannot be bypassed to ship a report. Research (Citations and Trust, AAAI 2025) shows citations raise trust even when fake — but trust collapses the moment a user actually checks. So we reward the opposite: citations that are real, and really accurate.
Citations raise trust even when they are fake. So we bet everything on one thing: a citation must click back to real evidence.
Literature map
Whose shoulders each scoring rule stands on. Engineering heuristics without a source are flagged in code as “assumption · tunable · not settled.”
- GEO: Generative Engine Optimization Formalizes visibility as coverage / position / influence. Measured lifts: adding statistics +33–41%, citing sources +30%, adding quotations +28–41% in citation rate. arXiv:2311.09735 ↗
- C-SEO Bench Infrastructure dominates content tweaks — our three-tier site audit weights basic access highest (45) on this basis. arXiv:2506.11097 ↗
- From Citation Selection to Citation Absorption Being listed as a source ≠ entering the answer body — two independent signals. Share of voice defaults to absorption; selection is reported separately. arXiv:2604.25707 ↗
- Don't Measure Once One measurement is not trustworthy; small samples need uncertainty → binary data use the Wilson 95% interval (E.B. Wilson, 1927). arXiv:2604.07585 ↗
- Recency Bias in LLM Reranking Changing only the timestamp shifts Top-10 years by up to 4.78 and flips pairwise preference up to 25% → freshness is a hard retrieval mechanism, weighted highly. arXiv:2509.11353 ↗
- GEO-SFE / SAGEO Arena Structural remodeling +17.3% citations; pure prose rewriting −9% vs structure +22% → document structure is the most stable citation signal. arXiv:2603.29979 ↗
- Cited but Not Verified Link validity >94% but factual accuracy only 39–77% → "cited ≠ accurately represented." Fidelity must be audited separately. arXiv:2605.06635 ↗