Product

From a single AI answer to an operable visibility asset

Four independent subsystems, each producing pure data; one engagement path that turns “being seen” into a long-term asset.

Four subsystems

recall

Recall pipeline

Playwright drives real web-interface evidence across the 6 must-test engines, scores each answer on five dimensions, and aggregates recall and Share of Voice.

  • The 6 must-test engines: DeepSeek, Doubao, Tongyi, Yuanbao, Kimi, Ernie.
  • Serial capture, resumable, multi-account failover; on rate-limits it rotates accounts or skips an engine.
  • Supports “first question + ≤2 follow-ups” chains that mirror real user behavior.
  • Deep-think dual track (normal / deep) captured separately, comparing two answers to the same question.
diagnostic

Site audit

Single-domain static-signal audit, three-tier weighted 0–100, plus a parallel agent-reachability L0–L3 and cloaking cross-validation.

  • 8 inspectors: crawler access, discovery, structured data, Chinese locale, Chinese entities, content, structure, agent reach.
  • Checks whether 37 AI crawlers are allowed: GPTBot, ClaudeBot, PerplexityBot, Baiduspider, Bytespider…
  • Chinese entity probes: are you on Baidu Baike / Wikipedia / Zhihu / WeChat / Toutiao?
  • Dual-UA diff to catch cloaking that feeds crawlers and browsers different content.
weights

Citation source weights

Aggregates existing evidence into a source × engine matrix of which source types each engine prefers to cite — turning “where to publish” into data.

  • Source taxonomy: portal / vertical / encyclopedia / social / UGC / e-commerce.
  • UGC citation-spike detection flags abnormal cross-batch jumps as anti-manipulation alerts (cf. PoisonedRAG).
  • Engine source preferences are only a starting point — corrected by the client’s measured matrix. Industry lore is not truth.
benchmark

Engine behavior benchmark

A brand-agnostic standard query set measuring response rate, citation rate and source diversity, all with Wilson intervals.

  • Append-only, reproducible, trendable over time.
  • Reuses the same recall pipeline for capture, emitting JSON + public HTML in one step.

Measure → Diagnose → Build → Operate

The baseline measurement stands alone; the next three steps turn “being seen” into a durable asset and practice.

  1. 01
    Measure

    Baseline measurement

    Where do you stand in AI answers right now? A health check that stands on its own.

    OutputGEO baseline report
  2. 02
    Diagnose

    Gap diagnosis

    Why not mentioned / not cited / described wrong? Attributed to fixable gaps.

    OutputGap inventory + priorities
  3. 03
    Build

    Asset building

    Translate brand facts into AI-readable sources: source of truth, knowledge base, site fixes, content.

    OutputSource of truth / KB / fixes
  4. 04
    Operate

    Continuous operation

    Re-measure monthly, attribute, update the knowledge base — because AI answers drift.

    OutputMonthly reports + actions

How BASELINE differs from “we’ll get you to the top” GEO

Typical GEO serviceBASELINE
Core deliverableOptimization promise / keyword listTraceable measurement + priorities
ScoringBlack box / subjective / LLM-judgedDeterministic rules, 4 dims zero-LLM
Public numbersPercentages, pretty curvesRaw counts X/Y + Wilson interval
EvidenceScreenshot decks, hard to verifyAppend-only batches, per-number traceable
Re-measurementOne-offCross-batch diff, monthly operation
Weight basisExperience / loreAnchored to public papers, flags assumptions