GEO · Generative Engine Optimization

Inside AI answers,
how is your brand described?

Buyers already make decisions inside DeepSeek, Doubao and Tongyi. BASELINE measures what 6 Chinese AI engines actually answer about you and your competitors — who gets mentioned, whose content gets cited, who ranks first. Deterministic rules, raw counts, and every number traces back to a real screenshot.

Not another “we’ll get you to the top” promise. An honest measurement.

6
Chinese AI engines, always tested
DeepSeek · Doubao · Tongyi · Yuanbao · Kimi · Ernie
5
deterministic scoring dimensions
4 zero-LLM rules · 1 human-verified to L1
100%
of numbers are traceable
every count clicks back to the raw answer

Why now

The decision layer is moving from the search box to the chat box

When users stop scanning ten blue links and simply ask AI “who’s good for this?”, your brand is either written into that answer — or it doesn’t exist. It’s a brand-new shelf, and you can’t see it.

01

Invisible until a customer screenshots it

You can still check rankings on a search engine. AI answers shift by person, phrasing and time — you have no dashboard, unless someone measures them for you, answer by answer.

02

Most “optimization” is ineffective or drifts

The academic benchmark C-SEO Bench (NeurIPS 2025) shows that "many GEO tactics are unstable in real conditions and drift with models and competition." The only durable value is trustworthy measurement.

03

Mentioned ≠ described correctly

PwC research found links in AI answers are valid >94% of the time, but factually accurate only 39–77%. AI may mention you — and still narrow you, get you wrong, or cast you as the fallback.

How we work

Measure before you optimize

We don’t sell magic. Everything BASELINE does rests on three rules — the kind that let a conclusion survive your legal team, your competitors, and your own re-checking.

01

Deterministic rules, not AI scoring

Four of the five dimensions are pure code: string matching, position tertiles, source-domain matching. Zero LLM, zero randomness. The same evidence always yields the same conclusion.

02

Raw counts, never percentages

“0 mentions out of 5,” not “0%.” "Percentages on small samples are noise that won’t survive scrutiny" — so we report raw counts with a Wilson 95% confidence interval.

03

Traceable to every single answer

Evidence is stored append-only by date, resumable, never overwritten. Every number in the report is wrapped with a data-ev anchor back to the screenshot of that exact answer — open it and see why the number is what it is.

"In a market where most tactics fail and results drift with time and competition, the only durable value is trustworthy measurement."

What we measure

Every AI answer, broken into five dimensions

We ask a question, fire follow-ups, run it five times — then break each answer into five dimensions. "Four are decided by rules; the only sentiment dimension is forced through human review to L1 confidence before it enters a report."

  1. 1

    Mention Mention

    Is the brand (and all its aliases) written into the answer body? Exact match, yes/no.

  2. 2

    Rank / first screen Rank

    If listed, in what position? Within the first 600 characters? Scored by ordinal and character position.

  3. 3

    Cited as a source Citation

    Does the brand appear in the AI’s source/reference list? Matched by URL, title, platform name.

  4. 4

    Sentiment / frame Sentiment

    “Indispensable” or “last resort”? The only non-automatic dimension — forced human review, never accuse unfairly.

  5. 5

    Source platforms Sources

    Which kinds of platform did this answer cite? Portal / vertical / encyclopedia / social / UGC — this sets the lever.

Four subsystems

From a single answer to a shareable white paper

Four independent subsystems, each producing pure data and no opinions; together they form a complete picture of your brand inside AI.

recall

Recall pipeline

Playwright drives real web-interface evidence across 6 engines → five-dimension scoring → recall aggregation and Share of Voice.

queries → evidence batches → white paper
diagnostic

Site audit

Single-domain static-signal audit: robots / llms.txt / JSON-LD / Chinese entities, three-tier weighted 0–100, plus agent reachability L0–L3.

domain → scorecard + JSON
weights

Citation source weights

Aggregates evidence into a source × engine matrix of which source types each engine actually prefers to cite.

evidence → weight matrix
benchmark

Engine behavior benchmark

A brand-agnostic standard query set measuring response rate, citation rate and source diversity — all with Wilson intervals.

query set → industry baseline

How we engage

Measure → Diagnose → Build → Operate

The baseline measurement stands alone as an honest health check; the next three steps turn “being seen” into a durable asset and an ongoing practice.

  1. Measure

    Baseline measurement

    Where do you stand in AI answers right now? A health check that stands on its own.

    Output GEO baseline report
  2. Diagnose

    Gap diagnosis

    Why not mentioned / not cited / described wrong? Attributed to fixable gaps.

    Output Gap inventory + priorities
  3. Build

    Asset building

    Translate brand facts into AI-readable sources: source of truth, knowledge base, site fixes, content.

    Output Source of truth / KB / fixes
  4. Operate

    Continuous operation

    Re-measure monthly, attribute, update the knowledge base — because AI answers drift.

    Output Monthly reports + actions

Standing on the literature

Every weight has a source

Our scoring constants aren’t guesses. The five-dimension frame, the three-tier weights, the high weight on freshness, the structure signals — each is anchored to published research. In the code, every cited paper carries its arXiv number, and every industry heuristic is explicitly flagged “assumption, not truth.”

Read the full methodology →
+33–41%
citation-rate lift from adding statistics
GEO foundational paper · KDD 2024 ↗
+22% vs −9%
structural remodeling vs pure rewriting
SAGEO Arena · arXiv ↗
45 / 30 / 25
three-tier site-audit weights
C-SEO Bench · NeurIPS 2025 ↗

Strategic technology partner

BASELINE × Tencent

We maintain a deep technical partnership with Tencent — from GEO expert support to the underlying technology of our benchmark platform.

GEO expert support

Tencent provides technical experts in GEO, helping refine our methods and definitions.

Benchmark platform support

The underlying technology of the engine-behavior benchmark platform involves Tencent’s engineering team.

Built by BASELINE

The assessment system and product are developed in-house by BASELINE; Tencent supports us as a deep technical partner.

Read about the partnership →

FAQ

You’re probably wondering

What is GEO (Generative Engine Optimization)?

GEO is the systematic work of getting a brand accurately mentioned, cited and ranked inside the generative answers of AI engines. The core difference from SEO: SEO competes for rank among ten blue links, while GEO competes for the single answer the AI writes directly — and users often decide right there, without clicking through to any website.

How does GEO relate to SEO and AEO?

SEO optimizes ranking on a results page; AEO (answer engine optimization) overlaps heavily with GEO — both care about the “direct answer.” We use “GEO” to stress that the target is the generated answer inside AI engines (DeepSeek, Doubao, etc.) and the sources it cites. They aren’t mutually exclusive: an AI-friendly website usually does well on SEO too.

Which AI engines does BASELINE actually test?

Six consumer Chinese AI engines, always: DeepSeek, Doubao, Tongyi, Yuanbao, Kimi and Ernie. Capture goes through real web-interface interaction (driven by Playwright), staying as close as possible to the answer a real user would get — not an idealized response from a public API.

Why report “0 out of 5” instead of a percentage?

Because a percentage on a small sample is noise. At N=5, “0%” and “20%” differ by a single answer yet read completely differently, and won’t survive scrutiny from a legal team or a competitor. We report only raw counts (X/Y) with a Wilson 95% confidence interval, stating the uncertainty honestly. It’s both integrity and self-protection.

See all FAQs →

Know where you stand before you talk about winning

Give us a brand and a few real questions. Within two weeks you’ll have an AI-visibility baseline you can send to your boss — and to your legal team.