Invisible until a customer screenshots it
You can still check rankings on a search engine. AI answers shift by person, phrasing and time — you have no dashboard, unless someone measures them for you, answer by answer.
GEO · Generative Engine Optimization
Buyers already make decisions inside DeepSeek, Doubao and Tongyi. BASELINE measures what 6 Chinese AI engines actually answer about you and your competitors — who gets mentioned, whose content gets cited, who ranks first. Deterministic rules, raw counts, and every number traces back to a real screenshot.
Not another “we’ll get you to the top” promise. An honest measurement.
Why now
When users stop scanning ten blue links and simply ask AI “who’s good for this?”, your brand is either written into that answer — or it doesn’t exist. It’s a brand-new shelf, and you can’t see it.
You can still check rankings on a search engine. AI answers shift by person, phrasing and time — you have no dashboard, unless someone measures them for you, answer by answer.
The academic benchmark C-SEO Bench (NeurIPS 2025) shows that "many GEO tactics are unstable in real conditions and drift with models and competition." The only durable value is trustworthy measurement.
PwC research found links in AI answers are valid >94% of the time, but factually accurate only 39–77%. AI may mention you — and still narrow you, get you wrong, or cast you as the fallback.
How we work
We don’t sell magic. Everything BASELINE does rests on three rules — the kind that let a conclusion survive your legal team, your competitors, and your own re-checking.
Four of the five dimensions are pure code: string matching, position tertiles, source-domain matching. Zero LLM, zero randomness. The same evidence always yields the same conclusion.
“0 mentions out of 5,” not “0%.” "Percentages on small samples are noise that won’t survive scrutiny" — so we report raw counts with a Wilson 95% confidence interval.
Evidence is stored append-only by date, resumable, never overwritten. Every number in the report is wrapped with a data-ev anchor back to the screenshot of that exact answer — open it and see why the number is what it is.
"In a market where most tactics fail and results drift with time and competition, the only durable value is trustworthy measurement."
What we measure
We ask a question, fire follow-ups, run it five times — then break each answer into five dimensions. "Four are decided by rules; the only sentiment dimension is forced through human review to L1 confidence before it enters a report."
Is the brand (and all its aliases) written into the answer body? Exact match, yes/no.
If listed, in what position? Within the first 600 characters? Scored by ordinal and character position.
Does the brand appear in the AI’s source/reference list? Matched by URL, title, platform name.
“Indispensable” or “last resort”? The only non-automatic dimension — forced human review, never accuse unfairly.
Which kinds of platform did this answer cite? Portal / vertical / encyclopedia / social / UGC — this sets the lever.
Four subsystems
Four independent subsystems, each producing pure data and no opinions; together they form a complete picture of your brand inside AI.
recall Playwright drives real web-interface evidence across 6 engines → five-dimension scoring → recall aggregation and Share of Voice.
diagnostic Single-domain static-signal audit: robots / llms.txt / JSON-LD / Chinese entities, three-tier weighted 0–100, plus agent reachability L0–L3.
weights Aggregates evidence into a source × engine matrix of which source types each engine actually prefers to cite.
benchmark A brand-agnostic standard query set measuring response rate, citation rate and source diversity — all with Wilson intervals.
How we engage
The baseline measurement stands alone as an honest health check; the next three steps turn “being seen” into a durable asset and an ongoing practice.
Where do you stand in AI answers right now? A health check that stands on its own.
Why not mentioned / not cited / described wrong? Attributed to fixable gaps.
Translate brand facts into AI-readable sources: source of truth, knowledge base, site fixes, content.
Re-measure monthly, attribute, update the knowledge base — because AI answers drift.
Strategic technology partner
We maintain a deep technical partnership with Tencent — from GEO expert support to the underlying technology of our benchmark platform.
Tencent provides technical experts in GEO, helping refine our methods and definitions.
The underlying technology of the engine-behavior benchmark platform involves Tencent’s engineering team.
The assessment system and product are developed in-house by BASELINE; Tencent supports us as a deep technical partner.
FAQ
GEO is the systematic work of getting a brand accurately mentioned, cited and ranked inside the generative answers of AI engines. The core difference from SEO: SEO competes for rank among ten blue links, while GEO competes for the single answer the AI writes directly — and users often decide right there, without clicking through to any website.
SEO optimizes ranking on a results page; AEO (answer engine optimization) overlaps heavily with GEO — both care about the “direct answer.” We use “GEO” to stress that the target is the generated answer inside AI engines (DeepSeek, Doubao, etc.) and the sources it cites. They aren’t mutually exclusive: an AI-friendly website usually does well on SEO too.
Six consumer Chinese AI engines, always: DeepSeek, Doubao, Tongyi, Yuanbao, Kimi and Ernie. Capture goes through real web-interface interaction (driven by Playwright), staying as close as possible to the answer a real user would get — not an idealized response from a public API.
Because a percentage on a small sample is noise. At N=5, “0%” and “20%” differ by a single answer yet read completely differently, and won’t survive scrutiny from a legal team or a competitor. We report only raw counts (X/Y) with a Wilson 95% confidence interval, stating the uncertainty honestly. It’s both integrity and self-protection.
Give us a brand and a few real questions. Within two weeks you’ll have an AI-visibility baseline you can send to your boss — and to your legal team.