Source Decoupling in Chinese AI Search Engines
A cross-industry empirical analysis of 2,307 category queries across six Chinese AI engines — citation sources are highly decoupled across engines, nearly disjoint from traditional search, with large industry differences in brand visibility.
A cross-industry empirical analysis across three industries and six Chinese AI engines.
As Chinese AI engines — DeepSeek, Doubao, Tongyi (Qwen), Yuanbao, Kimi and Ernie — become major gateways for how consumers find information, a brand’s “visibility” inside them has become a new battleground in digital marketing. Yet existing research focuses almost entirely on English engines; the citation behavior of Chinese AI engines remains an academic blank.
Using measured data from the BASELINE GEO (Generative Engine Optimization) platform, this study systematically analyzes 2,307 category-query answers across six Chinese AI engines and four industries (supplements, outdoor advertising, OTC pharma, men’s skincare). We strictly separate category queries (natural queries containing no brand name) from brand/competitor queries (which trivially force a mention), and measure visibility using category queries only, avoiding a methodological trap in prior work.
Three core findings:
- Citation-source domains overlap very little between Chinese AI engines — the average Jaccard coefficient (a set-overlap metric) is just 1.9%–8.7% across four industries, and about 90% of domains are cited by a single engine only.
- The overlap between Chinese AI engines and traditional search (Baidu) is even lower (Jaccard 0.7%), far below the 10.7%–38% seen in the English market.
- Under the category-query lens, brand visibility skews toward the head, but the degree varies by industry, and engines show systematic asymmetry in source preference.
Three research questions
- RQ1: How non-overlapping are citation sources across Chinese AI engines? Is the pattern robust across industries?
- RQ2: How decoupled are Chinese AI engines from traditional search (Baidu)?
- RQ3: After removing brand-query contamination, is there systematic asymmetry in a brand’s visibility across engines?
Data and method
The study uses cross-industry measured data from the BASELINE GEO platform across three independent industries. Brand names are anonymized for confidentiality.
| Industry | Brands | Category-query records | Engine coverage |
|---|---|---|---|
| Supplements | 5 (A–E) | 1,600 | All 6 engines |
| Outdoor advertising | 4 (F–I) | 320 | 2 engines (DeepSeek / Doubao) |
| OTC pharma | 1 (K) | 225 | 5 engines |
| Men’s skincare | 1 (J) | 162 | All 6 engines |
| Total | 11 brands | 2,307 | 6 engines |
Key methodological decision: we measure brand visibility using category queries only (e.g. “what vitamins help low immunity,” “which firms do high-speed-rail ad placement”) — natural queries with no brand name. Brand queries (e.g. “how good is Brand A”) contain the brand name in the question itself, so the mention rate is trivially near 100% and carries no signal; mixing them in systematically overestimates head brands.
In supplements, the gap is stark: Brand A’s mention rate is 52.9% under the all-query lens but only 4.7% under category queries — an 11× overestimate. More importantly, Brand C looks stronger than Brand B under the mixed lens, but is far weaker under category queries — a true ordering visible only under the category-query lens.
Results
RQ1 · Sources are highly decoupled across engines
| Industry | Engines | Avg. Jaccard | Single-engine domains |
|---|---|---|---|
| Supplements | 6 | 2.8% | 89.9% |
| Outdoor advertising | 2 | 8.7% | 91.3% |
| OTC pharma | 5 | 2.1% | 92.4% |
| Men’s skincare | 6 | 1.9% | 90.2% |
All four industries show average Jaccard below 10%, with ~90% of domains cited by a single engine only. The cross-industry consistency indicates this decoupling is a structural feature of the Chinese AI ecosystem, not an industry quirk. Engines also differ sharply in sourcing strategy — under supplements category queries, Kimi cites 135 unique domains (highly dispersed) while Ernie cites only 25 yet covers equivalent answers, relying heavily on a few authoritative sources.
RQ2 · AI engines are nearly disjoint from traditional search
Comparing all six engines’ citation domains (category queries, 424 total) with Baidu search-result domains yields a Jaccard of just 0.7%.
| Comparison | Overlap | Source |
|---|---|---|
| English AI vs Google organic | 10.7%–38% | Ahrefs, SE Ranking |
| Chinese AI engines, pairwise (this study) | 1.9%–8.7% | Four-industry measurement |
| Chinese AI vs Baidu results (this study) | 0.7% | This study |
DeepSeek and Kimi have zero overlap with Baidu results — they almost never cite Baidu’s top-ranked content. Ernie is the only engine with meaningful overlap (8.2%), likely tied to Baidu’s own search ecosystem.
RQ3 · Systematic asymmetry in brand visibility
Under the strict category-query lens, brand visibility looks completely different by industry. In outdoor advertising the gap between the head brand (F, 73.8%) and a zero-mention brand (I, 0%) is extreme; in supplements it is “mild and concentrated” — the head brand (B) is only 11.6%, with most brands under 5% natural recall. The strength of the winner-takes-all effect varies by industry.
A single brand’s visibility also varies widely across engines (spread up to 14.3 percentage points). Ernie remains the top mention engine for several brands under category queries, while Doubao is systematically lower. This asymmetry implies structural differences in each engine’s retrieval strategy — optimization must be engine-specific.
Discussion
Query-type contamination: an overlooked trap. The study’s key methodological contribution is exposing how mixing query types systematically distorts GEO conclusions. In supplements, the all-query lens overestimates the client brand by 11.3×. We recommend the category-query lens as the standard for visibility measurement.
Practical implications:
- “Ranking well on Baidu” ≠ “AI can see you.” With only 0.7% overlap, traditional SEO contributes very little to AI visibility.
- Engine-specific work is mandatory. With ~90% single-engine domains, being cited by one engine does not mean being cited by another.
- Industry matters. GEO strategy can’t be copied across industries; it must be tailored to competitive structure.
Conclusion
Across four industries of measured data, this study reveals three structural features of the Chinese AI citation ecosystem: high decoupling across engines (Jaccard 1.9%–8.7%, ~90% single-engine domains), near-total separation from traditional search (Jaccard 0.7%), and large industry differences in brand visibility. It is also the first to expose the query-type contamination trap in GEO — brand queries overestimate client visibility by 11–17×.
GEO in the Chinese market faces deeper decoupling than the English market, and cannot simply import English-market strategy or methods. Brands need engine- and industry-specific visibility strategies grounded in their own ecosystem.
References
- Aggarwal, P., et al. “GEO: Generative Engine Optimization.” KDD 2024. arXiv:2311.09735.
- Puerto, H., et al. “C-SEO Bench: Does Conversational SEO Work?” NeurIPS 2025 D&B Track. arXiv:2506.11097.
- Ahrefs. “AI Overviews vs. Organic Search Results.” 2025.
- SE Ranking. “AI Overviews Study: URL Overlap.” 2025.
- 5W PR. “AI Platform Citation Source Index.” Meta-analysis, 2026-05.
- Broder, A. “A Taxonomy of Web Search.” SIGIR Forum, 2002.
Data collected by the BASELINE GEO platform. Brand names anonymized for confidentiality. All figures computed from measured data.