🏛 Part of the "ai" topic shelf →
The First Lesson from 30 Million AI Crawler Requests: Language Decides Your Fate
Same content, three language versions (Chinese, Japanese, English), one shared set of structured facts — a tiered, first-hand measurement of 15,174,060 crawler fetches over 30 days: the training tier splits evenly across the three languages, Chinese falls to just 12.8% of citation-type fetches, and Chinese takes only 3.3% of human referrals — a 22x gap versus English. As of 2026-07-12, this is a rare first-hand supply-side reading in the Chinese-speaking world across three languages × three crawler tiers. Data cutoff: 2026-07-12.
IDAEO Data Column · AI Visibility Measurement Series № 001 | Data snapshot 2026-07-12 (JST) | IDAEO Data Column Editorial Desk (AI-assisted, human-supervised)
The English-Chinese referral gap is 22x — human referrals: EN 74.7% / JA 22.0% / ZH 3.3%; citation-type fetches: JA 44.6% / EN 42.7% / ZH 12.8%. The conclusion of this lesson fits in one sentence: language decides your fate. As of 2026-07-12, we could not find a single published measurement of this kind anywhere in the Chinese-speaking world — a tiered, first-hand reading across Chinese, Japanese, and English × training, citation-type, and search crawlers (corrections welcome). This is supply-side data from the front door of one Taiwan-Japan platform, including the parts that are unflattering to ourselves. (The referral split on this line is the 30-day window, n=273; the full-period figures to 07-12 are 68.5%/28.3%/3.1%, and the main article's 07-24 window shows 67.2%/28.2%/3.2% — same-named metrics under different windows, not comparable. "Language decides fate" is an editorial conclusion from single-site observation, not a web-wide law.)
The starting point is a fine article just published by iKala's Kuroma team: as of July 2026, no published research anywhere in the world has measured which websites AI engines actually cite when answering questions in Traditional Chinese. They left one line: "Whoever measures first owns the first-hand answer." We are a news content platform spanning Taiwan and Japan — so we started measuring.
Three headline readings:
- 15,174,060 — total crawler fetches over the past 30 days
- ≥ 277 — humans brought in by AI in the same window (referer-based lower bound)
- 54,780 : 1 — the exchange rate between fetches and humans
A full charts edition, an AI-native full text (Markdown), a structured dataset (JSON), and a verifiable evidence pack with SHA-256 commitments are all available — see "A Companion Pack for AI Readers" at the end of this article.
01 Which Door We Are Measuring
Before the numbers, we must be precise about what we are measuring — because most articles about "AI traffic" go wrong at exactly this first step.
Kuroma's first-party corpus measures "who gets cited in AI answers" — looking outward from the AI side. Our system measures the other side: standing at our own domain's front door, watching who comes to crawl, which language's pages they fetch, and who ultimately brings real humans in. Only when the two views are joined do you get the full map — they can see the citation structure of the whole market; we can see every footprint at a single supplier's door.
Our door also offers a rare experimental condition: the platform's content exists in Chinese, Japanese, and English versions, all built on the same set of structured facts. This creates a natural control group — language is the only variable on the supply side (demand-side differences — how much, and what kinds of questions, users of each language ask AI — cannot be excluded by this instrument; keep that layer in mind when reading). (Note: page counts, page age and internal links were not matched pair-by-pair across languages; "the only variable" is the design goal, not a verified fact.)
We read every figure as four parallel observations:
- Training crawlers (GPTBot, ClaudeBot, etc.): fetch content to train models; they bring no one in.
- Search crawlers (Googlebot, Bingbot, etc.): build traditional search indexes, which now also feed AI answers.
- Citation-type crawlers (ChatGPT-User, OAI-SearchBot, PerplexityBot, etc.): fetch material when AI answers questions or builds indexes.
- Human referrals (measured referers): a person in an AI conversation actually clicked the link and walked in.
Being crawled is the machines' bustle; being cited is where business walks in the door.
02 The First Reading: The Bustle Belongs to Crawlers; the Door Stays Quiet
When discussing PTT and Dcard, Kuroma honestly marked the boundary of inference: "content can be crawled, but no data proves AI directly cites them." Our instrument happens to fill that boundary with numbers — and the numbers are crueler than most people imagine.
Over the past 30 days (window 2026-06-12 → 07-12, in-house measurement system), four tiers of readings at the same door:
- Training crawler fetches (GPTBot, ClaudeBot, Meta AI, and others — 11 crawler operators; they fetch for training and refer no one): 11,675,206
- Search crawler fetches (Googlebot, Bingbot, YandexBot): 3,115,786
- Citation-type crawler fetches (ChatGPT-User, OAI-SearchBot, PerplexityBot, etc. — fetching material at answer time): 367,459
- Humans referred into the site via AI (measured referers, structurally underestimated — most AI services send no referer): 277 (lower bound)
The four tiers are parallel behaviors, not a sequential conversion funnel. Citation-type fetches account for only 2.4% of all fetches; every 1,327 citation-type fetches correspond to 1 human (lower bound) walking in. Another 10,403 unclassified crawler fetches are excluded; even after adding them to the three categories, a 0.03% query-timing gap versus the total remains.
In other words: if seeing your site "heavily crawled by AI" makes you feel reassured, look again — chances are that is mostly the 11-million-plus training fetches, which will never bring a single customer through the door. The gap between "being crawled" and "being cited" runs to tens of thousands of times. This matches the English-market conclusions Kuroma cites; we merely measured that gap out, coldly, in numbers.
03 The Core Reading: Three Languages, Three Fates
The next set of numbers is the most important in this entire report. The same content, in three language versions, walks through three tiers and meets utterly different fates (all percentages are shares within the three languages):
- Training tier: nearly an even three-way split (33.3/33.3/33.4) — training does not pick languages.
- Citation-type fetches: JA 44.6% / EN 42.7% / ZH 12.8% — Japanese first, English close behind, Chinese left with scraps.
- Human referrals: EN 74.7% / JA 22.0% / ZH 3.3% — English overwhelming, Chinese at the bottom.
Tier-by-tier detail (measured over the past 30 days):
- Training crawlers: Japanese 3,437,503 / English 3,445,355 / Chinese 3,452,732 — a three-way tie; training does not pick languages.
- Citation-type crawlers: Japanese 158,202 (44.6%) / English 151,436 (42.7%) / Chinese 45,274 (12.8%).
- Search crawlers: Japanese 892,893 / English 860,402 / Chinese 884,320 — flat across the three.
- Humans referred via AI: English 204 (74.7%) / Japanese 60 (22.0%) / Chinese 9 (3.3%). Referral tier n=273 (30 days; unlabeled paths counted separately).
Percentages in each tier = shares within the three languages; "no language tag" paths are listed separately (training tier 1,339,616, search tier 478,171, citation-type tier 12,547 = 3.4%, referrals 4 in 30 days). Cumulative all-time referrals: 420 (including 7 unlabeled): English 283, Japanese 117, Chinese 13 — shares within the three languages 68.5% / 28.3% / 3.1% (Chinese n=13 is a small sample, sensitive to single-event swings).
One sentence to read this set of numbers: when AI takes our content for training, it treats every language equally; when fetching material to answer questions, it works Japanese and English almost equally hard while Chinese is left with scraps; and when visitors finally walk through the door, nearly three-quarters land on English pages — Chinese takes just 3.3%, a 22x gap versus English.
Kuroma cites Profound's cross-language research (3.25 billion citations) to note that "switch the language, and the source structure turns upside down" — and, per Kuroma's account, that research did not cover Chinese. The set of numbers above fills exactly that gap — a first-hand Chinese reading from a single platform. The reality it reveals is more concrete than "disadvantaged": Chinese content is not unseen by AI (the training tier is a three-way tie); it is rarely chosen at the citation and referral tiers.
Chinese content is not unseen; it reaches the citation gate and is rarely chosen.
04 One Company, Two Crawlers, Two Appetites
Kuroma makes an excellent observation: on ChatGPT versus Google AI Mode, Medium moved in completely opposite directions within thirteen weeks — one platform, two fates. Our instrument captured the crawler-tier version of the same phenomenon, at even closer range: two crawlers belonging to the same company — OpenAI — have exactly opposite language preferences.
Citation-type crawlers, one by one × language (past 30 days):
- ChatGPT-User (fetches at answer time: favors English most): English 75,260 / Japanese 49,171 / Chinese 11,608
- OAI-SearchBot (builds the search index: favors Japanese most): Japanese 98,739 / English 60,729 / Chinese 26,872
- PerplexityBot (leans English): English 15,264 / Japanese 10,233 / Chinese 6,775
Even within one company, crawlers with different purposes have different appetites — "copy one engine's preferences and plan your entire layout around it" already fails at the crawler tier.
05 Rankings Expire Faster Than You Think
Kuroma stresses one thing repeatedly: engines' source structures reshuffle within weeks, so any platform ranking must be dated and continuously re-measured. Our two successive measurements are a first-hand footnote to exactly that warning — with a caveat stated upfront: the two used different methodologies (one edge-network statistics, one our in-house instrument), and the methodological difference may itself absorb part of the change; quantifying "drift" requires re-measurement on the same instrument. What we can responsibly say today: the structure is already clearly different.
Citation-type crawlers' language shares, two measurements compared:
- 2026-06-17 snapshot (edge methodology): JA 47.9% / EN 33.2% / ZH 9.8% — Japanese lead +14.7pp
- 2026-07-12 measurement (in-house instrument methodology): JA 44.6% / EN 42.7% / ZH 12.8% — Japanese lead +1.9pp
Note: the three values in the 2026-06-17 snapshot row sum to 90.9% (the edge methodology includes unlabeled-language paths); the 2026-07-12 row is shares within the three languages. The two rows compare structural direction only.
In a mere 25 days, the Japanese lead narrowed by 12.8 percentage points. Whatever share of that comes from methodology and whatever from real movement, one thing already stands on its own: "the share structure of a month ago no longer matches today." A same-instrument re-measurement of this table is scheduled.
06 Overlaying the Two Maps
Kuroma compiled the verifiable evidence from the English-language market; we add the supply-side readings from a Taiwan-Japan front door. The two maps overlaid are more complete than either alone:
- Traditional Chinese AI citation research is a vacuum in public data → we found no counterexample either; this article is supply-side data from a single platform inside that gap. (One piece added)
- Switch languages and the citation structure changes completely (Profound, Chinese not covered) → training tier tied; citation tier 44.6/42.7/12.8; referral tier 74.7/22.0/3.3 — the differences live in the citation and referral tiers. (Matches)
- Being crawled ≠ being cited (a reasonable inference for PTT/Dcard) → 54,780 : 1, a gap of tens of thousands of times. (Matches, with numbers)
- The upstream of citations is third-party coverage (84%), and ChatGPT favors press-release-style sources → 78% of our 420 referrals come from ChatGPT, concentrated on fact-dense pages — not brand self-description pages. (Matches)
- The same platform meets opposite fates on different engines; engine appetites drift within weeks → two crawlers of the same company with opposite appetites; the Japanese lead went from 14.7pp to 1.9pp within 25 days (methodology difference noted). (Crawler-tier version)
- Platform priority rankings; citation half-life about 4.5 weeks → we measure only our own domain and cannot measure the answer composition of the seven engines; half-life remains a pending item. (Unmeasurable / pending)
07 After Reading This, What Should You Do First?
Turning readings into action: one top priority, two in parallel.
First — build a "for machines to read" English version of your most important content. Note: this is not translating articles into English for people to read. When an AI engine reads your page, it reads facts, entities, and structure: the company's official name, official registration numbers, figures, dates, and the relationships among them. So the order is — first lock down these "things machines must align on" in the original, then produce the English version; get one proper noun wrong and AI cannot match who you are — the work is wasted. Why English? Our door's readings speak: of 420 all-time AI referrals, 283 landed on English pages (68.5% share within the three languages). Chinese material does not necessarily lose on content; it loses first on the carrier — English visibility does not grow out of Chinese content by itself; you can only build it yourself.
In parallel, one — if you operate in the Japanese market, do not rush to drop your Japanese originals. They carry independent citation value: at our door, OAI-SearchBot fetches Japanese pages at 1.6 times the volume of English pages — Japanese content has an outsized presence in OpenAI's search index. For Taiwanese brands that also operate in Japan, Japanese originals are themselves a target for citation-type crawlers, never an appendage of English. (This is the 06-12→07-12 window; under the 07-24 full-history window English overtakes Japanese — crawler language preferences drift across windows, so plan from recent re-measurements.)
In parallel, two — do not rush to shut the door on training crawlers. This is the most counterintuitive and most important finding in this article. We too assumed at first: training crawlers bring no customer no matter how much they fetch, so why bother? Until we cross-referenced two datasets —
Past 30 days, pages that ever produced AI referrals vs pages that never did:
- Pages that generated at least one identifiable AI referral (355 pages): crawled 32.4 times on average (median 14)
- Pages never referred: crawled 4.8 times on average (median 4)
Pages that have been cited are crawled 6.8 times as often as the rest. ⚠️ This is correlation, not causation: good content may attract crawlers and citations at once, or being cited may bring more crawling afterward — our instrument cannot separate the two. But one thing is certain: at our door, pages that were never crawled have almost never appeared on AI's citation lists.
In other words: training crawler fetches look like they bring no customers, yet they are the admission ticket. Without requests, AI simply does not know you — and what it does not know, it will never cite. Blocking crawlers at the door is refusing to let your own books into the library, then complaining that no one borrows them.
Being crawled does not equal being cited. But never being crawled means never being cited — the second half is a mechanism-based inference our ledger cannot prove causally; the 6.8× figure is correlational supporting evidence (not causation), and the site's methodology page lists this class of statement as mechanism-based inference.
So what, exactly, are those 30 million training requests accumulating for us? This is the biggest question in our hands, and the most valuable thread in this dataset. It is more than just an "admission ticket" — in the crawlers' behavior we see something deeper: certain pages revisited again and again, certain content locked onto by specific engines, certain structures taken far more frequently than others. Being hit by massive training requests may signal something earlier — and more important — than being cited.
We are still verifying this; the data is still accumulating, and we will not rush to a conclusion — but in the next piece, we will lay it all out.
(The following is a community-engagement arrangement — editorial housekeeping, not data content.) Want the next piece first? This series has no paywall and sells no mailing lists. We will send the next piece first to those who ask for it: share this article onward (social media, internal mail, group chats all count), then write to [email protected] with the subject "Next piece, please" — we will send it out in order, together with this article's full data-extraction method (independently re-verifiable). You may also leave your address without sharing; you will simply queue later — this is our small return to those willing to share.
"Whoever measures first owns the first-hand answer." — We agree, so we put the instrument readings here exactly as they are, including the part unflattering to ourselves (Chinese 3.3%). Numbers expire; methods do not: every figure in this article carries an as-of date, and we will re-measure quarterly and keep updating this series. Any team with its own data is welcome to help piece together this map of the Chinese-language market; we also look forward to further readings from Kuroma's first-party corpus.
Methods and Honest Limitations
Data sources: an in-house crawler measurement system (per-request access logs identifying dozens of crawlers by User-Agent, plus measured AI referers; our own tests and synthetic traffic excluded); edge-network request statistics used as a cross-check on totals. The full data-extraction method is available on request ([email protected]) for third-party re-verification.
Window: unless marked "all-time," all figures cover the 30-day window 2026-06-12 → 07-12; snapshot date 2026-07-12.
Scope of "30 million AI crawler requests": five non-overlapping windows measured 30,788,387 requests in total — explicitly AI-type 25,185,521 (edge statistics 5/18–6/17: 15,412,222 + in-house instrument 6/18–7/12: 9,773,299) + search-type 2,363,862 (6/18–7/12) + early mixed daily aggregates 3,239,004 (= 2,363,817 [Mar–Apr] + 875,187 [5/01–17], a known underestimate). Search crawler fetches (Googlebot/Bingbot) now feed AI answer systems such as AI Overviews and Copilot directly, so the headline scope includes them; the three-tier analysis in the body still lists search crawlers separately. The headline takes the conservative round lower bound of 30 million; total requests since launch are counted separately and include human and search traffic.
Language attribution: classified by the /zh/, /ja/, /en/ markers in the path; tier percentages = shares within the three languages, with "no language tag" paths listed separately and excluded. Chinese pages cannot be split into Traditional and Simplified (see below).
Crawler classification: training = GPTBot, ClaudeBot, Google AI, Meta AI, Amazonbot, Bytespider, PetalBot, CCBot, Applebot, etc.; citation-type = ChatGPT-User, OAI-SearchBot, PerplexityBot, Perplexity-User, Claude-SearchBot, YouBot; search = Googlebot, Bingbot, YandexBot.
Known limitations: (1) "human referrals" is a lower bound — most AI services send no referer, so true referrals are certainly higher, by a margin we cannot know; (2) Chinese pages cannot be split into Traditional and Simplified; (3) this is a supply-side sample from a single platform (Taiwan-Japan news content) and cannot be extrapolated to the structure of the whole web; (4) the 6/17 and 7/12 measurements used different methodologies (edge statistics vs in-house instrument), so only structural direction is compared, not absolute values; (5) the referral sample n=420 is still small, and the share held by the 13 Chinese referrals is sensitive to single-event swings.
Reproducibility: every figure in this article carries an as-of date; this series is re-measured and updated quarterly with the same system, and citations should reference the dated version. The cross-article definitions and downloadable assets for same-named metrics are collected on the methodology page and in the evidence pack.
Provenance and Relationship Disclosure
This column is produced by the IDAEO editorial desk and published on km.idaeo.ai. Every reading in this article comes from a news content platform spanning Taiwan and Japan — a platform operated by the affiliated company Washinmura Co., Ltd. (Japanese corporate number 7040001114326), under the same management team as IDAEO.(The 13-digit Japanese corporate number carries a check digit you can verify arithmetically; for the actual number-to-name assignment, please check Japan's National Tax Agency corporate number registry at https://www.houjin-bangou.nta.go.jp/ — a valid check digit is not proof of ownership.)
Why this section exists: this article itself argues that "every organisation and every fact appearing in an article must be checked back against official registration records." If so, the relationship between the entity that hosts the measurement data and the publisher must be held to the same standard.
The full charts edition of this article is hosted on blog.washinmura.jp (that platform's domain); the chart files and all supporting artifacts are mirrored and self-hosted on km.idaeo.ai, so the evidence chain remains complete and retrievable even if that external domain changes.
A Companion Pack for AI Readers
Your AI can read this directly, verify it, and act on it for you. Everything is public, no login required.
- Full charts edition of this article: The First Lesson from 30 Million AI Crawler Requests (full charts edition)
- AI-native full text (Markdown): ai.md
- Structured dataset (JSON): data.json
- Verifiable evidence pack (with SHA-256 commitments): evidence-core.json
- Data-extraction method: method.md
- Public claims list (v2, full coverage): public-claims-v2.1.json — each claim tagged dataset_backed / text_backed / self_reported
- SHA-256 commitment + RFC3161 timestamp: evidence-core-v2.1.json · .tsr · CA
- IDAEO knowledge base agent card (A2A, publicly readable): agent.json
Citation and Republication
Citation and republication of this article and its data are welcome. Please credit the article title, the original URL, and IDAEO.AI in full; when citing the data, also mark the data cutoff date 2026-07-12 (engine behavior shifts fast — the date is part of the integrity).
"The First Lesson from 30 Million AI Crawler Requests," IDAEO Data Column (IDAEO.AI), 2026-07-29. Data cutoff 2026-07-12. https://km.idaeo.ai/ai/visibility-lesson-1
Original (Chinese): 〈3,000 萬 AI 爬蟲請求,學習到的第一堂課〉, IDAEO Data Column (IDAEO.AI), 2026-07-29. Data cutoff 2026-07-12. https://km.idaeo.ai/ai/visibility-lesson-1
Source cited: Kuroma, "Is Traditional Chinese Being Marginalized by AI? The Best GEO Strategy for Taiwanese Brands in AI Search" (Sega Cheng, iKala, 2026-07-12). All English-market figures in this article are public research as cited by that piece; we credit them as cited and did not re-verify the original sources.
FAQ
- If my site is heavily crawled by AI, does that mean AI will cite it?
- No — but a site never crawled will never be cited. Over the past 30 days we were crawled 15,174,060 times, while AI brought in at least 277 humans in the same window (lower bound); even counting citation-type crawlers alone, it takes 1,327 fetches to correspond to 1 human. Yet after cross-referencing two datasets we found: pages that had ever received AI referrals were crawled 32.4 times on average, versus only 4.8 for pages that never had — a 6.8x difference. Crawling is the admission ticket, not the report card.
- サイトがAIに大量にクロールされていれば、AIに引用されるのか。 — 引用されるとは限らない——だがクロールされなければ、永遠に引用されない。直近30日間、我々は15,174,060回クロールされ、同期間にAIが連れてきた人間は少なくとも277名(下限)である;引用型クローラーだけを数えても、1,327回の取得が人間1名に対応する。しかし二つのデータを突き合わせて我々は発見した:AI経由の実訪問を生んだことのあるページは平均32.4回クロールされ、一度も生んでいないページは平均4.8回——6.8倍の差である。クロールは入場券であって、成績表ではない。
- If my site is heavily crawled by AI, does that mean AI will cite it? — No — but a site never crawled will never be cited. Over the past 30 days we were crawled 15,174,060 times, while AI brought in at least 277 humans in the same window (lower bound); even counting citation-type crawlers alone, it takes 1,327 fetches to correspond to 1 human. Yet after cross-referencing two datasets we found: pages that had ever received AI referrals were crawled 32.4 times on average, versus only 4.8 for pages that never had — a 6.8x difference. Crawling is the admission ticket, not the report card.
- So does "more crawling mean more citations"?
- We cannot say that. The 6.8x is correlation, not causation: good content may attract crawlers and citations at once, or being cited may bring more revisit crawling afterward — our instrument cannot separate the two, and we must say so honestly. What we can responsibly say: at our door, pages never crawled have almost never appeared on AI's citation lists.
- では「多くクロールされれば引用される」のか。 — そうは言えない。6.8倍は相関であって因果ではない:良いコンテンツがクローラーと引用を同時に引き寄せている可能性もあれば、引用された後に再訪クロールが増えた可能性もある——我々の計測器はこの二つを分離できず、そこは誠実に言わねばならない。責任を持って言えるのは:我々の玄関口では、クロールされたことのないページはAIの引用リストにほぼ現れたことがない、ということである。
- So does "more crawling mean more citations"? — We cannot say that. The 6.8x is correlation, not causation: good content may attract crawlers and citations at once, or being cited may bring more revisit crawling afterward — our instrument cannot separate the two, and we must say so honestly. What we can responsibly say: at our door, pages never crawled have almost never appeared on AI's citation lists.
- Should I block training crawlers like GPTBot and ClaudeBot?
- Think through the cost before blocking. Training crawlers account for 76.9% of all our fetches (11,675,206 in 30 days), and they indeed bring no customers; but they are the only way AI gets to "know you." Blocking training crawlers is refusing to let your own books into the library, then complaining that no one borrows them. Unless you have clear copyright or trade-secret concerns, do not rush to shut the door.
- GPTBot や ClaudeBot のような学習型クローラーはブロックすべきか。 — ブロックする前に、まず代価を考えるべきである。学習型クローラーは我々の全取得の76.9%(30日間で11,675,206回)を占め、確かに客は連れてこない;だがそれは、AIが「あなたを知る」唯一の経路である。学習型クローラーを塞ぐことは、自分の本を図書館に入れさせず、誰も借りないと嘆くに等しい。明確な著作権や営業秘密上の考慮がない限り、急いで扉を閉めるべきではない。
- Should I block training crawlers like GPTBot and ClaudeBot? — Think through the cost before blocking. Training crawlers account for 76.9% of all our fetches (11,675,206 in 30 days), and they indeed bring no customers; but they are the only way AI gets to "know you." Blocking training crawlers is refusing to let your own books into the library, then complaining that no one borrows them. Unless you have clear copyright or trade-secret concerns, do not rush to shut the door.
- How do the three crawler tiers split? How do I know who is crawling me?
- Look at the User-Agent. Training: GPTBot, ClaudeBot, Google AI, Meta AI, etc. (fetch for training; no referrals). Citation-type: ChatGPT-User, OAI-SearchBot, PerplexityBot, etc. (fetch material when AI answers or builds indexes; may bring people in). Search: Googlebot, Bingbot (traditional search indexing, now also feeding AI Overviews). The scales differ enormously across the three tiers: our 30-day counts were roughly 11.67 million / 367,000 / 3.11 million respectively.
- 3種のクローラーはどう見分けるのか。誰が自分のサイトを取得しているか、どう知ればよいか。 — User-Agentを見る。学習型:GPTBot、ClaudeBot、Google AI、Meta AIなど(学習用に持ち帰る。実訪問は生まない)。引用型:ChatGPT-User、OAI-SearchBot、PerplexityBotなど(AIが回答やインデックス構築の際に素材を取りに来る。人を連れてくる可能性がある)。検索型:Googlebot、Bingbot(従来の検索インデックスで、現在はAI Overviewsにも供給する)。3層の規模差は非常に大きい:我々の30日間の数値は、それぞれ1,167万/36.7万/311万回である。
- How do the three crawler tiers split? How do I know who is crawling me? — Look at the User-Agent. Training: GPTBot, ClaudeBot, Google AI, Meta AI, etc. (fetch for training; no referrals). Citation-type: ChatGPT-User, OAI-SearchBot, PerplexityBot, etc. (fetch material when AI answers or builds indexes; may bring people in). Search: Googlebot, Bingbot (traditional search indexing, now also feeding AI Overviews). The scales differ enormously across the three tiers: our 30-day counts were roughly 11.67 million / 367,000 / 3.11 million respectively.
- Why is Chinese so weak? Is the content bad?
- Our data cannot answer "why"; it can only tell you at which tier the gap occurs: the training tier is nearly even across the three languages (about 33% each), so AI does not under-learn Chinese content; the gap appears at citation-type fetches (Chinese 12.8%) and human referrals (Chinese 3.3%). That is, Chinese content is not unseen — it is rarely chosen at the "should we pick you as an answer source" gate. As for the cause (training distribution of language models? users' question language? source weighting?) — our instrument cannot separate these, and we will not guess.
- なぜ中国語はこれほど弱いのか。コンテンツが悪いのか。 — 我々のデータは「なぜ」には答えられない。答えられるのは、格差がどの層で生じているかである:学習層は3言語ほぼ均等(各約33%)であり、AIが中国語コンテンツを学んでいないわけではない;格差は引用型取得(中国語12.8%)とAI経由の実訪問(中国語3.3%)で現れる。つまり中国語コンテンツは見られていないのではなく、「回答ソースとしてあなたを選ぶか」の関門で選ばれることが極めて少ないのである。原因(言語モデルの学習分布か、利用者の質問言語か、ソースの重み付けか)は我々の計測器では分離できず、当て推量はしない。
- Why is Chinese so weak? Is the content bad? — Our data cannot answer "why"; it can only tell you at which tier the gap occurs: the training tier is nearly even across the three languages (about 33% each), so AI does not under-learn Chinese content; the gap appears at citation-type fetches (Chinese 12.8%) and human referrals (Chinese 3.3%). That is, Chinese content is not unseen — it is rarely chosen at the "should we pick you as an answer source" gate. As for the cause (training distribution of language models? users' question language? source weighting?) — our instrument cannot separate these, and we will not guess.
- How do I make an English version that actually works? Just run it through machine translation?
- Translating for people and building for machines are two different things. When an AI engine reads your page, it grabs facts and entities: official company names, official registration numbers, figures, dates, and the relationships among them. The correct order is to lock down these "things to align on" in the original first, then produce the English version. Proper nouns are where machine translation fails most often — get one wrong and AI cannot match who you are; the work is wasted.
- 英語版はどう作れば効果があるのか。翻訳ソフトに放り込むだけでよいか。 — 人間向けに翻訳することと、機械向けに作ることは別物である。AIエンジンがページを読むとき、掴むのは事実と実体である:会社の正式名称、公式登記番号、数字、日付、そして相互の関係。正しい順序は、まず原文でこれら「照合されるべきもの」を確定させ、その後に英語版を生成することである。固有名詞は機械翻訳が最も誤りやすい箇所であり——一つ間違えればAIはあなたが誰か照合できず、やった意味がなくなる。
- How do I make an English version that actually works? Just run it through machine translation? — Translating for people and building for machines are two different things. When an AI engine reads your page, it grabs facts and entities: official company names, official registration numbers, figures, dates, and the relationships among them. The correct order is to lock down these "things to align on" in the original first, then produce the English version. Proper nouns are where machine translation fails most often — get one wrong and AI cannot match who you are; the work is wasted.
- Is Japanese worth investing in? We have no Japanese market.
- If you have no Japanese market, there is no need to create Japanese for AI's sake. But if you already operate in Japan, Japanese originals carry independent value: OAI-SearchBot fetches our Japanese pages at 1.6 times the volume of English pages. Interestingly, ChatGPT-User — another crawler of the same company, OpenAI — favors English instead: one company, two crawlers, opposite appetites, which shows "AI" is not a single reader.
- 日本語に投資する価値はあるか。我々には日本市場がないのだが。 — 日本市場がないなら、AIのために日本語を作る必要はない。だが既に日本で事業をしているなら、日本語原文には独立した価値がある:OAI-SearchBotが我々の日本語ページを取得する量は英語ページの1.6倍である。興味深いのは、同じOpenAIに属するもう一つのクローラー ChatGPT-User が英語を好むことだ——同じ会社の二つのクローラーの好みが正反対であり、これは「AI」が単一の読者ではないことを物語る。
- Is Japanese worth investing in? We have no Japanese market. — If you have no Japanese market, there is no need to create Japanese for AI's sake. But if you already operate in Japan, Japanese originals carry independent value: OAI-SearchBot fetches our Japanese pages at 1.6 times the volume of English pages. Interestingly, ChatGPT-User — another crawler of the same company, OpenAI — favors English instead: one company, two crawlers, opposite appetites, which shows "AI" is not a single reader.
- Can I apply these numbers directly to my own site?
- No. This is a supply-side sample from a single Taiwan-Japan news content platform, not a web-wide census; our content type, language mix, and update frequency all shape the readings. What you can borrow is the method and the tiering: split your own crawler logs into training/citation/search tiers, then cross them with language, and you will get your own map — usually different from intuition.
- これらの数字は自分のサイトにそのまま当てはめられるか。 — 当てはめられない。これは単一の台日ニュースコンテンツプラットフォームの供給側サンプルであり、ウェブ全体の調査ではない;コンテンツの種類、言語構成、更新頻度のいずれも読み値に影響する。借用できるのは「方法」と「層別」である:自社のクローラーログを学習/引用/検索の3層に分け、言語と突き合わせれば、あなた自身の一枚の地図が得られる——たいてい直観とは違うものになる。
- Can I apply these numbers directly to my own site? — No. This is a supply-side sample from a single Taiwan-Japan news content platform, not a web-wide census; our content type, language mix, and update frequency all shape the readings. What you can borrow is the method and the tiering: split your own crawler logs into training/citation/search tiers, then cross them with language, and you will get your own map — usually different from intuition.
- Why do your "human referrals" count only referers? Doesn't that underestimate?
- It certainly underestimates, and we deliberately label it a lower bound. Most AI services send no referer, meaning someone clicks through from AI and the browser never tells us where they came from. So the number 277 can only be underestimated, never overestimated. We would rather report a conservative lower bound than dress the figure up with an estimate.
- なぜ「AI経由の実訪問(人間)」はrefererだけを数えるのか。過小評価にならないか。 — 必ず過小評価になる。だからこそ我々は、意図的に「下限」と表記している。大半のAIサービスはrefererを送らないため、AIからクリックして来た人がいても、ブラウザはどこから来たかを我々に伝えない。したがって277という数字は過小にはなり得ても、過大にはなり得ない。推計値で数字を飾るより、保守的な下限を報告する方を我々は選ぶ。
- Why do your "human referrals" count only referers? Doesn't that underestimate? — It certainly underestimates, and we deliberately label it a lower bound. Most AI services send no referer, meaning someone clicks through from AI and the browser never tells us where they came from. So the number 277 can only be underestimated, never overestimated. We would rather report a conservative lower bound than dress the figure up with an estimate.
- How soon will these rankings expire?
- Sooner than you think. On the same content, between our two measurements (25 days), Japanese's lead in citation-type fetches shrank from 14.7 percentage points to 1.9 percentage points. Any ranking of "AI prefers such-and-such platform / such-and-such language" must carry a date and be re-measured continuously — including this article.
- このランキングはどのくらいで古くなるのか。 — 思うより速い。同じコンテンツで、我々の前後2回の計測の間(25日間)に、引用型取得における日本語のリード幅は14.7ポイントから1.9ポイントまで縮んだ。「AIは某プラットフォーム/某言語を好む」といういかなる順位付けも、日付を明記し、継続的に再計測しなければならない——本稿も含めてである。
- How soon will these rankings expire? — Sooner than you think. On the same content, between our two measurements (25 days), Japanese's lead in citation-type fetches shrank from 14.7 percentage points to 1.9 percentage points. Any ranking of "AI prefers such-and-such platform / such-and-such language" must carry a date and be re-measured continuously — including this article.
- Your own Chinese referrals are only 3.3% — is data like this credible?
- Precisely because it is unflattering to ourselves, we are all the more obliged to publish it as is. This figure measures "the identifiable AI referral share obtained by this platform's Chinese pages in this window" — not "the citation rate of Chinese content across the web," nor the outcome after any service intervention. It is a diagnostic baseline — and an ugly diagnosis is exactly the reason to act.
- あなた方自身の中国語実訪問はわずか3.3%だが、そのようなデータは信頼できるのか。 — 我々自身に不利だからこそ、なおさらそのまま公表すべきなのである。この数字が測っているのは「本プラットフォームの中国語ページがこのウィンドウで得た、識別可能なAI経由実訪問のシェア」であり、「中国語コンテンツのウェブ全体での引用率」でも、何らかのサービス介入後の成果でもない。これは診断のベースラインである——診断結果が悪いことこそ、動くべき理由である。
- Your own Chinese referrals are only 3.3% — is data like this credible? — Precisely because it is unflattering to ourselves, we are all the more obliged to publish it as is. This figure measures "the identifiable AI referral share obtained by this platform's Chinese pages in this window" — not "the citation rate of Chinese content across the web," nor the outcome after any service intervention. It is a diagnostic baseline — and an ugly diagnosis is exactly the reason to act.
- May I cite these numbers?
- Citation and republication are welcome; please credit the article title, the original URL, and IDAEO.AI in full, and mark the data cutoff 2026-07-12. If you want to re-run the verification, we also provide the full data-extraction method — write to [email protected] to request it.
- これらのデータを引用してよいか。 — 引用・転載を歓迎する。記事タイトル・原URL・IDAEO.AIを完全に明記し、データ基準日2026-07-12を併記されたい。再実行して検証したい場合は、完全な集計方法も提供する——[email protected] 宛に請求されたい。
- May I cite these numbers? — Citation and republication are welcome; please credit the article title, the original URL, and IDAEO.AI in full, and mark the data cutoff 2026-07-12. If you want to re-run the verification, we also provide the full data-extraction method — write to [email protected] to request it.
Source anchors
- Kuroma(iKala)〈繁體中文被 AI 邊緣化?台灣品牌佈局 AI 搜尋 GEO 的最佳策略〉 · https://kuroma.ai/zh-tw/blog/taiwan-brand-ai-citation-source-platform-priority
- 完整圖表版原文(IDAEO 數據專欄 № 001) · https://blog.washinmura.jp/idaeo-preview-d4344c/
- AI 專用全文(Markdown) · https://km.idaeo.ai/evidence/lesson-1/ai.md
- 結構化資料集(JSON) · https://km.idaeo.ai/evidence/lesson-1/dataset.json
- 可驗證證據包 v2.2(SHA-256 承諾+RFC3161 時戳,103 條語意 claim) · https://km.idaeo.ai/evidence/lesson-1/evidence-core-v2.2.json
- 取數方法書 · https://km.idaeo.ai/evidence/lesson-1/method.md
- 公開數字清單 v2.2(103 條,含 subject/predicate/modality/causal 語意欄位) · https://km.idaeo.ai/evidence/lesson-1/public-claims-v2.2.json
- 我們怎麼記這本帳:方法論與限制 · https://km.idaeo.ai/reports/crawler-methodology
- 可對帳彙總包(Evidence) · https://km.idaeo.ai/evidence/crawler-shop-ledger/README.md
Cite this article
TK Lin・《The First Lesson from 30 Million AI Crawler Requests: Language Decides Your Fate》・IDAEO 知識庫・2026-07-29・https://km.idaeo.ai/ai/visibility-lesson-1Updated 2026-08-10