km.idaeo.ai · IDAEO 知識庫

🏛 Part of the "reports" topic shelf

How We Keep This Ledger: Methodology and Limits of the AI Crawler Series

The four AI-crawler data articles on this site (the "30 million hauls" series) share one server ledger. This page opens it up: how the numbers are counted, how crawler identities were verified, which figures carry known limits, and which statistics can no longer be traced — including three bookkeeping corrections we found ourselves. The verification commands are published below; you are welcome to rerun them.

This page is the methods appendix for our AI crawler series. We lay out how the ledger was kept — including three bookkeeping corrections discovered in our own audit. One principle throughout: where something can be measured, we publish the command; where it cannot, we say so plainly.

An auditable evidence chain turns the server ledger into explicit definitions, identity checks, corrections, stated limits, and conclusions that readers can rerun and reconcile.
The auditable-ledger method: publish the command when something can be measured, and say plainly when it cannot.

Where the data comes from

All request counts come from the server databases of our two websites (washinmura.jp and idaeo.ai), covering 2026-03-03 to 2026-07-24. The ledger splices two segments: before 2026-06-22, permanently retained daily aggregate tables; from 2026-06-23 to 07-24, request-level detail tables. The two segments split cleanly at 06-22/06-23, with no overlap and no gap.

One detail we owe you: the detail segment from 06-23 to 07-24 spans 32 calendar days inclusive. The "30-day" metrics in the articles follow the system's habitual name for its "30-day rolling window," not a precise 30.0 days.

Detail tables are retained for only 30 days by policy, then deleted, leaving daily aggregates. The cost: some questions that require going back into the details can no longer be answered after publication — section three below has two of them.

Permanently retained daily aggregates join request-level details retained for only 30 days, with no overlap or gap between 2026-06-22 and 06-23; expired detail cannot be reconstructed.
Data lineage and retention boundary: aggregates persist, while request-level detail expires after 30 days.

How crawler identities were verified

The name a crawler reports (its User-Agent) can be forged, so we reverse-DNS-checked every source IP, matched each against the network ranges the companies officially publish, and checked the registered owner of each range. All 565 of Apple's IPs resolve back to applebot.apple.com; all three OpenAI crawlers hit officially published ranges; Amazon, Microsoft and Meta pass the same way. Sampling caught 1 fake source impersonating Bytespider from a cloud data center.

Crawler identity evidence chain: a User-Agent claim is only the start; source IP, reverse DNS, official network ranges, and registered ownership must agree, otherwise the source is treated as an impostor.
AI crawler identity cannot be established from a self-reported name alone; it requires cross-checking the full evidence chain.

The "476 human clicks": we count only web-browser visits that carry referrer information, with our own test traffic removed. In-app clicks carry no referrer, and zero-click answers never visit, so 476 is a visible lower bound.

Self-audit corrections (the honest ledger, continuously updated)

First: the Common Crawl inclusion count. The original report said "1,515 records included, of which 1,089 were full-page originals," but the breakdown only added up to 1,087. We reran the public index query and settled the account: the correct figures are 1,515 records included, 1,100 originals with status 200 (ainews 734 / aeo 333 / www 29 / two further subdomains at 2 each), which sums to the total with zero discrepancy. Cause of the old error: the original query hit a service interruption and dropped rows, and two small subdomains were missed in classification. Verification command: `curl 'https://index.commoncrawl.org/CC-MAIN-2026-25-index?url=*.washinmura.jp&output=json'`.

Second: Applebot's volume includes redirects. After our brand moved from washinmura.jp to idaeo.ai, Applebot has kept crawling the old URLs — in current records, 100% of its requests to the old domain receive a 301 redirect response (all 1,173 of them) before it fetches the actual content on the new domain. This means the "Applebot 5 million" figure includes redirect requests to the old domain, and the same content may be counted once on the old side and once on the new. The redirect share during the surge period (from 6/28) can no longer be traced after log rotation; we state this limit as is.

Third: the sample size behind 4.7 vs 2.5 fetches per page. The comparison of "hand-crafted in-depth articles vs automatically generated pages" came from the 30-day details as of 07-24. The page-count denominators were not archived at the time, and the details have since been deleted under the retention policy — the sample size cannot be reconstructed, and we will not backfill a number for it. The same-shape statistic in the current window (real-time fetching overall at 2.3 per page, 86,950 requests across 38,530 URLs) is not directly comparable, because the URL structure has since been redesigned. In the articles, this figure keeps its "in our ledger" qualifier.

Three self-audit corrections: Common Crawl originals corrected from 1,089 to 1,100; Applebot volume includes 1,173 old-domain 301 responses; and the sample size behind 4.7 versus 2.5 fetches per page cannot be reconstructed.
Credibility comes from publishing corrections and clearly separating what remains measurable from what cannot be reconstructed.

Fourth (2026-08-11): "280,000 real user questions" corrected to "about 28,000." "GSC Isn't Broken — the Customers Moved" originally said the question bank came from "280,000 real user questions"; reconciliation against the database shows the actual table holds 28,363 questions (collected since April 2026). The original figure was a tenfold typo and has been corrected in all language versions. Discovered via: the site-wide article-by-article review cycle's reconciliation against source data.

Known limits

  • One operator, two websites, five months: an honest case observation, not an industry-wide law.
  • The "Google AI" count merges two labels and cannot be attributed to any single Google product; Google-side clicks blend into ordinary search referrers and are systematically undercounted.
  • Quality and read-count are co-occurring; confounders (page count, page age) are not fully separated, and we make no causal claim.
  • Reposts and zero-click reproduction are unobservable on the server side; feasible detection is sentence-probe checks and Common Crawl reverse lookups.
  • The figure "roughly 10–12% of crawled content eventually gets cited" is an approximation from reconciliation against historical citations: a directional estimate within the observable scope, with no precise denominator (the denominator's definition shifts with the statistical window), and it must not be used as a conversion rate. First mentions in the series link to this entry.
  • The "content quality tiers" (green/yellow/red) are an internal editorial rating: the dimensions include density of source anchoring, completeness of definitions, and coverage by auditable assets; the full rules are internal and not published, so numbers citing these tiers should be read as within-site relative comparisons, not externally recomputable metrics.
  • Same-named metrics across sister articles on this site may come from different statistical windows (e.g., the 30-day sitemap fetch counts 32,993 and 33,155 belong to two articles' different snapshots) — when numbers disagree, each article's stated window governs, and a discrepancy does not mean the ledger is wrong.
  • "Training crawlers are the foundation of AI traffic" is a mechanism-based inference, not a measured conclusion: the causal segment from training to later citations is unobservable on the server side. The article supports the direction with three measured signals (training crawlers arrived earliest; their crawling is breadth-first and most evenly spread across the three languages; the content flows into public training corpora), and does not claim proven causation.

Data Download (Auditable Aggregate Pack)

Rather than only publishing commands, we publish the ledger itself. Two de-identified aggregate files are available: daily requests per crawler (crawler-daily.csv, full history since 2026-03-03) and daily human clicks per AI source (referrals-daily.csv). Field definitions, SHA-256 hashes, the exact queries, and the known deltas against the frozen article figures are documented in the README — including why the aggregate view differs from the frozen total by roughly 3.54 million. The core numbers are no longer just "trust us"; you can check us. There is also a citation-tiers experiment pack: the full questions and anonymized verdicts of the three-round review experiment, with checksums.

Reality-Check Policy (from 2026-08-10)

However good a simulated review sounds, it does not count — only real-world citation does. From today, this site adopts reality-checking as a site-wide policy: 30 days and 90 days after each article is published, we return to the page and publish its real observed numbers (AI crawler reach and AI-platform referral clicks, from this site's server observation, recorded from 2026-08-10). First due dates: "The Five Tiers of Citation Sources" on 2026-09-09 (30 days) and 2026-11-08 (90 days). Honest boundary up front: observation starts 2026-08-10, and behavior before that date is not in the ledger; referral clicks count only web-side identifiable referrers and are a lower bound. Backfilled numbers will be written directly into the corresponding articles; if they are wrong, they get corrected publicly under the "self-audit corrections" rule.

Disclosure of interest

All data in this series comes from server records of sites operated by the IDAEO operator; IDAEO offers services for "getting content correctly cited by AI." Readers weighing this series' conclusions should be aware of this interest. The verification commands are published with the text — you are welcome to rerun them and prove us wrong; that is exactly how we found the three corrections above.

FAQ

Why do some numbers say "cannot be reconstructed" instead of being recalculated?
The details were deleted under the retention policy. A recalculation would use a different window and different definitions — a number that looks similar but is not comparable. Better to mark it unknowable than to publish a plausible fake.
「再構築できない」と言う数字を、なぜ計算し直さないのですか?明細は保持ポリシーで削除済みです。計算し直すと別のウィンドウ・別の口径になり、似て見えて比較できない数字になります。もっともらしい数を出すより、調べようがないと明記するほうが誠実です。
Why do some numbers say "cannot be reconstructed" instead of being recalculated?The details were deleted under the retention policy. A recalculation would use a different window and different definitions — a number that looks similar but is not comparable. Better to mark it unknowable than to publish a plausible fake.
1,089 became 1,100 — do the articles need fixing?
The article text already avoided citing those two figures, so it is unaffected. This page publishes the corrected account in full.
1,089 が 1,100 に変わりました。元の記事は直すべきですか?記事本文は当時からこの2つの数字の引用を避けており、影響はありません。このページは訂正後の帳簿を完全公開するものです。
1,089 became 1,100 — do the articles need fixing?The article text already avoided citing those two figures, so it is unaffected. This page publishes the corrected account in full.
Applebot's volume includes redirects — is "5 million" still credible?
The requests are real, but they include 301 redirect requests to old URLs, so not all of them are content fetches. We cannot retroactively split the share, so we state it as is: read it as "request volume," not "content-fetch volume."
Applebot の量がリダイレクトを含むなら、「500万回」はまだ信用できますか?リクエストの発生は事実ですが、旧URLへの 301 転送要求を含み、すべてが内容取得ではありません。比率は遡って分解できないため、そのまま注記します。「リクエスト量」であって「内容取得量」ではない、と読んでください。
Applebot's volume includes redirects — is "5 million" still credible?The requests are real, but they include 301 redirect requests to old URLs, so not all of them are content fetches. We cannot retroactively split the share, so we state it as is: read it as "request volume," not "content-fetch volume."
How can I verify what you claim?
The Common Crawl lookup command is in the text, and anyone can rerun it. The internal server data cannot be released raw, but the definitions and table structure are described on this page.
あなたたちの言うことは、どう検証できますか?Common Crawl の照会コマンドは本文に付してあり、誰でも再実行できます。サーバー内部データの生ファイルは公開できませんが、口径とテーブル構造はこのページで説明しています。
How can I verify what you claim?The Common Crawl lookup command is in the text, and anyone can rerun it. The internal server data cannot be released raw, but the definitions and table structure are described on this page.
What is a "preferred citation source," and how does it differ from "partial citation"?
These are tiers from our own review framework, which simulates how AI answer engines pick sources — they are not industry-standard terms. Partial citation = the AI is willing to cite you, but picks only the safest sentences and adds defensive qualifiers like "according to the site's own account." Preferred citation source = when multiple sources compete on the same topic, the AI ranks you first — because your definitions are clear, your limits are self-declared and your data can be audited, citing you requires no discount. Tier results are still a simulated review; real citation is verified against the referral-click ledger.
「優先引用ソース」とは何ですか?「部分引用」とどう違いますか?これは本サイトの評価フレームワークの等級で、AI答案エンジンがソースを選ぶ振る舞いを模擬したものです——業界標準の用語ではありません。部分引用=AIはあなたを引用するが、安全な文だけを選び、「同サイトの自称によれば」という防衛的な限定を付ける。優先引用ソース=同じテーマで複数のソースから選べるとき、AIがあなたを前に並べる——口径が明確で、限界を自ら記し、データが照合可能だから、引用時に割り引く必要がないのです。等級の結果はあくまで模擬評価であり、実際の引用は引用クリックの帳簿で検証します。
What is a "preferred citation source," and how does it differ from "partial citation"?These are tiers from our own review framework, which simulates how AI answer engines pick sources — they are not industry-standard terms. Partial citation = the AI is willing to cite you, but picks only the safest sentences and adds defensive qualifiers like "according to the site's own account." Preferred citation source = when multiple sources compete on the same topic, the AI ranks you first — because your definitions are clear, your limits are self-declared and your data can be audited, citing you requires no discount. Tier results are still a simulated review; real citation is verified against the referral-click ledger.
Will this page be updated?
Yes. New bookkeeping issues will be added to "self-audit corrections," and each update will be dated.
このページは更新されますか?はい。新しい帳簿の問題が見つかるたびに「自己監査訂正」に追記し、更新ごとに日付を記します。
Will this page be updated?Yes. New bookkeeping issues will be added to "self-audit corrections," and each update will be dated.

Cite this article

TK Lin・《How We Keep This Ledger: Methodology and Limits of the AI Crawler Series》・IDAEO 知識庫・2026-08-10・https://km.idaeo.ai/reports/crawler-methodology

更新 2026-08-10T15:26:19.302Z · server-rendered · four-language · IDAEO 知識庫