km.idaeo.ai · IDAEO 知識庫

🏛 Part of the "ai" topic shelf

The Invisible Readers: 12 AI Crawlers Revealed

More and more customers are skipping search and asking AI directly. Whether AI mentions you depends on whether it has “read” your website. We opened up five months of real server logs from our own site—more than thirty AI crawlers and roughly 30 million content fetches—and found that crawlers fall into three groups: training crawlers, search-indexing crawlers, and user-triggered fetchers. Only the last two bring customers back. Great content wins over people, not machines. And the crawlers that drive the most traffic are the slowest to respond—but site owners already hold the levers to speed them up. This article also presents the limitations honestly: it is an n=1 observation of a single site, the measured read-to-citation conversion rate is roughly 10–12%, and every boundary is disclosed.

The Invisible Readers: 12 AI Crawlers Revealed

The overview links the claim “No crawler request, no citation; no citation, no human traffic” to three crawler roles: training crawlers GPTBot, ClaudeBot, and Meta AI; indexing crawlers OAI-SearchBot, Bingbot, and Applebot; and the human-triggered ChatGPT-User. Volume narrows from about 30 million requests through a conversion of roughly one in ten to 476 observable referrals, a lower bound. The cards show OAI-SearchBot at 85.6% citation coverage, ClaudeBot at 6.44 million requests, Bingbot at 32,993 requests, and Applebot at 4.847 million requests in 30 days; Applebot is marked “🔭 Worth watching,” with the editorial observation that its volume is unusually high and something must be in the works behind the scenes.
A map of the whole AI-visibility story: distinguish crawler roles, then follow requests toward citations and human traffic; Applebot’s unusually high volume is a worth-watching editorial observation.

Start with one question. When someone wants what you sell—a restaurant meal, a guesthouse stay, a consulting service—will they open a search engine and compare options one by one, or will they simply ask AI, “What do you recommend?” More and more people are choosing the second route. AI gives them a direct answer, names a few businesses, and includes a handful of links. If your name comes up, customers arrive. If it does not, you never even make the shortlist. That is why we say visibility in the AI era is a new marketing channel.

This article is not about theory. We opened up five months of server logs from our own website—more than thirty AI crawlers and roughly 30 million content fetches—and examined them one by one: who came, what they took, how often they returned, and whether they brought customers back (more than thirty crawlers seen over five months; this article profiles 12 of them in detail). Every figure comes from our own server logs (directly measured), not estimates. The original full observation report was published on km.idaeo.ai.

Before going further, let us define the key term used throughout this article. In this article, "AI share of voice" means the requests AI crawlers make to your website: who comes, how often, and what they take. Every request is recorded by the server—countable and verifiable. The rest of this article simply reads those 30 million requests, group by group.

Horizontal bars rank twelve AI crawlers by all-time request volume and color-code training, indexing, human-triggered, and Google-family traffic.
High crawl volume does not mean high referral traffic: the largest crawlers bring no visitors, while OAI-SearchBot has much lower volume.

Meet a New Kind of Visitor: The AI Crawler

“Crawler” may sound technical, but its role is easy to understand. It is a machine visitor sent by an AI company to copy content from your website. Imagine a delivery app sending a representative to your restaurant to copy the menu. The representative does not order or spend money; they simply collect information so the platform can introduce your restaurant later. AI crawlers do the same thing, but they take the content back for different purposes: some use it to teach AI, some add it to a recommendation index, and some retrieve it for a real person asking a question right now.

Lesson One: If AI Has Never Read You, It Will Never Recommend You

No crawler request means no citation; no citation means no human traffic.

In plain English, AI must “read you” before it can recommend you—it must request one of your webpages at least once. Think of that as knocking on your door and opening your information. The AI company may collect it directly, or it may use a copy collected by someone else, such as the public database Common Crawl. Without that first request, you do not merely rank poorly in the AI world; you do not exist. You are like a restaurant absent from every guide and every platform: the food can be exceptional, but no one will mention it.

But this is only the price of admission, not a guarantee. We need to be honest about three things:

  • Being read ≠ being cited. A representative may copy your menu, but the platform may not recommend you right away. Our measured conversion rate is roughly 10–12% (source: reconciliation against historical citations; this is an approximation with no precise denominator, covering the observable scope and excluding zero-click citations for which the AI provides no link; see the methodology page for definitions).
  • Being cited ≠ receiving a click. AI often gives the entire answer directly, without a link. This is called a “zero-click” answer, and when it happens, we receive no data at all.
  • A visitor arriving ≠ us seeing the visit. Clicks inside mobile apps often carry no referral marker. The 476 visits we measured are the observable lower bound; the true number must be higher.

So there is only one viable order of operations: first make sure you are read, then improve the conversion from each reading.

A three-layer illustrative funnel shows AI crawler requests, conversion from crawled content to citations, and observable AI referrals.
Volume is the ticket in and conversion is the report card; the three layers use different scopes and cannot be multiplied directly.

Three Visitors, Three Futures

The logs divide crawlers neatly into three groups, and only the last two bring people back:

  • Training crawlers (GPTBot, ClaudeBot, Meta AI, and others): Think of them as encyclopedia editors who take your content back and write it into the AI’s “brain.” Results take months, sources are not credited, and customers do not arrive—but whether AI remembers you depends on them. First, let us set the record straight: they are not here to “take” anything—they are selecting material for AI’s knowledge base. If your content is selected, it becomes part of AI’s common knowledge. Across the entire AI crawler stack, this is the most foundational—and most important—layer.
  • Search-indexing crawlers (OAI-SearchBot, Bingbot, Applebot, and others): Think of them as travel-guide researchers who add your pages to a database that answer engines can consult at any time. That database is called an “index.” Results appear in hours to weeks, with cited links—and this is where customers come from.
  • User-triggered fetchers (ChatGPT-User and others): A real person is reading your page through AI right now, and the AI retrieves it on that person’s behalf. Results appear in seconds. It is like a customer already seated at your restaurant, asking the server to read the menu aloud. When this visitor appears, it means you got the first two steps right.

The one-line takeaway: if you want customers, feed search-indexing crawlers; if you want to be remembered over the long term, feed training crawlers; and when you see a user-triggered fetcher, it means the first two steps are working.

Three concept cards compare the purpose, time to impact, and traffic meaning of training, indexing, and human-triggered crawlers.
A crawler visit can signal three very different futures: model training, answer indexing, or a person reading now.

Finding One: Great Content Wins Over People, Not Machines

We originally assumed that “better content makes crawlers want to crawl more.” Our measured results showed otherwise: machines take everything, while people are selective. The representative copying the menu records every dish whether it tastes good or not; only the customer who sits down to eat makes a choice. Two completely independent measurements point in the same direction (both measured):

  • ChatGPT’s user-triggered fetcher read human-crafted deep-dive cards 4.7 times per page and automated pages 2.5 times per page—1.9 times as often.
  • Comparing quality grades green and red, user-triggered fetches were 3.14 versus 1.39—2.26 times as often.
  • When we ran the same test with the machine crawler ClaudeBot, the results were flat: green scored 2.56 and red scored 2.62, with red slightly higher.

What this means: better content will not make machines crawl you more, but once people arrive, they will read you more often—and AI will choose your content more often to show them. The return on content investment should therefore be judged at the “cited and traffic-driving” end of the funnel, not by crawler volume.

Grouped bars compare average crawls per page by OAI-SearchBot, GPTBot, and ClaudeBot across green, yellow, and red quality levels.
ClaudeBot is almost quality-agnostic, while OAI-SearchBot and ChatGPT human fetches respond more strongly to high-quality content.

Finding Two: The Visitors Most Likely to Bring Customers Are the Slowest to Arrive

After a new piece of content was published, how long did each crawler take to visit it for the first time (median, measured)?

  • ClaudeBot (training): 4.2 hours
  • GPTBot (training): 21.8 hours
  • OAI-SearchBot (search-indexing): 332.8 hours (13.9 days)
  • Bingbot (search-indexing): 374.6 hours (15.6 days)
Horizontal bars compare median first-visit lag after publication across AI crawlers; training crawlers arrive first, while indexing and human-triggered crawlers arrive later.
Training crawlers arrive early, but indexing and human-triggered crawlers that can produce citations and readers often arrive much later.

Launch a new dish and the menu-copying representative arrives the same day, but the guidebook researcher takes two weeks. During those two weeks, your hard-won new content may already be used for training while remaining absent from every list that could send customers your way. That gap is lost opportunity. And because the statistical window covers only 30 days, content published too late in the window never has time to await the slow crawlers. The slow crawlers are systematically undercounted, so the real gap can only be larger than these figures suggest.

The good news is that this problem has a solution, and the controls are already in your hands:

  • sitemap: the “complete menu directory” posted at the door, showing visitors every page available inside. Bingbot depends on it heavily, fetching it 32,993 times over 30 days (measured from our server logs; the sister article’s statistical window shows 33,155 — the difference comes from the timing of the snapshot).
  • llms.txt: a venue guide written specifically for AI. Amazonbot and GPTBot consume it heavily.
  • IndexNow: when you add a new dish, ring the bell and announce it instead of waiting for someone to arrive.

Deploying these three tools is the most direct way to accelerate the search-indexing crawlers that are the slowest to respond yet the most capable of bringing customers. To be candid, this is the logical next step inferred from “what they consume.” We have not yet run a before-and-after deployment experiment; we will report back when we do.

Grouped horizontal bars compare Bingbot, Applebot, GPTBot, and Amazonbot fetches of sitemap and llms.txt.
Bingbot lives on sitemap, while GPTBot and Amazonbot consume llms.txt; to accelerate a crawler, feed it what it likes.

A Reality Check: Is “Whenever I Launch a Major Site Overhaul, the Crawlers Arrive” Actually True?

The site team had long felt that every major integration project triggered a flood of crawlers. We pulled the database records to test that belief, and the answer was more interesting than a simple yes or no.

The sequence of events is real. On 2026-05-23, we created 550,807 llms summaries in a single day (measured; 5/23 = 550,807, 5/24 = 11,678, and every other day saw only a few thousand). Nine days later, on 6/01, three crawlers started up at once. ClaudeBot had visited only 669 times in the entire month of May, then surged to 53,468 on 6/01 and 354,414 on 6/02. GPTBot and Meta moved on the same day (measured). We also ruled out the alternative explanation of “a start-of-month schedule”: there were only 92 visits on 4/01 and 464 on 5/01. Only 6/01 exploded, so this was not a calendar effect (measured).

A logarithmic bar chart and event timeline show ClaudeBot daily volume surging after a large batch of llms summaries was created.
Three crawlers switched on together nine days after the large integration, but sequence does not establish causality.

But the data cannot tell us why they came. That project did two things at the same time: it improved the content and created 550,000 new URLs in one stroke. We verified that robots.txt had not changed—we did not open the door; we multiplied the visible surface area several times over. Imagine renovating a store and expanding it into an entire street on the same day. When the crowds arrive the following week, you cannot tell whether they came for the renovation or the new storefronts. To separate the effects, we need a clean experiment: improve only the quality of existing pages, add no new URLs at all, and then watch how the crawlers respond. Nothing like that has ever occurred in our records. It is the experiment we should deliberately run next.

Meet the Stars: Five Crawlers Worth Knowing

OAI-SearchBot—The Selective Guidebook Researcher, and the Most Important Player in This Study

Its volume is small (289,000 visits across all recorded history), but its value is unmatched: 85.6% of the URLs cited by ChatGPT appear in its crawl set (measured). What it selects, ChatGPT cites. This is a co-occurrence statistic; the two may simply favor the same types of pages, and causality has not been established. It is also the machine crawler most sensitive to quality in the entire study (green 5.94 versus red 3.54, measured). Its fatal weakness is speed: it takes 332.8 hours (13.9 days) to reach new content. Its slowness costs us, which makes speeding it up the first priority.

Two cards compare coverage of ChatGPT-cited URLs: OAI-SearchBot 85.6% versus GPTBot 34.9%.
Precision beats volume: the small, picky investigator decides who gets cited; co-occurrence only, causality unproven.

ClaudeBot—The Big Eater: Fast to Arrive, Hungry for Everything, but Brings No Customers

ClaudeBot ranks first by total volume (6.44 million visits across all recorded history), and it eats in bursts. Its median daily volume is only 97 visits, but its single-day high is 354,414 (measured). Most days it is nowhere to be seen; when it arrives, it clears the entire table. It responds to new content the fastest (a median of 4.2 hours), but is completely indifferent to quality and has brought back only 8 human visits across all recorded history (measured). Treat it as an investment in long-term presence, not sales performance.

Across its full history, ClaudeBot made 644 萬 crawls and brought back humans 8 times in measured results. Treat it as an investment in long-term presence, not performance.
644 萬 crawls versus 8 human visits: ClaudeBot is better framed as a long-term presence investment than a short-term performance engine.

Bingbot—The Old-School Researcher Who Checks the Store Every Day

Bingbot has the steadiest pulse, visiting daily at a stable volume (measured). It feeds on sitemap: 32,993 fetches over 30 days, the highest in the study. It also responds the slowest, taking 374.6 hours (15.6 days) to arrive. Yet Copilot, which it feeds, is our second-largest source of returning traffic (66 visits, 13.9%, measured). Using sitemap and IndexNow is most effective for Bingbot.

Log-scale bars compare daily max-to-median ratios of four crawlers, from Amazonbot's 4.74x to Meta AI's nominal 5,886x.
Some visit daily, some arrive once in months and sweep the table clean; Meta AI's small base is not directly comparable.

Applebot—The Mysterious Heavyweight: Crawling Like This Means Something Is Brewing

Applebot was July’s largest crawler (4.847 million visits over 30 days, measured), yet it has generated zero identifiable citation clicks so far (Siri is mostly zero-click). Its behavior is unusual: it almost never uses the AI training flag (Applebot-Extended appeared only twice in five months, measured), and it pays almost no attention to sitemap or llms.txt—it simply reads content quietly, at scale, and continuously.

Editor's observation (a judgment, not a measurement): One possible reading is that it is preparing for a product that has not yet been unveiled—a new Siri? A new search product? Another is that our brand migration triggered a full re-crawl (see the honest ledger). Until either is established, we watch without betting. So our strategy is to make no special investment, keep access fully open so it can read, and watch it every week. One further note: this volume includes 301 redirect requests to our old URLs (see the second correction on the methodology page), so read it as request volume, not content-fetch volume.

ChatGPT-User—The Customer Already Seated at the Table

ChatGPT-User represents user-triggered fetches. Its traffic peaks at 10 a.m. Japan time and follows a distinctly human daily rhythm (measured). Its appearance means someone is reading you right now. This is not future traffic; it is happening in the present. ChatGPT-User is a signal of success, not a target to optimize.

The dispersion spectrum runs from 0.05 to 0.42; higher means more human-like. GPTBot is 0.05, a pure machine schedule whose peak hour accounts for only 4.52%; Applebot is 0.06; OAI-SearchBot is 0.22 and peaks at 23 UTC; ChatGPT-User is 0.39 and peaks at 1 UTC, equal to 10 a.m. in Japan, matching human routines; PerplexityBot is 0.42, the most human-like overall, with its 13 UTC peak accounting for 10.18%. Data are measured.
Dispersion turns daily rhythm into a spectrum: GPTBot is the most clocklike, while PerplexityBot looks the most human.

The Other Players—Every One of Them Has a Reason to Be Here

Operating crawlers costs money, and every one that shows up has a purpose. Here is what each one is here to do:

  • GPTBot (OpenAI · primary training crawler): Its purpose is to bring content into the foundational knowledge used by the GPT family of models. It feeds on llms.txt (7,873 fetches over 30 days, measured; 8,618 is the combined llms.txt total across all crawlers — a different scope), so feeding llms.txt means feeding GPTBot. It brings no customers, but whether ChatGPT remembers you depends largely on it.
  • PerplexityBot (Perplexity · indexing crawler): Its purpose is to build the index for Perplexity’s answer engine. It is small but precise (59,000 visits across all recorded history), yet it delivered the first genuine citation in the history of our deep-dive cards (7/08, measured). Ranked by “citations earned per read,” it sits at the very top.
  • Amazonbot (Amazon · indexing crawler): Its purpose is to feed the retrieval ecosystem behind Alexa and the Rufus shopping assistant. It is quiet, steady, and shows up every day, with the highest share of llms.txt consumption in the field—for Amazonbot, llms.txt is especially effective.
  • Google family (blended labels): The 2.17 million visits recorded as “Google AI” are actually a combined count for the GoogleOther and Google-Extended labels. Each label serves a different purpose, but they are mixed together in our records. Until they are separated, any conclusion about Google should remain on hold.
  • CCBot (Common Crawl · public corpus): Its purpose is the most unusual: it builds a public corpus for the entire world, the shared water source for open-source models everywhere. washinmura.jp was added to the public corpus released in June 2026 (1,515 records, measured); once you enter it, you effectively enter every model that drinks from this source.
  • Bytespider (ByteDance · training crawler): Its purpose is to gather training data for ByteDance’s AI products. It logged 78,000 visits across all recorded history, with a clear tilt toward Chinese-language pages—its main battlefield is the Chinese-language ecosystem, and our Taiwan-Japan content is simply collected in passing.

The Honest Ledger: The Limits of This Record

Marketing content should be especially clear about its boundaries, because honesty is precisely what we want to give you:

  • n=1: This is a sample from one site, a five-month snapshot. It does not establish that these are universal crawler behaviors, much less laws.
  • Some of these “personalities” are our own doing: Applebot’s surge began the day after a brand migration, while Amazonbot’s high volume was amplified by falling into a black hole of junk paths. That is our architecture speaking, not their character.
  • Google crawlers cannot be attributed: “Google AI” is a combined count from two labels, so any conclusion based on it rests on blended data.
  • Page-level evidence from June has been lost: URL-level details are retained for only 30 days. Only daily totals remain from that wave, so no one should claim to know “which pages they actually crawled in June.”
  • Quality grades (green/yellow/red) are a custom three-tier labeling system used internally; the grading criteria are not disclosed in this article.
  • Quality testing has exposure bias: “better content” and “more pages” are highly correlated in the data (correlation coefficient 0.86, measured), and their effects cannot be completely separated.
  • The conversion rate of roughly one in ten comes from reconciliation against historical citations. It is an approximation within the observable scope, not exact visit-by-visit tracking. Time to first visit is reported only as a median, without sample size or distribution.
  • Zero-click answers and reposts are invisible: when AI gives the entire answer without a link, or another site republishes the content, neither leaves a trace on our server. This is a structural blind spot, not a failure to investigate.

Three Things to Take Away

  1. AI visibility is a real new marketing channel, and it begins with “being read.” Put sitemap, llms.txt, and IndexNow in place so the crawlers most likely to bring customers can find you faster.
  2. Measure the return on great content at the “human” layer: look at citations and referral traffic, not crawler volume. Machines take everything; people choose.
  3. Do not be fooled by big numbers. The crawler with the highest volume may bring no customers, while the ones that do bring customers are often small and slow. Focus your effort on the right few.

One final honest thought: no one can guarantee that AI will cite your content. Even our measured conversion rate is only roughly one in ten. But the reverse is certain: content that has never been read will never appear in an AI answer. Marketing in the AI era starts by making sure AI can see you.

To cite this article: IDAEO (2026). The Invisible Readers: 12 AI Crawlers Revealed. km.idaeo.ai/ai/ai-crawler-marketing. Data through 2026-07-24 (full-history window from 2026-03-03); source: our server-log measurements. Original full report: km.idaeo.ai/reports/ai-crawler-profiles-20260724

FAQ

I am not in the tech industry. What do AI crawlers have to do with my website?
If you want people to find you through AI, they have everything to do with it. Before AI recommends any website, a crawler must first read it. If your site has never been read, it simply does not exist in AI’s answers. This is not a question of ranking high or low; it is a question of whether you have the price of admission.
私はテクノロジー業界ではありません。AIクローラーは私のWebサイトとどんな関係があるのですか?AIを通じて誰かに見つけてもらいたいなら、関係があります。AIがWebサイトをおすすめする前には、必ずクローラーを送り、事前にそのサイトを読んでいます。読まれていなければ、あなたのWebサイトはAIの回答の中にそもそも存在しません。順位の高低ではなく、入場券を持っているかどうかの問題です。
I am not in the tech industry. What do AI crawlers have to do with my website?If you want people to find you through AI, they have everything to do with it. Before AI recommends any website, a crawler must first read it. If your site has never been read, it simply does not exist in AI’s answers. This is not a question of ranking high or low; it is a question of whether you have the price of admission.
If I write better content, will AI crawlers visit more often?
Our measurements say no. Machines take everything—ClaudeBot was almost entirely indifferent to the quality grades (green 2.56, red 2.62, with red slightly higher). But real people read good content through AI between 1.9 and 2.26 times as often (measured). Judge the return on good content by citations and referral traffic, not crawler volume.
良いコンテンツを書けば、AIクローラーはもっと頻繁に来るようになりますか?実測では、そうなりません。機械は選り好みしません。ClaudeBotは品質区分にほぼ反応せず(green 2.56、red 2.62で、redがわずかに上回りました)、一方で実ユーザーがAIを通じて良いコンテンツを読む回数は明らかに増えました(1.9倍から2.26倍、実測)。良いコンテンツの成果は引用と流入で評価し、クローラーの件数で評価しないでください。
If I write better content, will AI crawlers visit more often?Our measurements say no. Machines take everything—ClaudeBot was almost entirely indifferent to the quality grades (green 2.56, red 2.62, with red slightly higher). But real people read good content through AI between 1.9 and 2.26 times as often (measured). Judge the return on good content by citations and referral traffic, not crawler volume.
What exactly should I do to help AI discover my new content faster?
Prepare sitemap (a complete directory of your website’s pages) and llms.txt (a site guide written for AI), then use IndexNow to send proactive notifications. Bingbot fetched sitemap 32,993 times over 30 days (measured), while Amazonbot and GPTBot consume llms.txt heavily. For the search-indexing crawlers that respond the slowest yet bring the most customers, these are the most direct ways to accelerate discovery.
AIに新しいコンテンツをもっと早く見つけてもらうには、具体的に何をすべきですか?sitemap(Webサイト内の全ページ一覧)とllms.txt(AI向けのWebサイト案内)を用意し、IndexNowで能動的に通知してください。実測では、Bingbotは30日間でsitemapを32,993回クロールし、AmazonbotとGPTBotはllms.txtを重点的に利用していました。これは、反応が最も遅い一方で最も顧客を連れてくるインデックス型クローラーを加速させる、最も直接的な方法です。
What exactly should I do to help AI discover my new content faster?Prepare sitemap (a complete directory of your website’s pages) and llms.txt (a site guide written for AI), then use IndexNow to send proactive notifications. Bingbot fetched sitemap 32,993 times over 30 days (measured), while Amazonbot and GPTBot consume llms.txt heavily. For the search-indexing crawlers that respond the slowest yet bring the most customers, these are the most direct ways to accelerate discovery.
What is "AI share of voice"?
This article uses a concrete definition: **your AI share of voice is the requests AI crawlers make to your website** (as recorded in our 伺服器紀錄 server logs—countable and verifiable). The three kinds of requests map to three kinds of voice: training requests determine whether AI remembers you, indexing requests determine whether AI can cite you, and user-triggered requests mean AI is telling someone about you right now. Zero requests means zero share of voice—you will not appear in AI's answers.
「AIでの声量」とは何ですか?本記事の定義は具体的です。**AIクローラーがあなたのWebサイトに送るリクエストこそが、あなたのAIでの声量です**(当サイトの伺服器紀錄サーバーログに基づき、数えられ、検証できます)。3種類のリクエストは3種類の声量に対応します。学習型リクエストはAIがあなたを覚えているかどうかを決め、インデックス型リクエストはAIがあなたを引用できるかどうかを決め、実ユーザー代理取得のリクエストはAIがいままさに誰かにあなたを語っていることを意味します。リクエストがゼロなら、声量もゼロ——AIの答えにあなたは登場しません。
What is "AI share of voice"?This article uses a concrete definition: **your AI share of voice is the requests AI crawlers make to your website** (as recorded in our 伺服器紀錄 server logs—countable and verifiable). The three kinds of requests map to three kinds of voice: training requests determine whether AI remembers you, indexing requests determine whether AI can cite you, and user-triggered requests mean AI is telling someone about you right now. Zero requests means zero share of voice—you will not appear in AI's answers.
If I follow these steps, am I guaranteed to be cited by AI?
No. Being read is only a necessary condition. Our measured read-to-citation conversion rate is roughly 約一成, and this is an n=1 observation of a single site; the findings cannot be assumed to apply directly to other websites. What we do know for certain is the reverse: content that has never been read cannot appear in an AI answer.
そのとおりにすれば、AIに引用されることが保証されますか?保証されません。読まれることは必要条件にすぎません。私たちが実測した、読まれてから引用に至る転換率は約約一成です。しかも、これはn=1の単一サイト観測であり、結論をほかのWebサイトへそのまま適用できる保証はありません。確実なのは、その逆です。読まれていないコンテンツが、AIの回答に現れることはありません。
If I follow these steps, am I guaranteed to be cited by AI?No. Being read is only a necessary condition. Our measured read-to-citation conversion rate is roughly 約一成, and this is an n=1 observation of a single site; the findings cannot be assumed to apply directly to other websites. What we do know for certain is the reverse: content that has never been read cannot appear in an AI answer.
How do I find out which AI crawlers have visited my website?
Check the User-Agent field in your server access logs—names like GPTBot, ClaudeBot, and OAI-SearchBot appear there directly. To verify identities, cross-check IPs against each company's officially published ranges (crawler identities can be spoofed; every figure in this article passed that verification). All the numbers in this article were counted exactly this way.
どのAIクローラーが自分のサイトに来たか、どうすれば分かりますか?サーバーのアクセスログにあるUser-Agent欄を確認してください。GPTBot、ClaudeBot、OAI-SearchBotといった名前がそのまま現れます。なりすましがあるため、身元の確認はIPの逆引きと各社が公開している公式IPレンジとの照合で行います(本記事のデータはすべてこの検証を通過しています)。本記事の数字は、すべてこの方法で1件ずつ数えたものです。
How do I find out which AI crawlers have visited my website?Check the User-Agent field in your server access logs—names like GPTBot, ClaudeBot, and OAI-SearchBot appear there directly. To verify identities, cross-check IPs against each company's officially published ranges (crawler identities can be spoofed; every figure in this article passed that verification). All the numbers in this article were counted exactly this way.
Would blocking the training crawlers that bring no customers hurt me?
It is a trade-off. Training crawlers (GPTBot, ClaudeBot, and others) indeed bring no customers, but they determine whether AI remembers you—blocking them means giving up your place in AI's long-term memory. Our choice is to stay fully open and observe. Whether to block is each site's own strategy, but you should at least know which layer of your share of voice you are giving up.
顧客を連れてこない学習型クローラーをブロックしても大丈夫ですか?トレードオフがあります。学習型クローラー(GPTBot、ClaudeBotなど)は確かに顧客を連れてきませんが、AIがあなたを覚えているかどうかを決めるのはこのタイプです。ブロックすることは、AIの長期記憶の中の居場所を手放すことを意味します。当サイトの選択は「全面開放して観測する」です。ブロックするかどうかは各サイトの戦略ですが、どの層の声量を手放すのかは知っておくべきです。
Would blocking the training crawlers that bring no customers hurt me?It is a trade-off. Training crawlers (GPTBot, ClaudeBot, and others) indeed bring no customers, but they determine whether AI remembers you—blocking them means giving up your place in AI's long-term memory. Our choice is to stay fully open and observe. Whether to block is each site's own strategy, but you should at least know which layer of your share of voice you are giving up.
After publishing new content, how long until AI can cite it?
In our measurements, the two indexing crawlers most likely to produce citations are also the slowest: OAI-SearchBot takes a median of 332.8 hours (13.9 days) and Bingbot 374.6 hours (15.6 days) to first visit new content—you must be seen before you can be cited. Plan for at least a two-week lead time, and use sitemap, llms.txt, and IndexNow to speed things up (a before-and-after control experiment is still pending; treat this as a reasoned recommendation).
新しいコンテンツは、公開からどれくらいでAIに引用されますか?当サイトの実測では、引用につながりやすい2つのインデックス型クローラーが最も遅く、OAI-SearchBotは中央値332.8時間(13.9日)、Bingbotは374.6時間(15.6日)かかって初めて新コンテンツを見に来ます。見られて初めて、引用の可能性が生まれます。新コンテンツには最低2週間以上の待機期間を見込み、sitemap、llms.txt、IndexNowで加速してください(導入前後の比較実験は未実施のため、合理的な推奨として扱ってください)。
After publishing new content, how long until AI can cite it?In our measurements, the two indexing crawlers most likely to produce citations are also the slowest: OAI-SearchBot takes a median of 332.8 hours (13.9 days) and Bingbot 374.6 hours (15.6 days) to first visit new content—you must be seen before you can be cited. Plan for at least a two-week lead time, and use sitemap, llms.txt, and IndexNow to speed things up (a before-and-after control experiment is still pending; treat this as a reasoned recommendation).

Cite this article

TK Lin・《The Invisible Readers: 12 AI Crawlers Revealed》・IDAEO 知識庫・2026-08-09・https://km.idaeo.ai/ai/ai-crawler-marketing

Updated 2026-08-10

更新 2026-08-10T10:29:32.169Z · server-rendered · four-language · IDAEO 知識庫