km.idaeo.ai · IDAEO 知識庫

🏛 Part of the "ai" topic shelf →

Why a New Site Was Heavily Cited by AI Search Within Two Weeks — A Real crawl→index→cite Breakdown

A new site launched on 2026-09-20 saw its Google impressions rise within two weeks and, in one Perplexity citation probe, was heavily sourced and ranked first. But there is no control group, so the causes can't be separated. All I can offer is a hypothesis consistent with the data: the right parent domain, plus facts turned into a source machines can verify.

Infographic: crawl to index to cite — 457 crawler visits, Google impressions 4 to 265, #1 source in one test
Figure: A new website, live for just two weeks, became the most frequently chosen source in one AI search test. In its first few days, 6 AI/search tools came to read it 457 times in total; Google impressions jumped from 4 to 265 in 14 days (about 66 times); in this test, AI answers used it as a source more often than any other (ranked first, chosen about 91% of the time). Single website, one test, no control group.

A site that only went live on 2026-09-20 saw its Google impressions climb within two weeks; separately, a Perplexity citation probe pulled it heavily as a source and ranked it first. Why so fast? I have no control group and ran no item-by-item ablation tests, so I can't yet separate the causes.

What I can offer for now is a hypothesis consistent with the data: the site sits under a parent domain that AI crawlers already visited routinely, and its facts are organized into a source machines can easily use and trace back to check. The parent domain may have gotten it seen sooner; the evidence structure may have made it easier to pick. How much each contributed, this measurement can't answer.

This piece lays out the numbers, and how far the numbers let us infer. For the concept of AEO and how it differs from SEO, see 《AEO Is Not SEO》; for how to build it, see 《Personal IDA Methodology》.

Real traces at three nodes

Bar chart: AI and search crawler visits in the first days (ClaudeBot 183 down to PerplexityBot 2)
Figure: In the new website's first few days, all 6 major AI/search crawlers came to read it, but visit counts varied widely—Claude's bot was the most active (183 visits), while Perplexity came only 2 times. The numbers are the 'verified bot identities' recorded by Cloudflare, not what each bot claims to be.

This is my own fact page, tklin.washinmura.jp, sitting under the parent domain washinmura.jp, which has several years of history. It went live on 2026-09-20. The parent domain already had AI-crawler traffic. At the three stages of crawl, index, and cite, the site left three different kinds of signal.

Crawl. Within days of launch, Cloudflare's "verified bot identity" recorded these counts: ClaudeBot 183, bingbot 161, Googlebot 67, GPTBot 36, OAI-SearchBot 8, PerplexityBot 2. These are verified bot identities, not self-declared User-Agents. All six crawlers came, but their visit volumes differed widely.

Index. Google Search Console impressions went, in sequence, 4 → 32 → 89 → 171 → 265 over two weeks. Rising GSC impressions prove search visibility is growing; they are not the same as the number of indexed pages.

Cite. I ran one citation probe with Perplexity's Agent API: 100 questions × 5 model backends. All 5 backends went through Perplexity's own web search. What this measured is whether Perplexity search would pull this site as a source, not five independent AIs each making its own judgment.

The result: a retrieval success rate of retrieved 91% (identity / company questions). Source counts: tklin.washinmura.jp 439 (first), tklin.me 246, washinmura.jp 145. All 5 tested backends drew on the new site; on questions that describe him without naming him, 19/20 pointed to him. This is a strong signal that Perplexity search drew on the site heavily; it can't be rewritten as "every AI consistently cites it first in ordinary conversation."

The timing also needs to be stated precisely. "Two weeks" is the window over which GSC impressions were observed after launch. The citation probe was a separate run; the DATA source didn't record the exact execution date, so I can't say the citations happened precisely within those two weeks. What is certain: search visibility rose within two weeks, and at the time of that probe, Perplexity search was already pulling the site heavily as a source.

Why it was crawled fast, and why it was picked as a source

Bar chart: most-picked source in one AI-search test, 439 vs 246 vs 145, 91% retrieved
Figure: In the same AI search test, AI answers used this new website as a source most often (439 times, ranked first), ahead of the author's old website (246) and the parent website (145); for identity-type questions, it was used about 91% of the time. This is the result of one test and does not mean every AI does this every time.

Start with crawling. The parent domain's existing traffic is one possible cause. washinmura.jp has years of history, and AI crawlers already came often; a subdomain riding that traffic pipeline may have skipped the cold-start wait of a brand-new domain. But this supports only the claim that it "rode existing crawler traffic"; it can't be used to claim that the parent domain passed "authority" down to the subsite. The data doesn't prove inheritance.

The technical setup may have helped too. robots.txt allows everything, the HTTP header sets X-Robots-Tag: all, sitemap.xml lists the site's pages, and IndexNow actively pushes updates. When crawlers come, they can get in, and they have a chance of being notified about new pages. How much the parent domain and the technical setup each contributed can't be separated without a controlled experiment.

Then there is a seemingly contradictory number: PerplexityBot came only 2 times, yet the site was sourced 439 times in the probe. A reasonable inference is that Perplexity's live retrieval on this occasion piggybacked on general search indexes rather than relying only on its own crawler to build a corpus. As a meta-search, Perplexity calls on search results from Bing / Google and others, then hands them to the model to synthesize; once a general index has taken in a page, Perplexity may be able to find it.

This inference applies only to this one Perplexity probe. The crawler requests Cloudflare recorded and the sourcing during Agent API retrieval are also not the same metric; you can't extrapolate from 2 visits straight to 439. Still less can the observation be stretched into "dedicated crawlers don't matter": in the same records, ClaudeBot came 183 times and GPTBot 36 times, and some AI retrieval systems, OpenAI's native retrieval environment for example, rely quite heavily on their own dedicated crawlers to fetch and build their corpus. General search engines and dedicated crawlers: block neither.

As for why it was picked once indexed, I think the evidence structure may have played a role. This is a mechanism hypothesis; I ran no item-by-item removal (ablation) tests and can't write it up as proven causation.

The schema.org JSON-LD uses a single @id to bind Person and Organization to the same canonical identity, which in theory reduces same-name confusion. anchors.json holds 44 third-party sources, each with a Wayback snapshot and sha256, making it easy to trace back and check entry by entry. sha256 can only verify whether content is consistent and whether it has been tampered with; it cannot verify that the content itself is true. llms.txt, llms-full.txt, brief.txt, and the four language versions provide plain-text facts that machines can easily parse, and raise the chance of matching questions asked in different languages.

These arrangements may have jointly influenced the 91% retrieval rate on identity questions and the 19/20 result on the unnamed questions. Which item worked, and how much each accounted for, no evidence so far can separate.

What is method, what is luck

Two columns: four repeatable moves versus two case-specific advantages
Figure: In this case, four things can be replicated (full technical access, a unified identity, a third-party evidence chain, and plain-text fact pages), and anyone can do them; the two that cannot be copied (being hosted under a parent website with years of traffic, and low competition for the person's name) are luck specific to this case. Get the four replicable items solid first.

Some starting conditions can't be taken along by others.

One is the parent domain. washinmura.jp's history and existing AI-crawler traffic gave the subsite a ready-made visiting pipeline. That doesn't mean the parent domain's authority has been inherited by the subsite; the data doesn't prove that.

The other is low competition for the name and a distinctive record. There aren't large numbers of authoritative same-name pages online competing with him, so once the fact page was indexed, it may have been easier to lock onto as a source. This may also relate to the 19/20 result on the unnamed questions, but I have no measurement of same-name competition, so it remains conjecture.

Four things can be followed: opening everything technically (robots, X-Robots-Tag, sitemap, IndexNow); entity disambiguation with a single @id; using anchors.json to link a third-party evidence chain checkable via Wayback and sha256; and providing llms.txt and multilingual plain-text fact pages. The method can be copied; this case's speed and sourcing volume can't be guaranteed as-is.

To replicate it: three steps

First, find a node crawlers already visit. Prefer building a subsite under a domain crawlers come to often, or at least build backlinks from a site that is already indexed. Ping IndexNow at launch.

Second, get identity and evidence in order. Fix identity with Person / Organization JSON-LD carrying a single @id; publish anchors.json so third-party sources can be checked entry by entry via the Web Archive and sha256.

Third, give machines text that's easy to use. Serve llms.txt and multilingual plain-text fact summaries at the root, and cut redundant layout so retrievers can read and cite easily.

These three steps are the skeleton. For the full four conditions, the five steps, and the discipline of "feed only the verifiable," see 《Personal IDA Methodology》.

Honest limitations

First, the causes still can't be separated. There is no fresh bare-domain control group and no item-by-item removal (ablation) test, so the separate contributions of the parent domain's existing traffic, content quality, and technical setup can't be told apart. The same spec placed on a brand-new bare domain with no history would very likely be indexed noticeably slower; that is a reasonable expectation, not a verified certainty. This article records one site's trajectory; it is not a controlled experiment.

Second, 439 is a probe-specific value. It comes from the Perplexity Agent API with web_search deliberately turned on, at a scale of 100 questions × 5 model backends. All 5 backends used Perplexity's own web search. 439 means tklin.washinmura.jp ranked first by source count in this Perplexity search; it does not mean five independent AIs each chose it as their top pick, still less that every AI consistently cites it first in ordinary conversation. Likewise, rising GSC impressions only prove increased search visibility; they can't be read as the number of indexed pages, nor do they prove every engine has indexed the site completely.

Third, crawling and sourcing run on two separate ledgers. PerplexityBot came only 2 times, yet the probe sourced the site 439 times. Piggybacking on general search indexes is a reasonable inference for this Perplexity observation; Cloudflare's crawler counts can't be converted directly into Agent API source counts, nor can you conclude that all AI search takes the same path. Dedicated crawlers still matter a great deal for some engines.

Fourth, cold demand (L3) is not yet proven. The questions this time mostly revolved around a specific person's identity. For cold-demand (L3) queries that don't involve this person and compete purely on topic, this batch of data has no evidence yet.

Within two weeks of launch, the new site's search visibility rose markedly; a separate probe also did show Perplexity search drawing on it heavily. "The right parent domain, plus a fact source machines can check" can explain the data in front of us, but it remains a hypothesis. What can be learned first is to do the four replicable things solidly; how much time the parent domain actually won for it, I still can't calculate.

FAQ

Was it really cited within two weeks?
"Two weeks" is the observation window for GSC impressions after launch. The citation probe was a separately run single cross-section, and DATA didn't record the exact execution date, so I can't claim the citations happened precisely within the two weeks. What is certain: indexing signals were clearly rising within two weeks, and the probe shows the site was already being pulled heavily as a source by Perplexity's search.
本当に二週間で引用されたのですか? — 「二週間」は、公開後に GSC の表示回数を観察した期間です。引用プローブは別に実行した一回の断面で、DATA には正確な実行日が記録されておらず、引用がちょうど二週間のうちに起きたとは主張できません。確実に言えるのは、二週間のうちにインデックスのシグナルが明らかに伸びていたこと、そしてプローブが、このサイトがすでに Perplexity の検索で大量にソースとして取得されていたことを示していることです。
Was it really cited within two weeks? — "Two weeks" is the observation window for GSC impressions after launch. The citation probe was a separately run single cross-section, and DATA didn't record the exact execution date, so I can't claim the citations happened precisely within the two weeks. What is certain: indexing signals were clearly rising within two weeks, and the probe shows the site was already being pulled heavily as a source by Perplexity's search.
Which crawlers came, and how many times each?
Within days of launch, Cloudflare verified bot identity recorded: ClaudeBot 183, bingbot 161, Googlebot 67, GPTBot 36, OAI-SearchBot 8, PerplexityBot 2. These are verified bots, not self-declared User-Agents.
どのクローラーが、それぞれ何回来ましたか? — 公開から数日のうちに、Cloudflare の認証済みクローラー ID で記録されたのは、ClaudeBot 183、bingbot 161、Googlebot 67、GPTBot 36、OAI-SearchBot 8、PerplexityBot 2 です。検証済みのボットであり、User-Agent の自称ではありません。
Which crawlers came, and how many times each? — Within days of launch, Cloudflare verified bot identity recorded: ClaudeBot 183, bingbot 161, Googlebot 67, GPTBot 36, OAI-SearchBot 8, PerplexityBot 2. These are verified bots, not self-declared User-Agents.
What does 439 mean?
In one probe (100 questions × 5 backends, Perplexity Agent API with web_search on), tklin.washinmura.jp was sourced 439 times and ranked first. It means first place within Perplexity's search; it doesn't mean every AI consistently treats it as the top source in ordinary conversation.
439 とは何を意味しますか? — 一回のプローブ(100 問 × 5 つのバックエンド、web_search を有効にした Perplexity Agent API)で、tklin.washinmura.jp が 439 回ソースとして取得され、第一位だったという意味です。Perplexity の検索で第一位だったことを表すもので、一般的な対話でどの AI も安定してこのサイトを第一のソースにすることと同じではありません。
What does 439 mean? — In one probe (100 questions × 5 backends, Perplexity Agent API with web_search on), tklin.washinmura.jp was sourced 439 times and ranked first. It means first place within Perplexity's search; it doesn't mean every AI consistently treats it as the top source in ordinary conversation.
Why did PerplexityBot visit only 2 times when the site was sourced 439 times?
A reasonable inference is that a meta-search like Perplexity piggybacked on general search indexes instead of relying only on its own crawler to build a corpus. But this holds only for this one Perplexity run and can't be generalized to all AI search; some engines depend heavily on their own dedicated crawlers, so block neither side.
なぜ PerplexityBot は 2 回しか来ていないのに、439 回もソースとして取得されたのですか? — 合理的な推論は、Perplexity のような meta-search が汎用検索インデックスに相乗りしており、自社クローラーだけでインデックスを構築しているわけではない、というものです。ただし、これは Perplexity の今回の一回についてのみ成り立つことで、すべての AI 検索がそうだと一般化することはできません。自社の専用クローラーに大きく依存するエンジンもあるので、どちらもブロックしないでください。
Why did PerplexityBot visit only 2 times when the site was sourced 439 times? — A reasonable inference is that a meta-search like Perplexity piggybacked on general search indexes instead of relying only on its own crawler to build a corpus. But this holds only for this one Perplexity run and can't be generalized to all AI search; some engines depend heavily on their own dedicated crawlers, so block neither side.
What can be replicated, and what was luck?
Two things are hard to carry over: the parent domain's existing traffic, and low competition for the name. Four things can be replicated: opening everything technically, single-@id entity disambiguation, the anchors.json third-party evidence chain, and llms.txt with multilingual plain-text fact pages.
どこが複製でき、どこが運ですか? — そのまま真似しにくいのは二項目:親ドメインの既存トラフィック、人名の競合の少なさ。複製できるのは四項目:技術面の全開放、単一の @id によるエンティティの曖昧性解消、anchors.json による第三者の証拠の連鎖、llms.txt と多言語のプレーンテキスト事実ページです。
What can be replicated, and what was luck? — Two things are hard to carry over: the parent domain's existing traffic, and low competition for the name. Four things can be replicated: opening everything technically, single-@id entity disambiguation, the anchors.json third-party evidence chain, and llms.txt with multilingual plain-text fact pages.
Does this mean every AI will cite you?
No. There is no fresh bare-domain control group and no item-by-item ablation testing, so the contributions can't be separated; 439 is a probe-specific value; cold-demand queries are unproven. This is a hypothesis consistent with the data, not a conclusion confirmed by taking the factors apart.
これは、すべての AI があなたを引用するということですか? — そうではありません。新規のネイキッドドメインによる対照群はなく、項目を一つずつ外すテストも行っておらず、寄与を切り分けられません。439 は特定のプローブでの値です。コールド需要のクエリは実証されていません。これはデータと整合する仮説であり、分解して実証された結論ではありません。
Does this mean every AI will cite you? — No. There is no fresh bare-domain control group and no item-by-item ablation testing, so the contributions can't be separated; 439 is a probe-specific value; cold-demand queries are unproven. This is a hypothesis consistent with the data, not a conclusion confirmed by taking the factors apart.

Cite this article

TK Lin・《Why a New Site Was Heavily Cited by AI Search Within Two Weeks — A Real crawl→index→cite Breakdown》・IDAEO 知識庫・2026-10-04・https://km.idaeo.ai/post/ai/why-new-site-cited

更新 2026-10-04T10:13:00.123Z · server-rendered · four-language · IDAEO 知識庫