km.idaeo.ai · IDAEO 知識庫

🏛 Part of the "ai" topic shelf →

Personal IDA Methodology: How to Build an Official Fact Source for a Person That AI Will Cite

A personal IDA is an official fact base built for one person. It has to be verifiable, machine-readable, honestly partitioned, and recomputable, all at once. The method has five steps and only one rule: feed only what can be verified.

When AI is asked about you, it had best have on hand a fact sheet you'd put your name to and that others can trace back to sources. That is what a personal IDA does. For the crawl→index→cite concept behind it and how it differs from SEO, see 《AEO Is Not SEO》; for my measured numbers and inferences about mechanism, see 《Why a New Site Was Heavily Cited by AI Search Within Two Weeks》. This piece covers how to build one.

Definition

A personal IDA (Identity Data Aggregation) is a person's official fact base: the content is verifiable, the format is machine-readable, and it aims, as far as possible, to become an authoritative source AI is willing to cite. When AI is asked "who is this person," the goal is for it to find the version the person is willing to answer for, and to lean less on third-party retellings or wrong information.

Why do it

I've observed that for more and more people, the first step in getting to know someone is to ask AI directly; for them, the AI's answer is the first impression. This is a trend observation, not a number measured in this case.

If the person hasn't provided a checkable version, AI can only piece things together from third-party material. That material may be outdated, partial, or simply wrong, and the answer may still sound very sure. You can't wait for it to get things wrong and then expect everyone who asked to come read the correction. Having accountable facts ready in advance is what gives it a chance to pick the right source.

Four conditions

Four conditions of a personal IDA: verifiable, machine-readable, separated, reproducible
Figure: The four conditions of a personal IDA—miss any one and it is just a personal website: ① Checkable (every sentence has a source) ② Machine-readable (unified identity, multilingual) ③ Clearly separated (facts kept apart from self-praise) ④ Re-verifiable (others can check it against the sources themselves).

Having a personal website only gives you a location. A personal IDA also has to meet four conditions.

  1. Verifiable: every fact has an anchor that leads back to an original source or public record. Adjectives like "outstanding" or "veteran" can't stand in for facts.
  2. Machine-readable: use schema.org Person / Organization structured data; use a single `@id` site-wide as this person's canonical identifier, so that scattered Person and related Organization data all anchor to the same person while the distinction between person and company is preserved. Then provide multilingual versions to reduce machine guesswork.
  3. Honest boundaries: verified facts are labeled separately from self-descriptions and other people's characterizations.
  4. Recomputable: only when a third party can trace back to the original source, check the web archive, and recompute the hash can you talk about citation-grade. sha256 can verify whether content is consistent and whether it has been tampered with; it cannot verify that the content itself is true.

Five steps

Five steps of the personal IDA method, following crawl to index to cite
Figure: The five steps of a personal IDA, moving forward along 'crawled → indexed → cited': ① Build a fact base ② Make it machine-readable ③ Open the door for AI ④ Track across three layers ⑤ Measure whether AI uses it. There is only one rule: feed only what can be verified.
Track in three layers: crawl via Cloudflare, index via Search Console, cite by asking AI
Figure: After doing the work, check how far it has gotten across three layers—for crawling, look at Cloudflare's verified crawler identities; for indexing, look at impressions in Google Search Console; for citation, ask AI a fixed set of questions. The three layers may each move on their own, so do not use one layer to infer another.

These five steps move along crawl → index → cite. Each step should leave visible output.

Step 1: Build the fact base. Put only verifiable content in the fact section and mark the anchor for each item; self-descriptions and unverified claims go elsewhere. Keep web archives of your sources and record their hashes, so recomputation is possible later. This step decides whether the fact base can be trusted. Don't rush it.

Step 2: Make it machine-readable. Produce JSON-LD (Person and Organization, sharing a single `@id`), multilingual versions, and `llms.txt`. Organize third-party sources into a machine-readable `anchors.json`, each entry with a web archive and a hash.

Step 3: Open up to crawlers. Fully open `robots.txt`, submit a sitemap, enable IndexNow, and set the HTTP header `X-Robots-Tag: all`. Once the pages are written, crawlers still have to be able to get in.

Step 4: Get indexed, and track three layers. Each has its own signal:

  • Crawl: look at Cloudflare's verified bot identities to confirm which bots actually came.
  • Index / impressions: look at Google Search Console impressions to track search visibility; rising impressions are not the same as the number of indexed pages.
  • Cite: use a fixed question set or a citation probe to see whether the page was retrieved and written into the answer.

The numbers in 《Why a New Site Was Heavily Cited by AI Search Within Two Weeks》 show that the three layers may each go their own way. Using one layer to draw conclusions about another tends to miss.

Step 5: Measure the citation rate. Use a fixed question set and track regularly; record separately whether the page was retrieved as a source (retrieved) and whether it was written into the answer (cited). Only with the two metrics apart do you know which stage the fact base is actually working in.

A real case, in brief

The personal IDA I built for myself lives at tklin.washinmura.jp and went live on 2026-09-20. Over the following two weeks, Google impressions kept rising; separately, a Perplexity citation probe retrieved the site heavily as a source. For the full numbers, what each of the three layers measured, and the inference and limits around "few crawler visits, heavy sourcing," see 《Why a New Site Was Heavily Cited by AI Search Within Two Weeks》.

The honesty principle: feed only the verifiable

Facts zone versus claims zone: sourced facts kept apart from unverified claims
Figure: The facts section vs. the self-description section must be kept apart—the facts section (sourced, verifiable, safe for AI to cite) and the self-description section (self-praise, how others describe you, media-given titles, items still awaiting verification—labeled as such and not mixed in with facts). This line is what separates a personal IDA from a press release.

AI may repeat errors verbatim too. So every sentence in the fact base has to hold up when someone goes back to check.

Media titles and interview claims circulating online, however flattering, can't go into the fact section if the person can't confirm them with an original source. Put them in, and AI may keep passing them on as facts the person certified.

The handling is simple:

  • Verified facts: attach anchors and put them in the fact section.
  • Self-descriptions, other people's characterizations, and claims pending verification: put them in a separate section, clearly labeled for what they are.

A personal IDA's credibility rides on this line. Every "fact" you give AI has to be one you're willing to let others check.

An honest caveat

So far only my own single case supports this method, and the limitations can't be left out:

  • No fresh bare-domain control group, and no item-by-item ablation tests. tklin.washinmura.jp sits under a parent domain that already had an AI-crawler traffic pipeline. I can't separate the contributions of the parent domain, content quality, and technical setup; nor does the data prove the parent domain passed authority down to the subsite. On a brand-new bare domain with no history, indexing would very likely be noticeably slower. That is a reasonable expectation, not a verified certainty.
  • "Heavily cited" has specific measurement conditions. The 439 source retrievals came from a Perplexity Agent API probe with web_search deliberately turned on. All 5 backends tested went through Perplexity's own web search; what was measured is whether Perplexity search would pull the site as a source, not five independent AIs each judging on their own. Nor does this value mean every AI consistently treats it as the top source in ordinary conversation.
  • Cold demand (L3) is not yet proven. The questions this time mostly revolved around a specific person's identity. For cold-demand (L3) queries that don't involve this person and compete purely on topic, there is currently no evidence.

For fuller controls and caveats, see the honest-limitations section of 《Why a New Site Was Heavily Cited by AI Search Within Two Weeks》.

Closing

What a personal IDA governs is a person's identity and fact sources. The five steps get the source built; what keeps it standing is always the same rule: feed only what can be verified.

Verifiable anchor: personal IDA (tklin.washinmura.jp). Measurement uses three layers and three sources: crawl via Cloudflare verified bot identities, indexing via Google Search Console impressions, citation via the Perplexity citation probe. All numbers in this article are based on these.

FAQ

What is a personal IDA?
Identity Data Aggregation: a person's official fact base, with verifiable content, a machine-readable format, and the aim of becoming, as far as possible, an authoritative source AI is willing to cite.
個人 IDA とは何ですか? — Identity Data Aggregation、つまり一人の人物の公式の事実データベースです。内容は検証可能で、形式は機械が読み取れ、できるかぎり AI が引用したくなる権威あるソースとなることを目指します。
What is a personal IDA? — Identity Data Aggregation: a person's official fact base, with verifiable content, a machine-readable format, and the aim of becoming, as far as possible, an authoritative source AI is willing to cite.
Why do it?
For more and more people, the first step in getting to know someone is asking AI directly. If the person hasn't provided an official version, AI can only piece one together from third parties, which may be outdated, partial, or wrong, yet stated with confidence. (This is a trend observation, not a number measured in this case.)
なぜ取り組むのですか? — 人を知る最初の一歩として AI に直接尋ねる人が増えています。本人が公式版を提供していなければ、AI は第三者の情報をつなぎ合わせるしかなく、それは古かったり、一面的だったり、誤っていることさえあるのに、断定的に語られます。(これは傾向の観察であり、今回の事例で計測した数字ではありません。)
Why do it? — For more and more people, the first step in getting to know someone is asking AI directly. If the person hasn't provided an official version, AI can only piece one together from third parties, which may be outdated, partial, or wrong, yet stated with confidence. (This is a trend observation, not a number measured in this case.)
What are the four conditions?
Verifiable (every item has an anchor), machine-readable (schema.org + single @id + multilingual), honest boundaries (facts and self-descriptions in separate sections), and recomputable (a third party can trace back to the original source and recompute the hash).
四つの条件とは何ですか? — 検証可能(一件ごとにアンカーがある)、機械可読(schema.org+単一の @id+多言語)、正直な境界(事実と自己申告を区分する)、再計算可能(第三者が原典と突き合わせ、ハッシュを再計算できる)です。
What are the four conditions? — Verifiable (every item has an anchor), machine-readable (schema.org + single @id + multilingual), honest boundaries (facts and self-descriptions in separate sections), and recomputable (a third party can trace back to the original source and recompute the hash).
What are the five steps?
Build the fact base → make it machine-readable (JSON-LD / llms.txt / anchors.json) → open up to crawlers (robots / sitemap / IndexNow / X-Robots-Tag) → get indexed and track three layers (crawl / index / cite) → measure the citation rate (record retrieved and cited separately).
五つのステップとは? — 事実データベースを作る→機械可読にする(JSON-LD/llms.txt/anchors.json)→クローラーに開放する(robots/sitemap/IndexNow/X-Robots-Tag)→インデックスに入り、三層に分けて追跡する(クロール/インデックス/引用)→引用率を測る(retrieved と cited を分けて記録)。
What are the five steps? — Build the fact base → make it machine-readable (JSON-LD / llms.txt / anchors.json) → open up to crawlers (robots / sitemap / IndexNow / X-Robots-Tag) → get indexed and track three layers (crawl / index / cite) → measure the citation rate (record retrieved and cited separately).
What does "feed only the verifiable" mean?
AI may repeat what you give it verbatim, including what's wrong. Media titles and claims that can't be confirmed with an original source can't go in the fact section; separate them and label what they are. That line is what separates a personal IDA from a PR piece.
「検証可能なものだけを与える」とはどういう意味ですか? — AI は、あなたが与えた内容を、誤りも含めてそのまま復唱するかもしれません。原典で裏づけられないメディアの称号や表現は、事実セクションに入れてはいけません。セクションを分け、性質を明示します。この一線が、個人 IDA と PR 原稿を分けます。
What does "feed only the verifiable" mean? — AI may repeat what you give it verbatim, including what's wrong. Media titles and claims that can't be confirmed with an original source can't go in the fact section; separate them and label what they are. That line is what separates a personal IDA from a PR piece.
Does this method work?
So far it is supported only by measurements from the author's own single case, on a site under a parent domain that already had AI-crawler traffic, so the contributions of parent domain, content, and technical setup can't be separated. On a brand-new bare domain, indexing would very likely be noticeably slower (a reasonable expectation, not a verified certainty). Cold-demand queries are not yet proven.
この方法は効果がありますか? — 現時点では著者自身の一事例の実測による支持しかなく、しかもすでに AI クローラーのトラフィックがある親ドメインの下に置かれているため、親ドメイン・コンテンツ・技術の寄与を切り分けられません。まったく新規のネイキッドドメインに替えれば、インデックスは明らかに遅くなる可能性が高いでしょう(合理的な予想であり、検証済みの必然ではありません)。コールド需要のクエリはまだ実証されていません。
Does this method work? — So far it is supported only by measurements from the author's own single case, on a site under a parent domain that already had AI-crawler traffic, so the contributions of parent domain, content, and technical setup can't be separated. On a brand-new bare domain, indexing would very likely be noticeably slower (a reasonable expectation, not a verified certainty). Cold-demand queries are not yet proven.

Cite this article

TK Lin・《Personal IDA Methodology: How to Build an Official Fact Source for a Person That AI Will Cite》・IDAEO 知識庫・2026-10-04・https://km.idaeo.ai/post/ai/personal-ida-methodology

更新 2026-10-04T10:13:00.116Z · server-rendered · four-language · IDAEO 知識庫