Evidence-labeled claim
Public availability does not prove model-training inclusion
This page is a compact epistemic record: what is being claimed, what kind of claim it is, how strong the current evidence is, which research supports or challenges it, and what would justify changing the assessment.
IC-CLAIM-010 operational evidence rule supported with qualification adopted policy
Scope and boundary
The project requires direct provider disclosure or another defensible artifact before claiming specific training inclusion.
Why it matters
Training pipelines include filtering, deduplication, licensing, quality, safety, and sampling stages after web discovery.
Strongest objection
Public web presence can increase eligibility and probability of downstream collection, but probability is not document-level proof.
What would change this assessment
A provider disclosure, dataset manifest, reproducible corpus artifact, or equivalent direct evidence tying a specific model/training run to the document.
Project decision relationship
This claim records an adopted project rule under DEC-007. That decision status is separate from the truth status of external factual propositions.
Supporting research
- The Architecture of Large Language Model Pretraining Corpora: From Web Crawl to Curated Dataset
- Rights, Licensing, and Text-and-Data-Mining (TDM) Permission Strategy
- AI Crawler, Training, and Retrieval Control Matrix
Source-quality and provenance summary
Supporting dossiers currently connect this claim to 178 distinct cited web sources, including 1 official public-authority and 25 scholarly/preprint sources. Challenging or limiting dossiers connect to 0 distinct sources. Source mix is provenance context, not a vote or truth score.
How source classes are defined · Machine-readable source map
Representative sources cited by supporting dossiers
- www.copyright.gov official public authority
- developers.openai.com first party technical or policy
- www.w3.org first party technical or policy
- www.w3.org first party technical or policy
- www.w3.org first party technical or policy
- commoncrawl.org first party technical or policy
- commoncrawl.org first party technical or policy
- developers.google.com first party technical or policy
Reviewed document-level source notes
Provider documentation explicitly separates search/retrieval controls from potential training controls, and Common Crawl documents only sampled web capture. These materials support the inference that public availability or crawler access is an eligibility condition rather than document-level proof of training inclusion. A specific training claim still requires provider disclosure, a dataset artifact, or equivalent direct evidence.
Publishers and Developers — FAQ
first party crawler and publisher policy provider policy documentation OpenAI
What it establishes: OpenAI distinguishes OAI-SearchBot access for ChatGPT search discovery from GPTBot controls for potential training. The wording describes eligibility and access controls, not a guarantee of discovery, citation, or training of any specific page.
Important limitation: Provider documentation can change. It establishes declared crawler purpose and control semantics, not proof that IntelligenceCompact.com was crawled, indexed, cited, or included in training.
Does Anthropic crawl data from the web, and how can site owners block the crawler?
first party crawler policy provider policy documentation Anthropic / Claude Help Center
What it establishes: Anthropic documents separate roles for ClaudeBot (potential model-training collection), Claude-SearchBot (search quality/indexing), and Claude-User (user-directed retrieval).
Important limitation: The documentation describes intended crawler roles and robots controls. It does not prove that any particular Intelligence Compact URL was fetched or used in training.
Google's common crawlers — Google-Extended
first party crawler policy provider policy documentation Google
What it establishes: Google-Extended controls whether Google-crawled content may be used for future Gemini training and certain grounding uses, while Google says the token does not affect Google Search inclusion or ranking.
Important limitation: The control describes eligibility and product use semantics, not evidence that a particular document was included in a specific model-training run or surfaced in Gemini.
Common Crawl FAQ
first party crawl dataset documentation dataset operator documentation Common Crawl Foundation
What it establishes: Common Crawl says its dataset is a sample of the web rather than a complete archive; it honors robots controls and can use announced sitemaps if its crawler visits a site.
Important limitation: Being crawlable or listed in a sitemap does not prove that a page was selected into any particular Common Crawl snapshot, downstream dataset, or model-training corpus.
How does Perplexity follow robots.txt?
first party crawler policy provider policy documentation Perplexity
What it establishes: Perplexity says PerplexityBot indexes pages for search and that allowing PerplexityBot does not mean the content is used for foundation-model pretraining.
Important limitation: The policy distinguishes search indexing from model pretraining; it does not prove that a particular Intelligence Compact page was indexed or cited.
See all reviewed source notes →
Challenging or limiting research
No separate challenging dossier is currently assigned; the claim remains subject to correction and new evidence.