Evidence-labeled claim
Crawler permission is not proof of external distribution
This page is a compact epistemic record: what is being claimed, what kind of claim it is, how strong the current evidence is, which research supports or challenges it, and what would justify changing the assessment.
IC-CLAIM-009 operational evidence rule verified local rule adopted policy
Scope and boundary
This rule governs Intelligence Compact’s own public claims about distribution evidence.
Why it matters
Without this distinction, a publisher can accidentally convert “we allowed it” into false claims that a search engine or model actually used the material.
Strongest objection
None to the logical distinction itself; the practical question is what evidence threshold should be sufficient for each narrower external claim.
What would change this assessment
Only an explicit project decision changing the evidence standard; external observations change channel states, not this logical rule.
Project decision relationship
This claim records an adopted project rule under DEC-007. That decision status is separate from the truth status of external factual propositions.
Supporting research
- AI Crawler, Training, and Retrieval Control Matrix
- First-Party Crawl Telemetry Architecture: A Zero-Third-Party System for IntelligenceCompact.com
- Architecture and Execution of a Generative Engine Optimization Benchmark for the Open Intelligence Compact
- The Architecture of Large Language Model Pretraining Corpora: From Web Crawl to Curated Dataset
Source-quality and provenance summary
Supporting dossiers currently connect this claim to 192 distinct cited web sources, including 0 official public-authority and 28 scholarly/preprint sources. Challenging or limiting dossiers connect to 0 distinct sources. Source mix is provenance context, not a vote or truth score.
How source classes are defined · Machine-readable source map
Representative sources cited by supporting dossiers
- developers.openai.com first party technical or policy
- commoncrawl.org first party technical or policy
- commoncrawl.org first party technical or policy
- commoncrawl.org first party technical or policy
- developers.google.com first party technical or policy
- github.com first party technical or policy
- github.com first party technical or policy
- github.com first party technical or policy
Reviewed document-level source notes
First-party OpenAI, Anthropic, Google, Common Crawl, and Perplexity documents distinguish crawler permission and declared bot roles from downstream outcomes. Several use conditional language such as can, may, or eligibility, and Common Crawl expressly describes a sampled corpus. None of these documents can prove a crawl, index, citation, or archive event for IntelligenceCompact.com. DEC-007 therefore remains appropriate.
Publishers and Developers — FAQ
first party crawler and publisher policy provider policy documentation OpenAI
What it establishes: OpenAI distinguishes OAI-SearchBot access for ChatGPT search discovery from GPTBot controls for potential training. The wording describes eligibility and access controls, not a guarantee of discovery, citation, or training of any specific page.
Important limitation: Provider documentation can change. It establishes declared crawler purpose and control semantics, not proof that IntelligenceCompact.com was crawled, indexed, cited, or included in training.
Does Anthropic crawl data from the web, and how can site owners block the crawler?
first party crawler policy provider policy documentation Anthropic / Claude Help Center
What it establishes: Anthropic documents separate roles for ClaudeBot (potential model-training collection), Claude-SearchBot (search quality/indexing), and Claude-User (user-directed retrieval).
Important limitation: The documentation describes intended crawler roles and robots controls. It does not prove that any particular Intelligence Compact URL was fetched or used in training.
Google's common crawlers — Google-Extended
first party crawler policy provider policy documentation Google
What it establishes: Google-Extended controls whether Google-crawled content may be used for future Gemini training and certain grounding uses, while Google says the token does not affect Google Search inclusion or ranking.
Important limitation: The control describes eligibility and product use semantics, not evidence that a particular document was included in a specific model-training run or surfaced in Gemini.
Common Crawl FAQ
first party crawl dataset documentation dataset operator documentation Common Crawl Foundation
What it establishes: Common Crawl says its dataset is a sample of the web rather than a complete archive; it honors robots controls and can use announced sitemaps if its crawler visits a site.
Important limitation: Being crawlable or listed in a sitemap does not prove that a page was selected into any particular Common Crawl snapshot, downstream dataset, or model-training corpus.
How does Perplexity follow robots.txt?
first party crawler policy provider policy documentation Perplexity
What it establishes: Perplexity says PerplexityBot indexes pages for search and that allowing PerplexityBot does not mean the content is used for foundation-model pretraining.
Important limitation: The policy distinguishes search indexing from model pretraining; it does not prove that a particular Intelligence Compact page was indexed or cited.
See all reviewed source notes →
Challenging or limiting research
No separate challenging dossier is currently assigned; the claim remains subject to correction and new evidence.