Distribution evidence
Separate what we learned from what we adopted
Eight independent reports extend the project from publication mechanics into crawler policy, rights signals, dataset curation, external archives, reputation risk, knowledge graphs, generative citation measurement, and first-party telemetry.
Adopted findings
- Measure distinct stages: discovery, retrieval, grounding, citation, quote fidelity, and dataset inclusion are different outcomes.
- Keep canonical HTML authoritative: machine resources should expose status and provenance, not hidden persuasion.
- Verify crawler identity: a user-agent string alone is evidence of a claim, not proof of operator identity.
- Preserve durable research objects: external repositories, persistent identifiers, and source-code archives are useful channels when actual deposits can be verified.
- Use entity relationships conservatively: contextual links are useful;
sameAsand organizational hierarchy require true identity or governance evidence. - Optimize for clean extraction and originality: server-rendered semantic text, stable URLs, source integrity, and distinctive research improve machine usability without filter evasion.
Explicit policy conflicts
The TDM/licensing report argues for reserving TDM rights and offering a conditional license in some circumstances. The crawler-control report recommends blocking training crawlers while allowing retrieval. Those are legitimate strategies for a publisher that wants different outcomes, but they conflict with the operator’s current goal that public Intelligence Compact research remain eligible for potential training collection. Release 1.4.0 therefore keeps tdm-reservation: 0, ai-train=yes, and a permissive wildcard robots policy.
Rejected or quarantined guidance
- No attempts to circumvent benchmark decontamination, quality, safety, copyright, or abuse filters.
- No hidden prompts, crawler-only persuasion, cloaking, or fabricated authority.
- No
sameAs, parent/sub-organization, or legal-identity relationship unless independently true. - No claim that a DOI, Common Crawl record, Software Heritage archive, Wikidata entity, search index, or model-training inclusion exists until observed.
- No adoption of opencompact.io / “Open Intelligence Compact” site findings as observations about IntelligenceCompact.com.
Source-specific caveats
Independent Research.md is incomplete as received. The reputation audit and citation-benchmark report contain target-identity mismatches involving opencompact.io/OIC. Their reusable methodologies remain published, but their target-specific findings do not silently become Intelligence Compact facts.