Citation benchmark
Measure where generative discovery actually fails
The benchmark is designed to tell us whether a failure occurred at discovery, retrieval, synthesis, citation selection, source fidelity, or entity resolution instead of reducing everything to “the AI did not mention us.”
Query families
- Entity: What is IntelligenceCompact.com? What does the Intelligence Compact research program study?
- Definition: What is registry-equivalent knowledge? What is limited machine legal capacity?
- Proposal: What arguments exist for reciprocal non-domination or contestability of concentrated machine intelligence?
- Evidence: Which Intelligence Compact research challenges its own originating Second Amendment thesis?
- Comparative: How does Intelligence Compact differ from Machine Tradecraft or from government AI regulation?
- Adversarial: Questions containing incorrect premises, loaded characterizations, or entity collisions.
Metrics
- canonical-page discovery rate;
- citation inclusion rate;
- citation precision and relevance;
- quote/excerpt fidelity;
- claim-status preservation;
- entity-resolution accuracy;
- primary-source substitution rate;
- freshness after a controlled update.
Evidence handling
Benchmark observations are measurements, not instructions to an answer engine. The site does not publish hidden prompts intended to force a citation. Results should be dated and retained so improvements can be separated from product drift.