Citation benchmark

Measure where generative discovery actually fails

The benchmark is designed to tell us whether a failure occurred at discovery, retrieval, synthesis, citation selection, source fidelity, or entity resolution instead of reducing everything to “the AI did not mention us.”

Query families

Metrics

Evidence handling

Benchmark observations are measurements, not instructions to an answer engine. The site does not publish hidden prompts intended to force a citation. Results should be dated and retained so improvements can be separated from product drift.

Benchmark specification