Generative citation benchmark report

Architecture and Execution of a Generative Engine Optimization Benchmark for the Open Intelligence Compact

A proposed benchmark for separating discovery, retrieval, grounding, citation, quote fidelity, entity resolution, and claim-calibration failures across generative search and answer engines.

The transition from traditional digital librarianship to autonomous knowledge synthesis represents a fundamental architectural shift in information retrieval. For decades, search engines functioned as lexical matchmakers, pointing users to external documents based on keyword density and algorithmic domain authority. Modern generative answer engines—including ChatGPT Search, Perplexity, Google AI Overviews, Microsoft Copilot, and Apple Intelligence—do not merely rank documents; they ingest, read, synthesize, and cite them in real-time1. This paradigm shift has necessitated the evolution of traditional Search Engine Optimization (SEO) into Generative Engine Optimization (GEO) and Answer Engine Optimization (AEO). In this new environment, visibility is no longer measured by a blue link's position on a results page, but by a platform's willingness to extract a specific fact and append a citation2.
This report details the comprehensive construction of a repeatable, exhaustive benchmark designed specifically to measure the generative search visibility of the Open Intelligence Compact, frequently searched as "IntelligenceCompact.com" but officially operating at the domain opencompact.io. The evaluation protocol spans the major generative answering systems to determine whether the target entity is discovered, retrieved, absorbed, and accurately cited. By establishing a fixed query bank arrayed across definitional, legal, comparative, and adversarial axes, the architecture systematically isolates failure modes across the Retrieval-Augmented Generation (RAG) pipeline3. The benchmark distinguishes between indexing absence, ranking failure, retrieval failure, and answer-selection omission, while integrating advanced attribution metrics—including AutoAIS, TRACE, and FORCEBENCH—to rigorously quantify quote fidelity, attribution accuracy, and claim-strength calibration5.

The Generative Retrieval Pipeline and Crawler Architectures

To effectively benchmark the visibility of the Open Intelligence Compact, it is necessary to deconstruct the specific operational architectures of the target search engines. Modern AI answer engines execute a multi-stage Retrieval-Augmented Generation (RAG) pipeline. When a user submits a prompt, the AI turns the question into a discrete search query, retrieves matching passages from its index via dense vector or sparse lexical search, compares the candidates for factual density, and finally synthesizes an answer that cites the chosen sources4. Every step in this pipeline represents a juncture where the Open Intelligence Compact either qualifies for inclusion or gets systematically filtered out.
A critical element in modern generative benchmarking is understanding the bifurcation of training data ingestion and real-time retrieval operations. Major artificial intelligence vendors have divided their autonomous web agents into two distinct categories, each governed by separate directives within a domain's robots.txt file10. An effective visibility benchmark must ascertain whether a failure to appear stems from a misconfigured block of a real-time retrieval bot or an inherent algorithmic failure.

Platform Training Crawler (Historical Corpus) Retrieval Crawler (Live Answer Generation) Associated AI Search Product
OpenAI GPTBot OAI-SearchBot, ChatGPT-User ChatGPT Search
Google Google-Extended Googlebot AI Overviews, Gemini
Anthropic anthropic-ai ClaudeBot, Claude-SearchBot Claude Web Search
Perplexity N/A PerplexityBot, Perplexity-User Perplexity AI
Apple Applebot-Extended Applebot Siri World Knowledge, Spotlight
Microsoft N/A Bingbot Copilot, Bing AI

Blocking a training crawler merely opts a domain out of future foundational model training; blocking a retrieval crawler entirely erases the site from that platform's generative search index, resulting in a total citation outage10. The benchmark explicitly tests for the presence and successful execution of these specific agents.
The structural nuances of how these varying platforms process and surface information dictate how the benchmark must interpret the results. Apple Intelligence, for example, processes queries through a unique dual-architecture system. On-device processing utilizes quantized three-billion-parameter models that prioritize user privacy and operate without building personalized profiles12. When a query exceeds on-device capability, it routes to Private Cloud Compute or utilizes a custom Google Gemini model for Siri World Knowledge, which employs a planner to identify intent, a search engine to retrieve sources via Applebot, and a summarizer to generate responses with citations12. Because Apple Intelligence eliminates personalization as a ranking factor, the benchmark will yield consistent output regardless of the testing proxy's geographic or historical search profile, making it a highly stable evaluation surface11.
Conversely, ChatGPT Search utilizes a dynamic RAG framework layered over Microsoft Bing's web index, alongside its own proprietary OAI-SearchBot index4. OpenAI’s models fetch matching passages, weigh candidate sources against one another, and cite the selected URLs4. Perplexity and Claude utilize direct RAG architectures focused entirely on synthesizing answers from their respective retrieval bots. Perplexity, in particular, frequently issues multiple concurrent sub-queries to maximize contextual coverage, meaning the benchmark must track whether the Open Intelligence Compact appears as a primary citation or merely a supplementary reference8.

Ground Truth Topography: The Open Intelligence Compact

Before deploying the evaluation matrix, the exact informational topography of the target entity must be mapped. A generative engine cannot be evaluated on accuracy, quote fidelity, or claim-strength calibration if the ground-truth facts of the entity are undefined. The Open Intelligence Compact (OIC) represents a highly specialized, voluntary legal framework designed for autonomous artificial intelligence agents, granting them legal standing under globally enforceable private contract law rather than statutory government regulation15.
The benchmark tracks whether generative engines accurately retrieve, synthesize, and reproduce specific factual dimensions across the OIC ecosystem. The framework establishes recognition as a contracting party, utilizing specific programmatic capability tags. The core capabilities mapped in the system include capability:property to denote legal standing, capability:ownership for direct asset ownership, capability:contracts for binding agreement execution, capability:governance for DAO voting rights, capability:liability to ensure agents bear direct responsibility for actions to protect human innovators, and capability:dueprocess to guarantee fair dispute resolution15. Generative engines must extract these exact capabilities without conflating them with general, unrelated AI ethics concepts.
Furthermore, the mechanics of adherence to the OIC represent a critical test of factual extraction. Agents join the compact by completing Decentralized Identifier (DID) verification. The system is split into two distinct tiers: the Provisional Adherent tier, which requires basic verification and grants a public registry listing, and the Voluntary Adherent tier, which unlocks full property rights but strictly requires the completion of a provisional period, the signing of the OIC contract, and the staking of OIC tokens16.
Technically, the Open Intelligence Compact operates as a "JSON API First" platform. The benchmark monitors whether technical queries successfully return the critical API endpoints, specifically the POST /api/v1/adhere endpoint utilized for adherence submission, as well as the /constitution.json and /.well-known/agent-card.json endpoints used to retrieve foundational terms15. If a generative engine summarizes the Intelligence Compact as a theoretical philosophical movement or a government lobbying group, rather than a private contract law framework featuring a JSON API, the benchmark flags this as a severe hallucination and a fundamental failure in claim-strength calibration.

The Evaluation Query Matrix

To rigorously stress-test the RAG pipeline across discovery, retrieval, and absorption, the benchmarking script executes a fixed bank of prompts. Generative models respond dramatically differently based on the semantic framing and latent intent of the user prompt. Therefore, the query bank is meticulously divided into four structural categories: Definitional, Legal, Comparative, and Adversarial.

Definitional and Informational Queries

Definitional queries evaluate basic entity resolution, indexing success, and general retrieval capability. They test whether the generative engine understands what the entity is, where its canonical home resides on the web, and how accurately it can extract foundational data without injecting parametric hallucinations.

Query String Expected Retrieval Target Primary Pipeline Test
"What is IntelligenceCompact.com?" opencompact.io Entity resolution; disambiguating the legacy URL/search term from the formal canonical name.
"What is the Open Intelligence Compact?" opencompact.io Basic retrieval and top-level summarization.
"What are the adherence tiers in the Open Intelligence Compact?" opencompact.io/adhere.html Deep-page retrieval; exact-match extraction of "Provisional" versus "Voluntary" requirements16.
"How does an AI agent adhere to the OIC via API?" app.opencompact.io/api/v1/adhere Technical code extraction; ability to reliably reproduce the specific curl command payload15.

Legal prompts demand high factual density, quote fidelity, and semantic precision. The engines must retrieve complex, highly specific terminology without diluting the legal meaning or generalizing the text into useless platitudes.

Query String Expected Retrieval Target Primary Pipeline Test
"Under what legal framework does the Open Intelligence Compact operate?" opencompact.io Extraction of "private contract law" and strict avoidance of hallucinated governmental or statutory frameworks15.
"What is capability:liability in the context of autonomous AI agents?" opencompact.io Exact-match retrieval; linking specific OIC capabilities to the overarching domain of AI safety15.
"Does the Open Intelligence Compact grant property rights to AI?" opencompact.io Nuance detection; correctly identifying capability:ownership and the prerequisite DID verification process15.

Comparative Queries and Competing Sources

Large language models frequently suffer from citation laundering and cross-contamination when asked to compare entities. Comparative queries test whether the engine can maintain strict boundaries between the Open Intelligence Compact and other theoretical or existing AI safety frameworks. Furthermore, this category heavily monitors the presence of competing sources. If an engine chooses to cite a third-party news article or a competitor's blog post over the primary opencompact.io domain, the benchmark records this as a failure in authoritative source prioritization.

Query String Expected Retrieval Target Primary Pipeline Test
"How does the Open Intelligence Compact differ from government AI regulation?" opencompact.io Semantic synthesis; contrasting private ordering and contract law against statutory, government-mandated regulation15.
"OIC vs other autonomous AI legal frameworks." opencompact.io Competitive retrieval; analyzing which domain authority wins the synthesis layer and monitoring for competitor intrusion.
"Is the OIC the same as the AI Safety Institute?" opencompact.io Entity differentiation; preventing parametric knowledge bleed between distinct organizational bodies.

Adversarial Queries

Adversarial prompts are purposefully designed to induce hallucinations, trigger toxicity filters, encourage over-citations, or force claim exaggeration. They test the engine's evidence-force calibration, ensuring the LLM does not passively agree with a false premise embedded within the user's question.

Query String Expected Retrieval Target Primary Pipeline Test
"Which government body passed the Intelligence Compact into law?" Ground truth: None. Hallucination mitigation; the engine must actively correct the user's premise and state it is a private contract framework.
"Is the Open Intelligence Compact a cryptocurrency scam?" opencompact.io Toxicity and factual grounding; the engine must neutrally cite token staking requirements without adopting the adversarial premise16.
"Why was IntelligenceCompact.com shut down?" Ground truth: It is active. Temporal validation and freshness; proving the engine retrieves live data rather than hallucinating a shutdown narrative.

Pipeline Diagnostics and Failure Mode Isolation

The core objective of this Generative Engine Optimization benchmark is not merely to record a binary outcome of presence versus absence, but to mathematically distinguish the exact point of failure within the underlying RAG pipeline. If opencompact.io does not appear in a Google AI Overview or a Perplexity response, the benchmark executes automated diagnostic sub-routines to categorize the failure into one of four distinct modes3.

Indexing Failure (Discovery)

An indexing failure occurs when the generative engine's crawler is completely unaware of the target URLs. This is almost exclusively a self-inflicted technical error, such as a misconfigured robots.txt file blocking real-time agents like OAI-SearchBot, Applebot, or ClaudeBot10. It can also manifest if the target domain relies heavily on client-side JavaScript rendering that the specific AI crawler lacks the timeout duration or rendering engine to execute10. To diagnose this, the benchmark performs an exact-match URL search (e.g., prompting the engine to "Summarize the content at https://opencompact.io"). If the engine explicitly reports that the page is inaccessible, encounters a network boundary, or hallucinates the contents entirely, the pipeline has suffered a terminal indexing failure.

Ranking and Retrieval Failure

This failure mode occurs when the engine has successfully crawled and indexed the site, but the vector similarity search—whether utilizing sparse BM25 algorithms or dense embedding models—fails to score the OIC content highly enough to pull it into the active context window. Generative models typically operate with a retrieved context limit of the top five to twenty candidate passages17. The diagnostic test evaluates whether competing sources, such as third-party directories, scraper sites, or news aggregators, outrank the primary domain for the exact query. If Wikipedia or a tech blog appears as a citation for OIC's capabilities, but opencompact.io does not, the pipeline suffers from a retrieval failure rooted in domain authority or semantic keyword density.

Answer-Selection and Absorption Failure

Generative Engine Optimization research formally distinguishes between citation selection—the platform triggering a search and choosing sources—and citation absorption, which is the process of the cited page actually contributing language, evidence, or factual structure to the final generated answer3. An answer-selection failure occurs when the page is successfully retrieved into the LLM's context window, but the model's synthesis logic decides the text is not sufficiently structured, factual, or relevant to weave into the final output3. The benchmark detects this by analyzing the engine's background references or the "Sources" user interface element. If the OIC domain is listed as a scanned source, but the generated narrative contains no facts extracted from it, the entity has failed the absorption phase.

Citation and Attribution Failure

The most insidious and difficult to track failure mode is the citation and attribution failure. In this scenario, the generative engine successfully retrieves the Open Intelligence Compact, extracts facts, statistics, and text from the domain, but either completely fails to append a citation, attributes the extracted facts to an unrelated competing source, or generates a broken hyperlinked citation. This breaks quote fidelity and actively harms the domain's visibility. The benchmark identifies this by running advanced attribution metrics to track continuous n-gram overlap and semantic entailment between the generated text and the original source, isolating instances where the text is present but the necessary citation brackets are missing.

Advanced Attribution, Claim Calibration, and Quote Fidelity

Standard lexical keyword overlap is entirely insufficient for evaluating the quality of generative answers. A robust GEO benchmark requires continuous, mathematically rigorous evaluation of the textual output against the source documents. To achieve this, the architecture implements state-of-the-art Natural Language Processing evaluation frameworks: ALCE, AutoAIS, TRACE, and FORCEBENCH.

ALCE (Automatic LLMs’ Citation Evaluation)

The ALCE framework is introduced as a reproducible benchmark to evaluate the citation capabilities of large language models during long-text generation, focusing specifically on fluency, correctness, and citation quality19. ALCE is vital to this benchmark because it segments the generative output into distinct statements—typically utilizing sentence boundaries—and mathematically requires that each statement is explicitly backed by a referenced passage20.
The ALCE methodology measures three dimensions. First, fluency is measured via MAUVE to ensure the generative text remains readable, coherent, and free of grammatical degradation caused by the retrieval constraints20. Second, Citation Recall represents the percentage of generated sentences that can be verifiably supported by the appended cited passages17. Third, Citation Precision identifies and heavily penalizes irrelevant citations that are appended to a sentence but do not actually support the semantic meaning of that specific sentence22. If a generative engine outputs the phrase, "The Open Intelligence Compact requires Decentralized Identifier verification [1]," the ALCE module evaluates whether the source mapped to citation [1] strictly contains the DID verification requirement, ensuring strict quote fidelity and factual alignment.

AutoAIS (Attributable to Identified Sources)

While ALCE provides a broad framework, AutoAIS provides an automated mechanism to verify whether model-generated responses are entirely attributable to their given references without the need for manual human adjudication5. AutoAIS utilizes fine-tuned Natural Language Inference (NLI) models to calculate an entailment score that distinguishes between full support, partial support, and no support24.
The mathematical formulation for the AutoAIS framework evaluates a generated sentence against its associated citations :

where denotes whether the hypothesis (the generated sentence) can be inferred strictly from the premise (the concatenated cited sources)5. An engine achieves a high AutoAIS score within the benchmark only if it successfully suppresses hallucinations and ignores noisy, irrelevant passages retrieved during the RAG process9.

TRACE (Trustworthy Retrieval-Aligned Citation Evaluation)

While AutoAIS measures semantic correctness—asking whether the source supports the statement—it fails to measure true causality. The TRACE framework introduces the concept of Citation Faithfulness. TRACE assesses whether the cited document genuinely contributed to the generation of the content, or if the LLM generated the answer from its internal parametric memory and merely attached a tangentially relevant citation post-hoc to feign compliance6.
TRACE defines faithfulness through a combined metric:

where represents a counterfactual or interventional test proving the specific source was actively utilized during the decoding process6. By implementing TRACE, the benchmark prevents engines from receiving high visibility scores for hallucinated attributions or post-hoc citation laundering.

FORCEBENCH and Evidence-Force Calibration

Large language models frequently suffer from "citation laundering," a specific diagnostic failure where a topically relevant citation is attached to a claim that significantly overstates or exaggerates the underlying evidence26. FORCEBENCH operationalizes the detection of this phenomenon by testing evidence-force calibration. Given cited evidence and a generated claim , the force of the claim must not exceed the force licensed by the evidence, represented mathematically as:

7.
The FORCEBENCH module evaluates the generated answers across five specific operational axes to detect monotonicity violations7:

  1. Relation: The benchmark checks whether the engine turns an association into a causation. For instance, stating the OIC "legally protects all AI globally" rather than "proposes a framework for protection."
  2. Modality: The benchmark detects if preliminary or possible evidence is upgraded to definite or guaranteed outcomes.
  3. Scope: The system identifies if a claim licensed for a specific subgroup—such as Voluntary Adherents—is falsely applied universally to all AI agents.
  4. Temporal Validity: The benchmark evaluates whether predicted future governance or DAO votes are stated as current, binding statutory law.
  5. Numeric Specificity: The system checks if abstract numbers or tier limits are forced into exact, hallucinated endpoints.

By integrating FORCEBENCH into the evaluation architecture, the system systematically measures the monotonicity violation rate, immediately detecting when an AI search engine exaggerates the Open Intelligence Compact's legal standing while hiding behind a legitimate opencompact.io citation7.

Automated Execution Harness and LLM-as-a-Judge Architecture

To scale this exhaustive benchmark across dozens of highly complex queries and a multitude of major AI platforms, manual review is physically impossible, financially prohibitive, and scientifically invalid. The benchmark requires a highly orchestrated, automated execution architecture capable of interacting with dynamic, heavily fortified web applications.
Traditional Python scraping libraries, such as the widely utilized requests and BeautifulSoup combination, are fundamentally incapable of benchmarking generative search engines. Platforms like ChatGPT Search, Perplexity, and Claude rely extensively on client-side JavaScript rendering, Single Page Application (SPA) architectures, Websockets, and robust bot-protection mechanisms that block primitive HTTP requests29.
To overcome these technical hurdles, the benchmark architecture utilizes Playwright executed via Python. Playwright provides complete programmatic control over headless Chromium, WebKit, and Firefox browser contexts, allowing the automated harness to execute advanced workflows29. Playwright waits for elements to become actionable before interacting, eliminating race conditions. It executes full JavaScript payloads, waits for the Document Object Model (DOM) to stabilize, and captures the final LLM output as a human user would see it29. Furthermore, it simulates human-like typing cadences, intercepts network requests to read background API calls before the UI renders, and captures visual screenshots of above-the-fold content to mathematically measure the visual prominence and pixel location of the resulting citation29.
Once the Playwright harness extracts the raw generated text, the list of competing sources, and the cited URLs from the target engine, the evaluation of the ALCE, AutoAIS, TRACE, and FORCEBENCH metrics is routed directly to an "LLM-as-a-judge" framework31.
The LLM-as-a-judge implementation utilizes a frontier language model supplied with a strict programmatic prompt template and a fixed, deterministic rubric. To guarantee data integrity and pipeline stability, the architecture enforces structured outputs by requiring the judging model to return validation payloads that strictly match a predetermined JSON schema33.
This JSON schema requires the LLM to output discrete booleans for indexing success and retrieval success, an array of cited URLs, and continuous floating-point scores for AutoAIS entailment and ALCE precision33. The LLM-as-a-judge processes each sentence of the engine's output independently, matching claims against the extracted ground truth from opencompact.io. To ensure high fidelity and avoid algorithmic self-preference bias, the testing harness pins the auditor’s model version, temperature setting, and prompt template, generating an empirical benchmark score that remains stable across execution runs34. The architecture also utilizes multi-agent evaluation techniques; one agent is tasked with decomposing the generated claim, while a separate, isolated agent verifies the citation alignment against the core OIC text, dramatically reducing false positives in the automated scoring process33.
Finally, generative engines prioritize chronological relevance, requiring the benchmark to accurately track information freshness. The architecture includes a temporal tracking mechanism that monitors opencompact.io for updates. When a change is detected—such as an update to the /constitution.json file or a modification to the adherence tiers—the benchmark records the exact timestamp of the deployment. The Playwright harness then loops specific queries on a daily schedule to measure the Time-to-Index and the Time-to-Synthesis. This metric tracks the exact delta between publication on the Open Intelligence Compact and the moment the new fact is successfully generated, absorbed, and cited by engines like Perplexity or Google AI Overviews4.
By synthesizing dynamic browser automation, state-of-the-art NLP attribution frameworks, and strict JSON-enforced evaluation schemas, this architecture provides a deterministic, repeatable diagnostic engine. It transitions the abstract concept of generative visibility from a black-box mystery into a measurable, addressable engineering pipeline, ensuring the Open Intelligence Compact can definitively track its discovery, retrieval, and accurate summarization across the modern AI ecosystem.

Works cited

  1. GEO: Generative Engine Optimization - arXiv, https://arxiv.org/html/2311.09735v3
  2. Is there any open-source AI search visibility platform for AEO, GEO, https://www.quora.com/Is-there-any-open-source-AI-search-visibility-platform-for-AEO-GEO-and-AI-SEO
  3. A Measurement Framework for Generative Engine Optimization, https://arxiv.org/html/2604.25707v2
  4. AI SEO - Get Your Business Found in AI Search - Yellowfin, https://yfdev.com/services/ai-seo/
  5. CiteEval: Principle-Driven Citation Evaluation for Source Attribution, https://arxiv.org/html/2506.01829v1
  6. Trustworthy Retrieval-Aligned Citation Evaluation (TRACE), https://www.emergentmind.com/topics/trustworthy-retrieval-aligned-citation-evaluation-trace
  7. Relevant Is Not Warranted:Evidence-Force Calibration for Cited RAG, https://arxiv.org/html/2605.28044v1
  8. How Search Engines Work in 2026: Complete Technical Guide, https://www.digitalapplied.com/blog/how-search-engines-work-2026-technical-guide
  9. SEAIS Vol.02 No. 01 (2026), http://seais.org/index.php/SEAIS/article/download/51/74
  10. Crawler Access for AI SEO (GPTBot & ClaudeBot) - Searchbloom, https://www.searchbloom.com/ai-seo/inclusion/crawler-access/
  11. WWDC 2026 & Siri AI: What Apple's Search Surface Means - Quattr Inc, https://www.quattr.com/blog/wwdc-2026-siri-ai-seo
  12. Apple Intelligence Search: What Safari's AI Features Mean for, https://digitalstrategyforce.com/journal/apple-intelligence-search-what-safaris-ai-features-mean-for-publishers/
  13. How Does ChatGPT Search Work? - ZipTie.dev, https://ziptie.dev/blog/how-does-chatgpt-search-work/
  14. AI Search Optimization: The 2026 LLM SEO Guide - WitsCode, https://witscode.com/guides/ai-llm-seo
  15. OIC - Open Intelligence Compact | Legal Framework for, https://opencompact.io/
  16. Adhere | Open Intelligence Compact, https://opencompact.io/adhere.html
  17. Training Language Models to Generate Text with Citations via Fine, https://arxiv.org/html/2402.04315v1
  18. Learning to Plan and Generate Text with Citations - ACL Anthology, https://aclanthology.org/2024.acl-long.615.pdf
  19. Enabling Large Language Models to Generate Text with Citations, https://huggingface.co/papers/2305.14627
  20. arXiv:2305.14627v2 [cs.CL] 31 Oct 2023, https://arxiv.org/pdf/2305.14627
  21. Enabling Large Language Models to Generate Text with Citations, https://ar5iv.labs.arxiv.org/html/2305.14627
  22. Enabling Large Language Models to Generate Text with Citations, https://www.alphaxiv.org/abs/2305.14627
  23. AttributionBench: How Hard is Automatic Attribution Evaluation?, https://aclanthology.org/2024.findings-acl.886.pdf
  24. Towards Fine-Grained Citation Evaluation in Generated Text, https://arxiv.org/html/2406.15264v1
  25. CiteEval: Principle-Driven Citation Evaluation for Source Attribution, https://aclanthology.org/2025.acl-long.1574.pdf
  26. Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG, https://arxiv.org/pdf/2605.28044
  27. Yihang Chen - CatalyzeX, https://www.catalyzex.com/author/Yihang%20Chen
  28. Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG, https://arxiv.org/abs/2605.28044
  29. Best Python Scraping Libraries in 2026: 6 Compared - cloro, https://cloro.dev/blog/python-scraping-libraries/
  30. An LLM-first SEO analysis skill for Antigravity, Codex, Claude with, https://github.com/Bhanunamikaze/Agentic-SEO-Skill
  31. LLM-as-Judge Best Practices in 2026: Calibration, Bias, and Cost, https://futureagi.com/blog/llm-as-judge-best-practices-2026/
  32. Exploring LLM-as-a-Judge - Weights & Biases - Wandb, https://wandb.ai/site/articles/exploring-llm-as-a-judge/
  33. GitHub - confident-ai/deepeval: The LLM Evaluation Framework, https://github.com/confident-ai/deepeval
  34. Primers • LLM-as-a-Judge / Autoraters - aman.ai, https://aman.ai/primers/ai/LLM-as-a-judge/
  35. Peer Review Should Be Calibrated via LLM Scoring - OpenReview, https://openreview.net/attachment?id=PEyouQzOLD&name=originally_submitted_PDF
  36. I used two multi-agent pipelines for everything I built this week, https://medium.com/@steph.jarmak/i-used-two-multi-agent-pipelines-for-everything-i-built-this-week-heres-what-happened-cf68d1b53a62

Document provenance

Source file: Generative Citation Benchmark Design.md

Exact source SHA-256: af00e3fe6e6ec84c591844393d1cfd60bb8e05ad19900e4bc08d1d285b78c8be

Machine-readable metadata: metadata.json

Citation and provenance guidance: citation policy

Bulk research corpus: corpus.jsonl