The Evaluation Archive Becomes the Frontier Claim
An evaluation archive is a versioned trace of selected observations, not the frontier itself. Yanan Long's paper shows why repeated snapshots can support narrowly defined claims that a terminal leaderboard cannot, while its own candidate frontier model fails the paper's audit gates. The governance lesson is to preserve the path, pre-specify the test, and report an unsupported or indeterminate result without turning rank into authority.
The Paper
The paper is Yanan Long's Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations, arXiv:2606.17005 [cs.AI, stat.ME]. The reviewed record is version 1, submitted June 15, 2026. It reconstructs repeated public records from LiveBench and Open LLM Leaderboard v2 as the primary objective evidence, uses LMArena as a separate preference stress test, and treats GAIA and τ-bench as limited agentic pilots.
The target is a familiar failure mode: a public score becomes a social fact because a system sits near the top of a leaderboard. The number travels farther than the reporting rule, excluded systems, benchmark revisions, failed runs, and uncertainty that produced it. Long reframes that repeated public record as an inference-and-adjudication problem rather than treating the last ranking as sufficient evidence.
The paper makes two different contributions, and they should not be merged. Its archive contract and terminal-history counterexample clarify what evidence a longitudinal claim needs. Its proposed selection-aware candidate model is a stress-test object, not a winning estimator: the paper reports that it fails synthetic recovery, both primary-archive prediction checks, preference transfer, and uncertainty calibration. The protocol's ability to reject its own candidate is useful; it is not evidence that every archive-based method is valid.
Four Objects, Four Jobs
The title of this essay is shorthand. An archive does not literally become a frontier, and a frontier estimate is not a general measure of intelligence, safety, or deployment value. Four objects need separate names:
- Snapshot: the records exposed by one source at one time, under one benchmark version and visibility rule.
- Evaluation archive: an ordered set of snapshots plus provenance, extraction logic, corrections, exclusions, duplicate handling, and explicit missingness. A folder of undated CSV files is not enough.
- Frontier path: a convention-dependent upper-tail or best-observed score trajectory. In this paper it depends on a normalized score scale, a ceiling, a candidate-pool size, a top-k reporting rule, and a model family. It is not a directly observed latent truth.
- Frontier claim: a bounded proposition about that path—for example, that a named score under a named source and convention will approach a stated gap within a stated period. An audit gate is the pre-specified rule used to decide whether the available evidence supports that proposition for its stated use.
This vocabulary matters because an archive can be valid while a model fitted to it is poor, and a model can predict a benchmark while remaining irrelevant to a safety or procurement decision. Evidence should inherit no broader scope than the measurement target it actually tested.
Terminal Scores Lose the Path
A terminal leaderboard compresses a repeated history into one public cross-section. It may answer who was listed highest under the current table, but it is weak evidence for a different question: whether a score is saturating, how much headroom remains, or when a specified threshold will be reached.
The paper demonstrates the distinction with a constructed counterexample, not a forecast. It fixes a pool of 1,000 systems, reports the top 10, normalizes the ceiling to 1, and compares two paths that have the same selected terminal likelihood at time 10. Under that particular generalized-Pareto terminal model, the paths imply times of 23.03 and 75.13 to come within 0.05 of the ceiling. The paper explicitly says this is one admissible construction, not a general identified set or an empirical estimate of AI progress.
The exact numbers are therefore not the headline. The terminal observation cannot identify the pre-terminal path under the stated model and reporting convention. A saturation claim needs repeated versioned observations; a current-rank claim may not. This is the same distinction behind the site's critiques of leaderboards answering the wrong question, coding-agent benchmark claims, and benchmarks becoming curricula.
The Archive Contract
Long's archive contract treats metadata as evidence, not bookkeeping. For every source it fixes the public source, snapshot unit, timestamp field, score fields and direction, rank treatment, duplicate policy, missingness summary, and inclusion grade. A timestamp is source-native or marked as derived; it should not silently acquire precision during extraction.
The grades constrain the work each source may do. In the paper, a main source must be an objective archive with at least five validated snapshots, ten distinct canonical systems, three eligible one-step folds, and one eligible two-step fold. A stress-test or secondary source can test transfer or archive applicability but cannot support the primary objective headline without formal regrading. An excluded source remains in the manifest with a reason.
That is why LMArena preference records are not pooled with objective scores, while GAIA and τ-bench remain pilots. LiveCodeBench, HELM Capabilities, and SWE-bench Verified are excluded because the public histories used by this paper did not contain the versioned source tables required for its reconstruction. That is a limitation of this evidence baseline, not a finding that those benchmarks are generally invalid.
A reconstruction also needs chain of custody beyond the paper's minimum fields: retrieval time, immutable source URI or content hash, raw snapshot hash, parser and code revision, transformation log, validation failures, and correction history. Otherwise a future rerun may query a changed live surface and obtain a different past. The complementary internal proposals are an evaluation record schema and a benchmark update receipt: one preserves the run, the other preserves changes to the instrument.
Current Context
The paper's archives are not interchangeable with today's live leaderboards. Hugging Face's official Open LLM Leaderboard team announced the leaderboard's retirement on March 13, 2025, after evaluating more than 13,000 models. Its v2 records now function as historical evidence, which is exactly when immutable snapshots, documented exclusions, and retained evaluation details become more important than a live rank.
The official LiveBench repository takes a different approach: it exposes dated release options and a changelog, and it tells evaluators to match the question source and release option used for a run. Its documentation also warns that not every question in a stated release is necessarily public. That makes visibility part of the observation regime, not a footnote.
Official measurement guidance points in the same direction without endorsing Long's model. NIST's January 2026 initial draft on automated benchmark evaluations organizes practice around defining the measurement target, implementing the evaluation, and analyzing and reporting results. NIST AI 800-3 separately distinguishes performance on a fixed benchmark from performance generalized to similar possible items, and warns that uncertainty estimates inherit modeling assumptions. An archive strengthens traceability; it does not erase that measurement boundary.
The Audit Gates
The paper separates archive validity from model endorsement. Its candidate, S0, is a dynamic selection-aware model fitted to repeated snapshots. The paper reports that it passes none of three truth-known synthetic recovery regimes; a separate slow-frontier negative control behaves correctly. It passes neither primary objective-archive predictive check, trails the native comparison in the LMArena stress test, and fails the calibration audit. Future-snapshot prediction is evidence about future observations, not direct validation of a latent frontier.
This negative result is valuable because the protocol does not launder an attractive model into authority. But an audit gate is not a truth oracle. Failure can mean the proposition is false, the model is misspecified, the archive is insufficient, or the chosen threshold was not met. The defensible outcomes are therefore supported for the stated use, not supported under this gate, indeterminate because evidence is missing, and out of scope—not a universal binary of true and false.
The manuscript says its gate rules were fixed before reported summaries were inspected. Operational governance should go one step further: publish a timestamped protocol or registry record, including the claim text, comparator, admissible rows, metric direction, threshold, stopping rule, exception policy, and treatment of ties, before results are opened. Any amendment should create a new version rather than quietly rewriting the gate.
Agentic Evaluation Needs More Metadata
An agentic record identifies more than a base model. It identifies a configured system: model or API revision, prompts, scaffold, tools, memory policy, environment and simulator, task and retrieval state, budgets, retry and stopping rules, judge, randomness, failure handling, and human-intervention policy. Change one of those and the evaluated object may have changed.
The paper uses GAIA and τ-bench to show that aggregate agentic histories can be staged under an archive contract, but it does not claim full execution-trace observability. It characterizes GAIA's public scaffold and tool metadata as weak and keeps τ-bench outside the main objective claim. A score can therefore show that one aggregate configuration succeeded under one harness without revealing which component produced the difference or how it would behave in deployment.
Full traces also create safety and privacy duties. They may contain personal data, credentials, proprietary prompts, hidden test material, or exploitable environment details. A serious archive needs tiered disclosure: public manifests and hashes, access-controlled raw traces where justified, a redaction log, retention limits, reviewer access records, and a way to verify restricted evidence without publishing secrets. See AI Agent Observability and AI Audit Trails.
Governance Standard
A consequential evaluation claim should travel as three linked receipts:
- Claim receipt: exact proposition, intended decision, system boundary, population and task domain, score direction, threshold, time horizon, and explicit non-claims.
- Archive receipt: source and evaluator relationship, retrieval date, immutable snapshot identifier or hash, benchmark and harness versions, task slice, system configuration, visibility and rank rules, duplicate and missingness rules, inclusion grade, transformations, exclusions, and known corrections.
- Audit receipt: precommitted gate and comparator, admissible evidence, uncertainty method, calibration and sensitivity checks, outcome with reasons, reviewer identity and conflicts, exceptions, approval authority, and the event that triggers re-evaluation.
The archive custodian, evaluator, and decision owner should be named separately even when one organization holds all three roles. Corrections should append a new version; an appeal should preserve the original outcome; and a benchmark or system change should trigger a fresh decision rather than retroactively editing the old one. A leaderboard trend may inform research planning, but it should not by itself approve deployment, procurement, or a safety claim.
This standard extends the site's Claim Hygiene Protocol, public evaluation ledger, and capability-frontier evaluation gap. The concise rule is: preserve the observation path, bind it to one claim, and keep the decision authority visible.
Limits
The paper is candid about several limits. Real archives validate future observations, not latent frontier truth. Candidate-pool reconstruction is assumption-driven. The frontier family uses a simple monotone gap model and does not establish robustness to power-law, stretched-exponential, or non-monotone alternatives. The normalized score, ceiling, pool size, visibility rule, and target gap are reporting conventions. Its decision layer uses stylized losses and synthetic posterior draws rather than utilities elicited from an accountable decision owner.
There is also a reproducibility boundary in the reviewed release. The manuscript says that appendix release pointers make inclusion decisions reproducible, and its table names machine-readable companions, manifests, and audit artifacts. In the v1 HTML and v1 source bundle reviewed for this page, we found the manuscript tables and figures but no linked repository, persistent artifact identifier, released manifests, or machine-readable audit companions. We could therefore check the paper's definitions, source inventory, formulas, reported gate outcomes, and internal consistency; we could not independently rebuild its archive or rerun its gates from the released package.
That does not erase the paper's conceptual contribution, but it limits the assurance level of its empirical readout. A reproducibility map is an index of promised evidence, not the evidence itself. Until the named artifacts are publicly retrievable or made available to qualified reviewers, the results should be described as manuscript-reported rather than independently reproduced.
Sources
- Yanan Long, Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations, arXiv:2606.17005v1 [cs.AI, stat.ME], submitted June 15, 2026. The HTML, PDF, and source bundle were checked separately.
- NIST, Practices for Automated Benchmark Evaluations of Language Models, NIST AI 800-2 Initial Public Draft, January 2026.
- Andrew Keller et al., NIST, Expanding the AI Evaluation Toolbox with Statistical Models, NIST AI 800-3, February 2026.
- LiveBench, official repository and release documentation, checked for release selection, question-source matching, public-data availability, and changelog conventions.
- Hugging Face Open LLM Leaderboard, official retirement announcement, March 13, 2025.
Related internal references: AI Evaluations, Chatbot Arena and LMArena, GAIA Benchmark, Tau-bench, and Benchmark Contamination.