Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
Abstract
The Discovery Certification Protocol validates AI research agent outcomes through executable recovery tests, controlled audits, and deterministic verification with finite-sample recovery bounds.
AI research agents combine prior knowledge, public sources, and experimental feedback to produce useful results. The Discovery Certification Protocol (DCP) turns claims about these results into executable recovery and feedback tests. Gate 1 validates useful improvement on sealed evaluation. Gate 2 gives matched agents the registered starting information and observed Web content while withholding the target research history. Every valid method reaching the numerical target supplies a recovery witness and triggers the Core veto. DCP Core requires adequate controls, zero observed recoveries, and a finite-sample bound on recovery in one fresh registered episode. Optional Gate 3 measures the average effect of truthful feedback relative to a specified neutral policy from a shared checkpoint. DCP Evidence adds this effect after independent null calibration and a registered effect margin. Two controlled audits exercise the complete protocol in SQLite optimization and virtual catalyst control under different models. Each produced zero recoveries in 96 episodes, with an upper bound of 0.0468. Each paired study yielded 30 truthful recoveries and zero neutral recoveries, with passing 60-pair null studies. Additional cases exercise Core, recovered, and audit-incomplete decisions. A deterministic, LLM-free verifier reproduces the decisions from frozen evidence. DCP provides a common evidence language for useful outcomes, alternative routes, and feedback effects across AI research.
Community
The Discovery Certification Protocol turns AI research claims into testable evidence. Validate the gain. Challenge its recovery. Measure the contribution of feedback.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- ATOBench: Tracing How Autonomous Penetration-Testing Agents Verify Vulnerabilities When Target Evidence Lies (2026)
- When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs (2026)
- Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI (2026)
- What Does an Evaluation License? A Commit-Bound Census of Claim Replay in Inspect Evals (2026)
- Auditing and Decomposing Feedback-Driven Evolution in LLM Test Generation under the Oracle Problem (2026)
- ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations (2026)
- Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.09219 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper