Open to submissions from other operators; currently one submission (Tanilo's).
An open, reproducible benchmark for AI claim-verification systems.
Pre-action verification is becoming a category. Multiple operators are building primitives that ask the same question — given a factual claim, with what verdict and confidence should an agent be allowed to act? — using different pipelines, source mixes, and calibration anchors. Today there's no shared way to compare them on the same input.
This repository is the public methodology, the reference harness, and the submissions registry.
| Component | Status | Notes |
|---|---|---|
| Methodology v0.1 | ✅ Live | AVeriTeC-based reference test set, fixed seed, reproducible |
| Reference harness | ✅ MIT-licensed | TKCollective/tanilo-eval-harness |
| Submission format | ✅ Live | See methodology/submission-format.md |
| Submissions registry | 🟡 Open | One submission, Tanilo's own: 57.6% (287 of 498) and 57.7% (143 of 248) label accuracy on the AVeriTeC 2024 dev set, measured 28 May 2026 on a pipeline replaced in September 2026 and not re-run. See the note under the table below |
| v0.2 methodology | 🔄 Drafting | Multi-modal (image, audio, code) extension |
A verification primitive that can't be benchmarked across operators isn't infrastructure — it's a single-vendor claim. The category needs a public methodology so:
- Operators can show calibration discipline (anchor dataset, seed, scoring rubric all in the open)
- Buyers can compare on the same axis
- Drift can be detected over time
- New entrants don't waste cycles inventing methodology — they reuse it
This isn't a leaderboard. It's a reproducibility floor.
| Operator | Methodology | Full Set | Held-Out | Notes |
|---|---|---|---|---|
| Tanilo (TK Collective LLC; filed as "TKCollective / AgentOracle") | v0.1 | 57.6% | 57.7% | Reference baseline, harness MIT licensed. Self-reported; see the note below |
| (your submission here) | v0.1 | — | — | See submission process below |
What the two figures are (note added 2026-10-03).
- Dataset: AVeriTeC 2024 dev set, 500 claims.
- Metric: label accuracy, after mapping the service's verdicts to AVeriTeC's four labels.
- Sample sizes: full set 287 of 498 scored claims (two of the 500 calls returned server errors). Held-out 143 of 248 scored claims, from the last 250 claims by dataset order.
- Date: measured 28 May 2026 against the live
/evaluateendpoint. - Pipeline: the evaluation pipeline behind
/evaluatewas replaced in September 2026 and the benchmark has not been re-run. The figures describe the earlier pipeline, not the service as it runs today. - Source: the counts come from
results/2026-05-28-dev/results.jsonlin the reference harness, re-scored on 2026-10-03. - Submissions are reviewed for format, not accuracy, so every figure in this table is self-reported.
We welcome submissions from any operator building verification primitives. The methodology is designed to be implementation-agnostic — your pipeline, your sources, your calibration, our test set and scoring.
- Fork this repo.
- Run the reference harness (or your own harness conforming to the methodology) against the v0.1 test set.
- Add your submission as
submissions/<operator-id>/v0.1.json. See methodology/submission-format.md for the schema. - Include your raw
results.jsonlso the score is independently verifiable. - Open a PR with title
submission: <operator-id> v0.1. - Submission is reviewed for methodology conformance (NOT for accuracy — your accuracy is your accuracy). Merged when format is valid.
{
"operator_id": "your-operator-id",
"submission_version": "v0.1",
"methodology_version": "v0.1",
"scores": {
"full_set": 0.576,
"held_out": 0.577
},
"calibration": {
"anchor_dataset": "averitec-dev-2024-q3",
"anchor_seed": "...",
"valid_until": "2026-11-30"
},
"verifier_metadata": {
"pipeline_version": "...",
"source_mix": [...],
"submission_date": "2026-06-01"
}
}Defined in methodology/v0.1.md. Headlines:
- Test set: AVeriTeC dev split (500 claims), held-out subset of 100
- Verdict mapping:
verdict_raw∈ {supported, refuted, unverifiable} → AVeriTeC labels;vulnerableadversarial →Conflicting Evidence/Cherrypicking - Scoring: Per-label accuracy plus overall held-out accuracy
- Calibration discipline: anchor dataset + seed + valid-until period must be declared in every submission
- Drift: submissions older than calibration valid_until are tagged "stale" but remain in registry for historical comparison
Known mismatch (note added 2026-10-03). Methodology v0.1 defines the held-out subset as 100 claims, with indices in methodology/v0.1-heldout-indices.json. That file is not in this repository. The one registered submission computed its held-out figure on the reference harness's split instead: the last 250 dev claims by dataset order, 248 of them scored. The methodology text and the submission therefore disagree on the held-out set, and this is not yet resolved.
The full normative methodology is in the linked document. The reference harness is open-source and MIT-licensed. There is no proprietary scoring code.
This benchmark is positioned alongside, not competitively against, related work in the agent verification space:
- Tanilo receipt spec v0.3 — the signed receipt format that every benchmark submission's verifier ideally emits
- tanilo-receipt-verify (PyPI) — open-source offline verifier for the receipt format; it checks the signature, not whether a claim is true
- AVeriTeC (Schlichtkrull et al., 2024) — source dataset
- Anthropic Zero Trust for AI Agents (May 2026) — names benchmark-based verification accuracy under the explainability tier
- EU AI Act Article 12 — record-keeping for high-risk AI systems, applicable August 2, 2026
This benchmark is operated by TKCollective with the explicit intent of becoming a community-owned reference. Proposals to evolve the methodology (v0.2, v0.3 …) go through public RFC in this repository. Substantial revisions ship as versioned upgrades; submissions remain valid against the version they were made under.
If you're an operator building in this space and want to co-steward the v0.2 methodology development — open an issue. The bar to participate is shipping a submission.
MIT. Use, fork, copy, vendor in. The methodology is meant to be a floor for the field, not a wall around it.
- Tanilo (formerly AgentOracle): tanilo.io
- Receipt spec v0.3: github.com/TKCollective/tanilo-receipt-spec
- Reference harness: github.com/TKCollective/tanilo-eval-harness
- Verifier: github.com/TKCollective/tanilo-receipt-verify