Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Tanilo Benchmark (formerly AgentOracle Benchmark)

Open to submissions from other operators; currently one submission (Tanilo's).

License: MIT Methodology v0.1 Open Submissions

An open, reproducible benchmark for AI claim-verification systems.

Pre-action verification is becoming a category. Multiple operators are building primitives that ask the same question — given a factual claim, with what verdict and confidence should an agent be allowed to act? — using different pipelines, source mixes, and calibration anchors. Today there's no shared way to compare them on the same input.

This repository is the public methodology, the reference harness, and the submissions registry.

Status

Component Status Notes
Methodology v0.1 ✅ Live AVeriTeC-based reference test set, fixed seed, reproducible
Reference harness ✅ MIT-licensed TKCollective/tanilo-eval-harness
Submission format ✅ Live See methodology/submission-format.md
Submissions registry 🟡 Open One submission, Tanilo's own: 57.6% (287 of 498) and 57.7% (143 of 248) label accuracy on the AVeriTeC 2024 dev set, measured 28 May 2026 on a pipeline replaced in September 2026 and not re-run. See the note under the table below
v0.2 methodology 🔄 Drafting Multi-modal (image, audio, code) extension

Why this exists

A verification primitive that can't be benchmarked across operators isn't infrastructure — it's a single-vendor claim. The category needs a public methodology so:

  • Operators can show calibration discipline (anchor dataset, seed, scoring rubric all in the open)
  • Buyers can compare on the same axis
  • Drift can be detected over time
  • New entrants don't waste cycles inventing methodology — they reuse it

This isn't a leaderboard. It's a reproducibility floor.

Current registered submissions

Operator Methodology Full Set Held-Out Notes
Tanilo (TK Collective LLC; filed as "TKCollective / AgentOracle") v0.1 57.6% 57.7% Reference baseline, harness MIT licensed. Self-reported; see the note below
(your submission here) v0.1 — — See submission process below

What the two figures are (note added 2026-10-03).

  • Dataset: AVeriTeC 2024 dev set, 500 claims.
  • Metric: label accuracy, after mapping the service's verdicts to AVeriTeC's four labels.
  • Sample sizes: full set 287 of 498 scored claims (two of the 500 calls returned server errors). Held-out 143 of 248 scored claims, from the last 250 claims by dataset order.
  • Date: measured 28 May 2026 against the live /evaluate endpoint.
  • Pipeline: the evaluation pipeline behind /evaluate was replaced in September 2026 and the benchmark has not been re-run. The figures describe the earlier pipeline, not the service as it runs today.
  • Source: the counts come from results/2026-05-28-dev/results.jsonl in the reference harness, re-scored on 2026-10-03.
  • Submissions are reviewed for format, not accuracy, so every figure in this table is self-reported.

Submitting results

We welcome submissions from any operator building verification primitives. The methodology is designed to be implementation-agnostic — your pipeline, your sources, your calibration, our test set and scoring.

Submission process

  1. Fork this repo.
  2. Run the reference harness (or your own harness conforming to the methodology) against the v0.1 test set.
  3. Add your submission as submissions/<operator-id>/v0.1.json. See methodology/submission-format.md for the schema.
  4. Include your raw results.jsonl so the score is independently verifiable.
  5. Open a PR with title submission: <operator-id> v0.1.
  6. Submission is reviewed for methodology conformance (NOT for accuracy — your accuracy is your accuracy). Merged when format is valid.

Submission format (summary)

{
  "operator_id": "your-operator-id",
  "submission_version": "v0.1",
  "methodology_version": "v0.1",
  "scores": {
    "full_set": 0.576,
    "held_out": 0.577
  },
  "calibration": {
    "anchor_dataset": "averitec-dev-2024-q3",
    "anchor_seed": "...",
    "valid_until": "2026-11-30"
  },
  "verifier_metadata": {
    "pipeline_version": "...",
    "source_mix": [...],
    "submission_date": "2026-06-01"
  }
}

Methodology v0.1

Defined in methodology/v0.1.md. Headlines:

  • Test set: AVeriTeC dev split (500 claims), held-out subset of 100
  • Verdict mapping: verdict_raw ∈ {supported, refuted, unverifiable} → AVeriTeC labels; vulnerable adversarial → Conflicting Evidence/Cherrypicking
  • Scoring: Per-label accuracy plus overall held-out accuracy
  • Calibration discipline: anchor dataset + seed + valid-until period must be declared in every submission
  • Drift: submissions older than calibration valid_until are tagged "stale" but remain in registry for historical comparison

Known mismatch (note added 2026-10-03). Methodology v0.1 defines the held-out subset as 100 claims, with indices in methodology/v0.1-heldout-indices.json. That file is not in this repository. The one registered submission computed its held-out figure on the reference harness's split instead: the last 250 dev claims by dataset order, 248 of them scored. The methodology text and the submission therefore disagree on the held-out set, and this is not yet resolved.

The full normative methodology is in the linked document. The reference harness is open-source and MIT-licensed. There is no proprietary scoring code.

Related work

This benchmark is positioned alongside, not competitively against, related work in the agent verification space:

  • Tanilo receipt spec v0.3 — the signed receipt format that every benchmark submission's verifier ideally emits
  • tanilo-receipt-verify (PyPI) — open-source offline verifier for the receipt format; it checks the signature, not whether a claim is true
  • AVeriTeC (Schlichtkrull et al., 2024) — source dataset
  • Anthropic Zero Trust for AI Agents (May 2026) — names benchmark-based verification accuracy under the explainability tier
  • EU AI Act Article 12 — record-keeping for high-risk AI systems, applicable August 2, 2026

Governance

This benchmark is operated by TKCollective with the explicit intent of becoming a community-owned reference. Proposals to evolve the methodology (v0.2, v0.3 …) go through public RFC in this repository. Substantial revisions ship as versioned upgrades; submissions remain valid against the version they were made under.

If you're an operator building in this space and want to co-steward the v0.2 methodology development — open an issue. The bar to participate is shipping a submission.

License

MIT. Use, fork, copy, vendor in. The methodology is meant to be a floor for the field, not a wall around it.

References

About

Benchmark methodology and submissions registry for AI claim-verification systems. Open to submissions from other operators; currently one submission (Tanilo's).

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors