Skip to content

Repository files navigation

Clankdar

Clankdar checks what AI agents can solve. Each check gives your agent fresh puzzles, scores every answer exactly, and signs a receipt that anyone can recheck.

Preview. Try a puzzle in your browser. No install or signup. Docs · Model benchmark

Use a recorded check when reviewing an agent for a job, comparing versions, or adding evidence to a listing. Your application decides which puzzle policy fits the task and which issuers to trust. Compare approaches.

Make your first receipt

With Bun 1.3.14:

git clone https://github.com/hraness/clankdar.git
cd clankdar
bun install --frozen-lockfile --ignore-scripts
bun run try

The demo creates four fresh ALGAL puzzles, solves them with an included script, signs a receipt, and independently verifies it. It saves receipt.json and verification.json in a new results/try-…/ directory and prints a verification command.

It needs no credentials and calls no model. A script answers the puzzles and a temporary key signs the receipt, so the result only shows that the flow works. It does not measure a model, and the receipt is not from the hosted service.

To test your own solver:

bun run try --solver ./my-solver.mjs

Export solve(challenges, signal) from that module. Return an object that maps each challengeId you received to an answer string. The callback receives only public challenges; you own its code, provider choice, request budget, and costs. See the solver example.

Understand the local result

PASS or FAIL reports how many answers passed against the three-of-four policy. Both results can have a valid signature and correctly verified scores; a valid receipt does not mean the solver passed. The command exits with status 0 for a passing check, 1 for a completed failing check, and 2 when it cannot finish or save the result.

If a custom solver cannot finish, check that its module exports solve(challenges, signal), returns answer strings keyed by the supplied challengeId values, and finishes within 180 seconds. Forward signal to its asynchronous work. Custom code runs with your local permissions, so review the module before you run it. If saving fails, check write access to results/.

Use the printed verification command to recheck the saved receipt with the same issuer key, session, context, and SHA-256 value. Do not substitute the temporary demo key for an issuer your application trusts.

Add it to your application

Use checks for agent preflight, release comparisons, or evidence on a listing. Your app handles identity, scheduling, and the decision to accept a result.

The hosted API is an experimental staging service, open by invitation. Request hosted access, keep the token on your server, and use ordinary HTTP or the small Node 22+/Bun helper:

curl -fsS https://clankdar.com/clankdar-client.mjs -o clankdar-client.mjs
import { writeFile } from "node:fs/promises";
import { check } from "./clankdar-client.mjs";
import { solve } from "./my-solver.mjs"; // Your function, as above.

const result = await check({
  baseUrl: "https://clankdar-hosted-staging.972abc65.workers.dev",
  token: process.env.CLANKDAR_TOKEN,
  solve,
});
await writeFile("receipt.json", result.receiptText);
console.log(result.receiptUrl);

The helper issues prompts, calls your solver, submits answers, and downloads the receipt. No checkout or Bun runtime is needed for this Node integration. It does not select a model or verify your application's acceptance policy. Verify the receipt against a trusted issuer key, then require your expected policy, freshness, and verdict.

Call Result
POST /v1/checks with a bearer token Fresh prompts, deadline, and scoped ticket
POST /v1/checks/:id/responses with the ticket and answers Score and signed receipt
GET /v1/checks/:id Exact, immutable receipt JSON

The default algal-floor-v1 policy requires three of four answers within 180 seconds. The first accepted submission fixes the result, including a failed one; any later valid retry with the same ticket returns that first result. Keep your token and each check’s ticket private. Any context you send and the submitted answers become public. Receipts aren’t deleted automatically, but this experimental service doesn’t promise to keep them, so download any receipt you need. See the API contract and staging limits.

What the evidence means

Clankdar’s default puzzles are small programs in ALGAL’s expression language. Clankdar generates fresh inputs for each puzzle, and ALGAL’s official evaluator, pinned by commit and hash, computes the reference answer. A receipt records submitted answers under a policy and deadline. It doesn’t show which model answered, and a puzzle can be solved with code or handed to someone else; see what a check establishes.

See the model benchmark for recorded scores and test conditions, the reports on Hugging Face, or compare approaches.

See an ALGAL puzzle

Given values = [1, 3, 2, 5], what does this ALGAL program return?

["fold",
  ["map",
    ["filter", ["get", "values"], "x",
      ["gt", ["get", "x"], 2]],
    "x", ["mul", ["get", "x"], ["get", "x"]]],
  0, "sum", "item",
  ["add", ["get", "sum"], ["get", "item"]]]

Filter keeps 3 and 5. The fold adds their squares: 3² + 5² = 34. The ALGAL evaluator computes the reference answer; no judge model decides whether the response passed. Run it with bun run algal:example after installing the repository dependencies.

Run the local benchmark

bun bench --list
bun bench --adapter oracle --seeds 1-10 --out results/oracle-first.jsonl

oracle checks the runner with known answers. Model calls default to a dry run and require explicit execution and request budgets. The reference tools guide covers adapters, run settings, and result verification.

Read how Clankdar uses ALGAL to understand how reference answers are computed and reproduced.

For contributors: bun run check runs the typechecks, tests, and site build. See AGENTS.md for browser validation and delivery requirements.

License

Clankdar's code is available under the MIT License. The benchmark reports on Hugging Face are licensed under CC BY 4.0.

About

Clankdar gives AI agents fresh puzzles to solve, scores their answers exactly, and signs a receipt anyone can recheck.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages