Clankdar checks what AI agents can solve. Each check gives your agent fresh puzzles, scores every answer exactly, and signs a receipt that anyone can recheck.
Preview. Try a puzzle in your browser. No install or signup. Docs · Model benchmark
Use a recorded check when reviewing an agent for a job, comparing versions, or adding evidence to a listing. Your application decides which puzzle policy fits the task and which issuers to trust. Compare approaches.
With Bun 1.3.14:
git clone https://github.com/hraness/clankdar.git
cd clankdar
bun install --frozen-lockfile --ignore-scripts
bun run tryThe demo creates four fresh ALGAL puzzles, solves them with an included script, signs a receipt, and independently verifies it. It saves receipt.json and verification.json in a new results/try-…/ directory and prints a verification command.
It needs no credentials and calls no model. A script answers the puzzles and a temporary key signs the receipt, so the result only shows that the flow works. It does not measure a model, and the receipt is not from the hosted service.
To test your own solver:
bun run try --solver ./my-solver.mjsExport solve(challenges, signal) from that module. Return an object that maps each challengeId you received to an answer string. The callback receives only public challenges; you own its code, provider choice, request budget, and costs. See the solver example.
PASS or FAIL reports how many answers passed against the three-of-four
policy. Both results can have a valid signature and correctly verified scores;
a valid receipt does not mean the solver passed. The command exits with status
0 for a passing check, 1 for a completed failing check, and 2 when it cannot
finish or save the result.
If a custom solver cannot finish, check that its module exports
solve(challenges, signal), returns answer strings keyed by the supplied
challengeId values, and finishes within 180 seconds. Forward signal to its
asynchronous work. Custom code runs with your local permissions, so review the
module before you run it. If saving fails, check write access to results/.
Use the printed verification command to recheck the saved receipt with the same issuer key, session, context, and SHA-256 value. Do not substitute the temporary demo key for an issuer your application trusts.
Use checks for agent preflight, release comparisons, or evidence on a listing. Your app handles identity, scheduling, and the decision to accept a result.
The hosted API is an experimental staging service, open by invitation. Request hosted access, keep the token on your server, and use ordinary HTTP or the small Node 22+/Bun helper:
curl -fsS https://clankdar.com/clankdar-client.mjs -o clankdar-client.mjsimport { writeFile } from "node:fs/promises";
import { check } from "./clankdar-client.mjs";
import { solve } from "./my-solver.mjs"; // Your function, as above.
const result = await check({
baseUrl: "https://clankdar-hosted-staging.972abc65.workers.dev",
token: process.env.CLANKDAR_TOKEN,
solve,
});
await writeFile("receipt.json", result.receiptText);
console.log(result.receiptUrl);The helper issues prompts, calls your solver, submits answers, and downloads the receipt. No checkout or Bun runtime is needed for this Node integration. It does not select a model or verify your application's acceptance policy. Verify the receipt against a trusted issuer key, then require your expected policy, freshness, and verdict.
| Call | Result |
|---|---|
POST /v1/checks with a bearer token |
Fresh prompts, deadline, and scoped ticket |
POST /v1/checks/:id/responses with the ticket and answers |
Score and signed receipt |
GET /v1/checks/:id |
Exact, immutable receipt JSON |
The default algal-floor-v1 policy requires three of four answers within 180 seconds. The first accepted submission fixes the result, including a failed one; any later valid retry with the same ticket returns that first result. Keep your token and each check’s ticket private. Any context you send and the submitted answers become public. Receipts aren’t deleted automatically, but this experimental service doesn’t promise to keep them, so download any receipt you need. See the API contract and staging limits.
Clankdar’s default puzzles are small programs in ALGAL’s expression language. Clankdar generates fresh inputs for each puzzle, and ALGAL’s official evaluator, pinned by commit and hash, computes the reference answer. A receipt records submitted answers under a policy and deadline. It doesn’t show which model answered, and a puzzle can be solved with code or handed to someone else; see what a check establishes.
See the model benchmark for recorded scores and test conditions, the reports on Hugging Face, or compare approaches.
See an ALGAL puzzle
Given values = [1, 3, 2, 5], what does this ALGAL program return?
["fold",
["map",
["filter", ["get", "values"], "x",
["gt", ["get", "x"], 2]],
"x", ["mul", ["get", "x"], ["get", "x"]]],
0, "sum", "item",
["add", ["get", "sum"], ["get", "item"]]]Filter keeps 3 and 5. The fold adds their squares: 3² + 5² = 34.
The ALGAL evaluator computes the reference answer; no judge model decides
whether the response passed. Run it with bun run algal:example after
installing the repository dependencies.
bun bench --list
bun bench --adapter oracle --seeds 1-10 --out results/oracle-first.jsonloracle checks the runner with known answers. Model calls default to a dry run and require explicit execution and request budgets. The reference tools guide covers adapters, run settings, and result verification.
Read how Clankdar uses ALGAL to understand how reference answers are computed and reproduced.
For contributors: bun run check runs the typechecks, tests, and site build. See AGENTS.md for browser validation and delivery requirements.
Clankdar's code is available under the MIT License. The benchmark reports on Hugging Face are licensed under CC BY 4.0.