dbaai_bench runs one model against one task on servers you already have. This
runs many models against many tasks on servers it creates itself, marks the
results with its own commands, and destroys the servers afterwards.
uv sync # installs ../dbaai_bench as an editable checkout
uv run python dbrun.py suites # what tests exist
uv run python dbrun.py run -m sonnet,gpt-oss-120b -s mysql-installOne run leaves a directory like this:
runs/20260822-143512/
run.json what was asked for, the launch command, and its working directory
results.jsonl one line per finished cell, appended as it happens
leaderboard.md the table, the tables that explain it, and the tasks it scored
results.csv one row per stage, for a spreadsheet
events.jsonl droplets created and destroyed, with their ids
known_hosts this run's host keys, not yours
cells/<model>/<suite>/01-<stage>/ <suite>-<os>-<db> for whichever axis is wide
task.md the task exactly as the model received it
runbook.md the notes it was given with the task, if --runbook gave it any
grade.json every check, what it returned, and the score
bench.stdout.txt
bench/ the do-dba run directory: transcript.jsonl, report.md, secrets.json
run.json keeps command as the script argument list and adds launch_command
with the Python executable and interpreter flags, command_line as a quoted
string, and working_directory for resolving relative paths. The string uses
Windows process quoting on Windows and shell quoting on POSIX. These record the
Python process invocation; outer launchers such as uv run, shell aliases, and
the original shell quoting are not available to Python.
Use --agent-mode coordinated to let the model split each stage into subtasks
and choose how many workers to run concurrently:
uv run python dbrun.py run -m qwen/qwen3.8-flash \
-s mysql-replication-openbao --db percona-8.4 \
--agent-mode coordinated --max-workers 3 -j 1The coordinator chooses tasks and host assignments from that stage's inventory. It can work alone, give several hosts to one worker, wait for dependencies, or start a new worker to repair a failed subtask. This works across suites without hard-coded roles. Each worker uses the same model, reasoning effort, provider, and routing settings as its coordinator.
Workers exchange versioned JSON state through a shared board. Host reservations prevent concurrent dispatch to the same server; all agents share the stage's step and model cost budgets. The coordinator checks the complete task after workers finish, then the suite's existing independent grader runs.
--max-workers is the maximum concurrent subagents per stage, from 1 to 32
(default 4); the coordinator is additional. -j still limits simultaneous
matrix cells. Single-agent execution remains the default. Run metadata, JSONL
results, CSV, and leaderboard notes record the choice; resume requires matching
agent settings. See coordination details for the protocol,
logs, budgets, and cancellation behavior.
Add --validate-steps to have typesafe/jev-1.13 review every proposed step
with the working agent's full input context before execution. Jev uses
OpenRouter independently of the working model's provider. Rejected steps
are blocked and the classification is returned so the agent can revise;
validation errors stop execution.
This works in single and coordinated modes. Full requests and responses,
including complete script/file bodies, are saved in each agent's
step-validation.jsonl. Costs include validation; JSONL/CSV expose its
request count, rejections, and cost separately. Resume requires matching
validation settings. See step validation for
the decision policy, context limits, and log locations.
--max-reply-tokens N controls the output allowance for each model call,
including reasoning tokens. The default is derived from the model's context
window and capped at 16,384; a large context window does not imply a larger
reply allowance. Use an explicit value such as --max-reply-tokens 65536 to
compare reasoning-heavy models with more room. It applies to workers too and
is recorded in run.json; resume requires the same setting.
Repeated truncation is reported as output-limit, repeated empty answer text
as empty-response, and API error envelopes as api-error. These are distinct
from a model returning an invalid step. Full redacted responses and reasoning
metadata are recorded as model_reply events. Only kind=usage events count
toward cost totals; the reply events repeat response metadata for diagnosis.
Rejected replies are kept out of working history except for short failure notes
or invalid excerpts.
See the zero-score failure review for the observed failures, fixes, and validation.
- A cell is one model against one suite, on one operating system, for one database. It owns its own droplets, its own subprocess and its own directory, so cells run four at a time and a model that wedges costs one cell.
- A suite is one TOML file in
suites/. It says how many servers the work needs and lists the stages in order. - A stage is one task handed to the model, plus the checks that mark it. One stage means a fresh-droplet single task. Several stages mean the same droplets, in order - which is how you test work that builds on itself: set up replication, then upgrade both servers underneath it.
- A lease is the droplets a cell holds from its first stage to its last.
--cloud awsmakes them EC2 instances instead, which changes where they come from and nothing else on this page: this file says droplets throughout, and everything it says about them is true of both.
Each stage ends with two verdicts, and they are allowed to disagree:
status |
what do-dba concluded from the model's own VERIFY commands - done, unverified, exhausted, ... |
score |
the fraction of check weight the runner's own commands found satisfied |
A stage that reports done and scores 40% is the interesting case: the model
convinced the harness and not the suite. The leaderboard lists every one of them
under Said done, was not. The checks were written before any model saw the
task, and they run over the runner's own SSH connection after the model's process
has exited.
Two things are needed - a gateway, and something to make servers with - and the second one is optional in the sense that you can bring your own servers instead.
A model gateway. The same keys dbaai_bench uses - OPENROUTER_API_KEY, or
DIGITALOCEAN_INFERENCE_KEY with --provider digitalocean. The runner reads its
own .env first and the bench's second, so a key that already works there works
here - this project's .env only needs the names it is actually overriding. A name
left blank in it is dropped rather than read as an empty answer, so it falls
through to the bench's file instead of shadowing the real key there.
A third gateway needs no key at all. --provider selfhosted - self: for short -
talks to whatever OpenAI-compatible server $DBA_SELFHOSTED_BASE_URL points at,
LM Studio or vLLM or llama.cpp or Ollama, and nothing it serves is billed per
token: those cells count as $0 and the leaderboard says so rather than leaving a
row that looks free by accident. Two things are worth knowing before pointing a
matrix at one. A server that holds one model in memory at a time wants -j 1, or
the run spends its droplet hours swapping weights instead of working. And the
first request to each model waits while those weights load, which happens inside
the stage's own timeout rather than beside it.
doctl, to create droplets.
winget install DigitalOcean.doctl # or: scoop install doctl
doctl auth init # paste a personal access token
doctl account get # check itA token in $DIGITALOCEAN_ACCESS_TOKEN is used in preference to whatever
doctl auth init stored, so a run can be pointed at a different account without
touching your own login. It is passed to doctl through the environment and never
as an argument, because arguments are visible in the process list.
Or the aws CLI, to create EC2 instances instead. --cloud aws is the only
difference to the run:
winget install Amazon.AWSCLI # or: scoop install aws
aws configure # or: aws sso login --profile my-profile
aws sts get-caller-identity # check itThe runner shells out to the CLI you already configured rather than talking to the
API itself, so $AWS_PROFILE, $AWS_REGION, ~/.aws/config and an instance role
all work here exactly as they do in your own shell. --region overrides the region
for one run, $AWS_BINARY says where the CLI is if PATH does not reach it, and
$DBRUN_CLOUD=aws makes the flag the default so you stop typing it.
Four things EC2 makes the runner do differently, all of them on the plan before anything is created:
- No image lets root in. The key pair is installed for the image's own account -
ubuntu,ec2-user,admin,rocky, whoever the vendor chose - and root'sauthorized_keysholds the same key behind a forced command that says so. So the first thing done to a new instance is one login as that user, which copies the key over root's and takes the refusal off; everything after it, the bench and the suites and every check, is root over the runner's own connection exactly as on a droplet. The runner knows the names the big vendors use and the plan says which one it will try. A Marketplace or custom AMI is whoever built it, so--ssh-user NAMEnames it - without that the runner falls back to trying ten common names in turn, which works surprisingly often and costs a refused connection each time it does not. Whatever the account, it needs passwordless sudo; an image that does not give its own user that cannot be graded here, and the run says so rather than waiting. - One key pair, not every key. EC2 installs exactly one, so
--ssh-key NAMEpicks it and a region with a single key pair needs no flag. Its private half has to be somewhere the runner can reach - your agent, or-i ~/.ssh/that-key.pem, which the plan asks for if it cannot see one.-iis expanded and opened before anything is created: the tilde is yours, not the runner's - paramiko anddba.pynever see a shell - and a path that is not a file is refused up front rather than at the first login, where a file that will not open is indistinguishable from a machine that is still booting and gets waited on for the whole boot timeout, per node, on instances that are already billing. - Its own security group. A new instance is in nothing that would let SSH in,
so the run creates a tagged group of its own - port 22 from
--ssh-from(default0.0.0.0/0, narrow it to your own address where you can) plus everything between the group's own members, which is what the replication suites need - and gives it back in teardown.--security-group IDuses one of yours instead and leaves it alone.--subnetand--vpcplace the instances; the default is the roomiest subnet of the region's default VPC, and whatever you name needs a route to the internet, because the runner arrives over SSH and the model installs packages. - The disk is billed separately, and the vendors' images ship 8 GB, which will
not hold a database and a dump of it. Each instance gets a 100 GB gp3 root volume;
--disk GBchanges that.
Two more things read the same as on DigitalOcean and are answered differently
underneath. --os rocky-9 has no account-wide catalogue to resolve against, so it
asks each vendor's own published account for the newest image they released under
their own naming scheme - an ami- id is taken as given. And --size s-4vcpu-8gb,
which is what the suites here pin, is read as the capacity that slug names and
answered with the cheapest instance type that meets it, so a suite saying what the
task needs still gets it (the default is t3.xlarge).
Skip all of that to run against servers you already have:
uv run python dbrun.py run -m sonnet -s mysql-install --host root@203.0.113.10Those are never destroyed, and every cell gets the same machines one after another - fine for trying the runner out, not a matrix.
# every suite in suites/, three models, four cells at a time
uv run python dbrun.py run -m sonnet -m gpt-oss-120b -m qwen3-max
# one suite, models from a file, a hard cost ceiling
uv run python dbrun.py run --models-file models.txt -s mysql-replication --max-cost 20
# see what it would do and spend nothing
uv run python dbrun.py run -m sonnet --dry-run
# mix gateways in one run: provider:model overrides the default
uv run python dbrun.py run -m sonnet -m do:openai-gpt-oss-120b
# the box down the hall, one model in memory at a time
uv run python dbrun.py run -m self:qwen/qwen3.8-27b -s mysql-install -j 1
# one model at two thinking levels: two cells, two rows, one comparison
uv run python dbrun.py run -m qwen/qwen3.8-27b:low -m qwen/qwen3.8-27b:high \
-s mysql-install
# hold the routing still: only these two providers may serve the model
uv run python dbrun.py run -m qwen/qwen3.8-27b -s mysql-install \
--upstream AkashML,Reka
# the same models and suites on three operating systems, compared
uv run python dbrun.py run -m sonnet -s mysql-replication \
--os ubuntu-26 --os debian-13 --os rocky-9
# and the same again across three databases
uv run python dbrun.py run -m sonnet -s mysql-replication \
--db 9.7 --db percona-8.4 --db mariadb-11.4
# one task, three engines: the install suites are written check for check
uv run python dbrun.py run -m sonnet \
-s mysql-install -s postgres-install -s mongodb-install
# the other cloud: EC2 instances rather than droplets, in London
uv run python dbrun.py run -m sonnet -s mysql-install \
--cloud aws --region eu-west-2 -i ~/.ssh/laptop.pem
# hand every model the notes from a run that worked
uv run python dbrun.py run -m sonnet -s mysql-pxc-haproxy \
--runbook runbooks/mysql-pxc-haproxy.mdModel names are resolved against the gateway's catalogue before anything is
created, the same partial-name matching dba.py -m does. A typo fails in two
seconds instead of after forty droplets. A name can carry a gateway: prefix and a
thinking level - self:qwen/qwen3.8-27b:high - and both survive the resolution;
only the id in the middle is matched against the catalogue.
An id served by two gateways is two cells rather than one, which is the point of
running both: -m qwen/qwen3.8-27b -m self:qwen/qwen3.8-27b gets its own droplets,
its own directory and its own row on each side, and the rows are labelled or: and
self: so the table says which of them was which. Only ids that appear on more
than one gateway are labelled - everything else keeps its bare name, so a matrix
that names each model once reads exactly as it did before.
Suites are matched by name, unique prefix or path, so -s postgres-t finds
postgres-tune. A prefix that matches several suites is refused and says which,
which is also the quickest way to remember the names: -s mongodb lists the three
MongoDB suites rather than guessing between them. With no -s, every suite in
suites/ runs.
The matrix is built suite-major: every model gets tried on the first suite before any model sees the second. A run stopped early is then a comparison rather than one model's complete row.
Ctrl+C sets a flag every cell polls. Running stages are killed, pending cells
are cancelled, the droplets are destroyed, and the cells that already finished are
tabulated - then --resume picks up the rest:
uv run python dbrun.py run -m sonnet -m opus --resume 20260822-143512--resume skips the (model, gateway, suite, OS, database) cells that finished - the
same pair on the run's other image, or its other database, or the same weights on
another gateway, is still owed its droplets. A
cell the runner broke - a timeout, a droplet that never came up - is not counted
as finished, because there is nothing to learn from it. Nor is one the gateway
broke: a rate limit that outlasts the waiting ends the cell as api-error, and a
cell that spent one of its 120 steps before a 429 is a cell to run again rather
than a model that scored nothing. Those cells show their status in place of a
percentage and are in none of the means on the leaderboard - a gateway's bad
minute is not a benchmark result. --rate-limit-wait is how much of one a cell
will sit through before that happens - fifteen minutes per request by default,
which is a much longer fuse than do-dba's own two, because the servers are running
either way. Five minutes was the earlier default and was not enough: a recorded
four-server cell had executed all 26 steps asked of it when the gateway, asked for
step 27, still wanted another 60s at 290s of the 300s budget.
ssh-lost is counted the same way, and it is the network's turn to be the one at
fault. do-dba reopens a dropped connection and only gives up when the server will
not answer for a minute (--reconnect-wait there); a cell that ends this way lost
a machine and not an argument. One recorded three-node cell lost a server to a
reset socket in the middle of step 69 of 70 and the next stage reached the same
address on its first try minutes later - scored as the model's, that is 0% for a
dropped TCP connection with 69 steps of work already done. Unlike a 429 the
servers are worth keeping: --keep-failed holds them, because a link that died
mid-step is a machine to look at.
--os (spelled --image if you prefer) is repeatable, and each one it is given is
another axis of the matrix: cells become models x suites x OSes, one OS at a
time within each suite, and one leaderboard compares them. Given one --os, or
none (the default is ubuntu-26-04-x64), nothing changes - the OS is recorded but
stays out of the names and the tables, because one OS is not a comparison.
- Friendly names.
rocky-9resolves against your own account's image list, so it findsrockylinux-9-x64without you looking it up. A name that matches several images, or none, fails before a droplet exists and says which images it was choosing between. On--cloud awsthe same name is answered by asking Rocky's own published account for its newestRocky-9-EC2-Base-*, because an AMI is per region and there is no one catalogue to match against; every image the matrix may boot is resolved in preflight, so a third OS is not discovered hours in. Anami-id is taken as given, and a distribution neither vendor table knows is named that way. - A suite may decline.
os_familysays which package manager a suite's shell was written for, and a suite that cannot run on one of the run's images is left out of that column with anot run:line on the plan, rather than being graded on adnfit never had. If that empties the matrix, the run stops instead of starting. - A suite may pin. A suite with its own
imageruns once, on that image, whatever--ossays - it is testing something about that OS specifically. - The OS is part of a cell's identity: droplet names, directory names, the row
labels,
results.jsonl, the CSV and--resumeall carry it.
--os and --host are mutually exclusive: an operating system is something the
runner chooses when it creates a server, and your own servers already have one.
Seven products across three engines, and none of them is a drop-in for any other.
--db (spelled --mysql if you prefer - the axis was called that when MySQL was
all it had) is repeatable the same way --os is, and each value is another axis of
the matrix:
uv run python dbrun.py run -m sonnet -s mysql-install \
--db 9.7 --db percona-8.4 --db mariadb-11.4
# the same question, asked of the other two engines
uv run python dbrun.py run -m sonnet -s postgres-tune --db pg-18 --db pg-16
uv run python dbrun.py run -m sonnet -s mongodb-install --db mongodb-8.0 --db psmdb-8.0Whatever you write is what the models are asked for. There is no list of versions to be on and no list of names to be on, because a benchmark whose vocabulary lags its subject cannot ask the only question worth asking the week MariaDB 12.3 ships.
| written | means |
|---|---|
mysql-9.7, percona-8.4, mariadb-11.4 |
that flavour at that version |
pg-18, postgres-17, mongodb-8.0, psmdb-8.0 |
the other two engines, in the words their own documentation uses |
mariadb-12.3, mysql-8.0.36, pg-19 |
the same, at any version you like - nothing is checked against a list |
percona, mariadb, postgres, psmdb |
that flavour at its newest version here, which is the one thing the built-in versions are still for |
9.7, 8.4 |
that version of MySQL Community - a bare number is a MySQL, because that is where this axis started |
"Percona Server 8.4", "MariaDB 11.4 rc" |
your words, reaching the task exactly as written |
"FerretDB 2.0" |
a product this runner has never heard of |
One token is the shorthand: mysql (Community), percona (Percona Server for
MySQL), pxc (Percona XtraDB Cluster), mariadb, postgres, mongodb (Community),
percona-mongodb, the vendors' own words (percona-server-for-mysql,
percona-xtradb-cluster, maria-db, postgresql, pgsql, psmdb, mongodb-org)
and a version if you want one. It resolves to the vendor's own name
for the product, so --db psmdb-8.0 asks the model for "Percona Server for MongoDB
8.0".
Anything with a space in it is prose, kept word for word: --db "Percona Server 8.4" asks for exactly that and is recorded as percona-server-8.4. The runner reads
three things out of it and edits nothing - a trailing version if the last word looks
like one, which of the seven flavours it is, and which engine that flavour
belongs to, because grading has to know both what a server should look like and which
suites could honestly mark it. The words are read most-specific-first and the MySQL
ones last, since "Percona Server for MongoDB" contains percona and is not a MySQL -
and the product beats the vendor, since "Percona XtraDB Cluster" contains percona and
is not Percona Server. A name
with none of them in it is a flavour of its own, identified by grepping whatever the
server says about itself - VERSION(), version(), db.version() - for the name
itself, and offered to every suite on the axis, because which engine somebody meant
by "FerretDB 2.0" is not something a runner should be guessing.
Two things are still refused, before anything is created: a value with no name in
it, and two values that would be recorded as one row (percona-8.4 and "Percona 8.4" are both percona-8.4 on a leaderboard). Everything else the plan says
rather than refuses, because a typo and a product released last Tuesday are the
same string:
db myqsl-8.4
myqsl-8.4: 'myqsl' is not a flavour anything here knows, so every suite on the
--db axis - whichever database it is written for - will be asked for 'myqsl 8.4'
and will identify what turns up by grepping the version it reports for 'myqsl'
- did you mean mysql?
A version outside the built-in list gets the quieter version of that, naming the ones this runner has been kept up to date with. Both land above the confirmation prompt, which is the point: a database nobody can install grades every model at zero and looks exactly like every model failing.
One --db, or none, changes nothing about the matrix - the database is recorded on
every row but stays out of the cell names and the table headings, which have nothing
to distinguish. The default is one per engine (mysql-9.7, postgres-18,
mongodb-8.0), so a bare dbrun run asks each suite for its own engine's. The plan
still prints the db line whenever it is not the default, because one database is a
decision even when it is not an axis.
-
The axis is really one axis per engine.
--db percona-8.4 --db psmdb-8.0does not cross every suite with both: nothing can grade both, and a MySQL suite asked for MongoDB would score zero for having been asked the wrong question. So a suite takes the values on its own engine's line, and an engine the run never named falls back to that engine's default rather than dropping its suites ---db mariadb-11.4says which MySQL to compare and says nothing at all about MongoDB, and reading it as "no MongoDB today" would shorten the matrix by six suites without a word. That default is the engine's community build, except for a suite whosedb_flavourdoes not include it:mysql-pxc-haproxygrades Percona XtraDB Cluster, community MySQL is a different product rather than a lesser version of the same one, and so a baredbrun runasks it forpxc-8.4. -
A suite may decline.
db_flavourlists the flavours a suite's checks can honestly mark, and one the suite has not claimed is left out with anot run:line on the plan.mysql-tune's tuning stage rests onperformance_schema.variables_info, which MariaDB has not got, so grading it there would score a MariaDB tuner zero for MariaDB not being MySQL. A flavour from outside the seven is nobody's to decline - no suite could have named a product that did not exist when it was written - so every suite on the axis is asked for it, and the warning above is where the doubt goes. -
A suite may sit the axis out. No
db_flavourat all means the suite is not on this axis: it runs once however many databases the run compares, and its rows record no database. That ismysql-restore, whose cloud-init installs whatever the distribution callsmysql-server. A--dbthat no cell in the run ends up using - because every selected suite sits the axis out, or because they are all on another engine's line - is not a refusal and gets its own line, since the alternative is a run that costs what was asked for and answers a different question:not used: mariadb-11.4 - no suite in this run is on the --db axis for mysql, so nothing here will be asked to install itSitting the axis out is also not the way to write a suite for one product - a single-flavour
db_flavouris, and it is whatmysql-pxc-haproxydoes: on the axis, the series comes from the flag, and--db mysql-9.7is answered withnot run:instead of being ignored. Note the asymmetry withos_family, where saying nothing means runs on anything - a suite that mentions the product but declares no flavour, or declares flavours and never mentions the product, is refused at load time rather than run. -
The database is part of a cell's identity: droplet names, directory names, the row labels,
results.jsonl, the CSV and--resumeall carry it.
--db and --host are not mutually exclusive, unlike --os: the product is
what the task asks the model to install, not what the server already boots.
Two meters run at once, and the runner shows both.
- Models.
--max-cost USDis the run's ceiling: once spend reaches it no new cell starts, and the report says which were skipped. A cell already running is left alone - it is bounded by its own--max-cost-per-stage, and killing a cell halfway leaves a half-built server and a result worth nothing. Suites can set their own per-stage cap, and most of the ones here do. Self-hosted cells report no token cost, so a matrix made only of them never reaches any ceiling you set and the plan says as much instead of promising a guard that cannot fire. The droplets they drive are billed the same as anyone's. - Droplets. Billed by the hour whether or not anybody remembers them. Every
droplet is tagged
dbaai-runner,dbrun-<runid>andsuite-<name>, and its id is written toevents.jsonlbefore the model is handed anything. So whatever happens to the process:
uv run python dbrun.py reap # anything this runner ever left behind
uv run python dbrun.py reap --run 20260822-143512 # just that run's
uv run python dbrun.py reap --dry-run # list them and stopOn AWS the same tags are key and value - dbaai-runner=<runid>, dbrun-suite=<name>
and a Name you can read in the console - and a reap takes back the run's security
group as well as its instances. Two things differ enough to be worth typing:
uv run python dbrun.py reap --cloud aws --region eu-west-2 # EC2 tags are per region
uv run python dbrun.py reap --cloud both # both accounts, one passreap looks where --cloud says and nowhere else, and the default is DigitalOcean
(or $DBRUN_CLOUD) - so a reap that says nothing was found has answered about one
cloud. The rate on an EC2 plan is arithmetic rather than a quote: EC2 has no call
that will price an instance without the pricing endpoint and its own permissions, so
the figure is the on-demand list price in us-east-1 plus the volume, and the plan
says "about". Your region, your agreement and any savings plan all say otherwise.
--keep leaves the droplets up for a post-mortem. It is the one way this runner
spends money after it exits, so it says so on the way in and on the way out.
--keep-failed is that aimed at the cell that needs it. Some failures cannot be
diagnosed from the log - an sshd that refuses the login, a cloud-init that wedged, a
disk that filled - and destroying the server destroys the evidence, which is how the
same failure gets diagnosed twice. So a cell the runner broke keeps its servers and
everything else is destroyed as usual: a matrix of forty can be run for the price of
the one that needs looking at. A cell that ended because somebody stopped the run,
because the cap was reached, or because the gateway stopped answering keeps nothing
- there is no evidence on those servers, and two idle instances is a strange price to pay for a busy minute at a gateway. What it keeps, it tells you how to get into, as the cell fails and again as the last line of the run:
kept dbrun-20260825-074411-01-mysql-in 198.51.100.7 for diagnosis: could not get a root login
ssh -i ~/.ssh/id_ed25519 ubuntu@198.51.100.7
aws ec2 get-console-output --instance-id i-0a1b2c --region us-east-1 --output text
dbrun.py reap --cloud aws --region us-east-1 --run 20260825-074411 # when you are done
Three details are the point of it. The login offered is the image's own user, not
root - root is exactly what a boot failure on EC2 means the runner never managed to
arrange, so the lease's own user is the one that will not work. The console output
comes with it because it is the only thing that answers on an instance whose sshd
never came up, and on an instance that never even reached running it is the whole
of the diagnosis - so those are kept too, addressless as they are. And on AWS the
run's security group stays behind with them, because it is what lets port 22 through:
deleting it would leave servers that are still billing and can no longer be reached.
A cell the runner broke means a cell that never got servers, timed out, crashed,
failed its setup, or lost a server's connection for good (ssh-lost) - not a model
that scored badly. A model that fails its task is a result, and keeping servers for
every bad answer would leave a forty-cell matrix holding forty servers. reap is
how it all goes, and the exact command is printed.
--effort {off,minimal,low,medium,high,xhigh,max} is the one flag that changes what
a run costs without changing anything the plan counts. It reaches every stage as
dba.py --effort, asking the model to think before each step, and the gateway turns
that one word into whatever the model upstream wants - so it is worth comparing
models on rather than a knob only one of them has. The ends of the ladder are the
levels most worth asking for: off and minimal for a reasoning model that would
otherwise think its way through a systemctl check, xhigh and max for the ones
that can spend more when it matters. The rungs between are a request rather than a
promise - one measurement here, on qwen/qwen3.8-27b at a fixed seed, could not
tell low from high in reasoning tokens, while off was honoured clearly - so a
comparison whose whole point is the level should pin who serves it with --upstream
below and check the token counts in the transcripts. Thinking is billed as output
tokens: expect
several times the model spend of the same matrix without it, and expect the suites'
own max_cost ($1.50 to $5 a stage across the ones here) to be reached sooner,
which shows up as stages stopping on cost rather than on a verdict - so the plan
warns about it above the confirmation prompt, beside the cost caps. Unset is the
absence of the flag rather than a word meaning none: a reasoning model goes on
reasoning at the gateway's default, exactly as in every run recorded so far. The
effort is part of each stage's recorded command line in grade.json, so a cell
that thought is distinguishable from one that did not long after the run.
A thinking level can also be written into a model name - -m qwen/qwen3.8-27b:low,
or the same suffix on any line of a --models-file - which sets it for that model
alone and wins over --effort. :none is the other direction: it sits a run-wide
--effort out for one model, which is how a reasoning model and one without the
knob go into the same matrix. :none and :off are not the same thing and the
distinction is worth keeping: :none asks for nothing and that cell's command line
carries no --effort at all, while :off asks the model to stop thinking, which a
reasoning model will otherwise do by default. "We made no request" and "we asked for
none" are different runs and score differently. Like the gateway prefix, the suffix is part of what a
cell is, so -m qwen/qwen3.8-27b:low -m qwen/qwen3.8-27b:high is two cells with
their own droplets, their own directory and their own row - qwen/qwen3.8-27b:low
and qwen/qwen3.8-27b:high, spelled the way -m takes them back, so a row can be
re-run by copying its label. That is the comparison the suffix exists for: whether
the thinking was worth what it cost, on the same task and the same OS, rather than
across two runs a week apart. --resume counts the effort too, so resuming a matrix
that ran at :high does not skip the :low half of it.
--upstream NAME says which of the providers behind the gateway may serve the
models. It is OpenRouter's own idea: one id like qwen/qwen3.8-27b is offered by
half a dozen companies who have the weights, and every request is routed to
whichever of them looks best at that moment. They are not interchangeable. They
quantise differently, cap the context differently, honour a reasoning level or
ignore it, and come and go week to week - so an unpinned matrix can hand two cells
of the same "model" to two different machines and record the difference as a fact
about the model. Pinning is how a comparison holds still:
# only these two may serve it, AkashML first
uv run python dbrun.py run -m qwen/qwen3.8-27b -s mysql-install \
--upstream AkashML --upstream Reka
# same thing, typed once - commas split like -m does
uv run python dbrun.py run -m qwen/qwen3.8-27b:high -s mysql-install \
--upstream AkashML,RekaRepeatable, comma-separated lists accepted, and the order is the order the gateway
tries them in. The names are the gateway's own, spelled as its model page spells
them (AkashML, Reka, DeepInfra, Chutes) - nothing here validates them,
because who serves a model changes faster than any list in this repo could, and a
name that matches nobody is a cell that fails saying so.
Naming any upstream turns fallbacks off, which is the point rather than a side
effect: a pin that quietly fell through to whoever was up would not be holding
anything. The price is that a cell whose named providers are busy, rate-limited or
have dropped the model ends as api-error with its servers paid for, instead of
being answered by somebody else - so the plan warns about it above the confirmation
prompt, and --resume is how the cells that lost their provider get another go.
It reaches each stage as dba.py --upstream, one flag per name, so it is in the
recorded command line like the effort and the patience, and run.json carries an
upstream key naming who was allowed to answer - which is what makes two runs of
one id comparable afterwards. Only OpenRouter routes across several providers; cells
on DigitalOcean or a self-hosted server have one server behind each model and are
sent nothing, so a mixed matrix keeps them and the plan says how many are outside
the pin. A run where no model is on a routing gateway is refused up front rather
than per cell, because there is nothing there to pin.
An unpinned run still says when a provider is letting it down. The gateway picks an
upstream per request, and some of them stop replies at their own output cap - one
served deepseek-v4.1-flash for a whole stage and cut 32 of its 131 replies off at
about 2,200 of the 16,384 tokens asked for. Each of those is asked again, so no step
fails and the stage just runs out of time. dba.py warns the first time a reply stops
well short of the limit and again at the third from the same provider; the runner
reads the same thing off the transcript as the stage runs and prints it beside the
cell, repeats it under the stage's score, and gives the leaderboard a Providers that
cut replies short section naming the provider, how many replies it cut, and the
providers that served the same model cleanly - which is the --upstream to re-run
with. dbrun report finds these in old runs too, off the transcripts they kept.
--rate-limit-wait SECONDS (default 900, or $DO_INFERENCE_RATE_LIMIT_WAIT) is
how long one model request may spend waiting out 429s before the cell is given up
as api-error. It is far longer than do-dba's own two minutes because the trade is
different here: a cell that stops for a rate limit has servers running, an hour or
more of stage timeout in hand, and usually a pile of work that dies with it - one
recorded cell was refused at step 35 of 35, having spent $0.018 of its $1.50 cap,
because the gateway wanted another sixty seconds and had ten left to give. The
first fix for that was five minutes and a later run showed it was still short:
qwen/qwen3.8-flash, asked for step 27 on a four-server PXC stage, waited 5s, 15s,
30s and 60s four times over, and the gateway wanted another 60s at 290s of the
300s budget - 26 executed steps and four paid-for servers thrown away for one more
minute. Hence a quarter of an hour. Waiting costs nothing but wall clock: a 429 is
not billed, the stage timeout still stops a cell whose gateway never comes back,
and each wait is printed as it happens. A cell that waits this long on step after
step runs out of stage rather than of patience, and timeout is a runner fault
too, so neither ending is scored against the model. Free tiers are what the flag is
for - --rate-limit-wait 1800 for a matrix of them - and 0 restores giving up on
the first 429, which the plan warns about. It reaches every stage as dba.py --rate-limit-wait, so like the effort it is in the recorded command line, and a
cell can be re-run exactly as patiently as it was run.
--runbook [SUITE=]PATH hands the models notes along with the task - the repository
to add, the setting that turned out to be the one that mattered, the error that
means the config is not being read. It is the flag for asking a different question
than the rest of this page: not can these models work this out, but can they
follow what somebody already worked out.
# one suite in the run, one file
uv run python dbrun.py run -m sonnet -s mysql-pxc-haproxy \
--runbook runbooks/mysql-pxc-haproxy.md
# several suites: say which file is for which
uv run python dbrun.py run -m sonnet -s mysql-pxc-haproxy -s mysql-replication \
--runbook mysql-pxc-haproxy=runbooks/mysql-pxc-haproxy.md
# or point at the directory and let it match by name
uv run python dbrun.py run -m sonnet --runbook runbooks/The directory form is the one that scales: every suite in the run with a
runbooks/<suite>.md gets it, and the suites without one are the ordinary case
rather than an error. The bare-file form insists the run has exactly one suite
instead of guessing from the file's name, because a six-suite matrix given one
runbook is somebody who meant one of the other two spellings, and picking a suite
for them would put a stranger's notes in front of five models.
It reaches each stage as dba.py --runbook-file, which puts the text in the system
prompt in a RUNBOOK section under the task, framed as advice: the task is still
what the model is asked for and judged on, the rules still win, and where the notes
disagree with what a server actually says, the server is right. That framing is not
decoration - the notes were written about other machines, and a model that treats a
remembered wsrep_cluster_address as ground truth will confidently configure the
wrong cluster.
It is not another axis, and it changes what the percentage means. A runbook is a decision about what every model in the run was shown, like the wording of the task itself, so it applies run-wide and is not compared against anything. What it costs is that a suite run with a good runbook measures a model following instructions rather than a model that can do the work - and for a suite whose difficulty is a handful of traps, a detailed enough runbook is a transcription test. Both numbers are worth having. Pooling them into one table is not, so every place a score is recorded says which it is:
run.jsongets arunbookslist: the suite, the path, the character count and a short digest of the text.- every row of
results.jsonland therunbookcolumn ofresults.csvcarry<file>.md@<digest>for a helped cell and nothing for an unhelped one - because a spreadsheet pooled from several runs is exactly where that difference disappears. - the leaderboard header says it in words, above the matrix, including what it does to the column.
- each stage directory keeps the notes as
runbook.md, in full, beside thetask.md. Prose gets edited between runs - usually in the direction of giving more away - so a path alone could not say what this model was told, which is also why the digest is on the row.
--resume refuses to change it. The OS, the database and the effort fill in an
older record's silences; this one cannot, because the silence means the opposite - a
row with no runbook is a cell that was given none, which is a fact about the
measurement rather than a gap in the record. Resuming with different notes, or with
none where there were some, would build one table out of two experiments, so it stops
and prints what changed. Repeat the --runbook flags that run used, or start a run
of its own.
The notes have to be prose somebody stands behind, which in practice means runs that
worked: runbooks/mysql-pxc-haproxy.md here was written from a run where three models
each scored 100% on all three stages, quoting the configs that were on the servers when
the grader passed them - and naming, wherever those three disagreed, all of the shapes
that passed rather than flattening them into one right answer. A second run put the same
suite on Amazon Linux, where one cell of three took every mark and two lost 57 marks to
a /root/.my.cnf that outranked the grader's MYSQL_PWD; that became a section of the
same file, because the packaging changes and the cluster does not. Nothing in it is
marked untested, which was not true of the version written before the suite grew its
encryption and Synced checks, and it is the reason to rewrite the file when the suite
moves: notes that teach a model to turn encryption off are worse than no notes. An empty
file is refused rather than treated as no runbook, and the plan (--dry-run) prints each
file, its size and how many cells will be handed it, because 73,000 characters in the
system prompt is re-sent and re-paid at every step.
That last number is why the same notes exist twice. runbooks/brief/mysql-pxc-haproxy.md
is the same file with the narrative taken out - 26,000 characters instead of 73,000,
keeping only what decides a mark: the paths that differ between the two package managers,
the config that passed, the options that are not server options, one health check that can
see a desynced node, and the symptom-to-cause table. Because the directory form looks for
<suite>.md, the two are chosen by which directory you point at:
--runbook runbooks/ the long one, evidence and all
--runbook runbooks/brief/ the same advice, a third of the tokens
Both are recorded by digest, so a run says which of them a cell was given.
runbooks/brief/mongodb-replication.md was written the same way and exists only in the
short form - 18,000 characters out of thirty-five graded attempts, nineteen of which took
every mark. Most of it is one fact: MongoDB 8.x refuses to start on the kernels these
images ship, the packaged unit's GLIBC_TUNABLES=glibc.pthread.rseq=0 is what trips the
guard rather than the kernel itself, and clearing it in a drop-in is the whole fix. The
suites that have no file in a directory are not an error, so --runbook runbooks/brief/
on a matrix of both suites hands each cell the notes for its own suite and nothing to the
rest.
runbooks/brief/mysql-replication-openbao.md provides reusable instructions
for Percona 8.4 and 9.7 with OpenBao: version selection, scoped tokens,
native keyring startup, encrypted TLS replication, and restart verification.
Addresses and credentials are parameters; the brief contains no model names
or past-run references. Select it with
--runbook runbooks/brief/mysql-replication-openbao.md for this suite, or
--runbook runbooks/brief/ for directory matching across suites.
The detailed recipe remains at runbooks/mysql-replication-openbao.md.
A suite is a TOML file in suites/. The short form is one task:
name = "mysql-install"
description = "Install MySQL 9.7, a database, and a login user for it"
nodes = 1
max_steps = 25
max_cost = 1.5
task = '''
Install MySQL 9.7 from the vendor's own apt repository on this server.
Create a database called `app` and a login user called `app_rw`.
'''
[[check]]
name = "mysql is running and starts at boot"
weight = 2
command = "systemctl is-active --quiet mysql && systemctl is-enabled --quiet mysql"
[[check]]
name = "the app database exists"
command = "mysql -N -B -e \"SHOW DATABASES\" | grep -qx app"The long form is [[stage]] tables with [[stage.check]] under each, which is
what a suite that reuses its droplets looks like - see
suites/mysql-repl-upgrade.toml.
| key | meaning |
|---|---|
name |
the suite's name, and part of its droplets' hostnames |
description |
one line, shown in listings |
nodes = N |
N servers, unnamed. The harness calls them node1, node2, ... and the model works out the roles |
names = [...] |
N servers whose roles the suite has decided; the model is told these names and follows them |
size / image / region |
override the run's droplet settings for this suite. An image here takes the suite off the --os axis: it runs once, on that image. On --cloud aws a droplet slug is read as the capacity it names and a friendly image name as the vendor's newest, so a suite written for DigitalOcean still asks for what it meant |
os_family |
which OS families the suite's shell is written for - "debian" (apt, so Ubuntu and Debian both) or "rhel" (dnf), one or a list. --os leaves the suite out of an image it has not claimed |
db_flavour |
which databases the suite's checks can honestly mark - "mysql", "percona", "pxc", "mariadb", "postgres", "mongodb", "percona-mongodb", one or a list. All of them must be the same engine, since one suite's checks cannot mark two. Saying nothing takes the suite off the --db axis rather than putting it on every column of it; naming exactly one keeps it on the axis for the version and refuses every other product by name. Spelled mysql_flavour in a suite written before there was anything but MySQL, and still read |
cloud_init / cloud_init_file |
user-data for the droplets, for a fixture that should be in place before the model logs in |
stop_on_fail |
end the cell if a stage does not pass (off by default) |
| key | meaning |
|---|---|
id |
short, used as the directory name and in the tables |
task / task_file |
the prose the model is given. This is the specification: a vague task grades vaguely |
title |
one line for the progress display |
setup |
commands the runner runs before the model is told anything - the fixture the task starts from. A failure here is filed as setup-failed and never blamed on the model |
setup_host |
which servers the setup runs on (default: all) |
weight |
this stage's share of the cell's score |
max_steps |
steps the model gets (default 30) |
max_cost |
dollars this one task may spend |
timeout |
wall clock for the whole stage, in seconds (default 3600) |
command_timeout |
per command on the server, in seconds (default 300) |
| key | meaning |
|---|---|
name |
what the tables call it |
command |
shell, run by the runner over its own connection. Exit 0 is a pass |
host |
"*" every server must pass (the default), "one" exactly one must, "any" at least one must, or a node's name |
weight |
its share of the stage score (default 1) |
expect_exit |
when success is not exit 0 |
contains / not_contains |
when the exit code is not the answer - SHOW REPLICA STATUS exits 0 either way |
timeout |
seconds (default 60) |
host = "one" and host = "any" are what make a role-agnostic suite gradeable:
exactly one of these servers is a replica that is caught up, without the suite
dictating which one the model should have chosen. The nodes that satisfied it are
recorded, so the report can say which way round the model built it.
Use "one" for a role only one server can hold and "any" for one that several
can. The difference is worth getting right: two servers that both take writes and
both have binary logging on are two standalone databases, and under "any" they
would collect the marks for a replication pair. Both replication suites use "one"
for every check that names a role.
Every check and setup command is told about the lease it is running in:
DBRUN_SELF this node's label
DBRUN_NODES every label, space separated
DBRUN_PEERS the other labels
DBRUN_HOST_<LABEL> that node's public address
DBRUN_PRIVATE_<LABEL> its private address (the public one if it has none)
which is how a check can require that replication reads from the peer's private address rather than merely from somewhere.
A suite on the --db axis is told which database this cell asked for as well:
DBRUN_DB mysql-9.7, percona-8.4, pg-18, psmdb-8.0, ferretdb-2.0
DBRUN_DB_ENGINE mysql | postgres | mongodb | empty for a product from outside
DBRUN_DB_FLAVOUR the flavour: percona | mariadb | postgres | percona-mongodb | ...
DBRUN_DB_VERSION 9.7 empty if none was asked for
DBRUN_DB_NAME Percona Server for MySQL (the vendor's words, or yours)
DBRUN_DB_TITLE Percona Server for MySQL 8.4 (exactly what the task asks for)
DBRUN_DB_VERSION_RE ^9\.7 what to grep the reported version for, empty for "any"
DBRUN_DB_MATCH percona what the version and what it says about itself must match
DBRUN_DB_REJECT percona|mariadb what it must not, or empty
Every one of them except _ENGINE is exported under DBRUN_MYSQL_* as well, because
the axis had that name when MySQL was all it had and a suite written then still greps
$DBRUN_MYSQL_MATCH. Nine variables of duplication is a cheap price for an unset
$DBRUN_MYSQL_VERSION_RE being impossible: an empty pattern is a version check that
passes for every version.
A check written against those needs nothing else to grade a product this runner
has never heard of: an empty _REJECT and an empty _VERSION_RE both mean "no
constraint", which is what grep -q "" does anyway, so --db "MariaDB Enterprise"
grades as MariaDB at any version without a line of the suite changing.
Patterns and not only names, because every suite that cares asks the same two
questions - is this the version that was asked for, and is it the right vendor - and
neither is a string comparison. VERSION() is not enough on its own: Percona's is a
bare 8.4.3-3 and names the vendor only in @@version_comment, and Community is
the one product that has to be identified by what it is not. So a check greps both
fields together:
version=$(q "SELECT VERSION()"); comment=$(q "SELECT @@version_comment")
echo "version=${version:-none} comment=${comment:-none}"
[ -n "$version" ] || exit 1
printf '%s\n' "$version" | grep -q "$DBRUN_DB_VERSION_RE" || exit 1
printf '%s %s\n' "$version" "$comment" | grep -Eqi "$DBRUN_DB_MATCH" || exit 1
[ -z "$DBRUN_DB_REJECT" ] || ! printf '%s %s\n' "$version" "$comment" | grep -Eqi "$DBRUN_DB_REJECT"The same three lines carry to the other engines, and only the two fields change.
PostgreSQL is one project, so version() and server_version are all there is to
read. MongoDB is the awkward one: Percona tracks upstream's numbering exactly, so
db.version() cannot tell a community server from a Percona one and the mongodb
suites grep the version together with the installed package names - which is the
only place either vendor writes itself down plainly.
In the task - and only there - {{db}} writes the product in as it was asked
for, with {{db_name}}, {{db_version}}, {{db_flavour}} and {{db_engine}} for
the parts. {{mysql}}, {{mysql_name}}, {{mysql_version}} and {{mysql_flavour}}
still render the same things. Commands are given the variables above instead,
deliberately: braces in a check command belong to whatever is going to read them -
docker ps --format '{{.Names}}', a kubectl -o go-template, a jq filter - and
nothing here may touch them. A {{db_ver}} that nothing replaces is refused at
load time rather than handed to a model as literal braces.
Give a check a product-neutral name. Where the points went groups by check
name, and it answers as the database and version asked for is one comparable row
where it answers as MySQL 9.7 would be three that are not. It holds even for a
suite with one flavour on its axis, because the version is still a column:
mysql-pxc-haproxy grades the series and calls the check all three nodes answer
as the cluster that was asked for, which is one row across a run of pxc-8.0 and
a run of pxc-8.4. A suite off the axis altogether - mysql-restore - has one
product in every cell and may name it.
Checks are read-only by convention, and the starter suites break that convention deliberately twice: one check writes a marker row on whichever server accepts writes, and the next looks for that row on a server that refuses them. Checks run in the order they are written, which is what makes that pair work - and a write that arrives is the only real evidence that replication replicates.
mysql-pxc-haproxy breaks it further and on purpose: two of its checks break the
node HAProxy is serving, and the checks after each of them grade the failover. The
first desyncs it and leaves it running, which is the failure a health check that only
completes a MySQL handshake cannot see; the second stops the database outright. Both
are host = "one", so exactly one node may satisfy them - the suite picks the victim
by asking the proxy who it is talking to. The guards differ because the failures do:
the stop can require a whole cluster of three, while a desynced node is still a
member, so that one is guarded by a primary-key insert Galera certifies cluster-wide
and exactly one node wins. The suite puts its own desync back and waits for three
Synced nodes before it stops anything.
A stage with no checks is refused at load time. Grading a stage on the model's own account of itself is the one thing the harness under test already refuses to do.
| suite | servers | OS | db | what it is |
|---|---|---|---|---|
mysql-install |
1 | debian | all three MySQLs | install the database, a schema, a scoped user, not open to the world |
mysql-restore |
1 | debian | - | a dropped database and last night's dump - cloud-init and setup build the situation |
mysql-replication |
2 | debian, rhel | all three MySQLs | replication over the private network, roles left to the model |
mysql-replication-openbao |
3 | debian, rhel | percona, mariadb | primary, replica, and OpenBao: native remote keys, encrypted InnoDB tables and logs, verified TLS, and fresh key reads after database restarts |
mysql-repl-upgrade |
2 | debian, rhel | mysql | two stages on the same pair: replicate, then move to Percona without breaking it |
mysql-tune |
1 | debian | mysql, percona | two stages: install it, then size it to the machine - graded as ratios of the server's own memory, and on whether the settings would survive a restart |
mysql-group-replication |
3 | debian, rhel | mysql, percona | single-primary group replication across three servers: three ONLINE members as seen by every one of them, group traffic on the private network, a write crossing to both others |
mysql-pxc-haproxy |
4 | debian, rhel | pxc | Percona XtraDB Cluster on three servers behind HAProxy on a fourth, with the cluster and state-transfer traffic encrypted rather than switched off, then the benchmark desyncs the node HAProxy was serving and later stops it and grades where the traffic went both times |
mysql-pxc-clone |
4 | debian, rhel | pxc | three stages that grow a cluster rather than build one: one encrypted Percona XtraDB Cluster node, then HammerDB on a fourth server loading a hundred TPROC-C warehouses into it over TLS, then two more nodes added by clone SST - graded on the method, on the ten gigabytes arriving whole, and on all three nodes trusting one certificate authority kept out of the data directory |
postgres-install |
1 | debian | postgres | the same task as mysql-install, in PostgreSQL's vocabulary: a cluster, a database, a login role that is not a superuser |
postgres-tune |
1 | debian | postgres | install it, then size it - shared_buffers as a ratio of the machine, and pending_restart to catch an ALTER SYSTEM that was never restarted into |
postgres-replication |
2 | debian, rhel | postgres | streaming replication over the private network, roles left to the model |
postgres-patroni-ha |
5 | debian, rhel | postgres | three stages: PostgreSQL under Patroni with etcd on three nodes, then PgBouncer on each and two HAProxies sharing a keepalived address - the benchmark stops etcd on the leader's node, HAProxy on the address holder and Patroni on the leader, and grades that the leader stayed, the address moved and the proxies followed the new leader - then the stopped node brought back as a replica |
mongodb-install |
1 | debian | mongodb, psmdb | install it, turn access control on, a database with a collection in it, a user scoped to it |
mongodb-tune |
1 | debian | mongodb, psmdb | install it, then tune the server and its host: the WiredTiger cache, transparent huge pages, swappiness - and each of those in a way that outlives a reboot |
mongodb-replication |
2 | debian, rhel | mongodb, psmdb | a two-member replica set with a keyfile, over the private network, roles left to the model |
They are the tasks this benchmark was built for, and they are not cheap: the runs
they were written from took 40-70 steps and half an hour each. Start with
-s mysql-install and one model.
The OS column is each suite's os_family, and it is what --os rocky-9 picks
up: the install and tune suites are written around a vendor apt repository, so they
are left out of an rhel column rather than failed on it. The replication suites
grade a thing the database does rather than a way of installing it, and run on
either family.
The db column is each suite's db_flavour, and what --db mariadb-11.4 or
--db psmdb-8.0 picks up. Three suites per engine, deliberately the same three
tasks: install, tune and replication are written check for check against each
other, so a model's PostgreSQL score is readable beside its MySQL one and the
difference is the database rather than the marking. On the MySQL line,
mysql-install and mysql-replication mark all three flavours;
mysql-repl-upgrade is Community-only because becoming Percona is the task;
mysql-tune and mysql-group-replication need something MariaDB has not got. On
the MongoDB line every suite marks both builds, because what tells them apart is the
installed package name and not a check - which makes --db mongodb-8.0 --db psmdb-8.0 the cleanest comparison in the whole set. PostgreSQL has one flavour:
what Percona distributes is PostgreSQL, with the same version(), so a second
column would be noise - the axis worth comparing there is the major version, and
--db pg-18 --db pg-16 is two repositories, two data directories and one set of
checks. mysql-pxc-haproxy names one flavour for the same reason from the other
direction: pxc is a product of its own - not community MySQL, which has no wsrep,
and not Percona Server, which is the same vendor's other product - so the axis
there is the series, --db pxc-8.0 against --db pxc-8.4, and anything that is not
a cluster is refused on the plan rather than marked at zero. Only mysql-restore
has no db_flavour at all and runs once whatever the run compares, because the dump
is what it is.
mysql-replication-openbao installs a primary, a replica, and a dedicated OpenBao
server. It compares Percona Server's component_keyring_vault (or the legacy
keyring_vault plugin) with MariaDB's hashicorp_key_management. Use a MariaDB
build that supplies that plugin. Oracle MySQL Community is excluded because it
does not ship a native Vault backend; Oracle's keyring_hashicorp requires
Enterprise components.
uv run python dbrun.py run --models-file models.txt \
-s mysql-replication-openbao --db percona-8.4 --db mariadb-11.4Each cell uses three servers. OpenBao runs with persistent storage and verified HTTPS on its private address. Each database gets a separate KV-v2 mount and a scoped token. The task specifies the CA, audit log, and inspection-token paths that the checks need; the inspection token reads key metadata, not key values. An installation that only copies a local keyring into OpenBao does not satisfy the task.
The grader writes a marker into an encrypted InnoDB table and waits for it on the replica. It then restarts both database services, verifies that the marker remains readable, and requires successful key-read request/response pairs from both private database addresses in OpenBao's new audit records. Those requests must use non-root tokens and the correct per-node mounts. It also checks table and log encryption settings after restart, inspects persisted tablespace files for marker plaintext, and sends a second marker through replication. It does not restart or seal OpenBao; recovery of the key server itself is outside this suite.
Run the offline checks with uv run python run_tests.py suites openbao. They cover
both database flavors, legacy Percona keyrings, stale or failed audit responses,
deleted keys, plaintext tables and logs, and replication that fails after restart.
uv run python dbrun.py report # rebuild the last run's tables
uv run python dbrun.py report 20260822-143512results.jsonl is authoritative and append-only; the leaderboard and the CSV are
derived from it and can be rebuilt at any time, including after a scoring change.
The leaderboard has five sections: the score matrix, Said done, was not, Where the points went (every check and how many attempts satisfied it - the fastest way to find a check that is wrong rather than hard), Cells that did not get a fair run, Providers that cut replies short where any did, and The tasks - the prose the models were given, quoted once per distinct wording, because a suite's paragraphs get edited between runs and a percentage is not a result without them.
A run that compared operating systems gets a sixth, By operating system, and one
that compared databases gets By database. Either way every table that names a
cell grows that axis in its row labels: rows read `sonnet @ rockylinux-9` or
`sonnet / mariadb-11.4`, or both at once, and Where the points went counts
each check per OS and per database rather than pooling them - because a check that
passes on one image and not the other, or on MySQL and not on MariaDB, is the whole
reason to run both. The tasks splits the same way, naming the database whenever a
stage's {{db}} made two wordings out of one paragraph.
dbrun.py entry point
dba_runner/
cli.py subcommands, screens, preflight
matrix.py cells, the thread pool, the budget, cleanup
suites.py the TOML format and its loader
products.py the --db axis: engines, flavours, versions, what a check is told
provision.py doctl, leases, droplets, reap
aws.py the same on EC2: instance types, AMIs, security groups
harness.py running dba.py as a subprocess and reading it back
grade.py the checks, over the runner's own SSH connection
results.py results.jsonl, the leaderboard, the CSV
runbooks.py --runbook: which suite is handed which notes
suites/ the tests
runbooks/ notes for a suite, written from a run that worked
tests/ offline suites - no network, no token, no money
run_tests.py runs them all, one line per suite
do_dba is imported for its SSH transport and terminal handling, and dba.py is
run as a subprocess: the thing being measured is allowed to fail badly, and in
a subprocess a wedged model costs one cell and a timeout.
uv run python run_tests.py
uv run python run_tests.py suites gradeNothing there touches the network, a DigitalOcean account, an AWS account or a
model. doctl and aws are fakes that answer from a script, the bench is a fake
dba.py that writes a transcript, and the servers are fakes that answer the checks -
so the whole matrix, including a cell that times out and a droplet that leaks, runs
on a laptop with no token and no money.
ec2 is the suite for the other cloud, and it is mostly about the three things with
no DigitalOcean counterpart: a security group that outlives its instances (deleted
on the retry after termination catches up, and reported as leaked when it does not),
the root login a vendor image refuses on purpose, and a size and an image named in
this project's vocabulary rather than EC2's - s-4vcpu-8gb becoming the cheapest
type that holds it, rocky-9 becoming a lookup against Rocky's own account.
runbook is mostly about the record rather than the reading. Resolving
--runbook is a page of parsing with seven ways to get it wrong, and those are
checked; the longer half runs a two-suite matrix with one suite helped and then
looks for that difference in every artifact that outlives the run - the row, the
line in results.jsonl, the CSV column, the leaderboard's header - because a
helped cell and an unhelped one scoring the same 100% are not the same result, and
nothing in the number says so. It also pins the resume refusal, including that an
edited runbook is a different runbook and that a run which wrote no run.json
resumes anyway.
Six of them grade shell rather than Python. suites loads all fourteen shipped
suites and puts every check and setup script through sh -n, renders every task
once per flavour the suite claims and fails on a {{ that survived, and asserts
that each suite accepts its own engine's default database - a suite that refused it
would be silently skipped by an ordinary run. tune runs mysql-tune's tuning
checks against
nine fake servers - one tuned well, one untouched, one whose settings would not
survive a restart, one with performance_schema switched off - and asserts which
checks should fail on each. repl grades both MySQL replication suites against twelve
pairs of servers that are really directories with a stub mysql on their PATH: a
correct pair, the same pair the other way round, an older server that spells every
field Slave_, a pair replicating over the public internet, a replica that still
takes writes, a faultless pair of MariaDB, two standalone databases, a dead peer.
Three of the twelve are the --db axis: a Percona pair graded as a Community
run and a Community pair graded as a Percona one both fail on the version check
alone, and the Percona pair graded as the Percona run it is comes out clean. The
suite file is byte-identical across all three - only the exported environment
differs, which is the whole claim the axis makes.
That stub is an executable and not a shell function on purpose, because the client
is where a suite goes wrong: this one refuses \G when the statement arrives
through -e, exactly as the real one does, and a helper that asked that way
scored a working replica zero on every model of two whole runs. One scenario
grades a correct pair with that old helper put back and asserts it still comes out
at the 0.625 the run recorded - so the day the stub stops being awkward, the test
says so instead of passing everything.
group and pxc are the same idea on three and four servers. group grades
mysql-group-replication against sixteen groups - a correct single-primary one, a
tarball install with the client on nobody's $PATH, a group talking over the public
addresses, two primaries where there should be one, a member still joining. pxc
grades mysql-pxc-haproxy, and its fake cluster is real enough to fail over: the
shared directory is the cluster, a node's data is the cluster's data only while it
is a member, and the stub answers as HAProxy on the proxy's address. So stage two
breaks whichever node the proxy was serving - the suite chooses it, this file does
not - twice, in the two ways that are not the same: it desyncs the node and leaves it
running, then later stops it outright, and the failover checks pass because two other
machines really had the write and the proxy really moved. The stub proxy has both
health checks in it, one that reads the served node's state and one that only opens a
connection, which is the whole difference the desync exists to measure. Six clusters
are graded, eight more only through the first stage, and five are built to tempt a
check into passing with only the checks that could be fooled run against them: a
proxy balancing across all three, a cluster that had already lost a member, a cluster
with nothing wrong with it, a proxy whose health check proves nothing but that the
port answers, and one serving a node that is feeding a state transfer. Two more grade
a check by its evidence and not only its verdict: the axis - three nodes answering
8.0.46, marked twice, refused by a run that asked for 8.4 and accepted by a run that
asked for 8.0, because a check that greps a series written into the suite file would
pass the first and fail the second - and the encryption, where the three switches
that can put cluster traffic in clear are turned off one node apiece, so a node's
line has to name which of the three it was.
The scenario that keeps that one honest is bootstrapped: a cluster brought back
from the node that was behind comes up Synced, three strong, one backend served -
and has silently lost the write made while it was down. Everything in stage three
passes except the check that counts the rows, which is the only reason that check
exists.
clone grades mysql-pxc-clone the same way, and the shared directory being the
cluster is what makes it possible at all: a joiner that becomes a member of the
component sees the hundred warehouses that were loaded before it existed, which is
a clone modelled as the thing a clone actually achieves. Five arcs run all three
stages - one done properly, one done differently in every way the suite is not
allowed to mind, one joined by the vendor's default backup transfer instead, one
whose joiners kept the certificates their own installer generated, and one that
built the whole cluster in stage one before there was anything to clone - and
thirty-odd fixtures grade a single stage against servers built to tempt one check
into the wrong answer. The certificate stubs are the interesting ones: a
certificate here is a file naming the authority that signed it, so openssl verify
and openssl x509 -fingerprint both answer, and the three ways the vendor's
default layout fails are separable - files left in the data directory, the same
files named relatively so that nothing in the configuration says they are inside
it, and files moved out with the installer's self-signed certificate still in them.
The load has a fixture that lies: a hundred rows in warehouse on top of a single
warehouse of data, with the aggregates and the reported size made to agree. Every
check that counts in bulk passes it. The hundredth warehouse is the only one that
does not, and that is the only reason that check is in the suite.
patroni grades postgres-patroni-ha, and is the first stub-server grader that is
not MySQL: a stub psql that is PostgreSQL on a node and HAProxy-then-PgBouncer on
a proxy port, and a stub curl that is Patroni's REST API and etcd's gateway. The
shared directory is the world again, and systemctl is where it moves: stopping
Patroni on the leader elects the next member that can see etcd's quorum, stopping
etcd on a leader whose Patroni was only given its own node's member demotes it, and
stopping HAProxy on the address holder hands the virtual IP over if keepalived was
tracking it. Six arcs are marked through the failures - one done properly, one done
differently in every way the suite may not mind, and one each for a Patroni that
knows only its local etcd, keepalived left on multicast, keepalived that never
watches HAProxy, and HAProxy that bypasses PgBouncer. Two more recoveries start from
a copy of the correct arc's servers as the failures left them, one with the stopped
node left down and one with it brought back by pg_ctl. A handful of refusals check
that each injection refuses to break a second thing and that each failover check refuses a cluster
nobody broke. local-etcd is the one the etcd failure exists for: the failover it
causes is fast and clean, and the only thing wrong is that it happened.
The PostgreSQL and MongoDB tune and replication suites are still the gap in this
battery. They are loaded, rendered and parsed by suites, but nothing offline runs
their checks against a fake server yet, so their marking has been argued rather
than tested. The psql stub above is a start; mongosh would need one of the same
kind, awkward where the real client is awkward, which for mongosh means an
adminCommand that answers {ok: 0} instead of failing.
A suite that mis-marks does not look like a broken suite. It looks like a model that failed.
runs/holds real credentials. The bench'ssecrets.jsonand the servers' generated passwords end up under it. It is in.gitignoreand should stay there.- Its own
known_hosts. Per run, inside the run directory. Both clouds reissue addresses, and neither your~/.ssh/known_hostsfilling with dead droplets nor a recycled address looking like an attack is wanted. Error reading SSH protocol banneris not a failure. Ubuntu socket-activates sshd, so port 22 answers while the host keys are still being generated and the first connection closes where the banner should be. The runner retries for up to 7 minutes and says what happened either way, so paramiko's own traceback for it is kept off stderr -DBRUN_SSH_DEBUG=1puts it back when the connection is the thing you are debugging.- The retries back off, because sshd punishes the ones that do not. OpenSSH from
9.8 keeps a penalty per source address and refuses connections from one whose
attempts keep failing - 15 seconds at first, up to 10 minutes if it continues, and
every refusal inside the penalty extends it. A loop knocking every 5 seconds can
therefore lock itself out of a machine that is working, for longer than it was ever
willing to wait, and the run dies with
could not get a root login. So the wait doubles after each failure up to 45 seconds. On EC2 there is a better answer than waiting: the runner asks for the instance's own two status checks first and only knocks once they pass, so most of the time there is no penalty to lapse. That check is a courtesy and never the verdict - if EC2 has not decided within 2 minutes, or will not answer, the login goes ahead and decides. - Every wait says which wait it is and how long it has been. Booting a cell is
four waits in a row - the instance reaching
runningwith an address, port 22 answering, the root login going through, cloud-init finishing - and any of them can legitimately take minutes. Until they reported themselves the whole thing was one silent gap, which reads as a hung runner and got read that way: EC2 reachingrunningtakes about 8 seconds, but the 7 minutes of nothing afterwards looked like the cloud being slow to say so. Each wait now prints a line every 20 seconds naming what it is still waiting for and its elapsed seconds, rate-limited so four cells booting at once do not bury the log. The four share oneBOOT_TIMEOUTbetween them rather than starting a fresh one each, so a cell cannot spend 7 minutes per wait and a slow port no longer buys the login extra time. - cloud-init is waited for. A freshly booted image is still installing things
for a minute or two, holding the apt or dnf lock, and a model whose first step is
apt-getwould meet "Could not get lock" through no fault of its own. That is the difference between grading a model and grading a race. - Multi-stage suites rely on do-dba's server-side credential store. Stage one's
generated passwords stay on the server in
/etc/profile.d/dba-secrets.sh, which is how stage two can log in at all.--no-server-secretsin the bench would break that, so the runner does not pass it. - Nothing in a cell may reboot a node. The grader's connection to a node is opened
once and used for every stage, and nothing on this side of it survives the far end
going down - so do-dba's guard refuses
reboot,shutdown,poweroff,halt,init 6,systemctl rebootand a character into/proc/sysrq-triggeroutright instead of asking about them:--mode unattendedanswers every question yes, so a question is not a limit. The suites ask for the smaller thing that is actually being graded - a service that "comes up on its own when the machine boots", whichsystemctl is-enabledanswers without trying it - and where a failure is wanted, it is a process that goes down: the PXC failover stage stops the database on the routed node. One cell paid for this rule. Asystemctl rebooton db3 in stage 1 made every later check on db3 reportunreachable: the SSH connection is not open, and a PXC cluster that was up and Synced on all three nodes scored zero on three stages. - A connection that dies anyway is waited out. Plenty besides a reboot can take one away - sshd restarted, the firewall rewritten, a node that simply went. When one dies the runner waits up to 2.5 minutes for that node to answer again and logs back in, saying so in the cell's log; a node that really is gone costs the cell its checks, with its own name in the reason, rather than costing every check after it.
- A crash is recorded in the bench's words, not paramiko's. When
dba.pyleaves no transcript the runner quotes it, and what it quotes now is the last complaintdba.pyprinted -could not connect to 18.207.93.154:22 - Error reading SSH protocol banner- rather than the last six lines of stderr, which for anything SSH is a transport thread's frames from inside site-packages. The leaderboard prints the first line of that field, so it used to readcrashed: Traceback (most recent call last):, which names neither the host nor the reason. Both files are kept whole beside the cell either way.