Skip to content

Repository files navigation

dba-runner

dbaai_bench runs one model against one task on servers you already have. This runs many models against many tasks on servers it creates itself, marks the results with its own commands, and destroys the servers afterwards.

uv sync                                     # installs ../dbaai_bench as an editable checkout
uv run python dbrun.py suites               # what tests exist
uv run python dbrun.py run -m sonnet,gpt-oss-120b -s mysql-install

One run leaves a directory like this:

runs/20260822-143512/
  run.json          what was asked for, the launch command, and its working directory
  results.jsonl     one line per finished cell, appended as it happens
  leaderboard.md    the table, the tables that explain it, and the tasks it scored
  results.csv       one row per stage, for a spreadsheet
  events.jsonl      droplets created and destroyed, with their ids
  known_hosts       this run's host keys, not yours
  cells/<model>/<suite>/01-<stage>/     <suite>-<os>-<db> for whichever axis is wide
      task.md         the task exactly as the model received it
      runbook.md      the notes it was given with the task, if --runbook gave it any
      grade.json      every check, what it returned, and the score
      bench.stdout.txt
      bench/          the do-dba run directory: transcript.jsonl, report.md, secrets.json

run.json keeps command as the script argument list and adds launch_command with the Python executable and interpreter flags, command_line as a quoted string, and working_directory for resolving relative paths. The string uses Windows process quoting on Windows and shell quoting on POSIX. These record the Python process invocation; outer launchers such as uv run, shell aliases, and the original shell quoting are not available to Python.

Dynamic coordinator and parallel workers

Use --agent-mode coordinated to let the model split each stage into subtasks and choose how many workers to run concurrently:

uv run python dbrun.py run -m qwen/qwen3.8-flash \
  -s mysql-replication-openbao --db percona-8.4 \
  --agent-mode coordinated --max-workers 3 -j 1

The coordinator chooses tasks and host assignments from that stage's inventory. It can work alone, give several hosts to one worker, wait for dependencies, or start a new worker to repair a failed subtask. This works across suites without hard-coded roles. Each worker uses the same model, reasoning effort, provider, and routing settings as its coordinator.

Workers exchange versioned JSON state through a shared board. Host reservations prevent concurrent dispatch to the same server; all agents share the stage's step and model cost budgets. The coordinator checks the complete task after workers finish, then the suite's existing independent grader runs.

--max-workers is the maximum concurrent subagents per stage, from 1 to 32 (default 4); the coordinator is additional. -j still limits simultaneous matrix cells. Single-agent execution remains the default. Run metadata, JSONL results, CSV, and leaderboard notes record the choice; resume requires matching agent settings. See coordination details for the protocol, logs, budgets, and cancellation behavior.

Optional step validation

Add --validate-steps to have typesafe/jev-1.13 review every proposed step with the working agent's full input context before execution. Jev uses OpenRouter independently of the working model's provider. Rejected steps are blocked and the classification is returned so the agent can revise; validation errors stop execution.

This works in single and coordinated modes. Full requests and responses, including complete script/file bodies, are saved in each agent's step-validation.jsonl. Costs include validation; JSONL/CSV expose its request count, rejections, and cost separately. Resume requires matching validation settings. See step validation for the decision policy, context limits, and log locations.

Reply limits and response diagnostics

--max-reply-tokens N controls the output allowance for each model call, including reasoning tokens. The default is derived from the model's context window and capped at 16,384; a large context window does not imply a larger reply allowance. Use an explicit value such as --max-reply-tokens 65536 to compare reasoning-heavy models with more room. It applies to workers too and is recorded in run.json; resume requires the same setting.

Repeated truncation is reported as output-limit, repeated empty answer text as empty-response, and API error envelopes as api-error. These are distinct from a model returning an invalid step. Full redacted responses and reasoning metadata are recorded as model_reply events. Only kind=usage events count toward cost totals; the reply events repeat response metadata for diagnosis. Rejected replies are kept out of working history except for short failure notes or invalid excerpts.

See the zero-score failure review for the observed failures, fixes, and validation.

What it is doing

  • A cell is one model against one suite, on one operating system, for one database. It owns its own droplets, its own subprocess and its own directory, so cells run four at a time and a model that wedges costs one cell.
  • A suite is one TOML file in suites/. It says how many servers the work needs and lists the stages in order.
  • A stage is one task handed to the model, plus the checks that mark it. One stage means a fresh-droplet single task. Several stages mean the same droplets, in order - which is how you test work that builds on itself: set up replication, then upgrade both servers underneath it.
  • A lease is the droplets a cell holds from its first stage to its last. --cloud aws makes them EC2 instances instead, which changes where they come from and nothing else on this page: this file says droplets throughout, and everything it says about them is true of both.

Each stage ends with two verdicts, and they are allowed to disagree:

status what do-dba concluded from the model's own VERIFY commands - done, unverified, exhausted, ...
score the fraction of check weight the runner's own commands found satisfied

A stage that reports done and scores 40% is the interesting case: the model convinced the harness and not the suite. The leaderboard lists every one of them under Said done, was not. The checks were written before any model saw the task, and they run over the runner's own SSH connection after the model's process has exited.

Getting set up

Two things are needed - a gateway, and something to make servers with - and the second one is optional in the sense that you can bring your own servers instead.

A model gateway. The same keys dbaai_bench uses - OPENROUTER_API_KEY, or DIGITALOCEAN_INFERENCE_KEY with --provider digitalocean. The runner reads its own .env first and the bench's second, so a key that already works there works here - this project's .env only needs the names it is actually overriding. A name left blank in it is dropped rather than read as an empty answer, so it falls through to the bench's file instead of shadowing the real key there.

A third gateway needs no key at all. --provider selfhosted - self: for short - talks to whatever OpenAI-compatible server $DBA_SELFHOSTED_BASE_URL points at, LM Studio or vLLM or llama.cpp or Ollama, and nothing it serves is billed per token: those cells count as $0 and the leaderboard says so rather than leaving a row that looks free by accident. Two things are worth knowing before pointing a matrix at one. A server that holds one model in memory at a time wants -j 1, or the run spends its droplet hours swapping weights instead of working. And the first request to each model waits while those weights load, which happens inside the stage's own timeout rather than beside it.

doctl, to create droplets.

winget install DigitalOcean.doctl        # or: scoop install doctl
doctl auth init                          # paste a personal access token
doctl account get                        # check it

A token in $DIGITALOCEAN_ACCESS_TOKEN is used in preference to whatever doctl auth init stored, so a run can be pointed at a different account without touching your own login. It is passed to doctl through the environment and never as an argument, because arguments are visible in the process list.

Or the aws CLI, to create EC2 instances instead. --cloud aws is the only difference to the run:

winget install Amazon.AWSCLI             # or: scoop install aws
aws configure                            # or: aws sso login --profile my-profile
aws sts get-caller-identity              # check it

The runner shells out to the CLI you already configured rather than talking to the API itself, so $AWS_PROFILE, $AWS_REGION, ~/.aws/config and an instance role all work here exactly as they do in your own shell. --region overrides the region for one run, $AWS_BINARY says where the CLI is if PATH does not reach it, and $DBRUN_CLOUD=aws makes the flag the default so you stop typing it.

Four things EC2 makes the runner do differently, all of them on the plan before anything is created:

  • No image lets root in. The key pair is installed for the image's own account - ubuntu, ec2-user, admin, rocky, whoever the vendor chose - and root's authorized_keys holds the same key behind a forced command that says so. So the first thing done to a new instance is one login as that user, which copies the key over root's and takes the refusal off; everything after it, the bench and the suites and every check, is root over the runner's own connection exactly as on a droplet. The runner knows the names the big vendors use and the plan says which one it will try. A Marketplace or custom AMI is whoever built it, so --ssh-user NAME names it - without that the runner falls back to trying ten common names in turn, which works surprisingly often and costs a refused connection each time it does not. Whatever the account, it needs passwordless sudo; an image that does not give its own user that cannot be graded here, and the run says so rather than waiting.
  • One key pair, not every key. EC2 installs exactly one, so --ssh-key NAME picks it and a region with a single key pair needs no flag. Its private half has to be somewhere the runner can reach - your agent, or -i ~/.ssh/that-key.pem, which the plan asks for if it cannot see one. -i is expanded and opened before anything is created: the tilde is yours, not the runner's - paramiko and dba.py never see a shell - and a path that is not a file is refused up front rather than at the first login, where a file that will not open is indistinguishable from a machine that is still booting and gets waited on for the whole boot timeout, per node, on instances that are already billing.
  • Its own security group. A new instance is in nothing that would let SSH in, so the run creates a tagged group of its own - port 22 from --ssh-from (default 0.0.0.0/0, narrow it to your own address where you can) plus everything between the group's own members, which is what the replication suites need - and gives it back in teardown. --security-group ID uses one of yours instead and leaves it alone. --subnet and --vpc place the instances; the default is the roomiest subnet of the region's default VPC, and whatever you name needs a route to the internet, because the runner arrives over SSH and the model installs packages.
  • The disk is billed separately, and the vendors' images ship 8 GB, which will not hold a database and a dump of it. Each instance gets a 100 GB gp3 root volume; --disk GB changes that.

Two more things read the same as on DigitalOcean and are answered differently underneath. --os rocky-9 has no account-wide catalogue to resolve against, so it asks each vendor's own published account for the newest image they released under their own naming scheme - an ami- id is taken as given. And --size s-4vcpu-8gb, which is what the suites here pin, is read as the capacity that slug names and answered with the cheapest instance type that meets it, so a suite saying what the task needs still gets it (the default is t3.xlarge).

Skip all of that to run against servers you already have:

uv run python dbrun.py run -m sonnet -s mysql-install --host root@203.0.113.10

Those are never destroyed, and every cell gets the same machines one after another - fine for trying the runner out, not a matrix.

Running a matrix

# every suite in suites/, three models, four cells at a time
uv run python dbrun.py run -m sonnet -m gpt-oss-120b -m qwen3-max

# one suite, models from a file, a hard cost ceiling
uv run python dbrun.py run --models-file models.txt -s mysql-replication --max-cost 20

# see what it would do and spend nothing
uv run python dbrun.py run -m sonnet --dry-run

# mix gateways in one run: provider:model overrides the default
uv run python dbrun.py run -m sonnet -m do:openai-gpt-oss-120b

# the box down the hall, one model in memory at a time
uv run python dbrun.py run -m self:qwen/qwen3.8-27b -s mysql-install -j 1

# one model at two thinking levels: two cells, two rows, one comparison
uv run python dbrun.py run -m qwen/qwen3.8-27b:low -m qwen/qwen3.8-27b:high \
    -s mysql-install

# hold the routing still: only these two providers may serve the model
uv run python dbrun.py run -m qwen/qwen3.8-27b -s mysql-install \
    --upstream AkashML,Reka

# the same models and suites on three operating systems, compared
uv run python dbrun.py run -m sonnet -s mysql-replication \
    --os ubuntu-26 --os debian-13 --os rocky-9

# and the same again across three databases
uv run python dbrun.py run -m sonnet -s mysql-replication \
    --db 9.7 --db percona-8.4 --db mariadb-11.4

# one task, three engines: the install suites are written check for check
uv run python dbrun.py run -m sonnet \
    -s mysql-install -s postgres-install -s mongodb-install

# the other cloud: EC2 instances rather than droplets, in London
uv run python dbrun.py run -m sonnet -s mysql-install \
    --cloud aws --region eu-west-2 -i ~/.ssh/laptop.pem

# hand every model the notes from a run that worked
uv run python dbrun.py run -m sonnet -s mysql-pxc-haproxy \
    --runbook runbooks/mysql-pxc-haproxy.md

Model names are resolved against the gateway's catalogue before anything is created, the same partial-name matching dba.py -m does. A typo fails in two seconds instead of after forty droplets. A name can carry a gateway: prefix and a thinking level - self:qwen/qwen3.8-27b:high - and both survive the resolution; only the id in the middle is matched against the catalogue.

An id served by two gateways is two cells rather than one, which is the point of running both: -m qwen/qwen3.8-27b -m self:qwen/qwen3.8-27b gets its own droplets, its own directory and its own row on each side, and the rows are labelled or: and self: so the table says which of them was which. Only ids that appear on more than one gateway are labelled - everything else keeps its bare name, so a matrix that names each model once reads exactly as it did before.

Suites are matched by name, unique prefix or path, so -s postgres-t finds postgres-tune. A prefix that matches several suites is refused and says which, which is also the quickest way to remember the names: -s mongodb lists the three MongoDB suites rather than guessing between them. With no -s, every suite in suites/ runs.

The matrix is built suite-major: every model gets tried on the first suite before any model sees the second. A run stopped early is then a comparison rather than one model's complete row.

Ctrl+C sets a flag every cell polls. Running stages are killed, pending cells are cancelled, the droplets are destroyed, and the cells that already finished are tabulated - then --resume picks up the rest:

uv run python dbrun.py run -m sonnet -m opus --resume 20260822-143512

--resume skips the (model, gateway, suite, OS, database) cells that finished - the same pair on the run's other image, or its other database, or the same weights on another gateway, is still owed its droplets. A cell the runner broke - a timeout, a droplet that never came up - is not counted as finished, because there is nothing to learn from it. Nor is one the gateway broke: a rate limit that outlasts the waiting ends the cell as api-error, and a cell that spent one of its 120 steps before a 429 is a cell to run again rather than a model that scored nothing. Those cells show their status in place of a percentage and are in none of the means on the leaderboard - a gateway's bad minute is not a benchmark result. --rate-limit-wait is how much of one a cell will sit through before that happens - fifteen minutes per request by default, which is a much longer fuse than do-dba's own two, because the servers are running either way. Five minutes was the earlier default and was not enough: a recorded four-server cell had executed all 26 steps asked of it when the gateway, asked for step 27, still wanted another 60s at 290s of the 300s budget.

ssh-lost is counted the same way, and it is the network's turn to be the one at fault. do-dba reopens a dropped connection and only gives up when the server will not answer for a minute (--reconnect-wait there); a cell that ends this way lost a machine and not an argument. One recorded three-node cell lost a server to a reset socket in the middle of step 69 of 70 and the next stage reached the same address on its first try minutes later - scored as the model's, that is 0% for a dropped TCP connection with 69 steps of work already done. Unlike a 429 the servers are worth keeping: --keep-failed holds them, because a link that died mid-step is a machine to look at.

Operating systems

--os (spelled --image if you prefer) is repeatable, and each one it is given is another axis of the matrix: cells become models x suites x OSes, one OS at a time within each suite, and one leaderboard compares them. Given one --os, or none (the default is ubuntu-26-04-x64), nothing changes - the OS is recorded but stays out of the names and the tables, because one OS is not a comparison.

  • Friendly names. rocky-9 resolves against your own account's image list, so it finds rockylinux-9-x64 without you looking it up. A name that matches several images, or none, fails before a droplet exists and says which images it was choosing between. On --cloud aws the same name is answered by asking Rocky's own published account for its newest Rocky-9-EC2-Base-*, because an AMI is per region and there is no one catalogue to match against; every image the matrix may boot is resolved in preflight, so a third OS is not discovered hours in. An ami- id is taken as given, and a distribution neither vendor table knows is named that way.
  • A suite may decline. os_family says which package manager a suite's shell was written for, and a suite that cannot run on one of the run's images is left out of that column with a not run: line on the plan, rather than being graded on a dnf it never had. If that empties the matrix, the run stops instead of starting.
  • A suite may pin. A suite with its own image runs once, on that image, whatever --os says - it is testing something about that OS specifically.
  • The OS is part of a cell's identity: droplet names, directory names, the row labels, results.jsonl, the CSV and --resume all carry it.

--os and --host are mutually exclusive: an operating system is something the runner chooses when it creates a server, and your own servers already have one.

Databases

Seven products across three engines, and none of them is a drop-in for any other. --db (spelled --mysql if you prefer - the axis was called that when MySQL was all it had) is repeatable the same way --os is, and each value is another axis of the matrix:

uv run python dbrun.py run -m sonnet -s mysql-install \
    --db 9.7 --db percona-8.4 --db mariadb-11.4

# the same question, asked of the other two engines
uv run python dbrun.py run -m sonnet -s postgres-tune --db pg-18 --db pg-16
uv run python dbrun.py run -m sonnet -s mongodb-install --db mongodb-8.0 --db psmdb-8.0

Whatever you write is what the models are asked for. There is no list of versions to be on and no list of names to be on, because a benchmark whose vocabulary lags its subject cannot ask the only question worth asking the week MariaDB 12.3 ships.

written means
mysql-9.7, percona-8.4, mariadb-11.4 that flavour at that version
pg-18, postgres-17, mongodb-8.0, psmdb-8.0 the other two engines, in the words their own documentation uses
mariadb-12.3, mysql-8.0.36, pg-19 the same, at any version you like - nothing is checked against a list
percona, mariadb, postgres, psmdb that flavour at its newest version here, which is the one thing the built-in versions are still for
9.7, 8.4 that version of MySQL Community - a bare number is a MySQL, because that is where this axis started
"Percona Server 8.4", "MariaDB 11.4 rc" your words, reaching the task exactly as written
"FerretDB 2.0" a product this runner has never heard of

One token is the shorthand: mysql (Community), percona (Percona Server for MySQL), pxc (Percona XtraDB Cluster), mariadb, postgres, mongodb (Community), percona-mongodb, the vendors' own words (percona-server-for-mysql, percona-xtradb-cluster, maria-db, postgresql, pgsql, psmdb, mongodb-org) and a version if you want one. It resolves to the vendor's own name for the product, so --db psmdb-8.0 asks the model for "Percona Server for MongoDB 8.0".

Anything with a space in it is prose, kept word for word: --db "Percona Server 8.4" asks for exactly that and is recorded as percona-server-8.4. The runner reads three things out of it and edits nothing - a trailing version if the last word looks like one, which of the seven flavours it is, and which engine that flavour belongs to, because grading has to know both what a server should look like and which suites could honestly mark it. The words are read most-specific-first and the MySQL ones last, since "Percona Server for MongoDB" contains percona and is not a MySQL - and the product beats the vendor, since "Percona XtraDB Cluster" contains percona and is not Percona Server. A name with none of them in it is a flavour of its own, identified by grepping whatever the server says about itself - VERSION(), version(), db.version() - for the name itself, and offered to every suite on the axis, because which engine somebody meant by "FerretDB 2.0" is not something a runner should be guessing.

Two things are still refused, before anything is created: a value with no name in it, and two values that would be recorded as one row (percona-8.4 and "Percona 8.4" are both percona-8.4 on a leaderboard). Everything else the plan says rather than refuses, because a typo and a product released last Tuesday are the same string:

  db       myqsl-8.4
  myqsl-8.4: 'myqsl' is not a flavour anything here knows, so every suite on the
  --db axis - whichever database it is written for - will be asked for 'myqsl 8.4'
  and will identify what turns up by grepping the version it reports for 'myqsl'
  - did you mean mysql?

A version outside the built-in list gets the quieter version of that, naming the ones this runner has been kept up to date with. Both land above the confirmation prompt, which is the point: a database nobody can install grades every model at zero and looks exactly like every model failing.

One --db, or none, changes nothing about the matrix - the database is recorded on every row but stays out of the cell names and the table headings, which have nothing to distinguish. The default is one per engine (mysql-9.7, postgres-18, mongodb-8.0), so a bare dbrun run asks each suite for its own engine's. The plan still prints the db line whenever it is not the default, because one database is a decision even when it is not an axis.

  • The axis is really one axis per engine. --db percona-8.4 --db psmdb-8.0 does not cross every suite with both: nothing can grade both, and a MySQL suite asked for MongoDB would score zero for having been asked the wrong question. So a suite takes the values on its own engine's line, and an engine the run never named falls back to that engine's default rather than dropping its suites - --db mariadb-11.4 says which MySQL to compare and says nothing at all about MongoDB, and reading it as "no MongoDB today" would shorten the matrix by six suites without a word. That default is the engine's community build, except for a suite whose db_flavour does not include it: mysql-pxc-haproxy grades Percona XtraDB Cluster, community MySQL is a different product rather than a lesser version of the same one, and so a bare dbrun run asks it for pxc-8.4.

  • A suite may decline. db_flavour lists the flavours a suite's checks can honestly mark, and one the suite has not claimed is left out with a not run: line on the plan. mysql-tune's tuning stage rests on performance_schema.variables_info, which MariaDB has not got, so grading it there would score a MariaDB tuner zero for MariaDB not being MySQL. A flavour from outside the seven is nobody's to decline - no suite could have named a product that did not exist when it was written - so every suite on the axis is asked for it, and the warning above is where the doubt goes.

  • A suite may sit the axis out. No db_flavour at all means the suite is not on this axis: it runs once however many databases the run compares, and its rows record no database. That is mysql-restore, whose cloud-init installs whatever the distribution calls mysql-server. A --db that no cell in the run ends up using - because every selected suite sits the axis out, or because they are all on another engine's line - is not a refusal and gets its own line, since the alternative is a run that costs what was asked for and answers a different question:

      not used: mariadb-11.4 - no suite in this run is on the --db axis for mysql, so
      nothing here will be asked to install it
    

    Sitting the axis out is also not the way to write a suite for one product - a single-flavour db_flavour is, and it is what mysql-pxc-haproxy does: on the axis, the series comes from the flag, and --db mysql-9.7 is answered with not run: instead of being ignored. Note the asymmetry with os_family, where saying nothing means runs on anything - a suite that mentions the product but declares no flavour, or declares flavours and never mentions the product, is refused at load time rather than run.

  • The database is part of a cell's identity: droplet names, directory names, the row labels, results.jsonl, the CSV and --resume all carry it.

--db and --host are not mutually exclusive, unlike --os: the product is what the task asks the model to install, not what the server already boots.

Money

Two meters run at once, and the runner shows both.

  • Models. --max-cost USD is the run's ceiling: once spend reaches it no new cell starts, and the report says which were skipped. A cell already running is left alone - it is bounded by its own --max-cost-per-stage, and killing a cell halfway leaves a half-built server and a result worth nothing. Suites can set their own per-stage cap, and most of the ones here do. Self-hosted cells report no token cost, so a matrix made only of them never reaches any ceiling you set and the plan says as much instead of promising a guard that cannot fire. The droplets they drive are billed the same as anyone's.
  • Droplets. Billed by the hour whether or not anybody remembers them. Every droplet is tagged dbaai-runner, dbrun-<runid> and suite-<name>, and its id is written to events.jsonl before the model is handed anything. So whatever happens to the process:
uv run python dbrun.py reap                        # anything this runner ever left behind
uv run python dbrun.py reap --run 20260822-143512  # just that run's
uv run python dbrun.py reap --dry-run              # list them and stop

On AWS the same tags are key and value - dbaai-runner=<runid>, dbrun-suite=<name> and a Name you can read in the console - and a reap takes back the run's security group as well as its instances. Two things differ enough to be worth typing:

uv run python dbrun.py reap --cloud aws --region eu-west-2   # EC2 tags are per region
uv run python dbrun.py reap --cloud both                     # both accounts, one pass

reap looks where --cloud says and nowhere else, and the default is DigitalOcean (or $DBRUN_CLOUD) - so a reap that says nothing was found has answered about one cloud. The rate on an EC2 plan is arithmetic rather than a quote: EC2 has no call that will price an instance without the pricing endpoint and its own permissions, so the figure is the on-demand list price in us-east-1 plus the volume, and the plan says "about". Your region, your agreement and any savings plan all say otherwise.

--keep leaves the droplets up for a post-mortem. It is the one way this runner spends money after it exits, so it says so on the way in and on the way out.

--keep-failed is that aimed at the cell that needs it. Some failures cannot be diagnosed from the log - an sshd that refuses the login, a cloud-init that wedged, a disk that filled - and destroying the server destroys the evidence, which is how the same failure gets diagnosed twice. So a cell the runner broke keeps its servers and everything else is destroyed as usual: a matrix of forty can be run for the price of the one that needs looking at. A cell that ended because somebody stopped the run, because the cap was reached, or because the gateway stopped answering keeps nothing

  • there is no evidence on those servers, and two idle instances is a strange price to pay for a busy minute at a gateway. What it keeps, it tells you how to get into, as the cell fails and again as the last line of the run:
kept dbrun-20260825-074411-01-mysql-in 198.51.100.7 for diagnosis: could not get a root login
  ssh -i ~/.ssh/id_ed25519 ubuntu@198.51.100.7
  aws ec2 get-console-output --instance-id i-0a1b2c --region us-east-1 --output text
  dbrun.py reap --cloud aws --region us-east-1 --run 20260825-074411   # when you are done

Three details are the point of it. The login offered is the image's own user, not root - root is exactly what a boot failure on EC2 means the runner never managed to arrange, so the lease's own user is the one that will not work. The console output comes with it because it is the only thing that answers on an instance whose sshd never came up, and on an instance that never even reached running it is the whole of the diagnosis - so those are kept too, addressless as they are. And on AWS the run's security group stays behind with them, because it is what lets port 22 through: deleting it would leave servers that are still billing and can no longer be reached.

A cell the runner broke means a cell that never got servers, timed out, crashed, failed its setup, or lost a server's connection for good (ssh-lost) - not a model that scored badly. A model that fails its task is a result, and keeping servers for every bad answer would leave a forty-cell matrix holding forty servers. reap is how it all goes, and the exact command is printed.

--effort {off,minimal,low,medium,high,xhigh,max} is the one flag that changes what a run costs without changing anything the plan counts. It reaches every stage as dba.py --effort, asking the model to think before each step, and the gateway turns that one word into whatever the model upstream wants - so it is worth comparing models on rather than a knob only one of them has. The ends of the ladder are the levels most worth asking for: off and minimal for a reasoning model that would otherwise think its way through a systemctl check, xhigh and max for the ones that can spend more when it matters. The rungs between are a request rather than a promise - one measurement here, on qwen/qwen3.8-27b at a fixed seed, could not tell low from high in reasoning tokens, while off was honoured clearly - so a comparison whose whole point is the level should pin who serves it with --upstream below and check the token counts in the transcripts. Thinking is billed as output tokens: expect several times the model spend of the same matrix without it, and expect the suites' own max_cost ($1.50 to $5 a stage across the ones here) to be reached sooner, which shows up as stages stopping on cost rather than on a verdict - so the plan warns about it above the confirmation prompt, beside the cost caps. Unset is the absence of the flag rather than a word meaning none: a reasoning model goes on reasoning at the gateway's default, exactly as in every run recorded so far. The effort is part of each stage's recorded command line in grade.json, so a cell that thought is distinguishable from one that did not long after the run.

A thinking level can also be written into a model name - -m qwen/qwen3.8-27b:low, or the same suffix on any line of a --models-file - which sets it for that model alone and wins over --effort. :none is the other direction: it sits a run-wide --effort out for one model, which is how a reasoning model and one without the knob go into the same matrix. :none and :off are not the same thing and the distinction is worth keeping: :none asks for nothing and that cell's command line carries no --effort at all, while :off asks the model to stop thinking, which a reasoning model will otherwise do by default. "We made no request" and "we asked for none" are different runs and score differently. Like the gateway prefix, the suffix is part of what a cell is, so -m qwen/qwen3.8-27b:low -m qwen/qwen3.8-27b:high is two cells with their own droplets, their own directory and their own row - qwen/qwen3.8-27b:low and qwen/qwen3.8-27b:high, spelled the way -m takes them back, so a row can be re-run by copying its label. That is the comparison the suffix exists for: whether the thinking was worth what it cost, on the same task and the same OS, rather than across two runs a week apart. --resume counts the effort too, so resuming a matrix that ran at :high does not skip the :low half of it.

--upstream NAME says which of the providers behind the gateway may serve the models. It is OpenRouter's own idea: one id like qwen/qwen3.8-27b is offered by half a dozen companies who have the weights, and every request is routed to whichever of them looks best at that moment. They are not interchangeable. They quantise differently, cap the context differently, honour a reasoning level or ignore it, and come and go week to week - so an unpinned matrix can hand two cells of the same "model" to two different machines and record the difference as a fact about the model. Pinning is how a comparison holds still:

# only these two may serve it, AkashML first
uv run python dbrun.py run -m qwen/qwen3.8-27b -s mysql-install \
    --upstream AkashML --upstream Reka

# same thing, typed once - commas split like -m does
uv run python dbrun.py run -m qwen/qwen3.8-27b:high -s mysql-install \
    --upstream AkashML,Reka

Repeatable, comma-separated lists accepted, and the order is the order the gateway tries them in. The names are the gateway's own, spelled as its model page spells them (AkashML, Reka, DeepInfra, Chutes) - nothing here validates them, because who serves a model changes faster than any list in this repo could, and a name that matches nobody is a cell that fails saying so.

Naming any upstream turns fallbacks off, which is the point rather than a side effect: a pin that quietly fell through to whoever was up would not be holding anything. The price is that a cell whose named providers are busy, rate-limited or have dropped the model ends as api-error with its servers paid for, instead of being answered by somebody else - so the plan warns about it above the confirmation prompt, and --resume is how the cells that lost their provider get another go.

It reaches each stage as dba.py --upstream, one flag per name, so it is in the recorded command line like the effort and the patience, and run.json carries an upstream key naming who was allowed to answer - which is what makes two runs of one id comparable afterwards. Only OpenRouter routes across several providers; cells on DigitalOcean or a self-hosted server have one server behind each model and are sent nothing, so a mixed matrix keeps them and the plan says how many are outside the pin. A run where no model is on a routing gateway is refused up front rather than per cell, because there is nothing there to pin.

An unpinned run still says when a provider is letting it down. The gateway picks an upstream per request, and some of them stop replies at their own output cap - one served deepseek-v4.1-flash for a whole stage and cut 32 of its 131 replies off at about 2,200 of the 16,384 tokens asked for. Each of those is asked again, so no step fails and the stage just runs out of time. dba.py warns the first time a reply stops well short of the limit and again at the third from the same provider; the runner reads the same thing off the transcript as the stage runs and prints it beside the cell, repeats it under the stage's score, and gives the leaderboard a Providers that cut replies short section naming the provider, how many replies it cut, and the providers that served the same model cleanly - which is the --upstream to re-run with. dbrun report finds these in old runs too, off the transcripts they kept.

--rate-limit-wait SECONDS (default 900, or $DO_INFERENCE_RATE_LIMIT_WAIT) is how long one model request may spend waiting out 429s before the cell is given up as api-error. It is far longer than do-dba's own two minutes because the trade is different here: a cell that stops for a rate limit has servers running, an hour or more of stage timeout in hand, and usually a pile of work that dies with it - one recorded cell was refused at step 35 of 35, having spent $0.018 of its $1.50 cap, because the gateway wanted another sixty seconds and had ten left to give. The first fix for that was five minutes and a later run showed it was still short: qwen/qwen3.8-flash, asked for step 27 on a four-server PXC stage, waited 5s, 15s, 30s and 60s four times over, and the gateway wanted another 60s at 290s of the 300s budget - 26 executed steps and four paid-for servers thrown away for one more minute. Hence a quarter of an hour. Waiting costs nothing but wall clock: a 429 is not billed, the stage timeout still stops a cell whose gateway never comes back, and each wait is printed as it happens. A cell that waits this long on step after step runs out of stage rather than of patience, and timeout is a runner fault too, so neither ending is scored against the model. Free tiers are what the flag is for - --rate-limit-wait 1800 for a matrix of them - and 0 restores giving up on the first 429, which the plan warns about. It reaches every stage as dba.py --rate-limit-wait, so like the effort it is in the recorded command line, and a cell can be re-run exactly as patiently as it was run.

Runbooks

--runbook [SUITE=]PATH hands the models notes along with the task - the repository to add, the setting that turned out to be the one that mattered, the error that means the config is not being read. It is the flag for asking a different question than the rest of this page: not can these models work this out, but can they follow what somebody already worked out.

# one suite in the run, one file
uv run python dbrun.py run -m sonnet -s mysql-pxc-haproxy \
    --runbook runbooks/mysql-pxc-haproxy.md

# several suites: say which file is for which
uv run python dbrun.py run -m sonnet -s mysql-pxc-haproxy -s mysql-replication \
    --runbook mysql-pxc-haproxy=runbooks/mysql-pxc-haproxy.md

# or point at the directory and let it match by name
uv run python dbrun.py run -m sonnet --runbook runbooks/

The directory form is the one that scales: every suite in the run with a runbooks/<suite>.md gets it, and the suites without one are the ordinary case rather than an error. The bare-file form insists the run has exactly one suite instead of guessing from the file's name, because a six-suite matrix given one runbook is somebody who meant one of the other two spellings, and picking a suite for them would put a stranger's notes in front of five models.

It reaches each stage as dba.py --runbook-file, which puts the text in the system prompt in a RUNBOOK section under the task, framed as advice: the task is still what the model is asked for and judged on, the rules still win, and where the notes disagree with what a server actually says, the server is right. That framing is not decoration - the notes were written about other machines, and a model that treats a remembered wsrep_cluster_address as ground truth will confidently configure the wrong cluster.

It is not another axis, and it changes what the percentage means. A runbook is a decision about what every model in the run was shown, like the wording of the task itself, so it applies run-wide and is not compared against anything. What it costs is that a suite run with a good runbook measures a model following instructions rather than a model that can do the work - and for a suite whose difficulty is a handful of traps, a detailed enough runbook is a transcription test. Both numbers are worth having. Pooling them into one table is not, so every place a score is recorded says which it is:

  • run.json gets a runbooks list: the suite, the path, the character count and a short digest of the text.
  • every row of results.jsonl and the runbook column of results.csv carry <file>.md@<digest> for a helped cell and nothing for an unhelped one - because a spreadsheet pooled from several runs is exactly where that difference disappears.
  • the leaderboard header says it in words, above the matrix, including what it does to the column.
  • each stage directory keeps the notes as runbook.md, in full, beside the task.md. Prose gets edited between runs - usually in the direction of giving more away - so a path alone could not say what this model was told, which is also why the digest is on the row.

--resume refuses to change it. The OS, the database and the effort fill in an older record's silences; this one cannot, because the silence means the opposite - a row with no runbook is a cell that was given none, which is a fact about the measurement rather than a gap in the record. Resuming with different notes, or with none where there were some, would build one table out of two experiments, so it stops and prints what changed. Repeat the --runbook flags that run used, or start a run of its own.

The notes have to be prose somebody stands behind, which in practice means runs that worked: runbooks/mysql-pxc-haproxy.md here was written from a run where three models each scored 100% on all three stages, quoting the configs that were on the servers when the grader passed them - and naming, wherever those three disagreed, all of the shapes that passed rather than flattening them into one right answer. A second run put the same suite on Amazon Linux, where one cell of three took every mark and two lost 57 marks to a /root/.my.cnf that outranked the grader's MYSQL_PWD; that became a section of the same file, because the packaging changes and the cluster does not. Nothing in it is marked untested, which was not true of the version written before the suite grew its encryption and Synced checks, and it is the reason to rewrite the file when the suite moves: notes that teach a model to turn encryption off are worse than no notes. An empty file is refused rather than treated as no runbook, and the plan (--dry-run) prints each file, its size and how many cells will be handed it, because 73,000 characters in the system prompt is re-sent and re-paid at every step.

That last number is why the same notes exist twice. runbooks/brief/mysql-pxc-haproxy.md is the same file with the narrative taken out - 26,000 characters instead of 73,000, keeping only what decides a mark: the paths that differ between the two package managers, the config that passed, the options that are not server options, one health check that can see a desynced node, and the symptom-to-cause table. Because the directory form looks for <suite>.md, the two are chosen by which directory you point at:

--runbook runbooks/            the long one, evidence and all
--runbook runbooks/brief/      the same advice, a third of the tokens

Both are recorded by digest, so a run says which of them a cell was given.

runbooks/brief/mongodb-replication.md was written the same way and exists only in the short form - 18,000 characters out of thirty-five graded attempts, nineteen of which took every mark. Most of it is one fact: MongoDB 8.x refuses to start on the kernels these images ship, the packaged unit's GLIBC_TUNABLES=glibc.pthread.rseq=0 is what trips the guard rather than the kernel itself, and clearing it in a drop-in is the whole fix. The suites that have no file in a directory are not an error, so --runbook runbooks/brief/ on a matrix of both suites hands each cell the notes for its own suite and nothing to the rest.

runbooks/brief/mysql-replication-openbao.md provides reusable instructions for Percona 8.4 and 9.7 with OpenBao: version selection, scoped tokens, native keyring startup, encrypted TLS replication, and restart verification. Addresses and credentials are parameters; the brief contains no model names or past-run references. Select it with --runbook runbooks/brief/mysql-replication-openbao.md for this suite, or --runbook runbooks/brief/ for directory matching across suites. The detailed recipe remains at runbooks/mysql-replication-openbao.md.

Writing a suite

A suite is a TOML file in suites/. The short form is one task:

name = "mysql-install"
description = "Install MySQL 9.7, a database, and a login user for it"
nodes = 1
max_steps = 25
max_cost = 1.5

task = '''
Install MySQL 9.7 from the vendor's own apt repository on this server.
Create a database called `app` and a login user called `app_rw`.
'''

  [[check]]
  name = "mysql is running and starts at boot"
  weight = 2
  command = "systemctl is-active --quiet mysql && systemctl is-enabled --quiet mysql"

  [[check]]
  name = "the app database exists"
  command = "mysql -N -B -e \"SHOW DATABASES\" | grep -qx app"

The long form is [[stage]] tables with [[stage.check]] under each, which is what a suite that reuses its droplets looks like - see suites/mysql-repl-upgrade.toml.

Suite keys

key meaning
name the suite's name, and part of its droplets' hostnames
description one line, shown in listings
nodes = N N servers, unnamed. The harness calls them node1, node2, ... and the model works out the roles
names = [...] N servers whose roles the suite has decided; the model is told these names and follows them
size / image / region override the run's droplet settings for this suite. An image here takes the suite off the --os axis: it runs once, on that image. On --cloud aws a droplet slug is read as the capacity it names and a friendly image name as the vendor's newest, so a suite written for DigitalOcean still asks for what it meant
os_family which OS families the suite's shell is written for - "debian" (apt, so Ubuntu and Debian both) or "rhel" (dnf), one or a list. --os leaves the suite out of an image it has not claimed
db_flavour which databases the suite's checks can honestly mark - "mysql", "percona", "pxc", "mariadb", "postgres", "mongodb", "percona-mongodb", one or a list. All of them must be the same engine, since one suite's checks cannot mark two. Saying nothing takes the suite off the --db axis rather than putting it on every column of it; naming exactly one keeps it on the axis for the version and refuses every other product by name. Spelled mysql_flavour in a suite written before there was anything but MySQL, and still read
cloud_init / cloud_init_file user-data for the droplets, for a fixture that should be in place before the model logs in
stop_on_fail end the cell if a stage does not pass (off by default)

Stage keys

key meaning
id short, used as the directory name and in the tables
task / task_file the prose the model is given. This is the specification: a vague task grades vaguely
title one line for the progress display
setup commands the runner runs before the model is told anything - the fixture the task starts from. A failure here is filed as setup-failed and never blamed on the model
setup_host which servers the setup runs on (default: all)
weight this stage's share of the cell's score
max_steps steps the model gets (default 30)
max_cost dollars this one task may spend
timeout wall clock for the whole stage, in seconds (default 3600)
command_timeout per command on the server, in seconds (default 300)

Check keys

key meaning
name what the tables call it
command shell, run by the runner over its own connection. Exit 0 is a pass
host "*" every server must pass (the default), "one" exactly one must, "any" at least one must, or a node's name
weight its share of the stage score (default 1)
expect_exit when success is not exit 0
contains / not_contains when the exit code is not the answer - SHOW REPLICA STATUS exits 0 either way
timeout seconds (default 60)

host = "one" and host = "any" are what make a role-agnostic suite gradeable: exactly one of these servers is a replica that is caught up, without the suite dictating which one the model should have chosen. The nodes that satisfied it are recorded, so the report can say which way round the model built it.

Use "one" for a role only one server can hold and "any" for one that several can. The difference is worth getting right: two servers that both take writes and both have binary logging on are two standalone databases, and under "any" they would collect the marks for a replication pair. Both replication suites use "one" for every check that names a role.

Every check and setup command is told about the lease it is running in:

DBRUN_SELF              this node's label
DBRUN_NODES             every label, space separated
DBRUN_PEERS             the other labels
DBRUN_HOST_<LABEL>      that node's public address
DBRUN_PRIVATE_<LABEL>   its private address (the public one if it has none)

which is how a check can require that replication reads from the peer's private address rather than merely from somewhere.

A suite on the --db axis is told which database this cell asked for as well:

DBRUN_DB                mysql-9.7, percona-8.4, pg-18, psmdb-8.0, ferretdb-2.0
DBRUN_DB_ENGINE         mysql | postgres | mongodb | empty for a product from outside
DBRUN_DB_FLAVOUR        the flavour: percona | mariadb | postgres | percona-mongodb | ...
DBRUN_DB_VERSION        9.7          empty if none was asked for
DBRUN_DB_NAME           Percona Server for MySQL      (the vendor's words, or yours)
DBRUN_DB_TITLE          Percona Server for MySQL 8.4  (exactly what the task asks for)
DBRUN_DB_VERSION_RE     ^9\.7        what to grep the reported version for, empty for "any"
DBRUN_DB_MATCH          percona      what the version and what it says about itself must match
DBRUN_DB_REJECT         percona|mariadb   what it must not, or empty

Every one of them except _ENGINE is exported under DBRUN_MYSQL_* as well, because the axis had that name when MySQL was all it had and a suite written then still greps $DBRUN_MYSQL_MATCH. Nine variables of duplication is a cheap price for an unset $DBRUN_MYSQL_VERSION_RE being impossible: an empty pattern is a version check that passes for every version.

A check written against those needs nothing else to grade a product this runner has never heard of: an empty _REJECT and an empty _VERSION_RE both mean "no constraint", which is what grep -q "" does anyway, so --db "MariaDB Enterprise" grades as MariaDB at any version without a line of the suite changing.

Patterns and not only names, because every suite that cares asks the same two questions - is this the version that was asked for, and is it the right vendor - and neither is a string comparison. VERSION() is not enough on its own: Percona's is a bare 8.4.3-3 and names the vendor only in @@version_comment, and Community is the one product that has to be identified by what it is not. So a check greps both fields together:

version=$(q "SELECT VERSION()"); comment=$(q "SELECT @@version_comment")
echo "version=${version:-none} comment=${comment:-none}"
[ -n "$version" ] || exit 1
printf '%s\n' "$version" | grep -q "$DBRUN_DB_VERSION_RE" || exit 1
printf '%s %s\n' "$version" "$comment" | grep -Eqi "$DBRUN_DB_MATCH" || exit 1
[ -z "$DBRUN_DB_REJECT" ] || ! printf '%s %s\n' "$version" "$comment" | grep -Eqi "$DBRUN_DB_REJECT"

The same three lines carry to the other engines, and only the two fields change. PostgreSQL is one project, so version() and server_version are all there is to read. MongoDB is the awkward one: Percona tracks upstream's numbering exactly, so db.version() cannot tell a community server from a Percona one and the mongodb suites grep the version together with the installed package names - which is the only place either vendor writes itself down plainly.

In the task - and only there - {{db}} writes the product in as it was asked for, with {{db_name}}, {{db_version}}, {{db_flavour}} and {{db_engine}} for the parts. {{mysql}}, {{mysql_name}}, {{mysql_version}} and {{mysql_flavour}} still render the same things. Commands are given the variables above instead, deliberately: braces in a check command belong to whatever is going to read them - docker ps --format '{{.Names}}', a kubectl -o go-template, a jq filter - and nothing here may touch them. A {{db_ver}} that nothing replaces is refused at load time rather than handed to a model as literal braces.

Give a check a product-neutral name. Where the points went groups by check name, and it answers as the database and version asked for is one comparable row where it answers as MySQL 9.7 would be three that are not. It holds even for a suite with one flavour on its axis, because the version is still a column: mysql-pxc-haproxy grades the series and calls the check all three nodes answer as the cluster that was asked for, which is one row across a run of pxc-8.0 and a run of pxc-8.4. A suite off the axis altogether - mysql-restore - has one product in every cell and may name it.

Checks are read-only by convention, and the starter suites break that convention deliberately twice: one check writes a marker row on whichever server accepts writes, and the next looks for that row on a server that refuses them. Checks run in the order they are written, which is what makes that pair work - and a write that arrives is the only real evidence that replication replicates.

mysql-pxc-haproxy breaks it further and on purpose: two of its checks break the node HAProxy is serving, and the checks after each of them grade the failover. The first desyncs it and leaves it running, which is the failure a health check that only completes a MySQL handshake cannot see; the second stops the database outright. Both are host = "one", so exactly one node may satisfy them - the suite picks the victim by asking the proxy who it is talking to. The guards differ because the failures do: the stop can require a whole cluster of three, while a desynced node is still a member, so that one is guarded by a primary-key insert Galera certifies cluster-wide and exactly one node wins. The suite puts its own desync back and waits for three Synced nodes before it stops anything.

A stage with no checks is refused at load time. Grading a stage on the model's own account of itself is the one thing the harness under test already refuses to do.

The starter suites

suite servers OS db what it is
mysql-install 1 debian all three MySQLs install the database, a schema, a scoped user, not open to the world
mysql-restore 1 debian - a dropped database and last night's dump - cloud-init and setup build the situation
mysql-replication 2 debian, rhel all three MySQLs replication over the private network, roles left to the model
mysql-replication-openbao 3 debian, rhel percona, mariadb primary, replica, and OpenBao: native remote keys, encrypted InnoDB tables and logs, verified TLS, and fresh key reads after database restarts
mysql-repl-upgrade 2 debian, rhel mysql two stages on the same pair: replicate, then move to Percona without breaking it
mysql-tune 1 debian mysql, percona two stages: install it, then size it to the machine - graded as ratios of the server's own memory, and on whether the settings would survive a restart
mysql-group-replication 3 debian, rhel mysql, percona single-primary group replication across three servers: three ONLINE members as seen by every one of them, group traffic on the private network, a write crossing to both others
mysql-pxc-haproxy 4 debian, rhel pxc Percona XtraDB Cluster on three servers behind HAProxy on a fourth, with the cluster and state-transfer traffic encrypted rather than switched off, then the benchmark desyncs the node HAProxy was serving and later stops it and grades where the traffic went both times
mysql-pxc-clone 4 debian, rhel pxc three stages that grow a cluster rather than build one: one encrypted Percona XtraDB Cluster node, then HammerDB on a fourth server loading a hundred TPROC-C warehouses into it over TLS, then two more nodes added by clone SST - graded on the method, on the ten gigabytes arriving whole, and on all three nodes trusting one certificate authority kept out of the data directory
postgres-install 1 debian postgres the same task as mysql-install, in PostgreSQL's vocabulary: a cluster, a database, a login role that is not a superuser
postgres-tune 1 debian postgres install it, then size it - shared_buffers as a ratio of the machine, and pending_restart to catch an ALTER SYSTEM that was never restarted into
postgres-replication 2 debian, rhel postgres streaming replication over the private network, roles left to the model
postgres-patroni-ha 5 debian, rhel postgres three stages: PostgreSQL under Patroni with etcd on three nodes, then PgBouncer on each and two HAProxies sharing a keepalived address - the benchmark stops etcd on the leader's node, HAProxy on the address holder and Patroni on the leader, and grades that the leader stayed, the address moved and the proxies followed the new leader - then the stopped node brought back as a replica
mongodb-install 1 debian mongodb, psmdb install it, turn access control on, a database with a collection in it, a user scoped to it
mongodb-tune 1 debian mongodb, psmdb install it, then tune the server and its host: the WiredTiger cache, transparent huge pages, swappiness - and each of those in a way that outlives a reboot
mongodb-replication 2 debian, rhel mongodb, psmdb a two-member replica set with a keyfile, over the private network, roles left to the model

They are the tasks this benchmark was built for, and they are not cheap: the runs they were written from took 40-70 steps and half an hour each. Start with -s mysql-install and one model.

The OS column is each suite's os_family, and it is what --os rocky-9 picks up: the install and tune suites are written around a vendor apt repository, so they are left out of an rhel column rather than failed on it. The replication suites grade a thing the database does rather than a way of installing it, and run on either family.

The db column is each suite's db_flavour, and what --db mariadb-11.4 or --db psmdb-8.0 picks up. Three suites per engine, deliberately the same three tasks: install, tune and replication are written check for check against each other, so a model's PostgreSQL score is readable beside its MySQL one and the difference is the database rather than the marking. On the MySQL line, mysql-install and mysql-replication mark all three flavours; mysql-repl-upgrade is Community-only because becoming Percona is the task; mysql-tune and mysql-group-replication need something MariaDB has not got. On the MongoDB line every suite marks both builds, because what tells them apart is the installed package name and not a check - which makes --db mongodb-8.0 --db psmdb-8.0 the cleanest comparison in the whole set. PostgreSQL has one flavour: what Percona distributes is PostgreSQL, with the same version(), so a second column would be noise - the axis worth comparing there is the major version, and --db pg-18 --db pg-16 is two repositories, two data directories and one set of checks. mysql-pxc-haproxy names one flavour for the same reason from the other direction: pxc is a product of its own - not community MySQL, which has no wsrep, and not Percona Server, which is the same vendor's other product - so the axis there is the series, --db pxc-8.0 against --db pxc-8.4, and anything that is not a cluster is refused on the plan rather than marked at zero. Only mysql-restore has no db_flavour at all and runs once whatever the run compares, because the dump is what it is.

Encrypted replication with OpenBao

mysql-replication-openbao installs a primary, a replica, and a dedicated OpenBao server. It compares Percona Server's component_keyring_vault (or the legacy keyring_vault plugin) with MariaDB's hashicorp_key_management. Use a MariaDB build that supplies that plugin. Oracle MySQL Community is excluded because it does not ship a native Vault backend; Oracle's keyring_hashicorp requires Enterprise components.

uv run python dbrun.py run --models-file models.txt \
  -s mysql-replication-openbao --db percona-8.4 --db mariadb-11.4

Each cell uses three servers. OpenBao runs with persistent storage and verified HTTPS on its private address. Each database gets a separate KV-v2 mount and a scoped token. The task specifies the CA, audit log, and inspection-token paths that the checks need; the inspection token reads key metadata, not key values. An installation that only copies a local keyring into OpenBao does not satisfy the task.

The grader writes a marker into an encrypted InnoDB table and waits for it on the replica. It then restarts both database services, verifies that the marker remains readable, and requires successful key-read request/response pairs from both private database addresses in OpenBao's new audit records. Those requests must use non-root tokens and the correct per-node mounts. It also checks table and log encryption settings after restart, inspects persisted tablespace files for marker plaintext, and sends a second marker through replication. It does not restart or seal OpenBao; recovery of the key server itself is outside this suite.

Run the offline checks with uv run python run_tests.py suites openbao. They cover both database flavors, legacy Percona keyrings, stale or failed audit responses, deleted keys, plaintext tables and logs, and replication that fails after restart.

Reading the results

uv run python dbrun.py report                     # rebuild the last run's tables
uv run python dbrun.py report 20260822-143512

results.jsonl is authoritative and append-only; the leaderboard and the CSV are derived from it and can be rebuilt at any time, including after a scoring change.

The leaderboard has five sections: the score matrix, Said done, was not, Where the points went (every check and how many attempts satisfied it - the fastest way to find a check that is wrong rather than hard), Cells that did not get a fair run, Providers that cut replies short where any did, and The tasks - the prose the models were given, quoted once per distinct wording, because a suite's paragraphs get edited between runs and a percentage is not a result without them.

A run that compared operating systems gets a sixth, By operating system, and one that compared databases gets By database. Either way every table that names a cell grows that axis in its row labels: rows read `sonnet @ rockylinux-9` or `sonnet / mariadb-11.4`, or both at once, and Where the points went counts each check per OS and per database rather than pooling them - because a check that passes on one image and not the other, or on MySQL and not on MariaDB, is the whole reason to run both. The tasks splits the same way, naming the database whenever a stage's {{db}} made two wordings out of one paragraph.

Layout

dbrun.py              entry point
dba_runner/
  cli.py              subcommands, screens, preflight
  matrix.py           cells, the thread pool, the budget, cleanup
  suites.py           the TOML format and its loader
  products.py         the --db axis: engines, flavours, versions, what a check is told
  provision.py        doctl, leases, droplets, reap
  aws.py              the same on EC2: instance types, AMIs, security groups
  harness.py          running dba.py as a subprocess and reading it back
  grade.py            the checks, over the runner's own SSH connection
  results.py          results.jsonl, the leaderboard, the CSV
  runbooks.py         --runbook: which suite is handed which notes
suites/               the tests
runbooks/             notes for a suite, written from a run that worked
tests/                offline suites - no network, no token, no money
run_tests.py          runs them all, one line per suite

do_dba is imported for its SSH transport and terminal handling, and dba.py is run as a subprocess: the thing being measured is allowed to fail badly, and in a subprocess a wedged model costs one cell and a timeout.

Tests

uv run python run_tests.py
uv run python run_tests.py suites grade

Nothing there touches the network, a DigitalOcean account, an AWS account or a model. doctl and aws are fakes that answer from a script, the bench is a fake dba.py that writes a transcript, and the servers are fakes that answer the checks - so the whole matrix, including a cell that times out and a droplet that leaks, runs on a laptop with no token and no money.

ec2 is the suite for the other cloud, and it is mostly about the three things with no DigitalOcean counterpart: a security group that outlives its instances (deleted on the retry after termination catches up, and reported as leaked when it does not), the root login a vendor image refuses on purpose, and a size and an image named in this project's vocabulary rather than EC2's - s-4vcpu-8gb becoming the cheapest type that holds it, rocky-9 becoming a lookup against Rocky's own account.

runbook is mostly about the record rather than the reading. Resolving --runbook is a page of parsing with seven ways to get it wrong, and those are checked; the longer half runs a two-suite matrix with one suite helped and then looks for that difference in every artifact that outlives the run - the row, the line in results.jsonl, the CSV column, the leaderboard's header - because a helped cell and an unhelped one scoring the same 100% are not the same result, and nothing in the number says so. It also pins the resume refusal, including that an edited runbook is a different runbook and that a run which wrote no run.json resumes anyway.

Six of them grade shell rather than Python. suites loads all fourteen shipped suites and puts every check and setup script through sh -n, renders every task once per flavour the suite claims and fails on a {{ that survived, and asserts that each suite accepts its own engine's default database - a suite that refused it would be silently skipped by an ordinary run. tune runs mysql-tune's tuning checks against nine fake servers - one tuned well, one untouched, one whose settings would not survive a restart, one with performance_schema switched off - and asserts which checks should fail on each. repl grades both MySQL replication suites against twelve pairs of servers that are really directories with a stub mysql on their PATH: a correct pair, the same pair the other way round, an older server that spells every field Slave_, a pair replicating over the public internet, a replica that still takes writes, a faultless pair of MariaDB, two standalone databases, a dead peer. Three of the twelve are the --db axis: a Percona pair graded as a Community run and a Community pair graded as a Percona one both fail on the version check alone, and the Percona pair graded as the Percona run it is comes out clean. The suite file is byte-identical across all three - only the exported environment differs, which is the whole claim the axis makes.

That stub is an executable and not a shell function on purpose, because the client is where a suite goes wrong: this one refuses \G when the statement arrives through -e, exactly as the real one does, and a helper that asked that way scored a working replica zero on every model of two whole runs. One scenario grades a correct pair with that old helper put back and asserts it still comes out at the 0.625 the run recorded - so the day the stub stops being awkward, the test says so instead of passing everything.

group and pxc are the same idea on three and four servers. group grades mysql-group-replication against sixteen groups - a correct single-primary one, a tarball install with the client on nobody's $PATH, a group talking over the public addresses, two primaries where there should be one, a member still joining. pxc grades mysql-pxc-haproxy, and its fake cluster is real enough to fail over: the shared directory is the cluster, a node's data is the cluster's data only while it is a member, and the stub answers as HAProxy on the proxy's address. So stage two breaks whichever node the proxy was serving - the suite chooses it, this file does not - twice, in the two ways that are not the same: it desyncs the node and leaves it running, then later stops it outright, and the failover checks pass because two other machines really had the write and the proxy really moved. The stub proxy has both health checks in it, one that reads the served node's state and one that only opens a connection, which is the whole difference the desync exists to measure. Six clusters are graded, eight more only through the first stage, and five are built to tempt a check into passing with only the checks that could be fooled run against them: a proxy balancing across all three, a cluster that had already lost a member, a cluster with nothing wrong with it, a proxy whose health check proves nothing but that the port answers, and one serving a node that is feeding a state transfer. Two more grade a check by its evidence and not only its verdict: the axis - three nodes answering 8.0.46, marked twice, refused by a run that asked for 8.4 and accepted by a run that asked for 8.0, because a check that greps a series written into the suite file would pass the first and fail the second - and the encryption, where the three switches that can put cluster traffic in clear are turned off one node apiece, so a node's line has to name which of the three it was.

The scenario that keeps that one honest is bootstrapped: a cluster brought back from the node that was behind comes up Synced, three strong, one backend served - and has silently lost the write made while it was down. Everything in stage three passes except the check that counts the rows, which is the only reason that check exists.

clone grades mysql-pxc-clone the same way, and the shared directory being the cluster is what makes it possible at all: a joiner that becomes a member of the component sees the hundred warehouses that were loaded before it existed, which is a clone modelled as the thing a clone actually achieves. Five arcs run all three stages - one done properly, one done differently in every way the suite is not allowed to mind, one joined by the vendor's default backup transfer instead, one whose joiners kept the certificates their own installer generated, and one that built the whole cluster in stage one before there was anything to clone - and thirty-odd fixtures grade a single stage against servers built to tempt one check into the wrong answer. The certificate stubs are the interesting ones: a certificate here is a file naming the authority that signed it, so openssl verify and openssl x509 -fingerprint both answer, and the three ways the vendor's default layout fails are separable - files left in the data directory, the same files named relatively so that nothing in the configuration says they are inside it, and files moved out with the installer's self-signed certificate still in them. The load has a fixture that lies: a hundred rows in warehouse on top of a single warehouse of data, with the aggregates and the reported size made to agree. Every check that counts in bulk passes it. The hundredth warehouse is the only one that does not, and that is the only reason that check is in the suite.

patroni grades postgres-patroni-ha, and is the first stub-server grader that is not MySQL: a stub psql that is PostgreSQL on a node and HAProxy-then-PgBouncer on a proxy port, and a stub curl that is Patroni's REST API and etcd's gateway. The shared directory is the world again, and systemctl is where it moves: stopping Patroni on the leader elects the next member that can see etcd's quorum, stopping etcd on a leader whose Patroni was only given its own node's member demotes it, and stopping HAProxy on the address holder hands the virtual IP over if keepalived was tracking it. Six arcs are marked through the failures - one done properly, one done differently in every way the suite may not mind, and one each for a Patroni that knows only its local etcd, keepalived left on multicast, keepalived that never watches HAProxy, and HAProxy that bypasses PgBouncer. Two more recoveries start from a copy of the correct arc's servers as the failures left them, one with the stopped node left down and one with it brought back by pg_ctl. A handful of refusals check that each injection refuses to break a second thing and that each failover check refuses a cluster nobody broke. local-etcd is the one the etcd failure exists for: the failover it causes is fast and clean, and the only thing wrong is that it happened.

The PostgreSQL and MongoDB tune and replication suites are still the gap in this battery. They are loaded, rendered and parsed by suites, but nothing offline runs their checks against a fake server yet, so their marking has been argued rather than tested. The psql stub above is a start; mongosh would need one of the same kind, awkward where the real client is awkward, which for mongosh means an adminCommand that answers {ok: 0} instead of failing.

A suite that mis-marks does not look like a broken suite. It looks like a model that failed.

Notes

  • runs/ holds real credentials. The bench's secrets.json and the servers' generated passwords end up under it. It is in .gitignore and should stay there.
  • Its own known_hosts. Per run, inside the run directory. Both clouds reissue addresses, and neither your ~/.ssh/known_hosts filling with dead droplets nor a recycled address looking like an attack is wanted.
  • Error reading SSH protocol banner is not a failure. Ubuntu socket-activates sshd, so port 22 answers while the host keys are still being generated and the first connection closes where the banner should be. The runner retries for up to 7 minutes and says what happened either way, so paramiko's own traceback for it is kept off stderr - DBRUN_SSH_DEBUG=1 puts it back when the connection is the thing you are debugging.
  • The retries back off, because sshd punishes the ones that do not. OpenSSH from 9.8 keeps a penalty per source address and refuses connections from one whose attempts keep failing - 15 seconds at first, up to 10 minutes if it continues, and every refusal inside the penalty extends it. A loop knocking every 5 seconds can therefore lock itself out of a machine that is working, for longer than it was ever willing to wait, and the run dies with could not get a root login. So the wait doubles after each failure up to 45 seconds. On EC2 there is a better answer than waiting: the runner asks for the instance's own two status checks first and only knocks once they pass, so most of the time there is no penalty to lapse. That check is a courtesy and never the verdict - if EC2 has not decided within 2 minutes, or will not answer, the login goes ahead and decides.
  • Every wait says which wait it is and how long it has been. Booting a cell is four waits in a row - the instance reaching running with an address, port 22 answering, the root login going through, cloud-init finishing - and any of them can legitimately take minutes. Until they reported themselves the whole thing was one silent gap, which reads as a hung runner and got read that way: EC2 reaching running takes about 8 seconds, but the 7 minutes of nothing afterwards looked like the cloud being slow to say so. Each wait now prints a line every 20 seconds naming what it is still waiting for and its elapsed seconds, rate-limited so four cells booting at once do not bury the log. The four share one BOOT_TIMEOUT between them rather than starting a fresh one each, so a cell cannot spend 7 minutes per wait and a slow port no longer buys the login extra time.
  • cloud-init is waited for. A freshly booted image is still installing things for a minute or two, holding the apt or dnf lock, and a model whose first step is apt-get would meet "Could not get lock" through no fault of its own. That is the difference between grading a model and grading a race.
  • Multi-stage suites rely on do-dba's server-side credential store. Stage one's generated passwords stay on the server in /etc/profile.d/dba-secrets.sh, which is how stage two can log in at all. --no-server-secrets in the bench would break that, so the runner does not pass it.
  • Nothing in a cell may reboot a node. The grader's connection to a node is opened once and used for every stage, and nothing on this side of it survives the far end going down - so do-dba's guard refuses reboot, shutdown, poweroff, halt, init 6, systemctl reboot and a character into /proc/sysrq-trigger outright instead of asking about them: --mode unattended answers every question yes, so a question is not a limit. The suites ask for the smaller thing that is actually being graded - a service that "comes up on its own when the machine boots", which systemctl is-enabled answers without trying it - and where a failure is wanted, it is a process that goes down: the PXC failover stage stops the database on the routed node. One cell paid for this rule. A systemctl reboot on db3 in stage 1 made every later check on db3 report unreachable: the SSH connection is not open, and a PXC cluster that was up and Synced on all three nodes scored zero on three stages.
  • A connection that dies anyway is waited out. Plenty besides a reboot can take one away - sshd restarted, the firewall rewritten, a node that simply went. When one dies the runner waits up to 2.5 minutes for that node to answer again and logs back in, saying so in the cell's log; a node that really is gone costs the cell its checks, with its own name in the reason, rather than costing every check after it.
  • A crash is recorded in the bench's words, not paramiko's. When dba.py leaves no transcript the runner quotes it, and what it quotes now is the last complaint dba.py printed - could not connect to 18.207.93.154:22 - Error reading SSH protocol banner - rather than the last six lines of stderr, which for anything SSH is a transport thread's frames from inside site-packages. The leaderboard prints the first line of that field, so it used to read crashed: Traceback (most recent call last):, which names neither the host nor the reason. Both files are kept whole beside the cell either way.

About

Matrix runner for the dbaai_bench DBA harness: models x suites x operating systems x databases, one droplet per cell, graded over the runner's own SSH connection

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages