Run performance benchmarks on PRs with a per-platform verdict - #374
Conversation
|
An automated preview of the documentation is available at https://374.corosio.prtest3.cppalliance.org/index.html If more commits are pushed to the pull request, the docs will rebuild at the same URL. 2026-09-29 21:15:19 UTC |
|
GCOVR code coverage report https://374.corosio.prtest3.cppalliance.org/gcovr/index.html Build time: 2026-09-30 14:42:15 UTC |
Benchmark report
linux — full results (226 benchmarks, 5 iterations, 1.5s each)accept_churn
fan_out
http_server
local_socket_latency
local_socket_throughput
socket_latency
socket_throughput
windows — full results (77 benchmarks, 5 iterations, 1.5s each)accept_churn
fan_out
http_server
socket_latency
socket_throughput
macos — full results (113 benchmarks, 5 iterations, 1.5s each)accept_churn
fan_out
http_server
local_socket_latency
local_socket_throughput
socket_latency
socket_throughput
Run & raw JSON artifacts · flag rule: |median Δ| > max(2%, 3×CV of the noisier side) · advisory only |
94313cf to
aeb6cec
Compare
Add a benchmarks workflow that builds the PR head and its merge-base, runs the full bench suite in interleaved base/head passes on dedicated self-hosted runners (Linux epoll+uring, Windows iocp, macOS kqueue), and posts one advisory PR comment, updated in place on each push. A row is flagged only when the median of its paired per-iteration deltas exceeds the larger of a fixed minimum effect size and three times the sample spread of the noisier side, so the noise floor is measured live on the same machine rather than maintained as stored calibration. Iterations alternate starting side so linear drift cancels. A workflow_dispatch A/A mode benchmarks the merge-base against itself to audit that flag rule whenever the hardware or the suite changes. The comparison and report generation live in a stdlib-only script, .github/bench/compare.py. It matches benchmark suites dynamically: benchmarks present on only one side are reported as added or removed rather than compared, unrecognized metrics degrade to an unsupported list, and malformed input files are warned about and skipped — the report never fails a job over suite shape. Untrusted code never reaches the machines without review: the pull_request_target gate runs collaborator PRs automatically and everyone else's only when a maintainer applies the benchmark label, which is consumed on use and stripped on later pushes so each approval covers exactly the diff that was reviewed. Head and base SHAs are pinned at gate time, and all tooling (this script and the composite actions) executes from the trusted base commit, never from the pull request.
aeb6cec to
6ba2c7a
Compare
Implements #343: every PR touching
include/,src/, orbench/gets an advisory benchmark comment — head vs merge-base, per platform, per backend — posted by the workflow and updated in place on each push.How it works
workflow_dispatchoverrides.workflow_dispatch mode=aabenchmarks the merge-base against itself to audit the rule whenever hardware or the suite changes.bench-<platform>artifacts; the comment links the run.Security model (
pull_request_target)compare.py, and composite actions always execute from the trusted base commit — never from the PR.benchmarklabel, which is consumed on use and stripped on later pushes — each approval covers exactly the reviewed diff. The label test matches the label that fired the event (no stale-label replay), and a failed label-consume fails the gate closed.Validation
compare.py: 24 local unit tests; validated against real bench output, including a doctored −10 % regression that flags correctly at N=5.Reviewer notes
prdispatch input was deliberately dropped — a manual dispatch reports to the workflow step summary only; giving dispatch the power to write comments to arbitrary PRs seemed wrong..github/bench/**exercises the base version of the tooling; tooling changes get their live test via branch dispatch or post-merge.pull_request_targetreads the base branch, which gains the workflow on merge). The comment below is from the final dress-rehearsal run at the pre-squash head; the code is unchanged since, only consolidated.Post-merge follow-ups
mode=aaonce per platform from develop and note results in Run performance benchmarks on PRs and report the impact #343 before closing it.