Skip to content

greenbench: an energy study of three sorts in C, Python and Node.js (time measured, energy pending a bare-metal run) - #2

Merged
fasharif merged 51 commits into
mainfrom
feature/greenbench
Oct 2, 2026
Merged

fasharif merged 51 commits into
mainfrom
feature/greenbench

Conversation

@fasharif

@fasharif fasharif commented Oct 2, 2026

Copy link
Copy Markdown
Owner

What

This PR rebuilds Algorithm-Energy-C into greenbench, a small, reproducible study of one question: is run time a good stand-in for energy, and how do C, Python and Node.js compare?

  • Energy from RAPL in C. Reads package and DRAM energy from /sys/class/powercap and handles counter wraparound. Missing, unreadable, corrupt or frozen counters are recorded as "energy unavailable", never as zero. Each socket is checked on its own. A run stops if a counter stands still for 100 ms or more, or jumps by more than 1 kW per package, as a counter reset after a suspend would.
  • Measurement harness. One worker process per sorting run in every language, timed with CLOCK_MONOTONIC_RAW.
    • Sort loops and startup batches are calibrated to about 240 ms, so that they last at least 200 ms.
    • An idle baseline of the same length follows every run.
    • 30 runs per configuration, in a seeded random order, optionally pinned to a CPU set. The harness refuses a CPU the kernel would quietly drop.
    • Output is a tidy runs.csv plus meta.json, which records settings, limits and provenance.
  • Same work in every language. The C, Python and Node.js workers produce identical inputs and operation counts. Parity tests check this, and the harness checks it again on every run.
  • Statistics and report. The report decides the question with net power above idle: net energy per sort divided by time per sort.
    • Time, energy and power ratios between languages and between input orders, each with a 95% bootstrap interval.
    • Each power ratio is read against a 10% margin fixed before measuring: same power within 10%, differs, inconclusive, or undefined.
    • The report states how many intervals exclude 1 against the 5% expected by chance.
    • Rank agreement across configurations, including per language and algorithm, and a run-to-run noise check.
    • UAE and UK carbon estimates, net of idle and including idle.
    • docs/results.md, its charts and the results block at the top of the README are all generated from one run.
  • One command on bare-metal Linux. scripts/run-study.sh runs the study.
    • It pins to two physical cores, never CPU 0 or its sibling thread, and only performance cores on a hybrid processor. --show-cpus prints the choice.
    • It blocks suspend with systemd-inhibit while it measures.
    • docs/bare-metal.md gives the Ubuntu-on-a-USB-stick steps, including a checksum-verified Node.js 24 install.
  • Tests that cannot flake on a busy machine. build/greenbench-simulated is a test-only build whose RAPL readings are computed from the clock. It labels itself everywhere, and the report refuses to put its energy in the README.
  • Documentation. docs/study.md sets out the question, method and threats to validity, and docs/decisions.md holds 19 decision records.

Why

Developers measure time and assume energy follows it. The original repository named measuring energy with RAPL as its next step. This PR takes that step as a proper study, with:

  • identical inputs in every language;
  • idle-baseline subtraction and confidence intervals;
  • a stated rule for when time does or does not predict energy, which never reads "no difference found" as "the same".

How it was tested

Every CI command was run from a fresh clone of this branch, in containers:

  • C: GCC 14.2 and Clang 19.1 builds with -Werror, and 955 unit-test checks each as a non-root user. ASan and UBSan cover the unit tests and a harness run on a fake two-socket tree.
  • Integration: 124 tests: parity, the harness end to end, the study script's CPU choice on fake topologies, and the report pipeline with simulated energy. They passed three more times with four busy loops on 2 CPUs.
  • Statistics: ruff, mypy --strict and 103 pytest tests on Python 3.12.14 and 3.14.7.
  • Node.js: ESLint, tsc and 30 tests on 24.21.0, and the worker tests on 18.20.8 and 20.20.2.
  • Other checks: ShellCheck, actionlint, and the Docker quick start through the documented mirror build.
  • Study script: a rehearsal on fake counters, with and without a stub systemd-inhibit.

docs/results.md was regenerated from a full study-profile run in Docker Desktop (2 CPUs, 2 GiB, 5,760 runs, complete, timings withheld). The operation counts, the chart and the README block are byte-identical to the previous run. One correction: the commit message of 1295633 says that run was of a48caef. It was of 60ddc84, which has the same code, because a48caef changed only two documentation files.

What is pending

  • Energy results. RAPL is unavailable in Docker Desktop, where this was built. Every time, energy, power, ratio, correlation and carbon result is marked pending until ./scripts/run-study.sh runs on bare-metal Linux. Please merge before following docs/bare-metal.md, which clones the default branch.
  • GitHub Actions. CI has not run yet, because the branch is not pushed.

Measured results

Measured results: time in Docker, energy pending

docs/results.md, docs/img/time_per_sort.png and the README results block now come from a full study-profile run with energy off:

  • Run: 5,760 measured runs, complete, no retries, every sort loop at least 200 ms.
  • Command: ./scripts/run-docker-timing.sh, effectively --cpus 2,4 --memory 2g, at 9e55f7b.
  • Environment: Docker Desktop 4.93.0 on a Windows 11 laptop (AMD Ryzen 7 6800H, 16 logical CPUs, on mains power). The container ran on vCPUs 2 and 4, pinned there, with 2 GiB, no swap and no network. Workers: gcc 14.2.0, Python 3.12.14 and Node.js 24.21.0.
  • Raw data: data/runs/20261002T211256Z-docker/ holds the raw data, environment.txt (with the exact docker command) and the summary. Regenerating the report from it in a fresh clone reproduces the committed files byte for byte.

Time per sort, random input of 2,000 elements (median, 95% bootstrap interval):

Algorithm C Python Node.js
bubble 6.05 ms (5.94 to 6.08) 181 ms (179 to 183) 2.47 ms (2.45 to 2.50)
insertion 473 µs (468 to 477) 102 ms (102 to 103) 1.05 ms (1.03 to 1.06)
quick 22.6 µs (22.4 to 22.7) 2.75 ms (2.69 to 2.78) 113 µs (111 to 114)

C bubble sort is slower than Node.js. GCC 14.2 at -O2 turns the C bubble sort's swap into an 8-byte load and store that the next comparison partly overlaps, so every swap stalls. A harness check in the same setup measured 6.12 ms per sort on random input and 11.3 ms on reversed input with the study's build, against 1.34 ms and 1.40 ms with -fno-tree-slp-vectorize. The study keeps -O2; docs/study.md gives the details.

Energy is still pending. Energy, power, the proxy question and carbon stay pending until ./scripts/run-study.sh runs on bare-metal Linux. ADR 20 records why times measured in Docker are published while energy is not.

🤖 Generated with Claude Code

fasharif and others added 30 commits September 26, 2026 02:25
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The sorts keep their exact behaviour and counts (the tests now check the
counts Algorithm-Energy-C published for 4,000 random integers). Counters are
64-bit, and the xorshift32 generator, the input orders and an FNV-1a input
hash move out of main.c into src/inputs.c so the harness and the workers can
share them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Package and DRAM zones are found under /sys/class/powercap, read with pread,
and summed over sockets. One counter wrap is handled with
max_energy_range_uj. A missing, unreadable or corrupt counter makes energy
unavailable with a reason; it is never reported as zero. Fake powercap trees
are stored as manifests because zone names contain colons.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every worker takes the same commands (run, sequence, input) and prints the
same result line: counts for one sort, its own loop time and a hash of its
input, so the harness can check that every language did the same work.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
greenbench run starts one worker process per measured run and reads RAPL
and CLOCK_MONOTONIC_RAW just before and after it. It calibrates sorts per
process to a minimum loop time, measures an idle window of the same length
after each run, shuffles every run into one seeded random order, can pin to
a CPU, checks that all languages report identical inputs and counts, and
writes a tidy runs.csv with meta.json. greenbench probe checks the counters.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A line-for-line port of the C worker using only the standard library, so it
runs on the Python a Linux distribution ships (3.10 or later). pyproject.toml
configures ruff, mypy --strict and pytest; requirements files are compiled
with hashes from the .in files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Plain JavaScript with JSDoc types, checked by the TypeScript compiler in
strict mode and linted with ESLint. It has no runtime dependencies, so it
runs without a build step on Node.js 18 or later.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The parity tests compare xorshift32 sequences, generated inputs up to
100,000 elements and the operation counts of every algorithm and order in
C, Python and Node.js. The harness tests run the real binary with energy
off, with fake powercap trees, and with failing or disagreeing workers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Ember's 2024 lifecycle figures as published by Our World in Data (CC BY 4.0),
with the source, basis and retrieval date recorded next to each value.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
greenbench_analysis loads runs.csv and meta.json with strict validation,
computes medians with percentile-bootstrap confidence intervals and
Spearman's rank correlation between time and energy, converts energy to
CO2-equivalent, and writes docs/results.md with its charts. With
--withhold-timings it marks time, energy and carbon figures as pending.
It replaces visualize.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
scripts/run-study.sh refuses WSL and containers, checks the tools, offers to
make the RAPL counters readable until the next reboot, builds and tests,
records the environment, runs the study profile with energy required and
writes the report. --check and --rehearse test a machine first. The
Dockerfile gives Windows and macOS users the same tools with energy off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
GCC and Clang builds, sanitizers with a fake powercap tree, cross-language
parity, harness and pipeline tests, a short harness run with energy off,
statistics tests on Python 3.12 and 3.14, the Node.js worker, ShellCheck and
actionlint. Dependabot covers actions, pip, npm and Docker.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
All three sorts make the same number of comparisons on reversed input, so
each algorithm now has its own line width, dash pattern and marker. Values
that a log axis cannot show are listed under the chart instead of silently
dropped, and a time-energy chart with nothing to plot is not drawn.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A study run in Docker Desktop stopped after 2,227 of 5,760 runs when one
Python worker could not read its own directory (ENOMEM from the shared
file system). A failed process is now rerun up to --retries times (default
1) and the retries are counted in meta.json. Disagreeing results are never
retried: they stop the run as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The README explains that greenbench grew out of Algorithm-Energy-C and says
plainly that energy results are pending until the study runs on bare-metal
Linux. The old README's timings, measured on a CI runner, are not carried
over; timings will come from the measured study run. docs/study.md sets out
the question, method and threats to validity, docs/bare-metal.md gives the
steps for an Ubuntu USB stick, and docs/decisions.md records the design
choices. The licence year is 2026.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sizes get thousands separators, worker versions and the energy reason read
cleanly, the report command includes the PYTHONPATH it needs, and a
withheld report says why its raw data is not committed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Generated with --withhold-timings from a full run of the study profile
(5,760 measured runs, all three languages) in Docker Desktop, where RAPL is
unavailable. It shows the exact operation counts, identical in C, Python
and Node.js, and marks time, energy, correlation and carbon as pending
until the study runs on bare-metal Linux.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Counter files that exist but never change, as some virtual machines and
firmware leave them, used to pass as valid energy: the harness wrote
energy_status=ok with 0 uJ and the report printed zero energy and carbon.

The harness now reads the counters 200 ms apart before measuring and
treats a package counter that has not moved as unavailable (exit 3 under
--energy require). A measured interval or idle window of 100 ms or more
with no package change stops the run. probe exits 3 on frozen counters.

Tests that need energy now move the fake counters with CounterWriter
(tests/integration/powercap.py), which advances them in place and wraps
them at max_energy_range_uj; make sanitize runs the harness under it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cess

A startup trial used to be one process that exits at once: about 1 to 3 ms
for C, one to three RAPL updates, so its energy was mostly counter
quantisation, and a quarter of the study's runs were affected.

Each measured startup run now launches the worker back to back until the
batch lasts about --min-loop-ms, calibrated during warm-up like the sorts.
runs.csv gains a launches column (schema 2, harness 2.1.0), meta.json
records it per trial, and the analysis divides startup time and energy by
it before subtracting the cost per process. Every launch's result line is
still checked against the other languages.

probe now needs at least 200 ms to judge frozen counters, so a sleep that
ends early on CLOCK_MONOTONIC_RAW cannot hide them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… run

The study pinned the harness and every worker to one CPU. A Node.js 24
worker runs 7 threads (V8's compiler and garbage collector), a Python
worker 1, so on a single CPU Node's helper threads took turns with the
sorting thread and the cross-language comparison was biased against it.

The harness now takes --cpus LIST (for example 2,3 or 2-3, validated
against the size of a cpu_set_t, which --cpu did not do) and records the
set in meta.json. run-study.sh picks two CPUs on different physical cores
of one processor from sysfs, never CPU 0. ADR 15 and the threats to
validity describe the choice and what has not been compared.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…valid

- meta.json records the harness command as invoked (build/greenbench run
  ...) rather than a bare "greenbench", and quotes an embedded single
  quote as '\'' so the line can be pasted back from the same directory.
- JSON strings from the harness replace bytes that are not well-formed
  UTF-8 with U+FFFD, so a --label or path in another encoding can no
  longer produce an invalid meta.json.
- The analysis rejects nan and inf energy cells with the line and
  column, and reports a meta.json that is not UTF-8 as a DataError
  instead of a traceback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
meta.json said "16 CPUs online" for a run in a container limited to two,
and nothing recorded whether calibrated sort loops reached their target.

- host.allowed_cpus (the affinity mask before pinning) and the cgroup v2
  cpu.max and memory.max of the harness's own cgroup, when present.
- short_loops and capped_loops: measured sorting runs whose loop fell
  short of --min-loop-ms, and how many of those --max-reps held back.
  The harness prints a note when there are any.
- The help text describes --min-loop-ms as a calibrated target rather
  than a guaranteed minimum, which is what it is.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
makeInput built the array with new Array(n) and filled it, which V8 stores
as HOLEY_SMI_ELEMENTS, and data.slice() kept that kind, so every element
access in the sorts paid a hole check: a small, systematic cost against
Node.js in a comparison of languages. Array.from gives a packed array
(checked with %HasHoleyElements on Node.js 24.21 for every order, before
and after slice). Parity tests still pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
study.md decides the question by whether energy ratios match time ratios
and whether any language or algorithm draws a different power, but the
report computed neither. Its main test, a pooled Spearman rho between
process time and process energy, was described as "dominated by input
size" although calibration holds process time nearly constant.

The analysis now computes, from one set of bootstrap draws per
configuration (each run's time and energy kept together):

- power while sorting (energy per sort / time per sort) with a 95% CI;
- time, energy and power ratios with CIs between languages (Python/C,
  Node.js/C, Python/Node.js) and between each input order and random,
  with a plain reading: same power, or who draws how much more;
- rank agreement across configurations (median time per sort against
  median energy per sort) at each size, per language and overall;
- the within-configuration rho, kept and described as a noise check.

The report also gains a power column, an energy-against-time chart with
constant-power lines, net and including-idle carbon side by side, a
warning for unfinished runs or fewer than 30 runs per configuration, the
share of sort loops that fell short of the target, and the cgroup
limits. It refuses a run whose meta.json status is not complete unless
--allow-partial is given, and refuses energy that is all zeros. Tests
use synthetic data in which one language draws 50% more power.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every harness-level energy test used frozen counters, so non-zero energy
had never gone from the harness through the net-of-idle columns into the
report. The new pipeline test moves a fake package counter faster while a
worker runs than during the idle windows, wraps it several times, and
checks the CSV (net formula, no mishandled wrap) and every energy section
of the report: power, ratios, rank agreement, carbon and both charts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Checked on stock Ubuntu 24.04 packages (Node.js 18.19.1, Python 3.12.3):

- "python3 -m venv --help" exits 0 without python3-venv, so the check
  passed and the script failed later. It now creates a throwaway venv.
- A fake --powercap-root in study mode was refused only after make
  clean, the build, the tests, the probe and the pip install. It is now
  refused right after the arguments are parsed.
- The chmod o+r prompt now says why it matters: until reboot any local
  program can read the counters, reopening the PLATYPUS side channel.
- A probe that finds frozen counters now stops the script with a reason.
- Node.js older than the 24 LTS the guide installs gets a warning, since
  the results describe whichever version ran.
- The closing message tells Farah to copy the results to a USB stick and
  commit from her usual computer, not from the live session.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every FROM line used ${REGISTRY}/..., which Dependabot cannot resolve, so
its docker entry did nothing, and the images tracked floating tags. The
FROM lines are now literal and pinned by digest (the current Docker Hub
digests, which Amazon's mirror of the official images serves too). The
comment gives a command that builds from that mirror when Docker Hub
limits pulls, deriving each --build-context name from the FROM lines so
that it stays right when Dependabot changes a digest.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The source column now reads "Ember (2026) – with major processing by Our
World in Data", OWID's short citation for indicator 1295504 (checked in
the chart's metadata on 2026-09-26), and data/README.md quotes the full
citation. The values and the CC BY 4.0 links are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fasharif and others added 21 commits September 26, 2026 04:57
- study.md: "What would count as an answer" now points at power ratios,
  power while sorting and rank agreement, and the analysis section
  describes them; the pooled rho and its wrong "dominated by input
  size" description are gone. The method says loops are calibrated to
  about 200 ms (a target, not a minimum), which clocks the workers use
  and why, and how startup batches work. New threats: helper threads
  and pinning, warm starts in startup batches, calibration shortfalls,
  startup medians left out of the intervals, and net against
  including-idle carbon.
- decisions.md: ADR 5 reworded; ADR 16 (Node.js 24 LTS for the study)
  and ADR 17 (power ratios rather than correlation decide the question).
- bare-metal.md: install Node.js 24 from nodejs.org, checked against the
  published SHA-256 sums; copy the results to a USB stick and commit from
  the usual computer (or set a git identity first); the new duration
  estimate; the side-channel risk of chmod; rehearsing without RAPL.
  Followed step by step on Ubuntu 24.04 in a container.
- README: operations chart first, then the probe output; plain feature
  names; the new analyses, options and limitations.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Generated with --withhold-timings from a full study-profile run of
commit ca63170 in Docker Desktop (container limited to 2 CPUs and 2 GiB,
energy unavailable): 5,760 measured runs, status complete, no retries.
The operation counts are unchanged. The provenance now records the
cgroup limits, the run status and a harness command that pastes back
from the repository root; the grid table uses OWID's citation. The raw
CSV is not committed, because it holds the withheld timings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…umps

A frozen package counter on one socket of a two-socket machine passed the
frozen check, because the check looked at the sum over sockets. Each
package domain is now judged on its own, and the error names the zone.

A reading that rises by more than 1 kW per package over its interval now
stops the run: a counter that reset after a suspend would otherwise be
read as a wrap worth hundreds of kilojoules. meta.json also records how
many counter wraps fell inside measured runs and idle windows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The energy tests moved fake counters from a Python thread. On a busy machine
the thread was not always scheduled in time, so a short interval saw no energy
and a long stall tripped the frozen-counter check; the pipeline test failed in
4 of 11 runs on a loaded machine.

make simulated now builds build/greenbench-simulated with
-DGREENBENCH_SIMULATED_RAPL. It reads a fake tree as usual and adds energy
computed at each reading from CLOCK_MONOTONIC_RAW and from the processor time
of its finished workers, wrapping at max_energy_range_uj. The integration
tests, the sanitizer run and the study script's fake-tree rehearsal use it;
the frozen-counter tests keep using build/greenbench. The simulated build says
so on standard error, meta.json records "simulated": true, and the report
opens with a warning that such energy is not a measurement. The report also
shows how many counter wraps were handled. ADR 18 records the decision.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The study asks for Spearman's correlation between time and energy per
language and algorithm. The only figure at that level was the median of
within-configuration rho values, which is a noise check. Rank agreement across
configurations now also covers each language and algorithm, across input
orders and sizes, with the same bootstrap interval.

RankAgreement now holds its scope as language, algorithm and size, and the
report puts it into words.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The report called two configurations "same power" whenever the 95% interval
of their power ratio included 1, which reads a result that is not significant
as proof of equality: a rehearsal printed "same power" for an interval from
0.0006 to 15.7. It also printed about 250 such intervals with no word on how
many would exclude 1 by chance.

Each power ratio now gets one of four readings against a margin of 10% either
way (1/1.1 to 1.1), fixed in docs/study.md: same power within the margin
(the whole interval inside it), differs (the interval excludes 1),
inconclusive (includes 1 but reaches beyond the margin), or undefined. A ratio
is also undefined when a time or energy per sort fell to zero or below in more
than 2.5% of the resamples, because an interval from the rest would look too
narrow. The report tabulates the readings and states how many intervals
exclude 1 against the 5% that chance would give. --equivalence-margin sets
another margin.

Power is labelled "net power above idle" wherever it is net of the idle
baseline, in the tables, the text and the chart. study.md explains the margin,
why the question is decided on net power, and the many comparisons, and
ADR 19 records the decision.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Calibration set the sorts per process so that a loop would last exactly
--min-loop-ms. A measured run is as likely to be a little faster than the
warm-up that set it as a little slower, so about half of all loops fell just
short of the target on any machine, and the count of short loops meant
nothing. Reviews found 128 to 217 short loops per 216 to 360 in busy
containers.

Warm-up now aims 20% higher (240 ms for the study profile's 200 ms), for sort
loops and startup batches alike, so ordinary variation stays above the target
and a short loop points to a busy or throttled machine. meta.json records the
aim as settings.aim_loop_ns and the report describes it; reports of earlier
runs read as before. The harness is now version 2.2.0. ADR 5, study.md and
the bare-metal time estimate (about 46 minutes of runs) are updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
sched_setaffinity quietly drops CPUs that are offline, do not exist or lie
outside the process's cpuset, as long as one listed CPU remains. The harness
then pinned to fewer CPUs than meta.json recorded. It now reads the set back
with sched_getaffinity and stops with a usage error naming the CPU, and says
plainly when none of the listed CPUs can be used.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
After the documented bare-metal run, docs/results.md would show measured
energy while the README's status banner, its limitations and study.md still
said "pending". The report now takes --readme README.md and rewrites the block
between <!-- results:start --> and <!-- results:end --> from the same run: a
pending notice while timings are withheld, and otherwise the readings of the
power ratios and the main table of ratios between languages. It refuses
simulated energy and a README without exactly one pair of markers, before
writing anything.

run-study.sh passes --readme in study mode and adds README.md to the commit
line; bare-metal.md lists it with the files to take away. The block in this
README was generated by the report from a withheld run. The README's
limitations and study.md's status box now read correctly before and after the
study.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The study script took the first online CPU after 0 and the next CPU on another
core. On processors that number a core's two threads 0 and 1, as this
project's own Ryzen 7 6800H does, that pinned the study to CPU 0's other
thread, which shares CPU 0's interrupt load. The choice also ignored core
types on hybrid processors.

It now skips every CPU in CPU 0's thread_siblings_list, and uses only the
performance cores in /sys/devices/cpu_core/cpus when Linux lists them. It
prints the choice with the reason, and --show-cpus prints them without
running anything. New tests check the choice on fake topologies through
GREENBENCH_SYSFS: adjacent and split thread numbering, a hybrid processor,
offline CPUs, a second package, no suitable pair, and the overrides. ADR 15,
study.md and bare-metal.md describe the rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The full study runs for about an hour unattended. A suspend would stop
CLOCK_MONOTONIC_RAW and can reset the RAPL counters; the harness already
stops on the implausible jump that a reset causes, but the hour is lost.
run-study.sh now runs the harness under systemd-inhibit, blocking suspend,
idle sleep and the lid switch, when logind is reachable, and otherwise warns
that nothing stops the machine from suspending. bare-metal.md asks for
automatic suspend to be turned off as well.

Checked in a container with a stub systemd-inhibit on the PATH: the
rehearsal ran the harness through it with the expected options.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The quick start chained eight commands on five lines. make setup creates
.venv and installs the hashed analysis requirements, and make smoke already
builds what it needs, so the block now reads: git clone, cd, make setup,
make smoke, ./scripts/run-study.sh. Checked in the Docker image as a normal
user: make setup and make smoke write results/smoke/results.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…peat

- Every job has timeout-minutes, so a hung run cannot hold a runner for six
  hours.
- The parity tests ran twice in the workers job: make parity, then make
  integration, which includes them. The job now runs make integration once.
- package.json and the README say the workers run on Node.js 18 or later,
  which CI did not test. A new job runs the worker tests on Node.js 18 and 20
  (they pass on 18.20.8 and 20.20.2 in containers).
- The Python worker claimed Python 3.10 or later, which nothing tested; ADR 9
  and its docstring now say 3.12 or later, as CI tests with 3.12 and 3.14.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- ADR 2's decision paragraph is wrapped like the rest again.
- study.md's measurement section says that a run stops on a counter that
  stands still for 100 ms or jumps by more than 1 kW per package.
- bare-metal.md says what the SHA-256 check of the Node.js download does and
  does not cover.
- The README's features, limitations, roadmap and folder notes no longer
  assume that the study has not run yet; the design decisions list mentions
  the simulated test build.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
study.md justified the 10% margin by saying that the languages differ in time
by factors of ten or more, which this repository has not measured yet. The
margin now rests on its own reasoning. Two paragraphs are rewrapped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A full study-profile run of commit a48caef in Docker Desktop (2 CPUs, 2 GiB,
5,760 measured runs, status complete, no retries), reported with
--withhold-timings --readme README.md. The operation counts, the chart and
the README's results block are unchanged; the provenance lines now name
harness 2.2.0 and its calibration aim. The raw data is not committed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A run without energy, such as one in a container, now publishes what it
measured and marks the rest as pending a run on bare-metal Linux:

- time per sort and startup time per process, with their intervals, and
  no energy columns;
- a "Time ratios" section: between languages and between input orders,
  each with a 95% bootstrap interval. A time ratio is undefined when a
  time per sort is at or below zero, or is in more than 2.5% of the
  resamples (Comparison.time_reliable);
- a note at the top, and pending sections for the proxy question and
  carbon, saying why energy is missing and where it will come from;
- in the README block, a table of time per sort by algorithm and
  language and a table of time ratios between languages.

The report and the README block also give the run's limits, pinning,
versions and label, a --reproduce command when one is passed, and a
link to the run's raw data.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
scripts/run-docker-timing.sh runs the study profile in a container on
machines that cannot boot Linux directly, and writes the raw data to
data/runs/<stamp>-docker/, the report and the README's results block.

- It builds the image from the Dockerfile and, in study mode, refuses
  uncommitted changes to tracked files, so the recorded commit is the
  code that ran.
- The harness and workers run from the image's own file system, not a
  mounted folder, and the results are copied out afterwards.
- The container gets the two CPUs that run-study.sh --show-cpus would
  choose, the harness pins to them, and it has a memory limit without
  swap and no network.
- environment.txt records the host, Docker, the image and base-image
  digests, lscpu inside the container and the exact docker command.

--rehearse runs the smoke profile under results/. Its option handling
is tested; its runs need Docker. ADR 20 records the decision.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
docs/results.md, its time chart and the README's results block now come
from a full study-profile run in Docker Desktop on an otherwise idle
Windows 11 laptop (AMD Ryzen 7 6800H, on mains power), with energy off.
Run with ./scripts/run-docker-timing.sh at 9e55f7b:

- 5,760 measured runs (4,320 sorting, 1,440 startup batches), status
  complete, no retries, every sort loop at least 200 ms;
- a container on CPUs 2 and 4 of the Docker VM, pinned there, with
  2 GiB of memory, no swap and no network;
- gcc 14.2.0, Python 3.12.14 and Node.js 24.21.0.

The raw data, the environment record and the summary are in
data/runs/20261002T211256Z-docker/. Energy, power, the proxy question
and carbon stay pending a run on bare-metal Linux. study.md's status
box and ADR 12 now say so.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
In the Docker time run, C bubble sort took 6.05 ms per sort on random
input of 2,000 elements against 2.47 ms for Node.js, with identical
counts. GCC 14.2 at -O2 vectorises the swap into one 8-byte load and
store of both elements; the next comparison's load half overlaps that
store and cannot be forwarded, so every swap stalls.

A harness check in the same container setup measured C bubble sort at
2,000 elements with the study's build and with -fno-tree-slp-vectorize
added: 6.12 against 1.34 ms per sort on random input, and 11.3 against
1.40 ms on reversed input. The threats to validity give the commands.
The study keeps the plain -O2 build, and its C bubble sort results
describe that build.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@fasharif
fasharif merged commit c757f05 into main Oct 2, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant