Measures the energy, time and work of three sorting algorithms in C, Python and Node.js on identical inputs, to test whether run time is a good stand-in for energy.
greenbench grew out of Algorithm-Energy-C, a C benchmark that counted the comparisons and array writes of bubble sort, insertion sort and quicksort and timed them. Its README named measuring energy directly with Intel's RAPL counters as the next step. This is that step, done as a small study: an energy-measuring harness in C, the same algorithms in three languages, statistics, carbon estimates and a written method.
Results from docs/results.md: AMD Ryzen 7 6800H with Radeon Graphics, Linux 6.18.33.2-microsoft-standard-WSL2, 5,760 measured runs on 2026-10-02.
Energy results are pending. This run measured time and operation counts only; energy was not measured (turned off with --energy off). Whether run time is a good stand-in for energy needs the study on bare-metal Linux (how).
Time per sort, random input of 2,000 elements: the median of each configuration's measured runs, with a 95% bootstrap interval in brackets.
| Algorithm | C | Python | Node.js |
|---|---|---|---|
| bubble | 6.05 ms (5.94 to 6.08) | 181 ms (179 to 183) | 2.47 ms (2.45 to 2.50) |
| insertion | 473 µs (468 to 477) | 102 ms (102 to 103) | 1.05 ms (1.03 to 1.06) |
| quick | 22.6 µs (22.4 to 22.7) | 2.75 ms (2.69 to 2.78) | 113 µs (111 to 114) |
Time ratios between languages, random input of 2,000 elements:
| Algorithm | Order | Size | Languages | Time ratio |
|---|---|---|---|---|
| bubble | random | 2,000 | Python / C | 30 (29.5 to 30.6) |
| bubble | random | 2,000 | Node.js / C | 0.408 (0.404 to 0.419) |
| bubble | random | 2,000 | Python / Node.js | 73.5 (71.9 to 74.3) |
| insertion | random | 2,000 | Python / C | 216 (214 to 220) |
| insertion | random | 2,000 | Node.js / C | 2.21 (2.17 to 2.25) |
| insertion | random | 2,000 | Python / Node.js | 97.7 (96.4 to 99.6) |
| quick | random | 2,000 | Python / C | 122 (119 to 124) |
| quick | random | 2,000 | Node.js / C | 4.99 (4.89 to 5.07) |
| quick | random | 2,000 | Python / Node.js | 24.5 (23.8 to 25) |
Environment: 16 CPUs online, 2 of them allowed to the harness, a cgroup memory limit of 2.0 GiB; CPU pinning to CPUs 2, 4; C: gcc 14.2.0; Python: 3.12.14; Node.js: v24.21.0. Run label: run-docker-timing.sh study 20261002T211256Z: a Docker container limited to CPUs 2,4 and 2g of memory without swap; energy off, because containers cannot read RAPL; Windows 11 Home laptop (ASUS TUF Gaming A15 FA507RC) on mains power in the Performance power mode, with no other containers or jobs running; Docker Desktop runs Linux in a virtual machine whose CPUs are not tied to physical cores. Reproduce with ./scripts/run-docker-timing.sh --cpus 2,4 --memory 2g. Every setting and command: docs/results.md.
Operation counts are exact, so they are already in docs/results.md, generated from a full run of the study profile. They are identical in all three languages, which the harness checks on every run:
Energy is another matter. In Docker Desktop, where this version was developed,
greenbench probe reports that there is nothing to measure:
$ build/greenbench probe
powercap root: /sys/class/powercap
energy: unavailable
reason: no RAPL package zone (intel-rapl:N named package-N) under /sys/class/powercap; this CPU, kernel or virtual machine does not expose RAPL
The harness then measures time and operation counts only, and leaves every energy cell empty rather than writing zeros.
Developers measure time and assume that faster code uses less energy. That assumption is rarely tested, because measuring energy properly is fiddly: the counters are coarse, the processor never stops drawing power, and comparing languages fairly means feeding them exactly the same work. greenbench does the fiddly parts so that the question can be answered with data: is run time a good stand-in for energy, and how do C, Python and Node.js compare? The research question, method and threats to validity are in docs/study.md.
- Energy from RAPL in C. Reads package and DRAM zones from
/sys/class/powercap, handles counter wraparound withmax_energy_range_uj, and reports "energy unavailable" with a reason when the counters are missing, unreadable or frozen, never a fake value. - Identical inputs everywhere. xorshift32 and the four input orders are implemented in C, Python and JavaScript; tests check that all three produce identical sequences, inputs and operation counts, and the harness stops if a worker ever disagrees.
- Measurement harness. One process per measured run for every language, time from
CLOCK_MONOTONIC_RAW, sorts per process calibrated to about 240 ms so that each loop lasts at least 200 ms (far above the ~1 ms RAPL update interval), an idle baseline of the same length after every run, 30 measured runs per configuration after warm-up, randomised run order, optional pinning to a set of CPUs, startup trials measured over batches of launches, and tidy CSV output. - Statistics. Net power above idle, and time, energy and power ratios between languages and between input orders, each with a 95% bootstrap interval. Each power ratio is read against a 10% margin set before measuring: the same power (so time predicts energy), different, or inconclusive. Also medians with intervals, rank agreement across configurations and a run-to-run noise check, in a Python package with its own tests.
- Carbon estimates. Energy converted to grams of CO2-equivalent for the UAE and UK grids, net of idle and including idle, from a cited table (Ember data, CC BY 4.0), with any other region or value on request.
- A generated report.
docs/results.md, its charts and the results block at the top of this README come from a script, never by hand. It refuses a run that did not finish unless told otherwise, and says when a run has fewer than 30 runs per configuration or loops shorter than their target. - Runs anywhere, measures where it can. The same harness runs on WSL2, in Docker and in CI
with energy off.
scripts/run-docker-timing.shmeasures time for the whole study profile in a Docker container, andscripts/run-study.shruns the full study, energy included, on bare-metal Linux.
flowchart LR
subgraph harness["build/greenbench (C)"]
plan["plan: configurations x 30 runs,<br/>shuffled from a seed"] --> run["run one worker process"]
rapl["RAPL reader<br/>/sys/class/powercap"] --> run
run --> files[("runs.csv<br/>meta.json")]
end
run -- "argv: algorithm, order, n, seed, reps" --> workers
subgraph workers["workers: same commands, same output"]
c["C"]
py["Python"]
node["Node.js"]
end
workers -- "result line: counts, loop time, input hash" --> run
files --> analysis["greenbench_analysis (Python):<br/>medians, power, ratios, bootstrap CIs, carbon"]
grid[("data/grid_intensity.csv")] --> analysis
analysis --> report["docs/results.md<br/>docs/img/*.png"]
For each measured run the harness reads the energy counters and the clock, starts a worker (for a startup trial, a batch of them back to back), waits for it to exit, and reads them again. It then sleeps for the same duration and reads the counters once more to measure idle energy. Each worker prints one line; the harness checks that its input hash and counts match the other languages, and writes a CSV row.
| Part | Choice | Why |
|---|---|---|
| Harness and C worker | C11, POSIX (posix_spawnp, pread, clock_gettime) |
Small, predictable overhead around each measured process, and direct access to sysfs. |
| Energy | Linux powercap sysfs | No kernel modules, no per-model unit decoding; the kernel reports microjoules (ADR 1). |
| Python worker | Python standard library | Runs on the Python a Linux distribution ships, with nothing to install. |
| Node.js worker | JavaScript with JSDoc types, checked by tsc --strict; Node.js 24 LTS for the study |
Type-checked like TypeScript, but no build step between the source and what is measured (ADR 9, ADR 16). |
| Statistics and report | Python 3.12+, numpy, matplotlib | The bootstrap, the ratios and Spearman's rho are short enough to write out and test; no SciPy needed. |
| Quality | GCC and Clang with -Werror, ASan and UBSan, pytest, mypy --strict, ruff, ESLint, ShellCheck, actionlint |
Every language in the repository is linted, type-checked where possible, and tested in CI. |
| Portability | Dockerfile, GitHub Actions | Try it on Windows or macOS; the energy study itself needs bare-metal Linux. |
On Linux (including WSL2) with a C compiler, make, Python 3.12+ (with venv) and Node.js 18+:
git clone https://github.com/fasharif/Algorithm-Energy-C.git greenbench
cd greenbench
make setup # a virtual environment (.venv) with the analysis packages
make smoke # build, run all three languages with energy off, report in results/smoke/
./scripts/run-study.sh # the energy study, on bare-metal Linux onlyThe Makefile uses .venv/bin/python when it exists, and python3 otherwise. make test
runs the C unit tests; Running the tests lists the rest.
On Windows or macOS, the Docker image runs the same smoke test and writes the report to
results/smoke/ on your machine (in PowerShell, write ${PWD} instead of $PWD):
docker build -t greenbench .
docker run --rm -v "$PWD/results:/work/results" greenbenchThe base images are pinned by digest. If Docker Hub limits your pulls, the Dockerfile gives a build command that takes the same images from Amazon's mirror.
To measure time and operation counts for the whole study profile without bare-metal Linux,
run ./scripts/run-docker-timing.sh from bash (Git Bash on Windows) on a machine doing nothing
else. It builds the image, runs the harness with energy off in a container limited to two CPUs
and 2 GiB of memory, and writes the raw data to data/runs/, the report to
docs/results.md and the results block of this README
(ADR 20).
--rehearse runs the smoke profile instead and writes under results/; --help lists the
other options.
build/greenbench run takes these options; the defaults are the study profile.
build/greenbench --help lists them all.
| Option | Default | Meaning |
|---|---|---|
--profile study|smoke |
study |
Study: sizes 250, 500, 1,000, 2,000; 30 runs; 3 warm-up runs; loops of at least 200 ms. Smoke: sizes 64, 256; 3 runs; 1 warm-up; 2 ms. |
--languages, --algorithms, --orders, --sizes |
all | Comma-separated subsets, for example --languages c,node --sizes 100,1000. |
--runs N, --warmup N |
30, 3 | Measured and warm-up runs per configuration. |
--min-loop-ms N, --max-reps N |
200, 1,000,000 | Minimum length of each sort loop and startup batch (calibration during warm-up aims 20% higher); and the cap on sorts per process. |
--energy auto|require|off |
auto |
require fails with exit status 3 if RAPL cannot be read. |
--idle-baseline auto|on|off |
auto |
On whenever energy is measured. |
--cpus LIST |
off | Pin the harness and every worker to these CPUs, such as 2,3 or 2-3; each must be online and allowed to the harness. |
--seed N, --order-seed N |
2463534242, 1234567 | Input generator seed, and run-order seed. |
--no-startup |
Skip the startup trials. | |
--retries N |
1 | Rerun a worker process that fails (cannot start, crashes) up to N times; disagreeing results are never retried. |
--out DIR |
results/latest |
Where runs.csv and meta.json go. |
--python CMD, --node CMD |
python3, node |
Interpreters to use. |
--powercap-root DIR |
/sys/class/powercap |
Read RAPL zones from elsewhere, for tests. |
--dry-run |
Print the plan and exit. |
build/greenbench probe checks that the counters exist, are readable and move, and prints idle
power over one second; it exits with status 3 when energy cannot be measured.
The report is python -m greenbench_analysis report --runs DIR/runs.csv with
PYTHONPATH=analysis. It takes --out, --img-dir, --regions ARE,GBR (codes from
data/grid_intensity.csv), --intensity NAME=G_PER_KWH for any
other grid, --equivalence-margin PERCENT (10 by default; see
study.md), --withhold-timings for runs on a
busy machine, --allow-partial to report on a run that did not finish, --summary-csv,
--reproduce COMMAND to show the command that reproduces the whole run, and
--readme README.md, which rewrites the results block at the top of this file.
make test # C unit tests: sorts, inputs, RAPL parsing and wraparound, options, plan, protocol, CSV, spawn
make sanitize # the C tests and a harness run under AddressSanitizer and UBSan
.venv/bin/pip install --require-hashes -r requirements-dev.txt
make check-python # ruff, mypy --strict, and the statistics and Python worker tests
make check-node # ESLint, tsc --strict and the Node.js worker tests
make integration # cross-language parity, the harness, the scripts and the report pipeline
make lint-shell # ShellCheck on the scriptsThe RAPL tests build fake powercap trees from the manifests in
tests/fixtures/powercap, including a two-socket server, an AMD
processor without DRAM, a missing tree, corrupt values and an unreadable counter, and rewrite
counters between readings to test wraparound. The integration tests also run
build/greenbench-simulated, a test build whose readings are computed from the clock
(ADR 18).
With it, energy that is higher during runs than during idle windows goes from the harness
through to every section of the report, and a counter wraps inside a measured interval,
however busy the machine is. The real build must refuse counters that never move. Tests of
the study script check its choice of CPUs on fake CPU topologies; the Docker timing script's
runs need Docker, so its tests check only its options. CI runs all of the above
and lints the workflow with actionlint; the C build and tests run with both GCC and Clang, the
statistics tests on Python 3.12 and 3.14, and the Node.js worker tests also on Node.js 18
and 20.
src/ C: harness (greenbench.c), C worker (worker.c), RAPL reader, sorts,
input generator, options, run plan, worker protocol, CSV output
workers/python/ Python worker (standard library only) and its tests
workers/node/ Node.js worker (no runtime dependencies), tests, ESLint and tsc config
analysis/ greenbench_analysis: loading, statistics, carbon, charts, report; tests
tests/ C unit tests, fake powercap manifests, integration tests (pytest)
scripts/ run-study.sh: the one-command study for bare-metal Linux;
run-docker-timing.sh: time and operation counts in Docker, energy off
data/ grid carbon intensity with sources; the scripts write runs to data/runs/
docs/ study.md, bare-metal.md, decisions.md, results.md (generated), img/
Recorded in docs/decisions.md: why powercap rather than MSRs, why every language is measured as a whole process, how repetitions and startup batches are calibrated, why the idle baseline follows every run, why missing or frozen energy is never written as zero, why the study pins to two cores and measures Node.js 24, why power ratios rather than correlation decide whether time is a good stand-in for energy, why an interval that includes 1 is not enough to call two powers the same, why the energy tests use a simulated build of the harness, and why times measured in Docker are published while energy waits for bare metal.
- Energy needs bare-metal Linux. RAPL counters are hidden in containers, WSL2 and most virtual machines, so the energy study runs only on Linux on real hardware. The block at the top of this README, generated with docs/results.md, says whether the published results include energy.
- RAPL is not a power meter. It covers the processor package and DRAM, not the whole machine, and its accuracy varies between processors. See threats to validity.
- A small workload. Three single-threaded sorts on integer arrays of up to 2,000 elements in the study profile (100,000 at most).
- One setup per run. Results describe the recorded compiler, Python and Node.js versions
and the two pinned CPUs of one machine. Pinned and unpinned runs have not been compared.
GCC 14.2 at
-O2compiles the C bubble sort's swap into code that stalls on every swap, which made it slower than Node.js on random input in the time run in Docker; the threats to validity give the measurements.
Next:
- Repeat the study on a second machine with a processor from the other vendor (Intel or AMD), to see whether the answer holds across processors.
- Check RAPL against a plug-in power meter on the same machine.
- Add merge sort and a median-of-three quicksort, to show a worst case being avoided.
- Add hardware counters (instructions, cache misses) as explanations for energy differences.
MIT. Grid intensity data: Ember, via Our World in Data, CC BY 4.0 (see data/README.md).
