Skip to content

Fair benchmark: identical inputs, correct counts, tests and CI - #1

Merged
fasharif merged 4 commits into
mainfrom
fix/fair-benchmark-and-tests
Sep 25, 2026
Merged

fasharif merged 4 commits into
mainfrom
fix/fair-benchmark-and-tests

Conversation

@fasharif

Copy link
Copy Markdown
Owner

What was wrong

  • The comparison was not fair. Each algorithm sorted a different random array, once, at a single size of 1,000.
  • The counts were wrong. Insertion sort never counted the comparison that ends its inner loop. "Swaps" meant different things in different algorithms, and included swapping an element with itself.
  • Timing was too coarse. It used clock() for a single run of a millisecond-scale task.
  • An empty array broke bubble sort. n - 1 underflowed.
  • Sorted input could overflow the stack. Quicksort always recursed on both sides, and sorted input is its worst case.
  • The claims went beyond the measurements. The README talked about energy, but nothing measures energy.

What changed

  • Every algorithm gets an identical copy of each input: sizes 500, 1,000, 2,000 and 4,000, in random, sorted, reversed and nearly sorted order.
  • Each run is timed with a monotonic clock, the median of 5 runs is reported, and every result is checked to be sorted.
  • Comparisons and array writes are counted the same way in all three algorithms.
  • Quicksort recurses into the smaller side, so its stack depth stays O(log n).
  • Unit tests check each algorithm against qsort, and check exact counts on sorted and reversed input.
  • CI builds with GCC and Clang, runs AddressSanitizer and UBSan, runs the benchmark, and uploads the CSV and charts.
  • visualize.py saves charts to files instead of opening windows.
  • Adds an MIT licence.

A follow-up commit on this PR rewrites the README with the numbers CI measures.

🤖 Generated with Claude Code

fasharif and others added 4 commits September 25, 2026 23:34
Each algorithm used to sort a different random array, once, at a single
size, so the comparison was not fair. Now every algorithm gets an identical
copy of each input at several sizes and in four orders (random, sorted,
reversed, nearly sorted), and each is timed over five runs with a monotonic
clock, reporting the median. Every result is checked to be sorted.

Insertion sort never counted the comparison that ends its inner loop. All
three algorithms now count comparisons and array writes the same way, and
no longer count swapping an element with itself. bubble_sort no longer
underflows when given an empty array. quick_sort takes the same arguments as
the others and recurses into the smaller side, so sorted input cannot
overflow the stack.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The unit tests check each algorithm against qsort on random arrays with
duplicates, and check exact comparison and write counts on sorted and
reversed input. CI builds with GCC and Clang with warnings as errors, repeats
the tests under AddressSanitizer and UBSan, runs the benchmark, and uploads
the CSV and charts. visualize.py now saves PNG files instead of opening
windows, and prints a Markdown summary table.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Reports the comparisons, writes and times measured in CI, explains what they
show, and describes what the program measures without claiming to measure
energy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@fasharif
fasharif merged commit 0bf73c8 into main Sep 25, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant