Skip to content

Compact a quiet table's fresh files every minute, between sweeps (T-627) - #451

Open
chasers wants to merge 2 commits into
t-626-shared-file-cachefrom
t-627-quiet-table-small-files
Open

chasers wants to merge 2 commits into
t-626-shared-file-cachefrom
t-627-quiet-table-small-files

Conversation

@chasers

@chasers chasers commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

TL;DR: Compaction now merges a quiet table's newest files every minute, not every 5 minutes. A recent-data query opens a few files instead of a dozen tiny ones.

Tracker: T-627. Stacked on #450 (T-626).

Why

  • A table that seals on age, not size, leaves one sealed file per seal.
  • On the sandbox, logs.cluster (about 110 rows a minute, 3 write partitions) sealed 2-3 files a minute of 1-30 rows each.
  • The hour lane ran once per compact_interval_ms (5 min). So a dozen tiny files piled up in the current hour. 17 of 24 files held fewer than 1,000 rows.
  • Every query over recent data opened each one over S3.
  • A 60-minute histogram ending now took about 5x one ending 15 minutes ago, on the same row count (776 ms vs 146 ms median).

What changed

  • New fresh tick in Smolquery.StorageService.Scheduler. Every compact_fresh_interval_ms (60 s) it runs the hour lane again for the quiet tables only.
  • A quiet table has a file in the current or previous compact_bucket_ms under compact_fresh_below_bytes (4 MiB). New Planner.fresh_tables/3 decides this from the listing's bytes, with no footer reads.
  • Each sweep refreshes the quiet set. A table whose listing is no longer quiet leaves the set.
  • New env vars: SMOLQUERY_COMPACT_FRESH_INTERVAL_MS (0 = off) and SMOLQUERY_COMPACT_FRESH_BELOW_BYTES.
  • Sweeper gains two optional callbacks for a second timer: on_start/1 and handle_tick/2. The retention sweeper defines neither, so it does not change.
  • The lane timer reports lane="fresh".
  • Docs: docs/configuration.md rows, and a docs/deployment.md upgrade note.

How it works

  1. A sweep lists every table, as before, and computes the quiet set from the listings.
  2. 60 s later, the fresh tick lists only the quiet tables.
  3. For each one it plans the hour lane, the same planner as a sweep, but only over files under compact_fresh_below_bytes, and merges them into one group. A bigger file waits for the 5-minute sweep.
  4. The outcomes update the row caps, the quarantine and the backoff, as a sweep's do.
  5. The tick runs in the scheduler process, so it never races a sweep over the same files.

Why this option

You chose this option over:

  • sealing quiet tables less often, which keeps rows in the hot tier longer and gives more hot segments per query
  • per-table write partitions = 1, which is config only and does not fix the merge lag

Tests

  • ✅ PlannerTest: fresh_tables/3 picks a table with a small current- or previous-hour file. It skips a table whose recent files are big, and a table whose small files are older. It returns nothing when the tick is off.
  • ✅ SchedulerTest (integration, real DuckLake and merge): a sweep marks the quiet table; 2 new seals arrive; one fresh tick merges all 3 into one segment, and no rows are lost.
  • ✅ SchedulerTest: a table whose recent files are not small stays out of the set.
  • ✅ SchedulerTest (integration): a 20,000-row file over a 20 KB fresh floor stays; the tick merges only the 2 small seals.
  • ✅ Mutation checks: with the fresh tick as a no-op, the merge test fails. Without the fresh floor, the floor test fails.

Review

A Fable review found no row loss and no broken invariant. The ownership gate, backoff, caps and quarantine behave as in a sweep. Fixed in Review of T-627: ...:

  • Write amplification. The tick used the hour lane's 32 MiB floor, so it could rewrite a growing file every minute. Now it merges only files under compact_fresh_below_bytes (4 MiB).
  • Spill gating. The tick logged and counted a "span paused" every minute and, when gated, planned settled files. Now it uses the un-gated runtime and does nothing without a quiet table.
  • Tests. The integration test waited 5 s around a real merge; now 60 s. The fresh timer is off in test config, so no test gets a background merge.
  • Docs. The moduledoc now says caps, quarantine and conflict waits advance per tick. The cost line no longer says "small by definition".

Checks

  • ✅ mix precommit: 3,090 tests pass
  • ✅ mix ci
  • ✅ mix dialyzer

Watch out

  • ⚠️ More swaps. Each quiet table can get one merge and swap per minute, of at most 4 MiB. Each swap can conflict with a seal of the same table (T-573). Watch smolquery_compaction_conflicts_total.
  • ⚠️ More listings. One catalog listing per quiet table per minute. A table with many small files (like metrics.samples when backlogged) is listed each minute too.
  • ⚠️ Caps, quarantine and conflict waits now advance per tick as well as per sweep. A lost commit parks the table for compact_interval_ms, skipping its ticks until then.
  • ⚠️ The hot tier is unchanged. T-627's "<10 hot segments per Grafana selector" item is not addressed here.
  • ⚠️ T-627 sandbox acceptance is open: at most ~5 sealed files under 1,000 rows on logs.cluster, and the 60-min histogram ending now within 1.5x of one ending 15 min ago.

🤖 Generated with Claude Code

A table that seals on age, not size, leaves one sealed file per seal. On
the sandbox logs.cluster (about 110 rows a minute, three write
partitions) sealed two or three files a minute of 1-30 rows each, and the
hour lane ran once per compact_interval_ms (5 min), so a dozen piled up
in the current hour: 17 of its 24 files held fewer than 1,000 rows. Every
query over recent data opened each one over S3; a 60-minute histogram
ending now took about 5x one ending 15 minutes ago on the same rows.

Every compact_fresh_interval_ms (60 s) the scheduler now runs the hour
lane again for the quiet tables its last sweep found: those with a file
in the current or previous bucket under compact_fresh_below_bytes
(4 MiB), from Planner.fresh_tables/3. Busy tables, whose seals fill by
size, are not listed on the tick. The tick runs in the scheduler process,
so it never races a sweep over the same files, and its outcomes feed the
row caps, quarantine and backoff as a sweep's do. A table whose listing
is no longer quiet leaves the set until the next sweep. 0 turns it off.

Sweeper gains two optional callbacks, on_start/1 and handle_tick/2, for
a second timer; the retention sweeper defines neither and is unchanged.
The lane timer reports lane="fresh". SMOLQUERY_COMPACT_FRESH_INTERVAL_MS
and SMOLQUERY_COMPACT_FRESH_BELOW_BYTES set it; docs/deployment.md notes
the cost (a listing per quiet table per minute, and more swaps that can
conflict with seals, T-573).
@chasers
chasers added this pull request to stack #452 October 3, 2026 01:00
Fable review of PR 451. Nothing lost rows or broke ownership, backoff or
caps; the hour lane's ownership gate and the partial-set adjustments
behave as in a sweep.

- The tick ran the hour lane with compact_below_bytes (32 MiB), and a
  quiet table's merged current-hour file grows with every merge, so the
  tick could rewrite it every minute up to 32 MiB, carried across hours.
  The tick now plans with compact_below_bytes capped at
  compact_fresh_below_bytes (4 MiB); a bigger file waits for the sweep.
  A test seals a 20,000-row file over a 20 KB floor and two small ones;
  the tick merges the two and leaves the big file.
- The tick ran spill_gated, which logged and counted a span pause every
  minute and, when gated, planned settled files too. It now plans with the
  un-gated runtime, and does nothing when no quiet table is due.
- The integration test's :sys.get_state waited 5 s around a real merge;
  60 s now. The fresh timer is off in test config and in the scheduler
  tests' defaults, so no test gets a background merge.
- The moduledoc says caps, quarantine and conflict waits advance per tick;
  the stop log names the fresh tick; the upgrade note's cost line no
  longer claims the rewritten file is small by definition.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant