Conversation
A table that seals on age, not size, leaves one sealed file per seal. On the sandbox logs.cluster (about 110 rows a minute, three write partitions) sealed two or three files a minute of 1-30 rows each, and the hour lane ran once per compact_interval_ms (5 min), so a dozen piled up in the current hour: 17 of its 24 files held fewer than 1,000 rows. Every query over recent data opened each one over S3; a 60-minute histogram ending now took about 5x one ending 15 minutes ago on the same rows. Every compact_fresh_interval_ms (60 s) the scheduler now runs the hour lane again for the quiet tables its last sweep found: those with a file in the current or previous bucket under compact_fresh_below_bytes (4 MiB), from Planner.fresh_tables/3. Busy tables, whose seals fill by size, are not listed on the tick. The tick runs in the scheduler process, so it never races a sweep over the same files, and its outcomes feed the row caps, quarantine and backoff as a sweep's do. A table whose listing is no longer quiet leaves the set until the next sweep. 0 turns it off. Sweeper gains two optional callbacks, on_start/1 and handle_tick/2, for a second timer; the retention sweeper defines neither and is unchanged. The lane timer reports lane="fresh". SMOLQUERY_COMPACT_FRESH_INTERVAL_MS and SMOLQUERY_COMPACT_FRESH_BELOW_BYTES set it; docs/deployment.md notes the cost (a listing per quiet table per minute, and more swaps that can conflict with seals, T-573).
chasers
added this pull request to stack #452
October 3, 2026 01:00
Fable review of PR 451. Nothing lost rows or broke ownership, backoff or caps; the hour lane's ownership gate and the partial-set adjustments behave as in a sweep. - The tick ran the hour lane with compact_below_bytes (32 MiB), and a quiet table's merged current-hour file grows with every merge, so the tick could rewrite it every minute up to 32 MiB, carried across hours. The tick now plans with compact_below_bytes capped at compact_fresh_below_bytes (4 MiB); a bigger file waits for the sweep. A test seals a 20,000-row file over a 20 KB floor and two small ones; the tick merges the two and leaves the big file. - The tick ran spill_gated, which logged and counted a span pause every minute and, when gated, planned settled files too. It now plans with the un-gated runtime, and does nothing when no quiet table is due. - The integration test's :sys.get_state waited 5 s around a real merge; 60 s now. The fresh timer is off in test config and in the scheduler tests' defaults, so no test gets a background merge. - The moduledoc says caps, quarantine and conflict waits advance per tick; the stop log names the fresh tick; the upgrade note's cost line no longer claims the rewritten file is small by definition.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR: Compaction now merges a quiet table's newest files every minute, not every 5 minutes. A recent-data query opens a few files instead of a dozen tiny ones.
Tracker: T-627. Stacked on #450 (T-626).
Why
logs.cluster(about 110 rows a minute, 3 write partitions) sealed 2-3 files a minute of 1-30 rows each.compact_interval_ms(5 min). So a dozen tiny files piled up in the current hour. 17 of 24 files held fewer than 1,000 rows.What changed
Smolquery.StorageService.Scheduler. Everycompact_fresh_interval_ms(60 s) it runs the hour lane again for the quiet tables only.compact_bucket_msundercompact_fresh_below_bytes(4 MiB). NewPlanner.fresh_tables/3decides this from the listing'sbytes, with no footer reads.SMOLQUERY_COMPACT_FRESH_INTERVAL_MS(0= off) andSMOLQUERY_COMPACT_FRESH_BELOW_BYTES.Sweepergains two optional callbacks for a second timer:on_start/1andhandle_tick/2. The retention sweeper defines neither, so it does not change.lane="fresh".docs/configuration.mdrows, and adocs/deployment.mdupgrade note.How it works
compact_fresh_below_bytes, and merges them into one group. A bigger file waits for the 5-minute sweep.Why this option
You chose this option over:
Tests
PlannerTest:fresh_tables/3picks a table with a small current- or previous-hour file. It skips a table whose recent files are big, and a table whose small files are older. It returns nothing when the tick is off.SchedulerTest(integration, real DuckLake and merge): a sweep marks the quiet table; 2 new seals arrive; one fresh tick merges all 3 into one segment, and no rows are lost.SchedulerTest: a table whose recent files are not small stays out of the set.SchedulerTest(integration): a 20,000-row file over a 20 KB fresh floor stays; the tick merges only the 2 small seals.Review
A Fable review found no row loss and no broken invariant. The ownership gate, backoff, caps and quarantine behave as in a sweep. Fixed in
Review of T-627: ...:compact_fresh_below_bytes(4 MiB).Checks
mix precommit: 3,090 tests passmix cimix dialyzerWatch out
smolquery_compaction_conflicts_total.metrics.sampleswhen backlogged) is listed each minute too.compact_interval_ms, skipping its ticks until then.logs.cluster, and the 60-min histogram ending now within 1.5x of one ending 15 min ago.🤖 Generated with Claude Code