Skip the file cache for a cold large scan, and let each job choose (T-629) - #453
Merged
Merged
Conversation
added 3 commits
October 3, 2026 21:13
…-629) With the T-626 cache on, a cold read waited on writing every cache block to the node's local disk: on the sandbox onebrc_v6 (1.26 GB, 7 files) took 41 s cold through the cache against 15.5 s with none, and larger blocks or more fan-out made it slower. Its warm rerun took 1.7-4.5 s. Each job now decides, after planning, whether to read through the cache (Smolquery.QueryService.FileCache.decision/2): - unset (auto): a scan with more than file_cache.bypass_bytes (512 MiB, SMOLQUERY_QUERY_FILE_CACHE_BYPASS_BYTES) of sealed bytes not yet cached is :bypassed; anything else is :used. - file_cache: true reads through the cache whatever the size, so a run with it on warms a large table; file_cache: false is :off. A skipped job runs SET cache_httpfs_type = 'noop' before lockdown, which makes its engine read like plain httpfs (verified: same request count, nothing written). Scatter shards get the decision in their request. "Not yet cached" comes from an index the janitor rebuilds each sweep: block names carry the sealed file's name, so the plan's sealed bytes minus the cached bytes of its tables' files is the cold part. The index is an ETS table per instance, bounded by the blocks in the cache. The API takes "fileCache": true | false on POST /v1/queries and POST /v1/jobs; the job reports fileCache (used, bypassed, off, or null with no cache). The query page gets a File cache select (auto, on, off) kept in the URL, and a badge. smolquery_query_file_cache_jobs_total counts decisions. docs say when the cache is bypassed and how to warm it.
Fable review of PR 453. It verified the noop switch (no reads or writes of the cache, kept under lock_configuration), the block naming and the index parse (DuckLake names with '-' included), scatter propagation and the API and UI wiring. Three problems, all fixed: - A wide table could never be warmed. The decision compared whole-file sealed bytes with cached bytes, but the cache holds only the columns a query read: reading 2 of 21 columns cached 5.9% of the file, so the next auto run bypassed again, forever. A file now counts as cached once any block of it is in the cache, and the planner sums the sizes of the live files the cache has never read (plan.sealed_uncached_bytes), calling Catalog.segment_files only when the cache is on, the job is on auto and the scan passes the threshold. - registered_through includes retired files, so after compaction the old files' blocks made the new ones look warm, and the next large scan went cold through the cache. The planner now reads the live file list. - A shard worker read every sealed file's footer through the cache before its noop statement. The statement now runs first, and only on a worker that has a cache, so a peer without one keeps scattering. Tests: a table whose live file has a block is used at a 0-byte threshold; a block for a retired file leaves the live file bypassed. docs say a file counts once any block is cached, and that fileCache is null for a job read from history.
The T-626 kind test asserted the query node's file cache held blocks after `count(*)` totals over the sealed tables. DuckDB can answer a count from metadata without opening a data file, so whether any block was written depended on the plan, and the test failed once on PR 453 with an empty cache. It now reads sum(id) from a table with a sealed segment, which reads that file's column data through the cache.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR: A cold large scan now skips the file cache and runs at the no-cache speed. Each job can force the cache on or off, from the API or the query page.
Tracker: T-629. Follows #450 (T-626).
Why
onebrc_v6(1.26 GB, 7 sealed files) took 41 s cold through the cache. With no cache it took 15.5 s.What changed
Each job decides, after planning, what to do with the cache:
bypassedwhen the scan's live sealed files that the cache has never read add up to more thanbypass_bytes(512 MiB); elseusedfileCache: trueused, whatever the size, so it fills the cache for the next runfileCache: falseoffbypassedoroff) runsSET cache_httpfs_type = 'noop'before lockdown. Its engine then reads like plainhttpfs.FileCacheindexes the cache directory. Block names carry the sealed file's name, so the index knows which files have any block cached.plan.sealed_uncached_bytes). It reads the live file list (Catalog.segment_files) only when the cache is on, the job is on auto, and the scan is over the threshold."fileCache": true | falseonPOST /v1/queriesandPOST /v1/jobs. The job reportsfileCache:"used","bypassed","off", ornullwhen there is no cache. A non-boolean value is a 400.smolquery_query_file_cache_jobs_total{decision}.SMOLQUERY_QUERY_FILE_CACHE_BYPASS_BYTES.docs/api.md,docs/configuration.md, and adocs/deployment.mdupgrade note that says when the cache is bypassed and how to warm a big table.How to warm a big table
"fileCache": true, or with File cache on in the UI. It pays the cold cost and fills the cache.Evidence
The
noopswitch, checked locally on one engine:noop(bypass)on_disk, coldon_disk, warmTests
FileCacheTest:decide/3in every mode and at the threshold;decision/2from a plan, including an unsized one;statements/1; the index built from real block names (temp files ignored, unknown instance = 0).ClientIntegrationTest(integration, real engines withcache_httpfs):usedfile_cache: falseisoffbypassed, andfile_cache: trueforcesusednilJobsTest:fileCachetakes a boolean;"off"is a 400.QueryLiveTest: the select reads and writesfile_cachein the URL.ClientIntegrationTest: a table whose live file has a cached block isusedeven at a 0-byte threshold; a block of a retired file does not make the live file look warm.cached?always false, the warmed-table test fails.Review
A Fable review verified the
noopswitch (no cache reads or writes, kept underlock_configuration), the block naming and index parsing (including DuckLake names with-), scatter propagation and the API and UI. Fixed inReview of T-629: ...:noopstatement runs first, and only on a worker that has a cache, so a peer without one keeps scattering.fileCacheisnullfor a job read from history.Checks
mix precommit: 3,097 tests passmix cimix dialyzerWatch out
fileCache: truemarks its files as read.🤖 Generated with Claude Code