Skip to content

Skip the file cache for a cold large scan, and let each job choose (T-629) - #453

Merged
chasers merged 3 commits into
mainfrom
t-629-file-cache-per-query
Oct 3, 2026
Merged

chasers merged 3 commits into
mainfrom
t-629-file-cache-per-query

Conversation

@chasers

@chasers chasers commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

TL;DR: A cold large scan now skips the file cache and runs at the no-cache speed. Each job can force the cache on or off, from the API or the query page.

Tracker: T-629. Follows #450 (T-626).

Why

  • With the T-626 cache on, a cold read waited on writing every cache block to the node's local disk.
  • On the sandbox, onebrc_v6 (1.26 GB, 7 sealed files) took 41 s cold through the cache. With no cache it took 15.5 s.
  • Its warm rerun took 1.7-4.5 s.
  • Larger blocks or more fan-out made the cold read slower, not faster.

What changed

Each job decides, after planning, what to do with the cache:

option decision
unset (auto) bypassed when the scan's live sealed files that the cache has never read add up to more than bypass_bytes (512 MiB); else used
fileCache: true used, whatever the size, so it fills the cache for the next run
fileCache: false off
  • A skipped job (bypassed or off) runs SET cache_httpfs_type = 'noop' before lockdown. Its engine then reads like plain httpfs.
  • Scatter shards get the job's decision in their request.
  • Cache index: after each sweep, FileCache indexes the cache directory. Block names carry the sealed file's name, so the index knows which files have any block cached.
  • Counting: a file counts as cached once any block of it is cached, because the cache holds only the columns a query read. The planner sums the sizes of the live files with no cached block (plan.sealed_uncached_bytes). It reads the live file list (Catalog.segment_files) only when the cache is on, the job is on auto, and the scan is over the threshold.
  • API: "fileCache": true | false on POST /v1/queries and POST /v1/jobs. The job reports fileCache: "used", "bypassed", "off", or null when there is no cache. A non-boolean value is a 400.
  • Query page: a File cache select (auto, on, off), kept in the URL, and a badge (cached, cache bypassed, cache off).
  • Metric: smolquery_query_file_cache_jobs_total{decision}.
  • Env var: SMOLQUERY_QUERY_FILE_CACHE_BYPASS_BYTES.
  • Docs: docs/api.md, docs/configuration.md, and a docs/deployment.md upgrade note that says when the cache is bypassed and how to warm a big table.

How to warm a big table

  1. Run the query once with "fileCache": true, or with File cache on in the UI. It pays the cold cost and fills the cache.
  2. Within 30 s the janitor's index sees the blocks. Later runs read through the cache.

Evidence

The noop switch, checked locally on one engine:

run time requests cache files written
noop (bypass) 396-432 ms 99 0
on_disk, cold 674 ms 307 307
on_disk, warm 117 ms 0 307

Tests

  • ✅ FileCacheTest: decide/3 in every mode and at the threshold; decision/2 from a plan, including an unsized one; statements/1; the index built from real block names (temp files ignored, unknown instance = 0).
  • ✅ ClientIntegrationTest (integration, real engines with cache_httpfs):
    • a small scan is used
    • file_cache: false is off
    • a scan over a 0-byte threshold is bypassed, and file_cache: true forces used
    • every job still answers the right rows
    • a node with no cache reports nil
  • ✅ JobsTest: fileCache takes a boolean; "off" is a 400.
  • ✅ QueryLiveTest: the select reads and writes file_cache in the URL.
  • ✅ ClientIntegrationTest: a table whose live file has a cached block is used even at a 0-byte threshold; a block of a retired file does not make the live file look warm.
  • ✅ Mutation checks: with auto never bypassing, the bypass test fails. With cached? always false, the warmed-table test fails.

Review

A Fable review verified the noop switch (no cache reads or writes, kept under lock_configuration), the block naming and index parsing (including DuckLake names with -), scatter propagation and the API and UI. Fixed in Review of T-629: ...:

  • A wide table could never be warmed. The rule compared whole-file bytes with cached bytes, but the cache holds only the columns a query read: reading 2 of 21 columns cached 5.9% of the file, so auto bypassed forever. Now a file counts once any block is cached.
  • Compaction made new files look warm. The file list included retired files, so their blocks counted for the new files. Now the planner reads the live file list.
  • Shard workers read footers through the cache before the bypass. Now the noop statement runs first, and only on a worker that has a cache, so a peer without one keeps scattering.
  • Docs: fileCache is null for a job read from history.

Checks

  • ✅ mix precommit: 3,097 tests pass
  • ✅ mix ci
  • ✅ mix dialyzer

Watch out

  • ⚠️ No background warmer. A big table stays at the no-cache speed until someone runs it with the cache on. That is the choice made on T-629.
  • ⚠️ The index is up to 30 s old, so a just-warmed table may bypass once more.
  • ⚠️ The threshold weighs whole-file sizes, not the columns a query reads. A narrow query over a wide table that has never been read through the cache can be bypassed when it would have been cheap to cache. One run with fileCache: true marks its files as read.
  • ⚠️ T-629 acceptance on the sandbox is open: onebrc_v6 cold within 20% of 15.5 s; a run with the cache on, then a fast rerun; the Grafana gain from T-626 holds.

🤖 Generated with Claude Code

Chase Granberry added 3 commits October 3, 2026 21:13
…-629)

With the T-626 cache on, a cold read waited on writing every cache block
to the node's local disk: on the sandbox onebrc_v6 (1.26 GB, 7 files)
took 41 s cold through the cache against 15.5 s with none, and larger
blocks or more fan-out made it slower. Its warm rerun took 1.7-4.5 s.

Each job now decides, after planning, whether to read through the cache
(Smolquery.QueryService.FileCache.decision/2):

- unset (auto): a scan with more than file_cache.bypass_bytes (512 MiB,
  SMOLQUERY_QUERY_FILE_CACHE_BYPASS_BYTES) of sealed bytes not yet cached
  is :bypassed; anything else is :used.
- file_cache: true reads through the cache whatever the size, so a run
  with it on warms a large table; file_cache: false is :off.

A skipped job runs SET cache_httpfs_type = 'noop' before lockdown, which
makes its engine read like plain httpfs (verified: same request count,
nothing written). Scatter shards get the decision in their request.
"Not yet cached" comes from an index the janitor rebuilds each sweep:
block names carry the sealed file's name, so the plan's sealed bytes
minus the cached bytes of its tables' files is the cold part. The index
is an ETS table per instance, bounded by the blocks in the cache.

The API takes "fileCache": true | false on POST /v1/queries and
POST /v1/jobs; the job reports fileCache (used, bypassed, off, or null
with no cache). The query page gets a File cache select (auto, on, off)
kept in the URL, and a badge. smolquery_query_file_cache_jobs_total
counts decisions. docs say when the cache is bypassed and how to warm it.
Fable review of PR 453. It verified the noop switch (no reads or writes
of the cache, kept under lock_configuration), the block naming and the
index parse (DuckLake names with '-' included), scatter propagation and
the API and UI wiring. Three problems, all fixed:

- A wide table could never be warmed. The decision compared whole-file
  sealed bytes with cached bytes, but the cache holds only the columns a
  query read: reading 2 of 21 columns cached 5.9% of the file, so the
  next auto run bypassed again, forever. A file now counts as cached once
  any block of it is in the cache, and the planner sums the sizes of the
  live files the cache has never read (plan.sealed_uncached_bytes),
  calling Catalog.segment_files only when the cache is on, the job is on
  auto and the scan passes the threshold.
- registered_through includes retired files, so after compaction the old
  files' blocks made the new ones look warm, and the next large scan went
  cold through the cache. The planner now reads the live file list.
- A shard worker read every sealed file's footer through the cache
  before its noop statement. The statement now runs first, and only on a
  worker that has a cache, so a peer without one keeps scattering.

Tests: a table whose live file has a block is used at a 0-byte
threshold; a block for a retired file leaves the live file bypassed.
docs say a file counts once any block is cached, and that fileCache is
null for a job read from history.
The T-626 kind test asserted the query node's file cache held blocks
after `count(*)` totals over the sealed tables. DuckDB can answer a
count from metadata without opening a data file, so whether any block
was written depended on the plan, and the test failed once on PR 453
with an empty cache. It now reads sum(id) from a table with a sealed
segment, which reads that file's column data through the cache.
@chasers
chasers merged commit 26794eb into main Oct 3, 2026
11 of 12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant