Skip to content

Share sealed-tier reads across query jobs with an on-disk cache (T-626) - #450

Open
chasers wants to merge 2 commits into
mainfrom
t-626-shared-file-cache
Open

chasers wants to merge 2 commits into
mainfrom
t-626-shared-file-cache

Conversation

@chasers

@chasers chasers commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

TL;DR: An optional node-local disk cache lets every query job reuse the sealed-tier bytes and footers the previous job read. A dashboard's second query is no longer cold. Off by default.

Tracker: T-626. Bottom of the stack; T-627 goes on top.

Why

  • Each query job runs on its own DuckDB engine. The engine stops when the job ends.
  • DuckDB's caches stop with it.
  • So every job read again from S3 the files and Parquet footers that the previous job had just read.
  • Sandbox numbers (from T-626): a HyperDX page spent 1,981 ms in DuckDB with a fresh engine per query, and 548 ms with the caches kept. A Grafana page: 2,328 ms and 1,224 ms.

What changed

  • New file_cache setting on the query runtime: directory (nil = off) and max_bytes (2 GiB).
    • Env vars: SMOLQUERY_QUERY_FILE_CACHE_DIR, SMOLQUERY_QUERY_FILE_CACHE_MAX_BYTES.
  • When it is on, every job and shard engine:
    1. loads the cache_httpfs community extension, after the other extensions
    2. sets cache_httpfs_type = 'on_disk' and the shared cache directory
    3. turns off the cache reader's own memory cache (hits come from the page cache)
    4. caps a cold read at 8 parallel block requests, and sets an absolute 1 GiB disk floor
    5. turns DuckDB's external file cache back on (LOAD turns it off)
    6. excludes ^https?://, so the hot tier is never cached
  • The cache directory joins allowed_directories. Without that, lockdown lets the query run but caches nothing.
  • New Smolquery.QueryService.FileCache keeps the directory under max_bytes:
    • every 30 s it deletes the oldest blocks until the directory is at 90% of the cap
    • it never deletes a block younger than 10 s
  • New metrics: smolquery_query_file_cache_bytes, smolquery_query_file_cache_evicted_bytes_total, smolquery_query_file_cache_evicted_files_total.
  • Engine.Connection installs {name, :community} extensions with INSTALL ... FROM community.
  • The Dockerfile installs cache_httpfs at build time, so pods never download it.
  • The kind overlay turns the cache on.

Why this option

T-626 listed three options:

  1. A shared on-disk cache (this PR). Each job keeps its private engine and lockdown, and the bound lives outside DuckDB.
  2. Long-lived engines reused across jobs. This conflicts with per-job CREATE OR REPLACE VIEW and lock_configuration, and T-461's cache-entry leak returns.
  3. Metadata caches only. This does not help across jobs while engines are per job.

Evidence

The spike was local. S3 was not available, so it used an HTTP source with Range support and 20 ms added to each request. One 240 MiB Parquet file, and a fresh engine for each query:

requests time
httpfs (today), every query 98 385-443 ms
cache_httpfs, warm cache 1 139-223 ms
cache_httpfs, cold cache up to 308 532-750 ms
  • 6 engines that filled one cold cache at the same time returned identical answers.
  • Under the job lockdown, the cache fills only when its directory is in allowed_directories. This PR adds it.
  • Exclusion regexes are per engine, so adding one per engine does not grow anything.

Tests

  • ✅ FileCacheTest: no eviction under the cap; oldest-first eviction to 90%; young blocks kept; non-files ignored; the process creates the directory and reports the gauge.
  • ✅ RuntimeTest: off by default; the directory joins allowed_directories.
  • ✅ JobEngineTest: no cache without a directory; cache_httpfs loads last, and the four statements are right (including SQL quoting).
  • ✅ JobEngineTest (integration): a real engine installs cache_httpfs from community, and the setting and the exclusion are applied.
  • ✅ New kind test: after a sealed read over MinIO s3://, the api pod's cache directory holds blocks.

Review

A Fable review read the extension's source at the pinned build, and ran experiments. It found no correctness bug. It confirmed that:

  • sealed s3:// reads are cached, and SigV4 passes through
  • the hot-tier exclusion matches the original path, so sealed reads are not excluded
  • every setting and cache is per DuckDB instance; nothing grows across engines
  • deleting a block under a live reader falls back to S3, and two processes on one directory are safe

Fixed in Review of T-626: ...:

  • A cold read fanned out into unbounded 512 KiB sub-requests on its own threads. Now it is capped at 8 per read.
  • The extension's disk guard used 5% of the node disk. Now it is an absolute 1 GiB floor.
  • LOAD turned DuckDB's external file cache off for the whole engine. Now it is turned back on.
  • FileCache could fail the query service's boot and sat before EnginePool. Now it logs and stops, starts last, and an empty env var means off.
  • Docs: eviction is least recently read (the extension touches a block on each hit); the image pins the DuckDB version, not the extension build; watch file-backed memory apart from anon.

Checks

  • ✅ mix precommit: 3,088 tests pass
  • ✅ mix ci
  • ✅ mix dialyzer

Watch out

  • ⚠️ Cold reads may get slower. The extension fetches in 512 KiB blocks, so the first read of a file can make about 3x the requests. Measure a dashboard's first and second query on the sandbox (S3) before turning the cache on.
  • ⚠️ A community extension in the query path. DuckDB signs it, and the image pins it to this DuckDB version. A deployment that builds its own image must install it, or keep the cache off.
  • ⚠️ Give the cache its own volume (an emptyDir with a sizeLimit above the cap). The janitor runs every 30 s, so the directory can pass the cap for up to one interval.
  • ⚠️ T-626 acceptance is open: replay the two query sets on the sandbox, and watch RSS and cache disk over a day.

🤖 Generated with Claude Code

Chase Granberry added 2 commits October 3, 2026 00:21
Each query job runs on its own DuckDB engine, which stops when the job
settles, and DuckDB's caches stop with it. Every job re-read from S3 the
files and footers the previous job on the same page had just read. On the
sandbox a HyperDX page spent 1,981 ms in DuckDB with a fresh engine per
query and 548 ms with the caches kept; a Grafana page 2,328 ms and
1,224 ms.

With file_cache.directory set (SMOLQUERY_QUERY_FILE_CACHE_DIR), every job
and shard engine loads the cache_httpfs community extension last and
points its on-disk cache at that directory, which all engines on the node
share. Engines stay private; only cached byte ranges are shared. The hot
tier's http(s):// reads are excluded, the cache reader's process-wide
memory cache is off, and the directory joins allowed_directories so the
cache keeps working under lockdown (without it the query still answers
but nothing is cached).

The extension bounds its cache only by free disk, which on an emptyDir is
the node's disk, so Smolquery.QueryService.FileCache keeps the directory
under max_bytes (2 GiB, SMOLQUERY_QUERY_FILE_CACHE_MAX_BYTES): every 30 s
it deletes the oldest blocks to 90% of the cap, never one younger than
10 s. smolquery_query_file_cache_bytes and the evicted counters report it.

Measured locally against a Range-serving HTTP source with 20 ms a request
(S3 itself was not available): a fresh engine's group-by over a 240 MiB
file took 385-443 ms and 98 requests with httpfs, 139-223 ms and 1 request
once the blocks were cached. A cold first read took 532-750 ms and up to
308 requests, because the extension fetches 512 KiB blocks; measure that
on S3 before turning the cache on. Six engines filling one cold cache at
once returned identical answers.

Off by default. Engine.Connection installs {name, :community} extensions
FROM community; the image installs cache_httpfs at build time. The kind
overlay turns the cache on, and a kind test checks the api pod's cache
directory fills after a sealed read over MinIO.
…r block boot

Fable review of PR 450. It read the extension's source and confirmed
sealed s3:// reads are cached, the hot-tier exclusion matches the original
path, SigV4 passes through, every setting and cache is per DuckDB
instance, and deleting a block under a live reader falls back to the
object store.

- A cold read splits into 512 KiB block requests, each on a thread of its
  own with no cap, which bypassed the request ceiling read_engine_threads
  sets. cache_httpfs_max_fanout_subrequest caps them at 8 per read.
- The extension's disk guard was 5% of the filesystem's free space, which
  on an emptyDir is the node's disk; cache_httpfs_min_disk_bytes_for_cache
  makes it an absolute 1 GiB floor.
- LOAD switches DuckDB's external file cache off for the whole engine,
  the hot tier included; it is switched back on.
- FileCache raised in init on a directory it could not create, failing
  the query service's boot, and sat before EnginePool under rest_for_one.
  It now logs and stops, starts last, and an empty
  SMOLQUERY_QUERY_FILE_CACHE_DIR means off.
- Docs: the extension writes a block to a temp file and renames it, and
  touches a block on every hit, so eviction is least recently read first;
  its memory cache is per instance, not process-wide; the image pins the
  DuckDB version, not the extension build; watch the pod's file-backed
  memory apart from anon.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant