Conversation
added 2 commits
October 3, 2026 00:21
Each query job runs on its own DuckDB engine, which stops when the job
settles, and DuckDB's caches stop with it. Every job re-read from S3 the
files and footers the previous job on the same page had just read. On the
sandbox a HyperDX page spent 1,981 ms in DuckDB with a fresh engine per
query and 548 ms with the caches kept; a Grafana page 2,328 ms and
1,224 ms.
With file_cache.directory set (SMOLQUERY_QUERY_FILE_CACHE_DIR), every job
and shard engine loads the cache_httpfs community extension last and
points its on-disk cache at that directory, which all engines on the node
share. Engines stay private; only cached byte ranges are shared. The hot
tier's http(s):// reads are excluded, the cache reader's process-wide
memory cache is off, and the directory joins allowed_directories so the
cache keeps working under lockdown (without it the query still answers
but nothing is cached).
The extension bounds its cache only by free disk, which on an emptyDir is
the node's disk, so Smolquery.QueryService.FileCache keeps the directory
under max_bytes (2 GiB, SMOLQUERY_QUERY_FILE_CACHE_MAX_BYTES): every 30 s
it deletes the oldest blocks to 90% of the cap, never one younger than
10 s. smolquery_query_file_cache_bytes and the evicted counters report it.
Measured locally against a Range-serving HTTP source with 20 ms a request
(S3 itself was not available): a fresh engine's group-by over a 240 MiB
file took 385-443 ms and 98 requests with httpfs, 139-223 ms and 1 request
once the blocks were cached. A cold first read took 532-750 ms and up to
308 requests, because the extension fetches 512 KiB blocks; measure that
on S3 before turning the cache on. Six engines filling one cold cache at
once returned identical answers.
Off by default. Engine.Connection installs {name, :community} extensions
FROM community; the image installs cache_httpfs at build time. The kind
overlay turns the cache on, and a kind test checks the api pod's cache
directory fills after a sealed read over MinIO.
…r block boot Fable review of PR 450. It read the extension's source and confirmed sealed s3:// reads are cached, the hot-tier exclusion matches the original path, SigV4 passes through, every setting and cache is per DuckDB instance, and deleting a block under a live reader falls back to the object store. - A cold read splits into 512 KiB block requests, each on a thread of its own with no cap, which bypassed the request ceiling read_engine_threads sets. cache_httpfs_max_fanout_subrequest caps them at 8 per read. - The extension's disk guard was 5% of the filesystem's free space, which on an emptyDir is the node's disk; cache_httpfs_min_disk_bytes_for_cache makes it an absolute 1 GiB floor. - LOAD switches DuckDB's external file cache off for the whole engine, the hot tier included; it is switched back on. - FileCache raised in init on a directory it could not create, failing the query service's boot, and sat before EnginePool under rest_for_one. It now logs and stops, starts last, and an empty SMOLQUERY_QUERY_FILE_CACHE_DIR means off. - Docs: the extension writes a block to a temp file and renames it, and touches a block on every hit, so eviction is least recently read first; its memory cache is per instance, not process-wide; the image pins the DuckDB version, not the extension build; watch the pod's file-backed memory apart from anon.
chasers
added this pull request to stack #452
October 3, 2026 01:00
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR: An optional node-local disk cache lets every query job reuse the sealed-tier bytes and footers the previous job read. A dashboard's second query is no longer cold. Off by default.
Tracker: T-626. Bottom of the stack; T-627 goes on top.
Why
What changed
file_cachesetting on the query runtime:directory(nil = off) andmax_bytes(2 GiB).SMOLQUERY_QUERY_FILE_CACHE_DIR,SMOLQUERY_QUERY_FILE_CACHE_MAX_BYTES.cache_httpfscommunity extension, after the other extensionscache_httpfs_type = 'on_disk'and the shared cache directoryLOADturns it off)^https?://, so the hot tier is never cachedallowed_directories. Without that, lockdown lets the query run but caches nothing.Smolquery.QueryService.FileCachekeeps the directory undermax_bytes:smolquery_query_file_cache_bytes,smolquery_query_file_cache_evicted_bytes_total,smolquery_query_file_cache_evicted_files_total.Engine.Connectioninstalls{name, :community}extensions withINSTALL ... FROM community.cache_httpfsat build time, so pods never download it.Why this option
T-626 listed three options:
CREATE OR REPLACE VIEWandlock_configuration, and T-461's cache-entry leak returns.Evidence
The spike was local. S3 was not available, so it used an HTTP source with Range support and 20 ms added to each request. One 240 MiB Parquet file, and a fresh engine for each query:
allowed_directories. This PR adds it.Tests
FileCacheTest: no eviction under the cap; oldest-first eviction to 90%; young blocks kept; non-files ignored; the process creates the directory and reports the gauge.RuntimeTest: off by default; the directory joinsallowed_directories.JobEngineTest: no cache without a directory;cache_httpfsloads last, and the four statements are right (including SQL quoting).JobEngineTest(integration): a real engine installscache_httpfsfrom community, and the setting and the exclusion are applied.s3://, the api pod's cache directory holds blocks.Review
A Fable review read the extension's source at the pinned build, and ran experiments. It found no correctness bug. It confirmed that:
s3://reads are cached, and SigV4 passes throughFixed in
Review of T-626: ...:LOADturned DuckDB's external file cache off for the whole engine. Now it is turned back on.FileCachecould fail the query service's boot and sat beforeEnginePool. Now it logs and stops, starts last, and an empty env var means off.Checks
mix precommit: 3,088 tests passmix cimix dialyzerWatch out
emptyDirwith asizeLimitabove the cap). The janitor runs every 30 s, so the directory can pass the cap for up to one interval.🤖 Generated with Claude Code