Skip to content

Add SMOLQUERY_JOB_MEMORY_LIMIT to size query job engines (T-631) - #455

Merged
chasers merged 2 commits into
t-630-shard-file-cache-decisionfrom
t-631-job-memory-limit-env
Oct 4, 2026
Merged

chasers merged 2 commits into
t-630-shard-file-cache-decisionfrom
t-631-job-memory-limit-env

Conversation

@chasers

@chasers chasers commented Oct 4, 2026 •

Copy link
Copy Markdown
Owner

TL;DR: Add SMOLQUERY_JOB_MEMORY_LIMIT. It sets each query job engine's DuckDB memory limit from the environment.

Tracker: T-631. Stacked on PR 454 (T-630).

Why

  • job_memory_limit was set only in compiled config, at 1GB.
  • Every other engine limit already has an env var.
  • A deploy could not size query engines without a rebuild.

What changed

  • config/runtime.exs: SMOLQUERY_JOB_MEMORY_LIMIT sets Smolquery.QueryService job_memory_limit.
  • Unset, nothing changes. Job engines keep 1GB.
  • Scatter workers still take this value, unless SMOLQUERY_DISTRIBUTED_WORKER_MEMORY_LIMIT is set.
  • Docs: a row in docs/configuration.md, and a 0.22.0 note in docs/deployment.md.

Budget

Engine Limit
One job engine SMOLQUERY_JOB_MEMORY_LIMIT
One node, all jobs that × max_concurrent_jobs
One scatter worker SMOLQUERY_DISTRIBUTED_WORKER_MEMORY_LIMIT, else SMOLQUERY_JOB_MEMORY_LIMIT

Tests

  • ✅ runtime_config_test: the variable reaches job_memory_limit, and an empty value fails the boot.
  • ✅ RuntimeConfig.size!/2: accepts 4GB, 512 MiB, 1.5GB; refuses empty, words, negatives.

Checks

  • ✅ mix precommit, mix ci, mix dialyzer pass locally.

Review

/code-review high found one issue. Fixed in "Review of T-631":

  • ✅ The value went in unchecked. Now the boot refuses a value that is not a number and an optional unit (RuntimeConfig.size!/2).

Watch out

  • ⚠️ DuckDB still checks the unit. A number with a unit DuckDB refuses makes every job engine fail to start.
  • The other memory limit variables are still not checked. This PR does not change them.

🤖 Generated with Claude Code

@chasers
chasers added this pull request to stack #456 October 4, 2026 01:45
Chase Granberry added 2 commits October 4, 2026 02:01
job_memory_limit, each query job engine's DuckDB memory_limit, was set
only in compiled config at 1GB. Every other engine limit already has an
environment variable (SMOLQUERY_MEMORY_LIMIT, SMOLQUERY_STORAGE_MEMORY_LIMIT,
SMOLQUERY_DISTRIBUTED_WORKER_MEMORY_LIMIT, SMOLQUERY_WRITE_ENGINE_MEMORY_LIMIT),
so a deploy could not size the query side without a rebuild.

SMOLQUERY_JOB_MEMORY_LIMIT now sets it in config/runtime.exs. Unset,
nothing changes. A scatter worker engine still inherits it whole unless
SMOLQUERY_DISTRIBUTED_WORKER_MEMORY_LIMIT is set. Like its siblings, the
size string passes through as given; docs say a value DuckDB refuses
fails every job engine as it starts.

docs/configuration.md gets the row and points the worker limit's default
at it; docs/deployment.md gets a 0.22.0 upgrade note.
Code review of PR 455 (/code-review high). SMOLQUERY_JOB_MEMORY_LIMIT
passed through unchecked, so an empty value from an unfilled template, or
a word, booted a healthy-looking node whose every job engine then failed
to start, warm-pool rebuilds included.

RuntimeConfig.size!/2 now checks the shape at boot: a number and an
optional unit. DuckDB still owns the unit grammar, so the check refuses
only what can never be a size. The other memory limit variables keep
passing through as before.
@chasers
chasers force-pushed the t-631-job-memory-limit-env branch from f22e263 to 7dafe9b Compare October 4, 2026 02:06
@chasers
chasers merged commit 947dfa0 into main Oct 4, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant