feat: injectable redis client and idle-connection handling for state persistence - #1209
Open
maartenbreddels wants to merge 1 commit into
Open
maartenbreddels wants to merge 1 commit into
maartenbreddels wants to merge 1 commit into
Conversation
maartenbreddels
force-pushed
the
feat/state-redis-client-factory
branch
from
September 29, 2026 12:56
dc92f8e to
a8f689a
Compare
…persistence Managed Redis servers and NATs drop connections that sat idle for a few minutes. The built-in client had no health check and no retry, so the next takeover on such a connection failed with "connection reset by peer" and the user came back to a fresh session instead of the restored one. The built-in client now pings a connection that sat idle and retries once, immediately, on a connection error. It never retries a timeout: a takeover that completes after the connect deadline deletes the stored state, so a retry must not outlive that deadline. Deployments that need their own pool, keepalive or TLS setup can now point SOLARA_STATE_REDIS_CLIENT_FACTORY at a function that builds the client, instead of overriding a private method. A bad path stops the server at startup rather than failing every restore. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
maartenbreddels
force-pushed
the
feat/state-redis-client-factory
branch
from
September 29, 2026 13:21
a8f689a to
a9b5d08
Compare
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Managed Redis servers (and NATs in between) drop connections that sat idle for a few minutes. The Redis state backend built its client with
from_url(...)and no health check. Withfrom_url, redis-py gives each connection zero retries. So the next takeover on such a connection failed withConnectionError: Error while reading from ...: Connection reset by peer, and the user got a fresh session instead of the restored one.Built-in client
health_check_interval=30: a connection that sat idle is pinged before it is reused.Retry(NoBackoff(), 1, supported_errors=(redis.ConnectionError,)): one immediate retry on a connection error.connect_timeoutdeletes the stored state (the claim-or-delete path inkernel_context), so a retry must not outlive the deadline.supported_errorsis set explicitly because redis-py's default includesTimeoutError.retry_on_erroris set as well because redis-py 4.x ignoresretrywithout it.SOLARA_STATE_REDIS_CLIENT_FACTORYA dotted path to a
callable(settings) -> redis.Redis. It replaces the built-in client, for deployments that need their own pool, keepalive, TLS or Sentinel setup.validate_state_settings()imports it at startup, so a typo stops the server instead of failing every restore. The docs describe the contract: synchronous, thread-safe,decode_responses=False, bounded by the connect deadline, and dedicated to Solara.No
set_backend()setter: the backend is a lazy process-wide singleton, and swapping it while kernels hold the old one would leave two backends with separate breaker state.Known and accepted
The retry covers every command, so a lost reply also re-runs a delete. A fenced delete then returns
Falsealthough the first attempt deleted the key; that value only goes to a log line. For a retried delete to remove new state, another instance would have to recreate the same kernel's key inside the millisecond gap between the lost reply and the retry. Deletes happen on a tab close and on a poisoned state for that same kernel, so I think this is acceptable.Tests
tests/unit/state_redis_client_test.pyputs a small TCP proxy between the real redis-py client and a fakeredis TCP server:Error while reading from 127.0.0.1:...: (54, 'Connection reset by peer')).TimeoutError, with no reconnect and no second send.Note: a connection closed with a FIN before the next command is already handled by redis-py itself (
can_read()check on checkout). The failing case is a drop the client never saw, which is why the proxy sends the RST in reply to the request.I also checked end to end against a local Redis, with an app-side factory loaded through the env var: a takeover after
CLIENT KILLrestores the state, and the generation goes up by exactly one.🤖 Generated with Claude Code