Skip to content

Keep result-set filesystems out of the fsspec instance cache and read pandas results through one filesystem - #1001

Merged
laughingman7743 merged 6 commits into
masterfrom
fix/978-instance-cache-leak
Oct 3, 2026
Merged

laughingman7743 merged 6 commits into
masterfrom
fix/978-instance-cache-leak

Conversation

@laughingman7743

@laughingman7743 laughingman7743 commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

WHAT

Result sets and AioS3FileSystem no longer put the filesystems they create into fsspec's instance cache, and the pandas result set reads through a single filesystem:

  • The pandas result set creates its S3FileSystem with skip_instance_cache=True and reads through it:
    • The CSV output is opened with that filesystem and passed to pd.read_csv() as a file, for both the plain and the binary-column paths. When storage_options are given in the read options (execute(..., storage_options=...)), even None, the path and these options go to pd.read_csv() (or, for binary columns, fsspec.open()) as before.
    • pd.read_parquet() receives that filesystem as filesystem, with the unload location as bucket/key (pyarrow rejects the s3:// form with an fsspec filesystem: "GetFileInfo() yielded path ... which is outside base dir").
    • Previously pd.read_csv() and pd.read_parquet() each built another S3FileSystem from storage_options.
  • The S3FS result set creates its filesystem (filesystem_class) with skip_instance_cache=True.
  • The Polars result set passes skip_instance_cache=True in the storage_options of pl.read_csv(). The Parquet and scan_csv() paths use Polars' native object store, not fsspec, and are unchanged.
  • AioS3FileSystem always creates its internal S3FileSystem with skip_instance_cache=True. fsspec already caches the AioS3FileSystem itself unless skip_instance_cache=True is given to it, so the internal instance follows the lifetime of its owner.

A result set's filesystem, its S3 client and its dircache are now released with the result set, and a closed connection can be garbage-collected. Each result set builds one S3FileSystem (measured: one per query for pandas CSV, pandas unload=True, Polars and S3FS; pandas previously built two).

Behavior changes (release-note):

  • The filesystem a result set reads through now has its own S3 client and HTTPS connections, instead of the client cached for the connection and thread. The base result set already builds a new S3 client per result set for HeadObject/GetObject (pyathena/result_set.py:105), so a query now opens one more HTTPS connection to S3 than before. Measured from a laptop to us-west-2: the first request of a new client took a median of about 450 ms, a request on an open connection about 150 ms; building the client took about 2.3 ms. The cost inside the region was not measured.
  • AioS3FileSystem(skip_instance_cache=True) instances created with the same arguments no longer share their internal S3FileSystem and dircache.
  • Without user storage_options, the pandas result set owns the CSV file it reads for every CSV result, not only for binary columns, and closes it with the chunk iterator (PandasDataFrameIterator.close()).
  • storage_options given by the user are used as given; a connection in them is cached by fsspec unless they also contain skip_instance_cache: True (unchanged).

WHY

Closes #978.

S3FileSystem is cachable and fsspec keys its instance cache on the constructor arguments, which include the connection object. A program that opens a connection per query kept one cached filesystem, with its connection, boto3 session and clients, per connection; one long-lived connection reused a single cached instance whose dircache gained an entry for every result file. Closed #417 (memory growth with about 1,000 PandasCursor queries per hour) may have been this problem.

TEST

Tested commit: 75ef9c266b513ca16fe613929e74c87a85db4f88 (based on 4e7b55ff).

  • just lint: passed.
  • New tests:
    • TestAioS3FileSystem.test_internal_file_system_not_cached (offline): two skip_instance_cache=True instances do not share the internal filesystem or dircache, and no internal filesystem enters S3FileSystem._cache.
    • TestPandasCursor.test_result_set_file_system (AWS; CSV, chunked CSV, unload=True): one S3FileSystem is built per result set, no filesystem holding the connection is in the S3FileSystem or AioS3FileSystem instance cache, and the CSV file is closed after fetchall().
    • TestPandasCursor.test_csv_storage_options (AWS; plain/binary × None/options): with storage_options in the read options, the result set's filesystem never opens the output, given options reach the filesystem that does, and pandas opens the plain file itself.
    • test_result_set_file_system_not_cached in TestPolarsCursor and TestS3FSCursor (AWS): no filesystem holding the connection is in either instance cache.
  • Failing before the fix:
    • Polars, S3FS and the AIO test fail with the source changes reverted.
    • With the earlier skip_instance_cache-only pandas change, the pandas filesystem test fails with assert 2 == 1 in all three cases. test_csv_storage_options fails on d232af0's pandas source (none-binary, none-plain, options-plain) and on c42670b's (none-binary).
    • With TestAioS3FileSystem run first in the same process (it registers AioS3FileSystem for s3 and does not restore it), the Polars test fails when its skip_instance_cache option is removed.
  • uv run --env-file .env pytest -n 4 tests/pyathena/pandas/ tests/pyathena/aio/pandas/ tests/pyathena/filesystem/test_s3_async.py tests/pyathena/polars/test_cursor.py tests/pyathena/s3fs/test_cursor.py on e1f4d470 (same patches, previous base 9b2f0337): 471 passed, 1 skipped.
  • After rebasing onto 4e7b55ff (the only range-diff change is the context of the AIO test), pytest -n 4 tests/pyathena/filesystem/test_s3_async.py tests/pyathena/{pandas,polars,s3fs}/test_cursor.py: 339 passed, 1 skipped.
  • Not run locally: the remaining Polars and S3FS tests (async/aio cursors, result sets) and the SQLAlchemy suites (left to the AWS CI after Ready).

🤖 Generated with Claude Code

Comment thread tests/pyathena/util.py
return [q.strip() for q in template.render(**kwargs).split(";") if q and q.strip()]


def cached_file_systems(connection):

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round 1 (behavior and implementation) — FINDINGS, repaired

Scope: git diff 0b656b8d6de41b1f87bdb4d1e670e00a480df895..eebef4099242b968b0db6f565b5c679b219cc3d3 (full initial diff, 8 files).

Covered:

  • pyathena/pandas/result_set.py: _create_s3_file_system(), the pd.read_csv() storage_options (also used by the binary-column filesystem_open() stream) and the pd.read_parquet() storage_options. fsspec's _Cached.__call__ pops skip_instance_cache before __init__, and pandas passes storage_options to fsspec.open() / url_to_fs(), so the key never reaches S3FileSystem.__init__ or S3File.
  • pyathena/s3fs/result_set.py: filesystem_class is typed type[AbstractFileSystem], so the fsspec metaclass handles the flag for any accepted class.
  • pyathena/polars/result_set.py: only pl.read_csv() goes through fsspec (polars 1.44.2 io/csv/functions.py); Parquet and scan_csv() use object_store and are unaffected.
  • pyathena/filesystem/s3_async.py: the internal S3FileSystem follows its owner's lifetime; **kwargs cannot carry skip_instance_cache because fsspec already consumed it.
  • Sync, async and aio cursors build the same result-set classes, so they share the change. Arrow uses pyarrow's own filesystem and is not affected.

Finding: TestAioS3FileSystem registers AioS3FileSystem for s3/s3a with a class-scoped fixture (tests/pyathena/filesystem/test_s3_async.py:31-38) and never restores it. In an xdist worker that ran it first, pandas and Polars read through a cached AioS3FileSystem, which the new tests did not inspect (S3FileSystem._cache only), so a regression in the Polars storage_options could pass there.
Repair (28738a8): this cached_file_systems() helper checks both instance caches. Verified with pytest -n 0 running TestAioS3FileSystem::test_parse_path first: the Polars test fails with its skip_instance_cache line removed and passes with it.

Out of scope: the leaked registration itself is pre-existing test isolation behavior.

Validation after repair: just lint passed; pytest -n 2 tests/pyathena/{pandas,polars,s3fs}/test_cursor.py -k "file_system_not_cached or binary_null_vs_empty": 8 passed.

@laughingman7743 laughingman7743 changed the title Keep result-set and internal S3FileSystem instances out of the fsspec instance cache Keep result-set filesystems out of the fsspec instance cache and read pandas results through one filesystem Oct 3, 2026
Comment thread pyathena/pandas/result_set.py Outdated
read_csv_kwargs["converters"] = converters
return binary_columns

def _open_output_location(self, storage_options: dict[str, Any] | None, **kwargs: Any) -> Any:

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round 1 follow-up (behavior and implementation) for the round-two repair d5326d0 — CLEAN

Scope: git diff dcc9600ee756d9608953233999a4aa4cd09cec1c..d5326d096ee6115917b62a995b06095134b2ff7c (pyathena/pandas/result_set.py, tests/pyathena/pandas/test_cursor.py), traced into _read_csv(), _open_binary_csv_stream(), _read_parquet(), PandasDataFrameIterator.close() and AthenaPandasResultSet.close().

Checked:

  • Plain CSV path: self._fs.open(path, mode="rb") gives a binary S3File; pandas wraps a binary handle in text for the C/Python engines and passes it to pyarrow for the pyarrow engine, as it does for the file it opens from a path. Non-chunked reads close the file on ExitStack exit; chunked reads hand it to the iterator (stack.pop_all()), which closes it on exhaustion, error or close(). S3File.close() is idempotent, so the later close of an already closed file is harmless.
  • Binary path: AbstractFileSystem.open() wraps the binary file in TextIOWrapper(encoding="utf-8", newline="") for mode="rt", matching the previous fsspec.open() call. S3FileSystem does not override open().
  • User storage_options: None (absent) uses the result set's filesystem; any given dict, including {}, goes through fsspec.open() as pd.read_csv() did before. No other caller of _get_csv_read_options() reads storage_options.
  • Parquet: with an fsspec filesystem, pandas 3.0.6 _get_path_or_handle() keeps the path, and pyarrow rejected the s3:// form, so the path is bucket/key as in _read_parquet_schema(). A user filesystem or storage_options kwarg raised before (duplicate keyword or pandas ValueError) and still raises.
  • Errors from self._fs.open() are inside the try and still become OperationalError.

Validation: pytest -n 4 tests/pyathena/pandas/ tests/pyathena/aio/pandas/ on this source: 288 passed. New test_result_set_file_system: 3 passed; with dcc9600's pandas source, 3 failed (assert 2 == 1).

default_block_size=self._block_size,
default_cache_type=self._cache_type,
max_workers=self._max_workers,
skip_instance_cache=True,

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round 2 (claims, compatibility, operations) — FINDINGS, repaired

Scope: git diff 0b656b8d6de41b1f87bdb4d1e670e00a480df895..28738a8dc7c963d8c28393ad08772d34d219512e for the initial pass; the PR description, commit messages and code comments.

Claims checked:

  • "Result sets build one filesystem per query / pandas builds two": measured by wrapping S3FileSystem.__init__ around live queries. Before the repair it was pandas 2 per query (CSV and unload), Polars 1 and S3FS 1; after d5326d0 it is 1 for all four.
  • "fsspec caches per connection and thread": fsspec 2026.9.0 _Cached.__call__ adds threading.get_ident() to the token for sync filesystems. Confirmed.
  • Client build cost about 2.3 ms: measured locally.
  • Finding: the description presented per-query client creation as a pure CPU cost. It also gives up HTTPS connection reuse. Measured from a laptop to us-west-2: median 447 ms for the first request on a new client, 154 ms on an open connection. It also missed that pyathena/result_set.py:105 already builds a new S3 client per result set, so the real increase is one new connection per query. With the pandas second filesystem it was two.
    Repair: the maintainer chose to reduce pandas to one filesystem (d5326d0). The description now states the measured connection cost and its limits (not measured inside the region).
  • Finding: "An AioS3FileSystem no longer shares its dircache with an S3FileSystem created with the same arguments" was misleading. The internal instance's token contains every keyword AioS3FileSystem passes, so only an S3FileSystem with exactly those keywords shared it. Repair: the claim now names only skip_instance_cache=True instances with the same arguments.
  • Comment "fsspec caches this instance itself" in s3_async.py was ambiguous inside the S3FileSystem(...) call; it now names AioS3FileSystem (dcc9600).
  • Polars read_csv() goes through fsspec (polars 1.44.2 io/csv/functions.py) and scan_csv()/Parquet through object_store (pyathena/polars/result_set.py:619-626). Confirmed; docs/polars.md:15 stays accurate.
  • PyAthena - Memory Issue #417 is cited only as "may have been this problem".

Compatibility: no public signature changed. Users passing storage_options keep the fsspec path, and filesystem_class is typed type[AbstractFileSystem]. The docs (docs/pandas.md, docs/s3fs.md, docs/aio.md) make no claim about instance caching.

@laughingman7743
laughingman7743 force-pushed the fix/978-instance-cache-leak branch from d5326d0 to 4484e16 Compare October 3, 2026 10:24
self._open_binary_csv_stream(binary_columns, storage_options)
)
source = self._csv_stream
elif "storage_options" not in read_csv_kwargs:

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent review (relayed) — FINDINGS

Reviewer: Codex CLI 0.160.0, model gpt-6-astra, reasoning effort max, sandbox read-only, session 01a1013c-da33-7313-b62b-ed4961777a1f.
Scope: git diff 0b656b8d6de41b1f87bdb4d1e670e00a480df895..d5326d096ee6115917b62a995b06095134b2ff7c, reviewed in a detached snapshot at d5326d0. The prompt carried the issue's intended behavior but no PR number, description, commit messages or earlier findings. The reviewer was told not to edit, build, run tests, use GitHub or the network. The review snapshot and the PR worktree were both unchanged afterwards. This is a static review.

Reviewer output (verbatim)

Covered surfaces:

  • pandas CSV: plain/chunked/binary reads, C/Python/pyarrow engines, .txt, read kwargs, and stream cleanup.
  • pandas Parquet: filesystem=, scheme-less paths, and schema reads.
  • S3FS/Polars result sets and sync, Future-based async, and aio callers.
  • fsspec cache semantics, installed dependency sources, supported ranges, and new tests’ failure and ordering behavior.

FINDINGS

  1. P2 — Explicit CSV storage options still retain connections. pyathena/pandas/result_set.py:751
    With storage_options={"connection": conn, "default_cache_type": "none"}, this branch creates a filesystem without skip_instance_cache=True. Its cached storage_options retains conn, and its dircache survives disposal of the result set. The existing test_binary_as_pandas uses exactly these options but does not check cache membership. This is the original defect left unfixed for explicit options, rather than a new regression.

  2. P2 — Chunked CSV reads discard the owner responsible for closing underlying streams. pyathena/pandas/result_set.py:571
    For gzip-compressed content at a .csv output location, using chunksize=1 and storage_options={"connection": conn, "compression": "gzip", "skip_instance_cache": True} returns a GzipFile from fsspec’s OpenFile. The subsequent stack.pop_all() discards the OpenFile cleanup context. Iterator exhaustion, close, or a read exception closes only the gzip wrapper; GzipFile.close() does not close its supplied underlying S3File. That file and its executor remain open while the retained wrapper references them. Previously pandas retained and closed the OpenFile. New regression for ordinary CSV reads; analogous binary-path behavior predates this diff. Transfer the cleanup context alongside the stream.

  3. P2 — Ordinary CSV reads lose pandas’ anonymous-access fallback. pyathena/pandas/result_set.py:572
    Given a connection with explicit credentials, publicly readable results, no ambient AWS credentials, and storage_options={}, metadata lookup succeeds through the connection. Previously pandas retried the separate filesystem open with anon=True after NoCredentialsError. Opening outside pandas now raises OperationalError instead. New regression for ordinary CSV reads. Preserve that fallback when handling explicit storage options.

The new cache assertions would catch the original default-path defects. Fresh connection identities and checking both filesystem caches avoid the identified registry-order concern under pytest-xdist. They do not cover the explicit-options cases above.

Static review only; no edits, builds, tests, GitHub calls, or network access.

Author verification:

  1. Explicit storage_options with connection stay cached. Confirmed, but not acted on. The options are the caller's: pandas also cached them before this PR, and adding skip_instance_cache: True to them avoids it, as the new test_csv_storage_options does. Pre-existing and outside Cached S3FileSystem instances keep every connection alive and grow their dircache without bound #978, which covers the filesystems PyAthena creates.
  2. Chunked read with compression in the user storage_options loses the OpenFile owner. Confirmed for d5326d0: stack.enter_context(OpenFile) returns the GzipFile, pop_all() drops the OpenFile, and GzipFile.close() leaves the supplied fileobj open. This was a regression for non-binary reads. The same pattern in the binary-column path predates this PR and is unchanged.
  3. pandas' anonymous retry is lost for user storage_options. Confirmed: pandas 3.0.6 _get_filepath_or_buffer() retries fsspec.open() with anon=True after NoCredentialsError/PermissionError for s3:// paths, and opening the file ourselves skipped that. This was a regression for non-binary reads with user options. The default path never benefited, because S3FileSystem(connection=...) ignores anon.

Repair for 2 and 3 (4484e16, on this line): when storage_options are in the read options, the path and the options go to pd.read_csv() unchanged, as on master. Only without them is the output opened through the result set's filesystem. The binary path keeps master's behavior for given options; without them it uses the result set's filesystem. New test_csv_storage_options fails on d5326d0 (_csv_stream is the self-opened file) and passes on 4484e16.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review of the repair 4484e16 (rounds 1 and 2) — CLEAN

Scope: git range-diff 0b656b8d6de41b1f87bdb4d1e670e00a480df895..d5326d096ee6115917b62a995b06095134b2ff7c 55af09a2..4484e16b90110154637b22657082c33dbcb82cc1. Commits 1-4 are unchanged (=), commit 5 is new. Upstream between the bases (#998, #999) touches s3.py/s3_async.py range reads and transaction part limits, not filesystem construction or caching.

Round 1 (behavior):

  • Without user options, the CSV output is still opened through self._fs and the stream ownership is unchanged from d5326d0.
  • Non-binary reads with user storage_options are again exactly master's call: pd.read_csv(path, storage_options=...). pandas owns its handles there, and _csv_stream stays None as on master.
  • The binary path with given options calls fsspec.open() with them, as on master. The only difference is an explicit storage_options=None, which now uses the result set's filesystem instead of fsspec's default S3FileSystem without the connection.
  • _open_output_location() was removed, so no caller is left.

Round 2 (claims): the PR description now says that user options keep master's path, and that caching of a connection inside user options is unchanged (independent-review item 1). The test list and the tested commit are current.

Validation on 4484e16: just lint passed. pytest -n 4 tests/pyathena/pandas/ tests/pyathena/aio/pandas/ tests/pyathena/filesystem/test_s3_async.py tests/pyathena/polars/test_cursor.py tests/pyathena/s3fs/test_cursor.py: 468 passed, 1 skipped.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent follow-up review 1 (relayed) — FINDINGS

Reviewer: Codex CLI 0.160.0, model gpt-6-astra, reasoning effort max, sandbox read-only, session 01a10153-3899-75a2-a769-7b2709c32935.
Scope: the patch series 0b656b8d..d5326d09 compared with 55af09a2..4484e16b (range-diff, the new commit and the full diff), reviewed in a detached snapshot at 4484e16. The prompt carried no PR number, description or earlier findings. The snapshot and the PR worktree were both unchanged afterwards. This is a static review.

Reviewer output (verbatim)

Covered at 4484e16b:

  • Supplied range-diff and follow-up diff; commits 1–4 are unchanged.
  • Plain, chunked, and binary CSV paths; C, Python, and PyArrow engine selection.
  • storage_options absent, None, {}, and non-empty; pandas fallbacks and handle ownership.
  • Stream cleanup on completion, exhaustion, explicit close, cursor reuse, and errors.
  • New tests and existing binary lifecycle tests.
  • Parquet filesystem reuse, instance-cache isolation, and upstream filesystem changes. The rebase adds no conflicting cache or ownership behavior.

FINDINGS

  1. P2 — Binary CSV still conflates explicit storage_options=None with omission.
    pyathena/pandas/result_set.py:567 loses that distinction, so line 745 selects the connection-backed filesystem. Before the change, explicit None became {} and opened through fsspec defaults.

    Concrete failure: consume a large, chunked VARBINARY result with storage_options=None; the connection uses fixed temporary credentials that expire during consumption, while ambient credentials remain valid. Subsequent S3 reads now fail instead of continuing through the ambient credentials. Plain CSV preserves the previous behavior after this follow-up.

    Origin: introduced by commit 4 (d232af0a), pre-existing relative to commit 5 and still unresolved. Preserve key presence in the binary branch and add explicit-None coverage; the new test exercises only a non-empty options dictionary.

No newly introduced defect was found in commit 5. This was static review only; no files were changed, tests/builds run, or network accessed.

Author verification: confirmed. On master, an explicit storage_options=None became {} in the binary path (fsspec defaults), and plain reads after 4484e16 pass None to pandas (also fsspec defaults). The binary path at 4484e16 used the result set's filesystem instead. The regression came from d232af0.

Repair (dbab34f, rebased as e1f4d47 onto 9b2f033): the binary path uses the result set's filesystem only when the read options have no storage_options key; an explicit None becomes {} as on master. test_csv_storage_options now covers plain/binary × None/options:

  • it asserts that the result set's filesystem never opens the file, that given options reach the filesystem that does (default_cache_type == "none"), and that pandas opens the plain file itself (_csv_stream is None);
  • it fails on 4484e16 (none-binary) and on d232af0 (none-binary, none-plain, options-plain), and passes on the repair.

Self-review of the repair:

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent follow-up review 2 (relayed) — CLEAN

Reviewer: Codex CLI 0.160.0, model gpt-6-astra, reasoning effort max, sandbox read-only, session 01a10160-8741-7b70-8509-b66970b82e69.
Scope: the patch series 55af09a2..4484e16b compared with 9b2f0337..e1f4d470 (range-diff: commits 1-5 unchanged, commit 6 new), the new commit, and the full diff 9b2f0337..e1f4d470, reviewed in a detached snapshot at e1f4d47. The prompt carried no PR number, description or earlier findings. The snapshot was unchanged afterwards. This is a static review.

Reviewer output (verbatim)

Covered:

  • Supplied patch comparison and new commit; confirmed HEAD is e1f4d470 and commits 1–5 are unchanged.
  • _read_csv, _open_binary_csv_stream, and PandasDataFrameIterator: plain/binary CSV, full/chunked reads, and C/Python/PyArrow engine selection.
  • storage_options omitted, None, {}, and non-empty; compared opening, fallback, and ownership behavior with the pre-series source.
  • Stream cleanup on construction/read failures, exhaustion, early close, context exit, and cursor reuse.
  • New/changed tests in tests/pyathena/pandas/test_cursor.py and related existing lifecycle tests.
  • Default Parquet filesystem reuse, result-set cache isolation, and AioS3FileSystem internal isolation.
  • Upstream filesystem changes from 55af09a2 to 9b2f0337, including dircache operations and S3Object mapping behavior, against installed dependency sources.

CLEAN

No actionable findings in this follow-up scope. The explicit-None repair restores the pre-series binary CSV opening behavior without disrupting omitted-option filesystem reuse or stream ownership.

Static review only; no files changed, tests/builds run, or network access performed.

@laughingman7743
laughingman7743 force-pushed the fix/978-instance-cache-leak branch from dbab34f to e1f4d47 Compare October 3, 2026 10:47
laughingman7743 and others added 6 commits October 3, 2026 19:58
… cache

The result sets created their filesystems with the connection as a
constructor argument, so fsspec's instance cache kept every connection
and the shared dircache grew with every result file read. The internal
S3FileSystem of AioS3FileSystem was also cached, so instances created
with skip_instance_cache=True still shared it.

Create these filesystems with skip_instance_cache=True so that they are
released with the result set or the AioS3FileSystem that owns them.

Closes #978

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TestAioS3FileSystem registers AioS3FileSystem for "s3" and does not
restore the registration, so pandas and Polars read through
AioS3FileSystem in a worker that ran it first. Look for the connection
in both instance caches through a shared helper.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pd.read_csv() and pd.read_parquet() built a second S3FileSystem, with
its own S3 client and connections, from storage_options for every
result set. Open the CSV output through the result set's filesystem and
pass it to pd.read_parquet() as filesystem, so each result set builds
one S3 client for its reads. storage_options given in the read options
still open the CSV output through fsspec with these options.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Opening the output with fsspec.open() for user storage_options dropped
pandas' anonymous-access retry and, with a compression option, the
owner that closes the underlying file of a chunked read. Pass the path
and the storage_options to pd.read_csv() as before, and open the output
through the result set's filesystem only without them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An explicit storage_options=None opened binary-column results through
the result set's filesystem, while plain results and the code before
this change open them through fsspec's default filesystem. Use the
result set's filesystem only when the read options have no
storage_options key.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@laughingman7743
laughingman7743 force-pushed the fix/978-instance-cache-leak branch from e1f4d47 to 75ef9c2 Compare October 3, 2026 10:58
@laughingman7743
laughingman7743 marked this pull request as ready for review October 3, 2026 11:05
@laughingman7743
laughingman7743 merged commit 7947d94 into master Oct 3, 2026
12 checks passed
@laughingman7743
laughingman7743 deleted the fix/978-instance-cache-leak branch October 3, 2026 12:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Cached S3FileSystem instances keep every connection alive and grow their dircache without bound PyAthena - Memory Issue

1 participant