Skip to content

Share one S3 client per connection across result sets and filesystems #1011

Description

@laughingman7743

Use case

Every result set builds new S3 clients, so each query opens new HTTPS connections to S3 instead of reusing the connections of earlier queries.

  • AthenaResultSet.__init__() builds an S3 client from the connection's session for HeadObject and the data manifest (pyathena/result_set.py:105). It does so for every result set, including those of the default Cursor and DictCursor (and their aio versions), which fetch rows with GetQueryResults and never use it: only the pandas, Arrow, Polars and S3FS result sets call _get_content_length() or _read_data_manifest(). botocore opens no connection until the first request, so these result sets pay only for building the client (about 2.3 ms, see below) and its memory.
  • The result sets that read through S3FileSystem (pandas, S3FS; Polars read_csv()) create a filesystem with connection=..., and S3FileSystem.__init__() builds another S3 client (pyathena/filesystem/s3.py:181).
  • The Spark cursors build one more per cursor (pyathena/spark/common.py:110).

Until #978, the result-set filesystems sat in fsspec's instance cache, so the filesystem's client and its connection pool were reused per connection and thread. That cache never released them, which leaked connections and grew dircache without bound. PR #1001 fixes the leak by creating these filesystems with skip_instance_cache=True, so a result set's filesystem client now lives only as long as the result set. With #1001, a query builds two S3 clients (the result set's and its filesystem's), each starting with a new TCP/TLS connection.

Measured on 2026-10-03 from a laptop to us-west-2 (list_objects_v2(MaxKeys=1), 10 runs each):

  • First request on a new client: median about 450 ms.
  • Request on a client with an open connection: median about 150 ms.
  • Building the client from a session that has already built one: about 2.3 ms.

The difference is mostly round trips for the new connection, so it should be much smaller inside the region. That was not measured. Workloads that run many short queries pay this cost once or twice per query.

Proposed change

Let a Connection own one S3 client, created on first use from its session, region, config and client kwargs, as GlueMetadataClient creates its Glue client (pyathena/glue.py:75-94). Cursors that never access S3, such as the default Cursor, then build no S3 client at all. Share it with:

  • AthenaResultSet for HeadObject/GetObject;
  • S3FileSystem(connection=...), instead of building a client per filesystem;
  • the Spark cursors, where applicable.

botocore clients are thread-safe, so one client per connection can serve result sets created in different threads. The client would be released with the connection. Each filesystem keeps its own dircache, so #978 stays fixed.

Questions to settle before implementing:

  • The API: a public or private Connection attribute, or an S3FileSystem argument that accepts a client.
  • Whether Connection.close() should close the S3 client, and what happens to result sets and filesystems still in use after that.
  • Whether S3FileSystem(connection=...) created by users should also share the client. This changes nothing in its arguments, but two such filesystems would then share one connection pool.
  • max_pool_connections: one shared client limits concurrent S3 requests across all result sets of the connection to its pool size (botocore default 10). Today each result set has its own pool.

Alternatives considered:

  • Keep the filesystems in fsspec's instance cache and evict them when the connection closes: this depends on fsspec's cache tokens and only covers connections that are closed explicitly.
  • Do nothing: the extra connection per query is small compared with typical Athena query latency.

Validation plan (if implementing)

  • Count the S3 clients built per query (patching Session.client) for the default, pandas, Arrow, Polars, S3FS and Spark cursors, before and after; the default cursor should build none.
  • Measure the per-query S3 latency before and after, at least from a client outside AWS; optionally inside the region.
  • Check concurrent queries on one connection (threads and AsyncCursor) against the pool size.
  • Run just test pyathena, plus the SQLAlchemy suites if the connection API changes.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions