Skip to content

S3FileSystem does not route request parameters per S3 operation #969

Description

@laughingman7743

Problem

S3 request parameters are not routed per operation, which causes several failures that share one fix.

  1. requester_pays=True breaks bucket operations and sign(). _call() adds RequestPayer to every request (pyathena/filesystem/s3.py:2112-2120, set at :192-195).
    • HeadBucket, ListBuckets, CreateBucket, DeleteBucket and PutBucketAcl have no RequestPayer member, so these raise ParamValidationError: exists("s3://bucket"), info("s3://bucket"), isdir("s3://bucket"), ls(""), rmdir(bucket) and chmod(bucket, acl).
    • mkdir(bucket) raises ValueError("Bucket create failed").
    • sign() raises TypeError: generate_presigned_url() got an unexpected keyword argument 'RequestPayer'.
    • Passing RequestPayer= explicitly, as in metadata(path, RequestPayer="requester"), raises a duplicate-keyword TypeError.
  2. Write-only s3_additional_kwargs break every read. S3FileSystem(s3_additional_kwargs={"ServerSideEncryption": "AES256"}) is the usual s3fs example. With it, writes work, but open(p, "rb").read() raises ParamValidationError: Unknown parameter in input: "ServerSideEncryption" for GetObject. The same happens with StorageClass, ACL, ContentType and Metadata (s3.py:1969, :2494, :2505).
  3. open() mutates the caller's s3_additional_kwargs and lets the filesystem-level values win. _open() updates the caller's dict in place, and S3File keeps and mutates that same dict (s3.py:1968-1969, :2184, :2227, :2240; s3_async.py:362-363).
    • Read mode adds IfMatch=<etag> to the caller's dict, so reusing the dict for a later write sends another object's IfMatch.
    • Append mode adds the appended object's ContentType, ContentEncoding, StorageClass and Metadata, which leak into later writes of other keys.
    • Filesystem-level values override per-call ones: a per-call StorageClass="GLACIER_IR" is sent as the filesystem's STANDARD. put_file() loses the guessed ContentType the same way. pipe_file()'s single-request path and s3fs both let the per-call value win.
  4. S3 parameters passed as keyword arguments are dropped.
    • open(p, "wb", ServerSideEncryption="aws:kms", ContentType="text/csv") sends neither, because S3File ignores extra keywords (s3.py:1953-1984, :2131, :2179; s3_async.py:347). s3fs merges them into the request parameters.
    • pipe_file(p, data, ContentType=...) keeps the parameters on its single-request path. They are dropped when the data exceeds the block size or the call runs inside fs.transaction (s3.py:1309-1315). Before 1732653 they reached CreateMultipartUpload.

Expected:

  • each S3 operation receives only the parameters it accepts;
  • per-call parameters win over filesystem-level ones;
  • callers' dicts are never mutated;
  • keyword S3 parameters are applied on every write path.

Fix together: #946 (part requests miss the per-file parameters) and #967 (cp_file() sends block_size/max_workers to S3) are the same routing problem.

Reproduction

from unittest.mock import MagicMock
from pyathena.filesystem.s3 import S3FileSystem

calls = []
fs = S3FileSystem(connection=MagicMock(), skip_instance_cache=True,
                  s3_additional_kwargs={"ServerSideEncryption": "AES256"})
def call(method, **request):
    name = method._extract_mock_name().split(".")[-1]
    calls.append((name, request))
    if name == "head_object":
        return {"ContentLength": 3, "ETag": '"etag-of-a"'}
    return {"ETag": '"e"', "UploadId": "u", "Body": MagicMock(read=lambda: b"abc")}
fs._call = call

opts = {"ExpectedBucketOwner": "111122223333"}
with fs.open("s3://bucket/a", "rb", s3_additional_kwargs=opts) as f:
    f.read()
print(opts)
# {'ExpectedBucketOwner': '111122223333', 'ServerSideEncryption': 'AES256', 'IfMatch': '"etag-of-a"'}
with fs.open("s3://bucket/b", "wb", s3_additional_kwargs=opts) as f:
    f.write(b"x")
print(calls[-1])
# ('put_object', {..., 'IfMatch': '"etag-of-a"'})

fs.pipe_file("s3://bucket/large", b"x" * (5 * 2**20 + 1), ContentType="text/csv")
print(next(c for c in calls if c[0] == "create_multipart_upload"))
# ('create_multipart_upload', {'Bucket': 'bucket', 'Key': 'large', 'ServerSideEncryption': 'AES256'}): no ContentType

Other observed results:

  • With requester_pays=True and Stubber: exists("s3://bucket") raised ParamValidationError ... Unknown parameter in input: "RequestPayer", must be one of: Bucket, ExpectedBucketOwner, and sign() raised TypeError.
  • With Stubber validating GetObject, the read under filesystem-level ServerSideEncryption raised ParamValidationError ... Unknown parameter in input: "ServerSideEncryption".

Environment

  • PyAthena master e0e85da, fsspec 2026.9.0, botocore from uv.lock, Python 3.13.1.
  • Found in a full audit of pyathena/filesystem/. Unless noted, checked offline with mocked S3 responses (fs._call replaced) or botocore's Stubber with dummy credentials.

Proposed fix (optional)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions