Problem
S3 request parameters are not routed per operation, which causes several failures that share one fix.
requester_pays=True breaks bucket operations and sign(). _call() adds RequestPayer to every request (pyathena/filesystem/s3.py:2112-2120, set at :192-195).
- HeadBucket, ListBuckets, CreateBucket, DeleteBucket and PutBucketAcl have no
RequestPayer member, so these raise ParamValidationError: exists("s3://bucket"), info("s3://bucket"), isdir("s3://bucket"), ls(""), rmdir(bucket) and chmod(bucket, acl).
mkdir(bucket) raises ValueError("Bucket create failed").
sign() raises TypeError: generate_presigned_url() got an unexpected keyword argument 'RequestPayer'.
- Passing
RequestPayer= explicitly, as in metadata(path, RequestPayer="requester"), raises a duplicate-keyword TypeError.
- Write-only
s3_additional_kwargs break every read. S3FileSystem(s3_additional_kwargs={"ServerSideEncryption": "AES256"}) is the usual s3fs example. With it, writes work, but open(p, "rb").read() raises ParamValidationError: Unknown parameter in input: "ServerSideEncryption" for GetObject. The same happens with StorageClass, ACL, ContentType and Metadata (s3.py:1969, :2494, :2505).
open() mutates the caller's s3_additional_kwargs and lets the filesystem-level values win. _open() updates the caller's dict in place, and S3File keeps and mutates that same dict (s3.py:1968-1969, :2184, :2227, :2240; s3_async.py:362-363).
- Read mode adds
IfMatch=<etag> to the caller's dict, so reusing the dict for a later write sends another object's IfMatch.
- Append mode adds the appended object's ContentType, ContentEncoding, StorageClass and Metadata, which leak into later writes of other keys.
- Filesystem-level values override per-call ones: a per-call
StorageClass="GLACIER_IR" is sent as the filesystem's STANDARD. put_file() loses the guessed ContentType the same way. pipe_file()'s single-request path and s3fs both let the per-call value win.
- S3 parameters passed as keyword arguments are dropped.
open(p, "wb", ServerSideEncryption="aws:kms", ContentType="text/csv") sends neither, because S3File ignores extra keywords (s3.py:1953-1984, :2131, :2179; s3_async.py:347). s3fs merges them into the request parameters.
pipe_file(p, data, ContentType=...) keeps the parameters on its single-request path. They are dropped when the data exceeds the block size or the call runs inside fs.transaction (s3.py:1309-1315). Before 1732653 they reached CreateMultipartUpload.
Expected:
- each S3 operation receives only the parameters it accepts;
- per-call parameters win over filesystem-level ones;
- callers' dicts are never mutated;
- keyword S3 parameters are applied on every write path.
Fix together: #946 (part requests miss the per-file parameters) and #967 (cp_file() sends block_size/max_workers to S3) are the same routing problem.
Reproduction
from unittest.mock import MagicMock
from pyathena.filesystem.s3 import S3FileSystem
calls = []
fs = S3FileSystem(connection=MagicMock(), skip_instance_cache=True,
s3_additional_kwargs={"ServerSideEncryption": "AES256"})
def call(method, **request):
name = method._extract_mock_name().split(".")[-1]
calls.append((name, request))
if name == "head_object":
return {"ContentLength": 3, "ETag": '"etag-of-a"'}
return {"ETag": '"e"', "UploadId": "u", "Body": MagicMock(read=lambda: b"abc")}
fs._call = call
opts = {"ExpectedBucketOwner": "111122223333"}
with fs.open("s3://bucket/a", "rb", s3_additional_kwargs=opts) as f:
f.read()
print(opts)
# {'ExpectedBucketOwner': '111122223333', 'ServerSideEncryption': 'AES256', 'IfMatch': '"etag-of-a"'}
with fs.open("s3://bucket/b", "wb", s3_additional_kwargs=opts) as f:
f.write(b"x")
print(calls[-1])
# ('put_object', {..., 'IfMatch': '"etag-of-a"'})
fs.pipe_file("s3://bucket/large", b"x" * (5 * 2**20 + 1), ContentType="text/csv")
print(next(c for c in calls if c[0] == "create_multipart_upload"))
# ('create_multipart_upload', {'Bucket': 'bucket', 'Key': 'large', 'ServerSideEncryption': 'AES256'}): no ContentType
Other observed results:
- With
requester_pays=True and Stubber: exists("s3://bucket") raised ParamValidationError ... Unknown parameter in input: "RequestPayer", must be one of: Bucket, ExpectedBucketOwner, and sign() raised TypeError.
- With
Stubber validating GetObject, the read under filesystem-level ServerSideEncryption raised ParamValidationError ... Unknown parameter in input: "ServerSideEncryption".
Environment
- PyAthena master e0e85da, fsspec 2026.9.0, botocore from
uv.lock, Python 3.13.1.
- Found in a full audit of
pyathena/filesystem/. Unless noted, checked offline with mocked S3 responses (fs._call replaced) or botocore's Stubber with dummy credentials.
Proposed fix (optional)
Problem
S3 request parameters are not routed per operation, which causes several failures that share one fix.
requester_pays=Truebreaks bucket operations andsign()._call()addsRequestPayerto every request (pyathena/filesystem/s3.py:2112-2120, set at:192-195).RequestPayermember, so these raiseParamValidationError:exists("s3://bucket"),info("s3://bucket"),isdir("s3://bucket"),ls(""),rmdir(bucket)andchmod(bucket, acl).mkdir(bucket)raisesValueError("Bucket create failed").sign()raisesTypeError: generate_presigned_url() got an unexpected keyword argument 'RequestPayer'.RequestPayer=explicitly, as inmetadata(path, RequestPayer="requester"), raises a duplicate-keywordTypeError.s3_additional_kwargsbreak every read.S3FileSystem(s3_additional_kwargs={"ServerSideEncryption": "AES256"})is the usual s3fs example. With it, writes work, butopen(p, "rb").read()raisesParamValidationError: Unknown parameter in input: "ServerSideEncryption"for GetObject. The same happens with StorageClass, ACL, ContentType and Metadata (s3.py:1969,:2494,:2505).open()mutates the caller'ss3_additional_kwargsand lets the filesystem-level values win._open()updates the caller's dict in place, andS3Filekeeps and mutates that same dict (s3.py:1968-1969,:2184,:2227,:2240;s3_async.py:362-363).IfMatch=<etag>to the caller's dict, so reusing the dict for a later write sends another object'sIfMatch.StorageClass="GLACIER_IR"is sent as the filesystem'sSTANDARD.put_file()loses the guessed ContentType the same way.pipe_file()'s single-request path and s3fs both let the per-call value win.open(p, "wb", ServerSideEncryption="aws:kms", ContentType="text/csv")sends neither, becauseS3Fileignores extra keywords (s3.py:1953-1984,:2131,:2179;s3_async.py:347). s3fs merges them into the request parameters.pipe_file(p, data, ContentType=...)keeps the parameters on its single-request path. They are dropped when the data exceeds the block size or the call runs insidefs.transaction(s3.py:1309-1315). Before 1732653 they reached CreateMultipartUpload.Expected:
Fix together: #946 (part requests miss the per-file parameters) and #967 (
cp_file()sendsblock_size/max_workersto S3) are the same routing problem.Reproduction
Other observed results:
requester_pays=TrueandStubber:exists("s3://bucket")raisedParamValidationError ... Unknown parameter in input: "RequestPayer", must be one of: Bucket, ExpectedBucketOwner, andsign()raisedTypeError.Stubbervalidating GetObject, the read under filesystem-levelServerSideEncryptionraisedParamValidationError ... Unknown parameter in input: "ServerSideEncryption".Environment
uv.lock, Python 3.13.1.pyathena/filesystem/. Unless noted, checked offline with mocked S3 responses (fs._callreplaced) or botocore'sStubberwith dummy credentials.Proposed fix (optional)
{**filesystem-level, **per-call}in a new dict.self._client.meta.service_model.operation_model(name).input_shape.members, as s3fs does with_get_s3_method_kwargs. Skip non-API methods such asgenerate_presigned_url, or put the parameters intoParams._open()andS3File.__init__.open()/pipe_file()into the per-file parameters.