Skip to content

Copy metadata, tags and annotations in multipart copies - #1036

Merged
laughingman7743 merged 10 commits into
masterfrom
fix/973-multipart-copy-metadata
Oct 4, 2026
Merged

laughingman7743 merged 10 commits into
masterfrom
fix/973-multipart-copy-metadata

Conversation

@laughingman7743

@laughingman7743 laughingman7743 commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

WHAT

A copy of an object larger than 5 GiB (cp_file() / _cp_file() → _copy_object_with_multipart_upload(), sync and async) now gives the same result as CopyObject:

  • Directives. No multipart request accepts CopyObject's directives, so they are implemented in S3FileSystem._get_multipart_copy_kwargs() (shared with AioS3FileSystem through asyncio.to_thread):
    • MetadataDirective (default COPY): HeadObject on the source supplies CacheControl, ContentDisposition, ContentEncoding, ContentLanguage, ContentType, Expires and the user-defined Metadata for CreateMultipartUpload. Values given to the copy for these are ignored, as CopyObject ignores them. REPLACE uses the given values.
    • TaggingDirective (default COPY): GetObjectTagging on the source supplies Tagging. A given Tagging is ignored. REPLACE uses the given Tagging.
    • AnnotationDirective (default COPY): ListObjectAnnotations (all pages) on the source before CreateMultipartUpload, so a caller who cannot list them fails before anything is written; after CompleteMultipartUpload, GetObjectAnnotation + PutObjectAnnotation for each annotation onto the destination. Each write targets the completion's VersionId (in a versioned bucket) and has ObjectIfMatch set to the completed ETag, so it fails instead of attaching an annotation to an object written over the copy. Sync copies the annotations one by one; async uses asyncio.gather under the max_workers semaphore. EXCLUDE skips them.
    • Any other directive value raises ValueError before any request.
    • These steps are skipped when the source cannot have them: annotations for an SSE-C source (given by CopySourceSSECustomerAlgorithm), and tags and annotations for a source in a directory bucket (name ending in --x-s3), which supports neither API.
  • Parameter split (from the maintainer's comment on Multipart copies drop the source metadata, and the async multipart copy does not abort on failure #973):
    • CreateMultipartUpload receives only the CopyObject parameters that it accepts. Source conditions (CopySourceIf*), CopySourceSSECustomer* and ExpectedSourceBucketOwner used to fail CreateMultipartUpload's validation; now they reach only the part copies, which already received the parameters they accept.
    • A parameter that CopyObject does not accept either is still passed through, so botocore keeps rejecting typos.
  • Stale size. If HeadObject reports a size of at most 5 GiB (the multipart path was chosen from a stale cached size), the reported version is copied with a single CopyObject request, with the same parameters as a small copy. This includes an empty object.
  • Source version. HeadObject on the source is always sent. A source without a given version is copied from the version ID that HeadObject reports, unless that ID is null. The part copies (CopySource VersionId), the tag read and the annotation reads all use it, and the part ranges use HeadObject's size, so a source replaced during the copy is not mixed in.
  • Source lookups. HeadObject, GetObjectTagging, ListObjectAnnotations and GetObjectAnnotation on the source receive RequestPayer, ExpectedSourceBucketOwner as ExpectedBucketOwner, and (HeadObject only) CopySourceSSECustomer* as SSECustomer*. The destination's ExpectedBucketOwner goes only to the destination requests.
  • Async abort.
    • When a part copy or the completion fails, AioS3FileSystem now waits for the running parts and aborts the upload, as S3FileSystem does.
    • No queued part starts after a failure: the flag is set before the semaphore is released.
    • The abort is shared as S3FileSystem._abort_multipart_upload(), extracted from _finish_multipart_upload().
  • When a copy fails after the destination has been written (a failed annotation copy), cp_file() still invalidates the destination's cache entries.
  • docs/filesystem.md and the cp_file() docstring describe the behavior.

Release notes

  • 4.0.0: the minimum versions are now botocore>=1.43.31 and boto3>=1.43.31. botocore 1.43.31 is the first release with the annotation operations (bisected: 1.43.30 lacks list_object_annotations), and boto3 1.43.31 is the first release that requires botocore>=1.43.31 (PyPI metadata).
  • Copies over 5 GiB keep the source's content headers, user-defined metadata, tags and annotations by default, as CopyObject does. Before, they got S3's defaults and no metadata, tags or annotations. Given ContentType/Metadata/Tagging values are now ignored unless the matching directive is REPLACE; before, they were used.
  • These copies send extra requests: one HeadObject (always), one GetObjectTagging, ListObjectAnnotations, and a GetObjectAnnotation + PutObjectAnnotation per annotation. They need s3:GetObjectTagging on the source, plus s3:ListObjectAnnotations and s3:GetObjectAnnotation on the source and s3:PutObjectAnnotation on the destination. Use TaggingDirective="REPLACE" or AnnotationDirective="EXCLUDE" to skip these.
  • When HeadObject reports a non-null version ID, these copies read that specific source version, so they need s3:GetObjectVersion and s3:GetObjectVersionTagging instead of s3:GetObject and s3:GetObjectTagging (per the S3 User Guide's "Required permissions for Amazon S3 API operations" for UploadPartCopy and GetObjectTagging).
  • The destination exists without its annotations between the completion and the last annotation write. A failed annotation copy raises, and the destination is kept, not deleted.
  • CopySourceIf*, CopySourceSSECustomer*, ExpectedSourceBucketOwner and the directives now work for copies over 5 GiB. Before, CreateMultipartUpload rejected them.
  • An invalid MetadataDirective/TaggingDirective/AnnotationDirective raises ValueError for copies over 5 GiB.
  • A failed async multipart copy now aborts its upload. Before, the incomplete upload and its parts were left behind.

WHY

Closes #973. Before this PR, whether a copy kept its metadata depended on the object's size, and a failed async multipart copy left an incomplete upload behind.

TEST

Tested commit e8d3f66, then 2154780 after the review repairs, the maintainer's decisions and the rebase onto master ec5323e. On 2154780: just lint, just docs lint and uv lock --check pass, the offline copy tests pass, and live pytest -n 4 tests/pyathena/filesystem/ → 639 passed. An earlier one-off failure of the sync failed-annotation test was traced to a test helper overridden by a same-named one from #1016 (parallel part copies against the Stubber) and fixed in the tests (see the review thread).

  • just format, just lint, just docs lint: pass.
  • New offline tests (botocore Stubber with dummy credentials, max_workers=1 for a deterministic request order), in tests/pyathena/filesystem/test_s3.py and test_s3_async.py:
    • test_copy_object_with_multipart_upload_copies_source (sync + async): every request of a 2-part copy, using the shared helper tests/pyathena/util.py:stub_multipart_copy. It covers metadata and tags from the source with the given ContentType/Tagging ignored, StorageClass from the call rather than the source, the source condition only on the parts, source and destination bucket owners, two pages of annotations listed before the upload, the version from HeadObject on the parts, tags and annotation reads, and VersionId + ObjectIfMatch on the annotation writes.
    • _small_head_object_size[0, 10] (sync + async): a stale size over 5 GiB with HeadObject reporting 0 or 10 bytes copies the pinned version with CopyObject, without reading tags.
    • _head_object_size (sync): the ranges follow HeadObject's size, not a cached size, and a null version is not pinned.
    • _failed_listing (sync + async): an AccessDenied from ListObjectAnnotations happens before CreateMultipartUpload; nothing is written.
    • _failed_part (sync + async): abort, no completion.
    • _failed_annotation (sync + async): PermissionError, with no abort or delete.
    • Async _waits_for_running_parts: the abort waits for the running part, and the queued part never starts. Passed 6 of 6 repeated runs.
    • Sync: _replace_directives, _invalid_directive (3 cases), _unknown_parameter (botocore still rejects ContentTyp), _sse_c_source, _directory_bucket_source.
    • With pyathena/ reverted, 15 of the 26 selected multipart/copy tests fail.
  • Existing tests that mock the multipart requests now pass REPLACE/EXCLUDE directives, so they still check part sizes and parameter routing without reading the source.
  • Live S3, CI bucket (us-west-2):
    • uv run --env-file .env pytest -n 4 tests/pyathena/filesystem/ → 531 passed.
    • A one-off measurement called _copy_object_with_multipart_upload() directly on a 10 MiB + 1 B source (5 MiB parts) with content headers, metadata, two tags and two annotations. Sync and async copies matched a plain CopyObject of the same source: headers, metadata, tags and annotation payloads, with ContentType="text/plain" ignored. REPLACE/EXCLUDE gave the given values and no annotations. A failing CopySourceIfMatch raised OSError and left no multipart upload. The objects were deleted.
    • Measured for the design: CopyObject's default copies content headers, metadata, tags and annotations, but not WebsiteRedirectLocation, storage class or encryption. Under COPY it ignores a given ContentType, Metadata and Tagging. The annotation APIs and ObjectIfMatch work in the CI bucket. urlencode tags with spaces round-trip through CreateMultipartUpload.
  • Not tested live: objects over 5 GiB (cost), SSE-C, directory buckets, requester-pays buckets (offline only).

Limits:

  • A null version (an object from before versioning, or with versioning suspended) is not pinned.
  • An Expires value that botocore cannot parse is not copied.
  • The CI bucket is not versioned, so the version pinning is covered offline only.
  • Async cancellation of a multipart copy still leaves the upload behind (as before): Cancelling an async multipart copy leaves its multipart upload behind #1046.

🤖 Generated with Claude Code

Comment thread pyathena/filesystem/s3.py
}
return {k: v for k, v in source_kwargs.items() if v is not None}

def _get_multipart_copy_kwargs(

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round 1 (implementation behavior): base de8cc52ac2a43cdba72bb4571c185883287741fc, head e8d3f669e218d26b19915ca6b17d9ad13ae1a221 (plus a42788c8, docstring only). Result: FINDINGS. One wording finding is repaired; the other two are pre-existing or deferred.

Covered: _copy_object_with_multipart_upload (sync/async), _get_multipart_copy_kwargs, _get_copy_source_kwargs, _copies_annotations, _list_object_annotations, _copy_object_annotation, _abort_multipart_upload and its caller _finish_multipart_upload, the cp_file/_cp_file callers through _copy_file, the docs, and the tests.

  • Failure paths: an invalid directive raises before any request. A failure in HeadObject or GetObjectTagging happens before CreateMultipartUpload, so there is nothing to abort. A failed part or completion aborts, sync and async. On the async side the flag is set inside the failing part before it releases the semaphore, so no queued part starts (verified by the start 3 race that the first version of the test caught). After the completion, a failure to list, read or write an annotation raises, and the destination is kept (as specified).
  • Parameter routing: CopyObject parameters that CreateMultipartUpload does not accept are dropped. Unknown parameters still pass through, so botocore's validation error is preserved (_unknown_parameter). Source lookups get the source's owner and SSE-C key; destination requests get the destination's.
  • Finding (repaired in a42788c8): the docstring said that ObjectIfMatch keeps an annotation off "an object written over the copy". It only guards against a different ETag, and the wording now says so.
  • Deferred, listed as limits in the PR body: (a) annotations are listed after the completion, so an unpinned source replaced during the copy can supply the new object's annotations; (b) a caller without s3:ListObjectAnnotations now fails after the completion even when the source has no annotations. The maintainer specified the ordering after the completion, so I did not change it.
  • Pre-existing, not changed: an Expires value that botocore cannot parse is lost (as in S3Metadata).

Comment thread docs/filesystem.md
source are copied, and the values given for them are ignored, as CopyObject does. A
`REPLACE` directive uses the given values instead, and `AnnotationDirective="EXCLUDE"`
skips the annotations. Copying the tags needs `s3:GetObjectTagging` on the source, and
copying the annotations needs `s3:ListObjectAnnotations` and `s3:GetObjectAnnotation` on

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round 2 (claims, callers, AWS effects): base de8cc52ac2a43cdba72bb4571c185883287741fc, head a42788c8ad454869c9cf24a17b33b2520f025ba8. Result: CLEAN after the PR-body corrections (the ObjectIfMatch wording and the added Limits).

Claims checked:

  • "CopyObject ignores given ContentType/Metadata/Tagging under COPY, copies content headers/metadata/tags/annotations, not WebsiteRedirectLocation/storage class/SSE": measured live in the CI bucket on 2026-10-03 (default, explicit COPY, COPY with ContentType/Metadata/Tagging, REPLACE, TaggingDirective=REPLACE without Tagging).
  • "Sync and async multipart copies now match CopyObject": measured live on a 10 MiB + 1 B source with 5 MiB parts, compared with a CopyObject of the same source. Headers, metadata, tags and annotation payloads matched. A failed CopySourceIfMatch left no multipart upload.
  • "Before, CreateMultipartUpload rejected CopySourceIf*/CopySourceSSECustomer*/ExpectedSourceBucketOwner/directives": in the botocore 1.43.102 model these are not members of CreateMultipartUpload; the maintainer's comment on Multipart copies drop the source metadata, and the async multipart copy does not abort on failure #973 reported the same.
  • Existing callers: a ContentType/Metadata/Tagging given for a copy over 5 GiB is now ignored by default. This is a behavior change, release-noted. mv() raises before removing anything if the copy fails, so a failed annotation copy keeps both source and destination.
  • AWS operator: extra requests per copy over 5 GiB are 1 HeadObject + 1 GetObjectTagging + ≥1 ListObjectAnnotations + 2 per annotation, all through _call's retries. Async annotation copies are bounded by max_workers. The new permission requirements are release-noted in the PR body and stated in docs/filesystem.md.
  • Evidence: the offline Stubber tests and the live measurement are kept separate in the PR body, and >5 GiB, SSE-C, directory and requester-pays cases are marked as not tested live.

@laughingman7743
laughingman7743 force-pushed the fix/973-multipart-copy-metadata branch from a42788c to f0e374e Compare October 3, 2026 15:06
Comment thread pyathena/filesystem/s3.py
"AnnotationName": name,
"AnnotationPayload": response["AnnotationPayload"].read(),
}
if completed.version_id:

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent review (relayed): reviewer Codex CLI 0.160.0, model gpt-6-astra, reasoning effort max (local config), codex exec -s read-only --ephemeral on a detached snapshot at head a42788c8ad454869c9cf24a17b33b2520f025ba8, base de8cc52ac2a43cdba72bb4571c185883287741fc. Static review only; the snapshot was unchanged. The prompt had no PR number, PR description or author conclusions.

Covered: the full diff, the sync/async multipart paths, the cp_file/_cp_file/mv/copy callers, directive semantics, parameter routing, SSE-C, owner/payer headers, botocore 1.43.102 shapes, annotation pagination/versioning/failure handling, gather/semaphore behavior, aborts, cancellation, the tests and the docs.

Result: FINDINGS. Introduced by this change:

  1. P1: pyproject.toml allows botocore>=1.41.2, but the annotation operations first appear in botocore 1.43.31 (bisected locally: 1.43.30 lacks list_object_annotations, 1.43.31 has it). With an older botocore, a default copy over 5 GiB completes and then raises AttributeError. Open, needs a maintainer decision (raise the floor, or skip annotation copying when the client lacks the operation).
  2. P2: the annotations of an unpinned versioned source are listed after the copy, so they can come from a newer version. Deferred; it is in the PR's Limits. Pinning would need the source version from HeadObject for the parts too, which is beyond Multipart copies drop the source metadata, and the async multipart copy does not abort on failure #973's scope; this needs a maintainer decision.
  3. P2: ObjectIfMatch alone could annotate a newer destination version with the same ETag. Fixed: PutObjectAnnotation now also gets the completion's VersionId.
  4. P2: when an annotation copy fails, the destination's cache entries are left stale. Fixed: _copy_file invalidates path2 in a finally (sync and async).
  5. P2: in the async annotation-failure test, gather(return_exceptions=True) still ran a2 and hid its UnStubbedResponseError. Fixed: async annotation copies stop starting after one fails (the same flag pattern as the parts, matching the sync loop), so the stubbed sequence is now complete.

Pre-existing (reported by the reviewer):

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Repair (f0e374eb, rebased onto fe21250a; the only conflict was the typing/urllib.parse import lines, kept from both sides):

  • PutObjectAnnotation gets VersionId=completed.version_id when present, as well as ObjectIfMatch (finding 3). stub_multipart_copy expects both.
  • S3FileSystem._copy_file and AioS3FileSystem._copy_file invalidate path2 in a finally (finding 4), with new tests test_cp_file_failed_multipart_copy_invalidates_cache (sync and async).
  • Async annotation copies stop starting after a failure (finding 5).
  • Self-review of the repair. Round 1: _finish_multipart_upload after the rebase catches BaseException and calls the extracted _abort_multipart_upload, so behavior matches master. Invalidating on a failure before completion only evicts cache entries and sends no request. Round 2: the PR body claims are unchanged apart from the open items.
  • Validation: just lint passes; the offline copy tests pass (99; the 5 failures are live-only tests run without credentials); live pytest -n 4 tests/pyathena/filesystem/ → 546 passed.
  • Still open: findings 1 and 2 and async cancellation parity. The PR stays Draft until they are decided, and an independent follow-up review of the repairs is still pending.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Repairs for the maintainer's decisions (794b62c1..0e755234, rebased onto master 8446e872):

  • 794b62c1: botocore>=1.43.31 and boto3>=1.43.31, with uv lock changing only the two specifiers in [package.metadata]. botocore 1.43.31 is the first release with the annotation operations (bisected: 1.43.30 lacks list_object_annotations). PyPI metadata: boto3 1.43.31 requires botocore<1.44.0,>=1.43.31, and 1.43.30 requires >=1.43.30. There are no hasattr branches.
  • a879b6ef: the annotations are listed before CreateMultipartUpload (new _failed_listing tests, sync and async: nothing is written). HeadObject is always sent. Without a given version, the reported version pins the part copies (CopySource VersionId), the tag read and the annotation reads.
  • f0c189d9: the ranges use HeadObject's ContentLength for the pinned version, not a possibly cached info() size. A null version is not pinned. The async failed-annotation test now spies on _copy_object_annotation and asserts ["a1"]; it fails with the stop guard removed (verified).
  • 32f6e3b4, 0e755234: the docstring and docs wording for pinning.

Self-review of the repairs (affected scope). Round 1: sync and async paths are in parity; the directive check still runs before any request; REPLACE copies now send one HeadObject (release-noted); a null, absent or given version is not overridden. Round 2: the permission claims were checked against the S3 User Guide's "Required permissions for Amazon S3 API operations": UploadPartCopy with versionId needs s3:GetObjectVersion, and GetObjectTagging with versionId needs s3:GetObjectVersionTagging.

Independent follow-ups (Codex CLI, same configuration, read-only, static):

  1. f0e374eb..9ab9caf3: FINDINGS. A null version is mutable, so the docs overclaimed (fixed in f0c189d9/32f6e3b4). The cached size could truncate the pinned version (fixed in f0c189d9). The async failure test did not prove the stop (fixed in f0c189d9). SSE-C on the public cp_file path is pre-existing and left alone by instruction (S3File does not send its request parameters with the lookups made while opening #1004/Send the lookup parameters of a file with its lookups #1024).
  2. 9ab9caf3..64dfddae: FINDINGS. The wording promised pinning whenever versioning is enabled, but objects from before versioning keep null (fixed in 32f6e3b4). A replacement with ContentLength=0 makes _get_copy_ranges(0) raise ValueError before anything is written. Deferred: this is an error, not silent corruption, and it would need a new zero-byte path; it is listed in Limits.
  3. 64dfddae..eaa7c825: FINDINGS, P3: "a write replaces a null version" was not true with versioning enabled (fixed in 0e755234).
  4. eaa7c825..c7f83b36: CLEAN.
  5. Rebase c7f83b36 → 0e755234 (range-diff; only the constants conflicted, both kept; upstream Keep system metadata in setxattr() and reject version paths #1016/Use the compression table properties that Athena applies on write #1025 checked for interaction): CLEAN.

Validation on 0e755234: just lint, just docs lint and uv lock --check pass. Offline copy tests pass. Live pytest -n 4 tests/pyathena/filesystem/ → 553 passed in 3 of 4 runs. The first run after the rebase had one failure of the offline test_copy_object_with_multipart_upload_failed_annotation (sync), with an unexpected abort logged. It did not reproduce in 30 serial runs of the copy tests, 15 parallel runs of the copy tests (-n 4) or 3 more full live runs. The cause is not identified; recorded in the PR body.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Flaky test root cause and the 0-byte race (598edd8d, 13e45c4d, rebased onto master abc99b09):

  • Root cause of the one unexplained failure of test_copy_object_with_multipart_upload_failed_annotation (sync): after the rebase onto 8446e872, TestS3FileSystem defined _stubbed_fs() twice. The copy tests' helper (max_workers=1) came first; Keep system metadata in setxattr() and reject version paths #1016's setxattr helper, with the same name further down the class, overrode it. With that one, fs.max_workers was 50 (checked), so the two part copies ran in parallel and could reach the Stubber out of order. A part then failed with StubAssertionError, and _finish_multipart_upload attempted an unstubbed abort, which produced the logged "Failed to abort multipart upload u to s3://bucket/dst". I reproduced it deterministically by delaying part 1 with a botocore provide-client-params hook: the overridden helper gives StubAssertionError plus the abort log, the renamed _stubbed_copy_fs() (max_workers=1) gives the expected PermissionError with no pending responses. The fix is test-only: the helper is renamed. An AST check finds no other duplicate method names in either test class.
  • 0-byte race: if HeadObject reports a size of at most MULTIPART_UPLOAD_MAX_PART_SIZE (a stale cached size over 5 GiB chose the multipart path), the reported version (non-null VersionId pinned) is copied with _copy_object and the same kwargs as the small-object path, sync and async. _get_multipart_copy_kwargs returns before reading the tags in that case, because CopyObject applies the directives itself. New tests test_copy_object_with_multipart_upload_small_head_object_size[0, 10] (sync and async) fail without the fix. The stubbed multipart copy is now MAX_PART_SIZE + MIN_PART_SIZE bytes in two parts, so it stays on the multipart path.
  • Self-review (affected scope). Round 1: directive validation still runs first; the fallback sends exactly CopyObject's kwargs, as cp_file does for small objects; cache invalidation still runs through _copy_file's finally. Round 2: the docs sentence matches the code.
  • Independent follow-up (Codex CLI, same configuration, read-only, static) on both commits plus the rebase range-diff and the interacting upstream changes: CLEAN.
  • Validation on 13e45c4d: just lint and just docs lint pass; the offline copy tests pass (75 before the rebase); live pytest -n 4 tests/pyathena/filesystem/ → 639 passed.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Rebased 13e45c4d → 2154780a onto master ec5323ea (#1033). The only conflict was the import lines of tests/pyathena/util.py (master added timedelta, timezone and dateutil.tz.gettz; this branch adds UTC and botocore.response.StreamingBody), and both sides are kept. The range-diff shows the other nine commits unchanged. Upstream changes are in the converters and their tests only, not in pyathena/filesystem. On 2154780a: just lint passes; live pytest -n 4 tests/pyathena/filesystem/ → 639 passed; tests/pyathena/test_util.py and test_converter.py offline → 195 passed.

laughingman7743 and others added 10 commits October 4, 2026 02:09
A copy of an object larger than 5 GiB goes through a multipart upload,
which started without the source's content headers, user-defined
metadata and tags, and never got its annotations. Implement CopyObject's
directives for it (default COPY): read the content headers and metadata
with HeadObject and the tags with GetObjectTagging for
CreateMultipartUpload, ignoring the values of the copy as CopyObject
does, and copy the annotations after the completion with
ListObjectAnnotations, GetObjectAnnotation and PutObjectAnnotation. REPLACE
and EXCLUDE use the values of the copy or skip the annotations, and an
invalid directive raises ValueError.

CreateMultipartUpload now receives only the CopyObject parameters that it
accepts, so source conditions reach the part copies instead of failing
validation, and the source's lookups get the source's SSE-C key and
expected bucket owner. A failed part or completion of the async multipart
copy now aborts the upload after the running parts finish, as the sync
one does.

Closes #973

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Write each annotation to the version that the multipart copy created,
remove the cached entries of the destination when a copy fails after the
completion, and stop starting async annotation copies after one fails,
as the sync loop does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The multipart copy uses ListObjectAnnotations, GetObjectAnnotation and
PutObjectAnnotation, which botocore has since 1.43.31. boto3 1.43.31 is
the first release that requires it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The multipart copy reads the source with HeadObject in any case, and
without a given version it copies the parts, the tags and the
annotations from the version that HeadObject reports in a versioned
bucket, so a source replaced during the copy is not mixed in. The
annotations are listed before CreateMultipartUpload, so a caller that
cannot list them fails before anything is written.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The ranges of a multipart copy now cover the size of the version that
HeadObject reports instead of a size from a cached listing, and the
mutable null version of a bucket with versioning suspended is not
pinned.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The copy tests used the _stubbed_fs() helper of the setxattr tests,
which master added under the same name in the same class after the
copy tests were written. It overrode theirs and dropped max_workers=1,
so the part copies could reach the Stubber out of order and fail,
which aborted the upload in a test that expects no abort.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A multipart copy chosen from a stale size over 5 GiB now copies the
version that HeadObject reports with a single CopyObject request when
its size fits, including an empty object, which no multipart copy can
split into ranges.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@laughingman7743
laughingman7743 force-pushed the fix/973-multipart-copy-metadata branch from 13e45c4 to 2154780 Compare October 3, 2026 17:11
@laughingman7743
laughingman7743 marked this pull request as ready for review October 3, 2026 17:16
@laughingman7743
laughingman7743 merged commit 587d689 into master Oct 4, 2026
15 checks passed
@laughingman7743
laughingman7743 deleted the fix/973-multipart-copy-metadata branch October 4, 2026 00:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Multipart copies drop the source metadata, and the async multipart copy does not abort on failure

1 participant