From 19cc33466cfdbd0eb26ed2f0f08769a0cd792f80 Mon Sep 17 00:00:00 2001 From: Louis Choquel Date: Thu, 24 Sep 2026 12:33:34 +0200 Subject: [PATCH 1/3] =?UTF-8?q?feature/Artifact-field-names=20=C2=B7=20L-2?= =?UTF-8?q?60924-27d936=20(#43)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit `download_artifacts` now names each saved file after the field it fills, following the rule `pipelex-sdk-js` states in its `docs/artifact-download.md`, and every verdict item carries `found_at` on both arms. The new `locate_artifacts` returns each reference with every path it sits at, and `artifact_filename` now takes that location, so a verdict item predicts its own saved name. A cross-SDK run of one corpus through both implementations found no difference in the paths or the names. Closes L-260924-27d936 Advances L-260924-ec45a1 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01KAUFtFH3dukuR2CkwoH9wv --- ## Summary by cubic `download_artifacts` now names each saved file after the field it fills, matching the rule `pipelex-sdk-js` documents, and every verdict item carries `found_at` on both arms. A new `locate_artifacts` returns each `pipelex-storage://` reference together with every `$`-rooted path it sits at, in the same discovery order as `collect_artifacts`, so a consumer can predict a saved filename before downloading. A cross-SDK run of one corpus through both implementations found no difference in the paths or the names. `artifact_filename` now takes an `ArtifactLocation` and scope instead of a URI and index, and files are named from the first path at which their reference sits (for example, `$.rooms[3].staged_photo.url` is saved as `rooms-3-staged_photo.png`). **Breaking changes** - `artifact_filename(location, content_type, scope)` replaces the old URI-based signature and raises `ArtifactOperationError` for anything that is not a valid location. - `DownloadedArtifact` now extends `ArtifactLocation` and requires `found_at`; code that builds verdict items, such as test fakes, must supply it. Closes L-260924-27d936. Advances L-260924-ec45a1. Written for commit 8068c4111efd1b4caae9a11ddd665a6a11021158. Summary will update on new commits. Review in cubic Co-authored-by: Claude Opus 5.5 --- CHANGELOG.md | 11 + docs/architecture.md | 4 +- docs/artifact-download.md | 60 ++++- docs/run-results.md | 2 +- pipelex_sdk/artifact_models.py | 29 +- pipelex_sdk/artifacts.py | 461 ++++++++++++++++++++++++-------- pipelex_sdk/client.py | 8 +- tests/e2e/test_artifacts_e2e.py | 5 +- tests/unit/test_artifacts.py | 335 +++++++++++++++++++---- 9 files changed, 730 insertions(+), 185 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index c5893fb..8a871eb 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,16 @@ # Changelog +## [Unreleased] + +### Added + +- **`locate_artifacts` and `ArtifactLocation`**: `pipelex_sdk.artifacts.locate_artifacts(value)` is the artifact walk with its paths — every `pipelex-storage://` reference in a JSON value, deduplicated in discovery order exactly as `collect_artifacts` returns them, each as an `ArtifactLocation` whose `found_at` lists every `$`-rooted path at which it sits (`$.rooms[3].staged_photo.url`, `$.items[0].url`, `$["a key"].url`), written the way `@pipelex/sdk`'s `locateArtifacts` writes them. + +### Changed + +- **`download_artifacts` names each file after the field it fills, and `artifact_filename` takes a location (Breaking)**: a saved file is named after the first path at which its reference sits — `$.rooms[3].staged_photo.url` is saved as `rooms-3-staged_photo.png`, and an output that is one image as `main_stuff.png` — instead of after the last segment of its storage key, which now supplies only the extension, so the same run saves under the same names from either SDK. `artifact_filename(location, content_type, scope)` replaces `artifact_filename(uri, content_type, index)` and raises `ArtifactOperationError` for anything but an `ArtifactLocation` whose first path is in the walk's notation; a field whose name Windows reserves for a device (`aux`, `nul`, `com1` and the like) is saved with a trailing `_` (`aux_.png`); the full rule is on `docs/artifact-download.md`. +- **`DownloadedArtifact` carries a required `found_at` (Breaking)**: every item of a `download_artifacts` verdict carries its reference's `found_at`, on the saved arm and the error arm alike, so a file that was not saved still says which field it would have filled. `DownloadedArtifact` is now an `ArtifactLocation`, so `artifact_filename` takes a verdict item as it is, and code that builds `DownloadedArtifact` values — a test fake standing in for `download_artifacts` — must now supply the field. + ## [v0.11.0] - 2026-09-23 ### Added diff --git a/docs/architecture.md b/docs/architecture.md index 0577399..347ad3d 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -257,9 +257,9 @@ The wire models are snake_case Pydantic v2. Response models are extension-open ( ## Artifact stack (`pipelex_sdk/artifacts.py` + `pipelex_sdk/artifact_models.py`) -The download twin of input preparation, and the Python twin of `@pipelex/sdk`'s `src/artifacts.ts`: `collect_artifacts` walks a value for `pipelex-storage://` references, `resolve_artifacts` mints fresh links for a whole list through the bulk route (chunked at its bound), `fetch_artifact` yields one bounded stream, and `download_artifacts` saves a run's produced files under a directory as a produced verdict. The operations take a `Protocol` rather than the client, so `PipelexAPIClient` satisfies them structurally and a test injects a fake; the client also carries all four as methods, the way it carries `upload_file` and `prepare_inputs`. The whole contract, the options and the error taxonomy are in [artifact-download.md](./artifact-download.md). +The download twin of input preparation, and the Python twin of `@pipelex/sdk`'s `src/artifacts.ts`: `locate_artifacts` walks a value for `pipelex-storage://` references and says every `$`-rooted path each one sits at, `collect_artifacts` is the same walk's bare references, `resolve_artifacts` mints fresh links for a whole list through the bulk route (chunked at its bound), `fetch_artifact` yields one bounded stream, and `download_artifacts` saves a run's produced files under a directory as a produced verdict, each file named after the field it fills by `artifact_filename`'s rule. The operations take a `Protocol` rather than the client, so `PipelexAPIClient` satisfies them structurally and a test injects a fake; the client also carries `resolve_artifacts`, `fetch_artifact` and `download_artifacts` as methods, the way it carries `upload_file` and `prepare_inputs`. The whole contract, the options and the error taxonomy are in [artifact-download.md](./artifact-download.md). -**The shapes are owned by `pipelex_sdk/artifact_models.py`** — the scope enum, the wire items, the options and the verdict, plus the public defaults — beside `product_models` and `crate_models`. That split is not cosmetic: `pipelex_sdk/errors.py` types `ScopeUnavailableError.scope` and `ArtifactAuthenticationError.verdict` with two of them and the operations module imports those errors, so keeping the shapes in the operations module would be an import cycle (`reportImportCycles` is an error here). The JS twin needs no such split, TypeScript tolerating the cycle. +**The shapes are owned by `pipelex_sdk/artifact_models.py`** — the scope enum, the wire items, the location a walk answers, the options and the verdict, plus the public defaults — beside `product_models` and `crate_models`. That split is not cosmetic: `pipelex_sdk/errors.py` types `ScopeUnavailableError.scope` and `ArtifactAuthenticationError.verdict` with two of them and the operations module imports those errors, so keeping the shapes in the operations module would be an import cycle (`reportImportCycles` is an error here). The JS twin needs no such split, TypeScript tolerating the cycle. **The object store is fetched on its own httpx client**, never the API client's: no `Authorization` header of ours can ride along to the store, redirects are refused rather than followed, and the call's timeout is both a budget over the whole exchange and httpx's per-operation timeout. `_new_storage_client` is the one seam the module opens, which is where the unit suite injects an `httpx.MockTransport`. diff --git a/docs/artifact-download.md b/docs/artifact-download.md index 05f0477..20aa486 100644 --- a/docs/artifact-download.md +++ b/docs/artifact-download.md @@ -1,8 +1,8 @@ -# Artifact download (`collect_artifacts` / `resolve_artifacts` / `fetch_artifact` / `download_artifacts`) +# Artifact download (`locate_artifacts` / `collect_artifacts` / `resolve_artifacts` / `fetch_artifact` / `download_artifacts`) -> **Status: implemented** (`pipelex_sdk/artifacts.py`, with its shapes in `pipelex_sdk/artifact_models.py`). This is the download twin of [input preparation](./input-preparation.md): where `prepare_inputs` turns local files into `pipelex-storage://` references before a run, these four operations turn the references a run produced back into bytes on disk, or into a bounded stream, afterwards. They are layered so that each is usable without the next, and they are the Python twins of `@pipelex/sdk`'s `collectArtifacts` / `resolveArtifacts` / `fetchArtifact` / `downloadArtifacts`. +> **Status: implemented** (`pipelex_sdk/artifacts.py`, with its shapes in `pipelex_sdk/artifact_models.py`). This is the download twin of [input preparation](./input-preparation.md): where `prepare_inputs` turns local files into `pipelex-storage://` references before a run, these operations turn the references a run produced back into bytes on disk, or into a bounded stream, afterwards. They are layered so that each is usable without the next, and they are the Python twins of `@pipelex/sdk`'s `locateArtifacts` / `collectArtifacts` / `resolveArtifacts` / `fetchArtifact` / `downloadArtifacts`. > -> **They need a platform that serves the bulk resolve route.** Everything below `collect_artifacts` mints its links through `POST /v1/resolve-storage-url/bulk`, a hosted-platform route. The public bare runner (`pipelex-api`) has no resolve route at all, single or bulk, and a hosted deployment that predates the route answers a `404`; in both cases the operation raises the existing `ApiResponseError` and nothing is downloaded. `resolve_storage_url`, the single-reference primitive, stays for the callers that have one link to mint. +> **They need a platform that serves the bulk resolve route.** Everything below the pure walk mints its links through `POST /v1/resolve-storage-url/bulk`, a hosted-platform route. The public bare runner (`pipelex-api`) has no resolve route at all, single or bulk, and a hosted deployment that predates the route answers a `404`; in both cases the operation raises the existing `ApiResponseError` and nothing is downloaded. `resolve_storage_url`, the single-reference primitive, stays for the callers that have one link to mint. ## Why this exists @@ -12,20 +12,28 @@ Downloading is **explicit and separate from running**, the input-preparation rul Nothing here reads the embedded `public_url`. Every link is minted fresh by the platform, which is what makes a download work long after the embedded link died, and what keeps the tenant boundary where the platform enforces it. -## The four operations +## The operations -### `collect_artifacts(value)` — the pure walk +### `locate_artifacts(value)` and `collect_artifacts(value)` — the pure walk ```python -from pipelex_sdk.artifacts import collect_artifacts +from pipelex_sdk.artifacts import collect_artifacts, locate_artifacts + +locations = locate_artifacts(results.main_stuff) +# [ArtifactLocation(uri="pipelex-storage://org/runs/01J…/outputs/2325fcfe.png", +# found_at=["$.rooms[0].staged_photo.url"]), …] uris = collect_artifacts(results.main_stuff) -# ["pipelex-storage://org/runs/01J…/outputs/illustration.png", …] +# ["pipelex-storage://org/runs/01J…/outputs/2325fcfe.png", …] — the same walk, references only ``` -Walks any JSON-shaped value and returns every string that **is** a `pipelex-storage://` reference — the whole string, scheme first, with something after the scheme. A string that merely contains a reference does not count; the bare scheme does not count; nothing else is looked at. The result is deduplicated and kept in discovery order. It is a contract rather than a heuristic, because the scheme is unambiguous: the runtime serializes a produced file as content carrying its reference in `url`, and nothing else on the wire starts that way. Mappings, lists and pydantic models are walked alike, so a parsed `main_stuff` and a whole `RunResults` both work. +Walks any JSON-shaped value and returns every string that **is** a `pipelex-storage://` reference — the whole string, scheme first, with something after the scheme. A string that merely contains a reference does not count; the bare scheme does not count; nothing else is looked at. The result is deduplicated and kept in discovery order, the order of each reference's first sighting. It is a contract rather than a heuristic, because the scheme is unambiguous: the runtime serializes a produced file as content carrying its reference in `url`, and nothing else on the wire starts that way. Mappings, lists and pydantic models are walked alike, a model by its `model_dump()` keys, so a parsed `main_stuff` and a whole `RunResults` both work. + +`locate_artifacts` also says where each reference sits. It answers one `ArtifactLocation` per reference (in `pipelex_sdk.artifact_models`), whose `found_at` lists every path at which the reference occurs, in walk order, so a reference the output repeats is one entry with several paths, and `found_at[0]` is where it was first seen — the path a saved file is named after. `collect_artifacts` is the same walk's references alone. -It is pure — no network, no key — so a consumer can count or list a result's files without resolving any of them. +**Path notation.** A path is rooted at `$`, the walked value itself. An object key matching `^[A-Za-z_][A-Za-z0-9_]*$` is written `.key`, any other key `["…"]` in JSON string escaping, and an array index `[n]`. So a nested field reads `$.rooms[3].staged_photo.url`, a list member `$.items[0].url` (a list output arrives as the `{"items": [...]}` envelope), a key that is not an identifier `$["a key"].url`, and an output that is itself a reference `$`. A path is the exact location of the string, the final `url` of a content object included, so a consumer can follow it into the JSON without guessing. The escaping is `JSON.stringify`'s — the quote, the backslash and the control characters are escaped, any other character is kept as typed, and a surrogate standing alone is written `\uXXXX` — so the same output yields the same paths from either SDK. + +Both are pure — no network, no key — so a consumer can count, list or place a result's files without resolving any of them. ### `resolve_artifacts(uris)` — fresh links for a whole list @@ -93,11 +101,38 @@ if not verdict.all_saved: **How it downloads.** The whole set is resolved through the bulk route ahead of the work, then an `asyncio.Semaphore` bounds how many references are in flight at once (`concurrency`, default 4) over the whole per-reference pipeline: fetch, create the file exclusively, stream the body in. Resolution is just-in-time where it matters: a link that has expired by the time its task reaches it — a large set downloaded a few at a time can outlive the fifteen-minute link — is resolved again for that reference alone, so no fetch ever runs on a stale signature. -**Filenames.** Each file is named by `artifact_filename` (exported): the last segment of the storage key, reduced to `[A-Za-z0-9._-]` with leading dots stripped so it can never name anything outside the directory, capped in length with the extension preserved, given an extension from the content type when the key has none, and falling back to `artifact-N`. Files are **never overwritten**: a name already on disk gets a numeric suffix (`report-1.pdf`, `report-2.pdf`), through exclusive creation (`os.O_EXCL`) rather than an exists-check, so two tasks cannot race for one name. The directory is created if missing. +**Filenames: each file is named after the field it fills.** The name comes from the first path in the reference's `found_at`, by the rule `artifact_filename(location, content_type, scope)` (exported) applies: + +1. A final `url` key is dropped, since the runtime's image and document contents carry their reference there. A reference under any other key keeps that key, and a `url` key that is not final is kept. +2. Each key is reduced to `[A-Za-z0-9_]`, every other character becoming `_` — `-` and `.` included, since one is the separator and the other would fake an extension. An index stays its digits. +3. The segments are joined with `-`. A reference that is the walked value itself, or its `url`, has no segment left and takes the scope's name, so an output that is one image is saved as `main_stuff.png`. +4. A name over the length cap (128 characters, extension included) keeps its tail: whole leading segments are dropped first, since the last ones are the specific ones, and a single segment still too long is cut to fit. +5. A stem Windows reserves for a device — `con`, `prn`, `aux`, `nul`, `com0` to `com9` or `lpt0` to `lpt9`, in any case — gets a trailing `_`, so a field named `aux` is saved as `aux_.png`. Windows reserves those names whatever the extension, and a field name is the method author's to choose. +6. The extension is the one the storage key's last segment carries, reduced to `[A-Za-z0-9]`, when it has a short one; otherwise the content type's, for the types a run produces (`image/png` gives `.png`, `application/pdf` gives `.pdf`); otherwise there is none. The segment is percent-decoded the way `decodeURIComponent` decodes it, and read as typed where that would fail, so both SDKs find the same extension in the same key. + +For example, a home-staging method whose output is + +```json +{ + "rooms": [ + { + "original_photo": { "url": "pipelex-storage://org/assets/53174b03.png", "public_url": "…" }, + "staged_photo": { "url": "pipelex-storage://org/runs/01J…/outputs/2325fcfe.png", "public_url": "…" } + }, + { "original_photo": { "url": "…" }, "staged_photo": { "url": "…" } } + ] +} +``` + +is saved as `rooms-0-original_photo.png`, `rooms-0-staged_photo.png`, `rooms-1-original_photo.png` and `rooms-1-staged_photo.png`, where the storage keys alone (`53174b03.png`, `2325fcfe.png`) would not say which picture is which. Each verdict item's `found_at` carries the unreduced path, `$.rooms[0].staged_photo.url`. + +The result is always a bare filename — ASCII letters, digits, `_` and the `-` joins, then an optional extension — never empty, never starting with a dot and never a device name, so it can name nothing but a regular file directly inside `dir_path`. `artifact_filename` takes an `ArtifactLocation` from `locate_artifacts`, which is how a consumer predicts a name before downloading; a `DownloadedArtifact` is an `ArtifactLocation` too, so a verdict item can be passed as it is. It raises `ArtifactOperationError` for anything that is not an `ArtifactLocation`, for a location whose `found_at` is empty or whose `found_at[0]` is not a path in the notation above, and for an unknown scope. + +Files are **never overwritten**: a name already on disk gets a numeric suffix (`report-1.pdf`, `report-2.pdf`), through exclusive creation (`os.O_EXCL`) rather than an exists-check, so two tasks cannot race for one name. Two references whose paths reduce to one name (`"staged photo"` and `"staged-photo"`) are told apart by the same suffix, and the verdict says which file is which. A reference found at several paths is saved once, under the name of the first. The directory is created if missing. **Cleanup.** A failed or cancelled download unlinks its partial file; nothing truncated is ever left under a final name. -**The verdict.** A `DownloadArtifactsResult`, one entry per reference in discovery order: `scope`, `artifacts` (each a `DownloadedArtifact` with `uri`, `path`, `content_type`, `size` and `error`), `saved_paths` (the absolute paths of the saved ones, same order) and `all_saved`. `len(verdict.artifacts)` is the count of references walked, errors included. An empty walk over a present scope — an output that references no stored file — is a verdict with empty lists and `all_saved=True`, not an error, and it touches neither the network nor the disk. `content_type` is the platform's guess from the reference, known before the fetch, on both arms. +**The verdict.** A `DownloadArtifactsResult`, one entry per reference in discovery order: `scope`, `artifacts` (each a `DownloadedArtifact` with `uri`, `found_at`, `path`, `content_type`, `size` and `error`), `saved_paths` (the absolute paths of the saved ones, same order) and `all_saved`. `len(verdict.artifacts)` is the count of references walked, errors included. An empty walk over a present scope — an output that references no stored file — is a verdict with empty lists and `all_saved=True`, not an error, and it touches neither the network nor the disk. `found_at` is the reference's paths in the walked scope, exactly as `locate_artifacts` reports them, and `content_type` is the platform's guess from the reference, known before the fetch; both are on both arms, so an item that was not saved still says which field it would have filled. Per-item `error.code` is the fetch vocabulary above plus the download's own: `resolve_failed` (an expired link could not be re-resolved, for a reason that is not the credential), `total_limit_exceeded` (the item that would take the call past `max_total_bytes`; an item refused only because files still in flight hold the room stops nothing else, since one of them may yet fail and give it back), `write_failed` (the file could not be created, written or closed) and `aborted` (not yet started when a credential failure stopped the call). @@ -107,7 +142,7 @@ Per-item `error.code` is the fetch vocabulary above plus the download's own: `re - `FieldNotIncludedError` — the results read never carried the scope's key, so it is absent from `results.model_fields_set`. This is the Python reading of the JS `undefined`: ask for the key and read again; - `ScopeUnavailableError` — the key WAS relayed and its value is `None`, which is the platform saying it has no such artifact for this run (`scope` and `run_id` on the error). Reading by `run_id`, a null `main_stuff` is already `MissingMainStuffError` from `get_run_result`; - `ArtifactAuthenticationError` — the resolve route refused the credential (`401` / `403`), on the first resolve or on a re-resolve part-way through. It carries `verdict`, the result as it stood: the refusal stops the remaining references being taken but lets the fetches already running finish, since they are on presigned links that do not carry the credential, so every file saved is real and listed and the rest are marked `aborted` with a detail naming the credential failure; -- `ArtifactOperationError` — an unusable directory, both selectors or neither, or nonsense bounds; +- `ArtifactOperationError` — an unusable directory, both selectors or neither, an unknown `scope` (one that reached the call unvalidated, since `DownloadArtifactsOptions` refuses it at construction), or nonsense bounds; - and the transport and lifecycle errors of the reads it makes, unchanged: `ApiResponseError` for a deployment without the bulk route, `RunLifecycleUnavailableError` for a bare runner asked by `run_id`, `ApiUnreachableError`. Everything else that can go wrong with one reference is that reference's `error`. @@ -134,6 +169,7 @@ The contract is the JS one — the same operations, the same defaults, the same - **An async context manager, not a returned response**, because an httpx stream is only live inside its own block. - **`dir_path`, not `dir`**, which is a Python builtin. - **Options as pydantic models**, so a caller passes `DownloadArtifactsOptions(...)` where the JS twin spreads keys into one request object. +- **A location is a model, not a shape.** Where the JS `artifactFilename` takes any object carrying `uri` and `found_at`, the Python one takes an `ArtifactLocation`, and `DownloadedArtifact` subclasses it so a verdict item passes as it is. - **The shapes live in `artifact_models.py`**, beside `product_models` and `crate_models`, because `pipelex_sdk.errors` types two of its artifact errors with them and the operations module imports those errors — one home for the shapes keeps that from being an import cycle. ## Round trip diff --git a/docs/run-results.md b/docs/run-results.md index f652c16..25a215b 100644 --- a/docs/run-results.md +++ b/docs/run-results.md @@ -177,7 +177,7 @@ The usage pair reports what each inference call consumed and cost — one `Token A run that produces an image, a PDF or a document does not embed the bytes. The content inside `main_stuff` carries the file's durable reference — a `pipelex-storage://` URI, in the content's `url` — beside a `public_url` the storage provider signed when the run wrote the file. **That signed link is short-lived and must not be stored**: it expires on the provider's own schedule, so a link persisted in a database or rendered into a cached page stops working without warning, while the `pipelex-storage://` reference beside it is permanent and is what belongs in your records. -**The whole download direction is [`artifact-download.md`](./artifact-download.md)**, and it is where a consumer should start: `collect_artifacts(results.main_stuff)` lists the references without touching the network, `resolve_artifacts` mints a fresh link for each through the platform's bulk route, `fetch_artifact` streams one within bounds, and `download_artifacts` saves a whole run's files under a directory and answers a produced verdict. None of them reads the embedded `public_url`. +**The whole download direction is [`artifact-download.md`](./artifact-download.md)**, and it is where a consumer should start: `locate_artifacts(results.main_stuff)` lists the references without touching the network, each with the paths of the fields it fills (`collect_artifacts` gives the references alone), `resolve_artifacts` mints a fresh link for each through the platform's bulk route, `fetch_artifact` streams one within bounds, and `download_artifacts` saves a whole run's files under a directory, each named after the field it fills, and answers a produced verdict. None of them reads the embedded `public_url`. For the single reference you already hold, the raw primitive is still there: diff --git a/pipelex_sdk/artifact_models.py b/pipelex_sdk/artifact_models.py index 9be6a06..d68e99b 100644 --- a/pipelex_sdk/artifact_models.py +++ b/pipelex_sdk/artifact_models.py @@ -133,14 +133,35 @@ class DownloadArtifactsOptions(FetchArtifactOptions): max_total_bytes: int = DEFAULT_DOWNLOAD_MAX_TOTAL_BYTES -class DownloadedArtifact(BaseModel): +class ArtifactLocation(BaseModel): + """Where one `pipelex-storage://` reference sits in a walked value — what `locate_artifacts` + answers per reference. `found_at` lists every path at which the reference occurs, in walk order, + and is never empty: `found_at[0]` is where it was first seen, and the path a saved file is named + after. + + A path is `$`-rooted: `$` is the walked value itself, an object key matching + `^[A-Za-z_][A-Za-z0-9_]*$` is `.key`, any other key is `["…"]` in JSON string escaping, and an + array index is `[n]` — `$.rooms[3].staged_photo.url`, `$.items[0].url`, `$["a key"].url`. It is + the exact path of the string, the final `url` of a content object included. + """ + + #: The reference, exactly as it appears in the walked value. + uri: str + #: Every `$`-rooted path at which the reference occurs, in walk order. + found_at: list[str] + + +class DownloadedArtifact(ArtifactLocation): """One reference's outcome in a download verdict — one shape with nullable fields, like `ResolvedArtifact`: either `path` and `size` are set and `error` is `None`, or `error` is set and - both are `None`. `content_type` is the platform's guess from the reference's extension, known - before the fetch, on both arms. + both are `None`. `found_at` says where the reference sits in the walked scope, the first path + being the one that named the file, and `content_type` is the platform's guess from the + reference's extension, known before the fetch; both are on both arms, so an item that was not + saved still says which field it would have filled. + + It is an `ArtifactLocation`, so `artifact_filename` takes a verdict item as it is. """ - uri: str #: Absolute path of the written file. path: str | None = None content_type: str | None = None diff --git a/pipelex_sdk/artifacts.py b/pipelex_sdk/artifacts.py index cb1e0bf..0b2faa7 100644 --- a/pipelex_sdk/artifacts.py +++ b/pipelex_sdk/artifacts.py @@ -1,15 +1,16 @@ """The artifact stack — the download twin of `prepare_inputs`, in layers so each operation is usable without the next: -- `collect_artifacts(value)` — a pure walk of any JSON-shaped value for the strings that ARE - `pipelex-storage://` references. No network, no key. +- `locate_artifacts(value)` — a pure walk of any JSON-shaped value for the strings that ARE + `pipelex-storage://` references, each with every `$`-rooted path it sits at. `collect_artifacts` + is the same walk's bare references. No network, no key. - `resolve_artifacts(client, uris)` — the platform's bulk resolve route over a whole list, chunked at the route's bound, one verdict per reference. - `fetch_artifact(client, uri)` — an async context manager yielding a bounded stream for one reference: resolved fresh, timed out, redirects refused, the byte cap enforced mid-stream, no credentials forwarded, headers neutral. - `download_artifacts(client, ...)` — a run's produced files saved under a directory by a bounded - pool of tasks, as a produced verdict. + pool of tasks, each named after the field it fills (`artifact_filename`), as a produced verdict. A produced file is never embedded in a run's results: the content carries its durable `pipelex-storage://` reference beside a signed `public_url` that expires on the provider's @@ -25,6 +26,7 @@ from __future__ import annotations import asyncio +import json import math import os import re @@ -32,7 +34,7 @@ from dataclasses import dataclass from datetime import UTC, datetime from pathlib import Path -from typing import TYPE_CHECKING, Any, BinaryIO, Protocol, cast +from typing import TYPE_CHECKING, Any, BinaryIO, Protocol, TypeAlias, cast from urllib.parse import unquote, urlsplit import httpx @@ -44,6 +46,7 @@ BULK_RESOLVE_MAX_URIS, PIPELEX_STORAGE_SCHEME, ArtifactItemError, + ArtifactLocation, ArtifactScope, BulkResolvedStorageUrls, DownloadArtifactsOptions, @@ -65,7 +68,7 @@ from pipelex_sdk.runs import RunResultCompleted, RunResultFailed, RunResultRunning, RunResults if TYPE_CHECKING: - from collections.abc import AsyncGenerator, AsyncIterator + from collections.abc import AsyncGenerator, AsyncIterator, Sequence from contextlib import AbstractAsyncContextManager from pipelex_sdk.runs import RunResultState @@ -80,6 +83,30 @@ #: Longest filename `download_artifacts` writes, extension included. _MAX_FILENAME_LENGTH = 128 +#: Longest extension taken from a storage key, dot excluded. Anything longer after the key's last +#: dot is read as part of a name rather than as an extension, and the content type's extension is +#: used instead. +_MAX_EXTENSION_LENGTH = 10 + +#: Stems Windows reserves for a device, in any case and whatever the extension: `aux.png` there +#: names the auxiliary device, not a file. A stem is a field name the method author chose, so +#: `$.aux.url` would otherwise reach one. +_WINDOWS_DEVICE_STEM = re.compile(r"con|prn|aux|nul|com[0-9]|lpt[0-9]", re.IGNORECASE) + +#: An object key rendered as `.key` in a path; any other key is rendered as `["…"]`. +_IDENTIFIER_KEY = re.compile(r"[A-Za-z_][A-Za-z0-9_]*") + +#: The three steps a rendered path is read back from, each matched where the last one ended. +_KEY_STEP = re.compile(r"\.([A-Za-z_][A-Za-z0-9_]*)") +_INDEX_STEP = re.compile(r"\[([0-9]+)\]") +_QUOTED_STEP = re.compile(r'\[("(?:[^"\\]|\\.)*")\]', re.DOTALL) + +#: A UTF-16 surrogate standing alone in a Python string, which `JSON.stringify` escapes as `\uXXXX`. +_LONE_SURROGATE = re.compile(r"[\ud800-\udfff]") + +#: A `%` that does not open a two-digit escape, which makes `decodeURIComponent` refuse the segment. +_BARE_PERCENT = re.compile(r"%(?![0-9A-Fa-f]{2})") + #: Ceiling on collision suffixes before the never-overwrite rule gives up. _MAX_UNIQUE_ATTEMPTS = 10_000 @@ -130,7 +157,22 @@ class ArtifactCapableClient(BulkResolveClient, Protocol): async def get_run_result(self, run_id: str) -> RunResultState: ... -# ── collect_artifacts ──────────────────────────────────────────────── +# ── locate_artifacts / collect_artifacts ───────────────────────────── + +#: One step of a path into a walked value: an object key, or an array index. +_PathSegment: TypeAlias = str | int + + +@dataclass +class _LocatedReference: + """A reference as the walk records it: the raw segments of every path it was found at. + + `download_artifacts` names its files from these segments directly, so it never parses a rendered + path back; `found_at` is only their rendering. + """ + + uri: str + paths: list[tuple[_PathSegment, ...]] def is_storage_reference(value: str) -> bool: @@ -138,9 +180,26 @@ def is_storage_reference(value: str) -> bool: return value.startswith(PIPELEX_STORAGE_SCHEME) and len(value) > len(PIPELEX_STORAGE_SCHEME) +def locate_artifacts(value: Any) -> list[ArtifactLocation]: + """Every `pipelex-storage://` reference inside a JSON-shaped value, each with every path at which + it occurs. + + The references are deduplicated and kept in discovery order (the order of their first sighting), + and each one's `found_at` lists its paths in walk order, so `found_at[0]` is where it was first + seen. The string test is `collect_artifacts`'s: a string counts only when it IS a reference. A + path is rooted at `$`, the walked value itself; an object key matching `^[A-Za-z_][A-Za-z0-9_]*$` + is written `.key`, any other key `["…"]` in JSON string escaping, and an array index `[n]`. The + runtime serializes a produced image or document as content carrying its reference in `url`, so a + typical path ends there: `$.rooms[3].staged_photo.url`, `$.items[0].url`, or `$.url` for an + output that is one image. Pure — no network, no key. Mappings, sequences and pydantic models are + walked alike, a model by its `model_dump()` keys. + """ + return [_render_location(located) for located in _walk_references(value, with_paths=True)] + + def collect_artifacts(value: Any) -> list[str]: """Every `pipelex-storage://` reference inside a JSON-shaped value, deduplicated, in discovery - order. + order — the references of `locate_artifacts`, without their paths. A string counts only when it IS a reference — the whole string, scheme first, with something after the scheme; text that merely contains one does not. The scheme is unambiguous, so this walk @@ -150,72 +209,229 @@ def collect_artifacts(value: Any) -> list[str]: any of them. Mappings, sequences and pydantic models are walked alike, so `results.main_stuff` (a parsed JSON value) and a whole `RunResults` both work. """ - found: dict[str, None] = {} - _walk_for_references(value, found) - return list(found) - - -def _walk_for_references(value: Any, found: dict[str, None]) -> None: - """Depth-first walk collecting every string that is a storage reference, in discovery order.""" - if isinstance(value, str): - if is_storage_reference(value): - found[value] = None - return - if isinstance(value, BaseModel): - _walk_for_references(value.model_dump(), found) - return - if isinstance(value, dict): - for entry in cast("dict[str, Any]", value).values(): - _walk_for_references(entry, found) - return - if isinstance(value, (list, tuple)): - for item in cast("list[Any]", value): - _walk_for_references(item, found) + return [located.uri for located in _walk_references(value, with_paths=False)] -# ── artifact_filename ──────────────────────────────────────────────── +def _walk_references(value: Any, *, with_paths: bool) -> list[_LocatedReference]: + """The walk both public functions share: depth first, keys in mapping order. + + With `with_paths` off it records each reference once and no path at all, which is what + `collect_artifacts` needs: copying the trail for every occurrence costs memory in proportion to + occurrences times depth, where the deduplicated list needs only one entry per unique reference. + """ + # A dict keeps insertion order, which is the order of first sighting. + by_uri: dict[str, _LocatedReference] = {} + trail: list[_PathSegment] = [] + + def visit(node: Any) -> None: + if isinstance(node, str): + if not is_storage_reference(node): + return + known = by_uri.get(node) + if known is None: + known = _LocatedReference(uri=node, paths=[]) + by_uri[node] = known + if with_paths: + known.paths.append(tuple(trail)) + return + if isinstance(node, BaseModel): + visit(node.model_dump()) + return + if isinstance(node, dict): + for key, entry in cast("dict[Any, Any]", node).items(): + # A JSON object's keys are strings; any other key is read as the string it prints as. + trail.append(key if isinstance(key, str) else str(key)) + visit(entry) + trail.pop() + return + if isinstance(node, (list, tuple)): + for index_item, item in enumerate(cast("list[Any]", node)): + trail.append(index_item) + visit(item) + trail.pop() + + visit(value) + return list(by_uri.values()) + + +def _render_location(located: _LocatedReference) -> ArtifactLocation: + return ArtifactLocation(uri=located.uri, found_at=[_render_path(segments) for segments in located.paths]) + + +def _render_path(segments: Sequence[_PathSegment]) -> str: + """Segments to the `$`-rooted notation `found_at` carries.""" + rendered = ["$"] + for segment in segments: + if isinstance(segment, int): + rendered.append(f"[{segment}]") + elif _IDENTIFIER_KEY.fullmatch(segment): + rendered.append(f".{segment}") + else: + rendered.append(f"[{_json_string(segment)}]") + return "".join(rendered) -def artifact_filename(uri: str, content_type: str | None, index: int) -> str: - """The bare filename a storage reference is saved under. +def _json_string(key: str) -> str: + r"""A key as `JSON.stringify` writes it, so a path reads the same from either SDK. - The last segment of the storage key, reduced to a conservative character set so it can never name - anything but a regular file directly inside the target directory. Path separators are the split - point, so no traversal survives; leading dots are stripped, so no hidden file and no `..`; - everything outside `[A-Za-z0-9._-]` becomes `_`; an empty result falls back to a numbered - `artifact-N`. Length is capped with the extension preserved, and an extension is added from the - content type when the key carries none. A collision on disk is not this function's concern: - `download_artifacts` suffixes the stem (`name-1.ext`) on exclusive creation, so a file is never - overwritten. + `json.dumps` with `ensure_ascii` off escapes exactly what `JSON.stringify` does — the quote, the + backslash and the control characters, the latter as lowercase `\u00XX` — and keeps every other + character as typed, except a surrogate standing alone: `JSON.stringify` escapes it, and left raw + it would make the path a string no UTF-8 encoder accepts. + """ + dumped = json.dumps(key, ensure_ascii=False) + return _LONE_SURROGATE.sub(lambda match: f"\\u{ord(match.group()):04x}", dumped) + + +def _parse_path(path: str) -> list[_PathSegment] | None: + """The `$`-rooted notation back to segments, for a location that reaches `artifact_filename` from + outside the walk. The rendering is lossless, so this is exact for every path the walk produced; + anything else is `None`. + """ + if not path.startswith("$"): + return None + segments: list[_PathSegment] = [] + position = 1 + while position < len(path): + key_step = _KEY_STEP.match(path, position) + if key_step is not None: + segments.append(key_step.group(1)) + position = key_step.end() + continue + index_step = _INDEX_STEP.match(path, position) + if index_step is not None: + segments.append(int(index_step.group(1))) + position = index_step.end() + continue + quoted_step = _QUOTED_STEP.match(path, position) + if quoted_step is None: + return None + try: + decoded: Any = json.loads(quoted_step.group(1)) + except json.JSONDecodeError: + return None + if not isinstance(decoded, str): + return None + segments.append(decoded) + position = quoted_step.end() + return segments + + +# ── artifact_filename ──────────────────────────────────────────────── + + +def artifact_filename(location: ArtifactLocation, content_type: str | None, scope: ArtifactScope) -> str: + """The bare filename a reference is saved under, named after the field it fills: the path in + `location.found_at[0]`, where the reference was first seen. + + 1. A final object key `url` is dropped, since the runtime's image and document contents carry + their reference there; a reference under any other key keeps that key, and a `url` key that + is not final is kept. + 2. Each key is reduced to `[A-Za-z0-9_]`, every other character (`-` and `.` included) becoming + `_`; an index stays its decimal digits. + 3. The segments are joined with `-`. An empty result — the reference is the walked value itself, + or its `url` — becomes the scope's name. + 4. Over the filename length cap (128 characters, extension included), the tail is kept: whole + leading segments are dropped first, since the last ones are the specific ones, and a single + segment still too long is cut to fit. + 5. A stem Windows reserves for a device (`con`, `prn`, `aux`, `nul`, `com0` to `com9`, + `lpt0` to `lpt9`, in any case) gets a trailing `_`, so `$.aux.url` is saved as `aux_.png`. + 6. The extension is the storage key's own, reduced to `[A-Za-z0-9]`, when it has a short one; + otherwise the content type's, for the types a run produces; otherwise there is none. + + So `$.rooms[3].staged_photo.url` is saved as `rooms-3-staged_photo.png`, and `$.url` in + `main_stuff` as `main_stuff.png`. The stem is ASCII letters, digits, `_` and the `-` joins, never + empty and never a device name, so the name can only ever be a regular file directly inside the + target directory. A collision on disk is not this function's concern: `download_artifacts` + suffixes the stem (`name-1.ext`) on exclusive creation, so a file is never overwritten. A + `DownloadedArtifact` is an `ArtifactLocation`, so a verdict item can be passed as it is. + + Raises `ArtifactOperationError` for a location that is not an `ArtifactLocation`, one whose + `found_at[0]` is not a path in the notation `locate_artifacts` writes, or an unknown scope. + """ + checked_scope = _require_scope(scope) + # A caller can hand anything here — the old signature's bare uri string among them — and every + # shape must reach the documented refusal rather than an AttributeError. + loose = cast("object", location) + if not isinstance(loose, ArtifactLocation): + msg = f"artifact_filename needs an ArtifactLocation, as locate_artifacts answers; got a {type(loose).__name__}." + raise ArtifactOperationError(msg) + first = loose.found_at[0] if loose.found_at else None + segments = _parse_path(first) if first is not None else None + if segments is None: + msg = f'artifact_filename needs a location whose first "found_at" entry is a path such as "$.items[0].url"; got {first!r}.' + raise ArtifactOperationError(msg) + return _filename_for(segments, loose.uri, content_type, checked_scope) + + +def _filename_for(segments: Sequence[_PathSegment], uri: str, content_type: str | None, scope: ArtifactScope) -> str: + """The naming rule of `artifact_filename`, over the walk's own segments.""" + named = segments[:-1] if segments and segments[-1] == "url" else segments + words: list[str] = [] + for segment in named: + word = str(segment) if isinstance(segment, int) else re.sub(r"[^A-Za-z0-9_]", "_", segment) + # Only the empty key reduces to nothing, and it says nothing about the field. + if word: + words.append(word) + extension = _extension_for(uri, content_type) + stem = _fit_stem(words or [scope], _MAX_FILENAME_LENGTH - len(extension)) + # A device stem is at most four characters, so the `_` cannot overrun the cap. + if _WINDOWS_DEVICE_STEM.fullmatch(stem): + stem += "_" + return stem + extension + + +def _fit_stem(words: Sequence[str], budget: int) -> str: + """The words joined with `-` within `budget` characters, keeping the tail.""" + start = 0 + length = sum(len(word) for word in words) + len(words) - 1 + while length > budget and start < len(words) - 1: + length -= len(words[start]) + 1 + start += 1 + return "-".join(words[start:])[:budget] + + +def _extension_for(uri: str, content_type: str | None) -> str: + """`.ext` for the saved file — the storage key's own, else the content type's — or `""`.""" + from_key = _storage_key_extension(uri) + if from_key: + return f".{from_key}" + if content_type is None: + return "" + return _EXTENSION_BY_CONTENT_TYPE.get(content_type.split(";")[0].strip().lower(), "") + + +def _storage_key_extension(uri: str) -> str: + r"""The extension the storage key's last segment carries, without its dot, reduced to + `[A-Za-z0-9]` — or `""` when it has none, or none that short. + + The segment is what follows the last `/` or `\` once the scheme, query and fragment are gone, + percent-decoded when it decodes; a leading dot is not an extension. """ key = uri.removeprefix(PIPELEX_STORAGE_SCHEME) key = re.split(r"[?#]", key, maxsplit=1)[0] - segments = [part for part in re.split(r"[\\/]", key) if part] - # `unquote` keeps a malformed escape as typed, where the JS twin's `decodeURIComponent` throws - # and falls back to the same thing; sanitization below handles either. - decoded = unquote(segments[-1]) if segments else "" - - name = re.sub(r"[^A-Za-z0-9._-]", "_", decoded) - name = re.sub(r"^[._-]+", "", name) - name = re.sub(r"[._-]+$", "", name) - if not name: - name = f"artifact-{index + 1}" - - if not _extension_of(name) and content_type is not None: - guessed = _EXTENSION_BY_CONTENT_TYPE.get(content_type.split(";")[0].strip().lower()) - if guessed is not None: - name += guessed - - # The cap comes last, so a guessed extension is inside it like any other. - if len(name) > _MAX_FILENAME_LENGTH: - extension = _extension_of(name) - # The extension is kept only if there is room left for a stem; a pathological extension is - # dropped rather than preserved. - if len(extension) < _MAX_FILENAME_LENGTH: - name = name[: _MAX_FILENAME_LENGTH - len(extension)] + extension - else: - name = name[:_MAX_FILENAME_LENGTH] - return name + parts = [part for part in re.split(r"[\\/]", key) if part] + decoded = _decode_uri_component(parts[-1]) if parts else "" + dot = decoded.rfind(".") + if dot <= 0: + return "" + extension = re.sub(r"[^A-Za-z0-9]", "", decoded[dot + 1 :]) + return extension if len(extension) <= _MAX_EXTENSION_LENGTH else "" + + +def _decode_uri_component(segment: str) -> str: + """`decodeURIComponent`, or the segment as typed where the JS twin's call would throw. + + `unquote` alone is lenient twice over — it keeps a `%` that opens no escape and replaces bytes + that are not UTF-8 — where `decodeURIComponent` refuses the whole segment, so the same key would + otherwise yield a different extension in each SDK. + """ + if _BARE_PERCENT.search(segment): + return segment + try: + return unquote(segment, errors="strict") + except UnicodeDecodeError: + return segment def _extension_of(name: str) -> str: @@ -224,6 +440,15 @@ def _extension_of(name: str) -> str: return name[dot:] if dot > 0 else "" +def _require_scope(scope: ArtifactScope) -> ArtifactScope: + """The scope as the enum, or the refusal naming the two there are.""" + try: + return ArtifactScope(scope) + except ValueError as exc: + msg = f'"scope" must be "main_stuff" or "working_memory", got {scope!r}.' + raise ArtifactOperationError(msg) from exc + + # ── resolve_artifacts ──────────────────────────────────────────────── @@ -529,24 +754,26 @@ async def download_artifacts( """Save a run's produced files under `dir_path`, and answer a produced verdict. Takes exactly one of `run_id` (the results are re-read, so a run is downloadable days later) or - `results` (a `RunResults` in hand); walks the requested scope with `collect_artifacts`; resolves + `results` (a `RunResults` in hand); walks the requested scope with `locate_artifacts`; resolves the whole set through the bulk route ahead of the tasks; then a bounded number of tasks (`concurrency`, held by an `asyncio.Semaphore`) each fetch, create their file exclusively and stream the body in, re-resolving any link that has expired by the time a task reaches it. The - embedded `public_url` is never used. Files are named by `artifact_filename` and never - overwritten; a failed or cancelled download unlinks its partial file. + embedded `public_url` is never used. Each file is named after the field it fills, by + `artifact_filename`'s rule, and never overwritten; a failed or cancelled download unlinks its + partial file. Returns one entry per reference, errors as values. It raises only when no verdict can be produced: `RunStillRunningError` or `RunFailedError` for a run that has not completed, `FieldNotIncludedError` when the results read did not carry the scope's key, `ScopeUnavailableError` when the key was relayed as `None`, `ArtifactAuthenticationError` (carrying the verdict so far) when the resolve - route refuses the credential, `ArtifactOperationError` for an unusable directory or nonsense - bounds, and the transport and lifecycle errors of the reads it makes (`ApiResponseError` for a - deployment without the bulk route, `RunLifecycleUnavailableError` for a bare runner asked by id, - `ApiUnreachableError`). Cancelling the awaiting task raises `asyncio.CancelledError` out of here, - with every partial file unlinked first. + route refuses the credential, `ArtifactOperationError` for an unusable directory, an unknown + scope or nonsense bounds, and the transport and lifecycle errors of the reads it makes + (`ApiResponseError` for a deployment without the bulk route, `RunLifecycleUnavailableError` for a + bare runner asked by id, `ApiUnreachableError`). Cancelling the awaiting task raises + `asyncio.CancelledError` out of here, with every partial file unlinked first. """ opts = options or DownloadArtifactsOptions() + scope = _require_scope(opts.scope) bounds = _validated_bounds(opts) if opts.concurrency < 1: msg = f'"concurrency" must be a positive integer, got {opts.concurrency}.' @@ -559,11 +786,14 @@ async def download_artifacts( msg = "download_artifacts takes exactly one of `run_id` (the results are re-read) or `results` (a RunResults in hand)." raise ArtifactOperationError(msg) read_results = results if results is not None else await _read_completed_results(client, cast("str", run_id)) - walked = _scope_value(read_results, opts.scope) + walked = _scope_value(read_results, scope) - uris = collect_artifacts(walked) - if not uris: - return _assemble_verdict(opts.scope, []) + # The walk's own record names the files; `locations` is what the verdict reports. + located = _walk_references(walked, with_paths=True) + if not located: + return _assemble_verdict(scope, []) + locations = [_render_location(reference) for reference in located] + uris = [reference.uri for reference in located] target_dir = Path(dir_path).resolve() try: @@ -578,7 +808,7 @@ async def download_artifacts( except ApiResponseError as exc: if not _is_credential_refusal(exc): raise - verdict = _assemble_verdict(opts.scope, [_item_error(uri, None, "aborted", _SKIPPED_CREDENTIAL) for uri in uris]) + verdict = _assemble_verdict(scope, [_item_error(location, None, "aborted", _SKIPPED_CREDENTIAL) for location in locations]) msg = f"The resolve route refused the credential ({exc.status}); no artifact was downloaded." raise ArtifactAuthenticationError(msg, status=exc.status, verdict=verdict) from exc @@ -594,17 +824,18 @@ async def download_artifacts( budget=budget, bounds=bounds, target_dir=target_dir, - index=index, - uri=uri, + scope=scope, + location=location, + name_path=located[index].paths[0], entry=resolved[index], ) - for index, uri in enumerate(uris) + for index, location in enumerate(locations) ] ) finally: await storage_client.aclose() - verdict = _assemble_verdict(opts.scope, list(outcomes)) + verdict = _assemble_verdict(scope, list(outcomes)) if budget.credential_failure is not None: status = budget.credential_failure.status msg = f"The resolve route refused the credential ({status}) part-way through the download; the verdict so far is on this error." @@ -647,9 +878,16 @@ def _is_credential_refusal(exc: ApiResponseError) -> bool: return exc.status in {401, 403} -def _item_error(uri: str, content_type: str | None, code: str, detail: str) -> DownloadedArtifact: - """One reference's failure, as a value on the verdict.""" - return DownloadedArtifact(uri=uri, path=None, content_type=content_type, size=None, error=ArtifactItemError(code=code, detail=detail)) +def _item_error(location: ArtifactLocation, content_type: str | None, code: str, detail: str) -> DownloadedArtifact: + """One reference's failure, as a value on the verdict — still saying where the reference sits.""" + return DownloadedArtifact( + uri=location.uri, + found_at=location.found_at, + path=None, + content_type=content_type, + size=None, + error=ArtifactItemError(code=code, detail=detail), + ) def _is_expired(expires_at: str | None) -> bool: @@ -673,41 +911,45 @@ async def _process_one( budget: _DownloadBudget, bounds: _FetchBounds, target_dir: Path, - index: int, - uri: str, + scope: ArtifactScope, + location: ArtifactLocation, + name_path: Sequence[_PathSegment], entry: ResolvedArtifact, ) -> DownloadedArtifact: - """One reference's whole pipeline — the semaphore's slot, a re-resolve on expiry, then the save.""" + """One reference's whole pipeline — the semaphore's slot, a re-resolve on expiry, then the save + under the name its first path gives it. + """ + uri = location.uri async with semaphore: if budget.credential_failure is not None: - return _item_error(uri, entry.content_type, "aborted", _SKIPPED_CREDENTIAL) + return _item_error(location, entry.content_type, "aborted", _SKIPPED_CREDENTIAL) if entry.error is None and _is_expired(entry.expires_at): try: entry = (await resolve_artifacts(client, [uri]))[0] except ApiResponseError as exc: if _is_credential_refusal(exc): budget.credential_failure = exc - return _item_error(uri, entry.content_type, "aborted", _SKIPPED_CREDENTIAL) + return _item_error(location, entry.content_type, "aborted", _SKIPPED_CREDENTIAL) msg = f"The expired link could not be re-resolved: {exc}." - return _item_error(uri, entry.content_type, "resolve_failed", msg) + return _item_error(location, entry.content_type, "resolve_failed", msg) except (PipelineRequestError, ValidationError, ValueError) as exc: # Anything else the re-resolve can fail with — an unreachable host, a malformed # answer, a body that does not parse — is this one reference's error, never the # whole download's: the other references already have their links. msg = f"The expired link could not be re-resolved: {exc}." - return _item_error(uri, entry.content_type, "resolve_failed", msg) + return _item_error(location, entry.content_type, "resolve_failed", msg) if entry.error is not None: - return _item_error(uri, None, entry.error.code, entry.error.detail) + return _item_error(location, None, entry.error.code, entry.error.detail) if entry.url is None: msg = "The bulk resolve route answered an item with neither a link nor an error." - return _item_error(uri, entry.content_type, "resolve_failed", msg) + return _item_error(location, entry.content_type, "resolve_failed", msg) return await _save_one( storage_client=storage_client, budget=budget, bounds=bounds, target_dir=target_dir, - index=index, - uri=uri, + location=location, + filename=_filename_for(name_path, uri, entry.content_type, scope), download_url=entry.url, content_type=entry.content_type, ) @@ -769,12 +1011,15 @@ async def _save_one( budget: _DownloadBudget, bounds: _FetchBounds, target_dir: Path, - index: int, - uri: str, + location: ArtifactLocation, + filename: str, download_url: str, content_type: str | None, ) -> DownloadedArtifact: - """Fetch one resolved link and write it under the download directory, within the total budget.""" + """Fetch one resolved link and write it under the download directory as `filename`, suffixed on a + collision, within the total budget. + """ + uri = location.uri target: _TargetFile | None = None written = 0 reserved = 0 @@ -792,19 +1037,19 @@ def share() -> int: f"Saving this {_format_mib(reserved)} artifact would take the download past its " f"{_format_mib(budget.max_total_bytes)} total limit." ) - return _item_error(uri, content_type, "total_limit_exceeded", msg) + return _item_error(location, content_type, "total_limit_exceeded", msg) budget.committed += reserved share_taken = True try: # Synchronous on purpose: a cancellation landing inside a worker thread would leave a file # the handler below cannot see, and an exclusive create is not worth a thread. - target = _open_unique_file(target_dir, artifact_filename(uri, content_type, index)) + target = _open_unique_file(target_dir, filename) except OSError as exc: budget.committed -= share() share_taken = False msg = f"The file could not be created: {exc}." - return _item_error(uri, content_type, "write_failed", msg) + return _item_error(location, content_type, "write_failed", msg) async for chunk in stream.aiter_bytes(): # Only the bytes past this file's reservation are new to the total. @@ -814,7 +1059,7 @@ def share() -> int: budget.committed -= share() share_taken = False msg = f"This artifact took the download past its {_format_mib(budget.max_total_bytes)} total limit." - return _item_error(uri, content_type, "total_limit_exceeded", msg) + return _item_error(location, content_type, "total_limit_exceeded", msg) budget.committed += growth written += len(chunk) try: @@ -824,20 +1069,20 @@ def share() -> int: budget.committed -= share() share_taken = False msg = f"The file could not be written: {exc}." - return _item_error(uri, content_type, "write_failed", msg) + return _item_error(location, content_type, "write_failed", msg) except ArtifactFetchError as exc: if target is not None: target.remove() if share_taken: budget.committed -= share() - return _item_error(uri, content_type, exc.code, str(exc)) + return _item_error(location, content_type, exc.code, str(exc)) except (httpx.HTTPError, OSError) as exc: if target is not None: target.remove() if share_taken: budget.committed -= share() msg = f"The artifact could not be read: {exc}." - return _item_error(uri, content_type, "network", msg) + return _item_error(location, content_type, "network", msg) except asyncio.CancelledError: # A cancelled download leaves nothing truncated behind, then lets the cancellation through. if target is not None: @@ -852,11 +1097,11 @@ def share() -> int: target.remove() budget.committed -= share() msg = f"The file could not be closed: {exc}." - return _item_error(uri, content_type, "write_failed", msg) + return _item_error(location, content_type, "write_failed", msg) # A body shorter than it declared gives the unused reservation back. budget.committed -= share() - written - return DownloadedArtifact(uri=uri, path=str(target.path), content_type=content_type, size=written, error=None) + return DownloadedArtifact(uri=uri, found_at=location.found_at, path=str(target.path), content_type=content_type, size=written, error=None) def _assemble_verdict(scope: ArtifactScope, artifacts: list[DownloadedArtifact]) -> DownloadArtifactsResult: diff --git a/pipelex_sdk/client.py b/pipelex_sdk/client.py index 13d44f5..d053c6a 100644 --- a/pipelex_sdk/client.py +++ b/pipelex_sdk/client.py @@ -1158,7 +1158,8 @@ async def resolve_artifacts(self, uris: list[str]) -> list[ResolvedArtifact]: """Resolve a whole list of `pipelex-storage://` references through the bulk route, chunked at its bound, answering one `ResolvedArtifact` per reference in request order with per-reference failure as a value. The reading layer of the artifact stack: pair it with `collect_artifacts` - to mint fresh links for everything a run produced. See `docs/artifact-download.md`. + (or `locate_artifacts`, which also says where each reference sits) to mint fresh links for + everything a run produced. See `docs/artifact-download.md`. """ return await _resolve_artifacts_impl(self, uris) @@ -1182,8 +1183,9 @@ async def download_artifacts( Keyed on a `run_id` (the results are re-read, so it works days after the run) or a `RunResults` in hand; walks the `main_stuff` scope by default, `working_memory` on request; resolves every - link fresh (never the embedded `public_url`); and returns a produced verdict, one entry per - reference, errors as values. See `docs/artifact-download.md`. + link fresh (never the embedded `public_url`); names each file after the field it fills; and + returns a produced verdict, one entry per reference with the paths it sits at, errors as + values. See `docs/artifact-download.md`. """ return await _download_artifacts_impl(self, dir_path=dir_path, run_id=run_id, results=results, options=options) diff --git a/tests/e2e/test_artifacts_e2e.py b/tests/e2e/test_artifacts_e2e.py index 1410939..6cb05f8 100644 --- a/tests/e2e/test_artifacts_e2e.py +++ b/tests/e2e/test_artifacts_e2e.py @@ -27,7 +27,7 @@ import pytest from pipelex_sdk.artifact_models import ArtifactScope, DownloadArtifactsOptions -from pipelex_sdk.artifacts import collect_artifacts, download_artifacts, fetch_artifact +from pipelex_sdk.artifacts import artifact_filename, collect_artifacts, download_artifacts, fetch_artifact from pipelex_sdk.client import PipelexAPIClient from pipelex_sdk.crate_models import MthdsFileItem @@ -108,6 +108,9 @@ async def _round_trip() -> tuple[str, DownloadArtifactsResult]: assert echoed.size == len(_PDF_BYTES) assert echoed.path in verdict.saved_paths assert echoed.path is not None + # The file is named after the working-memory field the echoed input sits in. + assert echoed.found_at[0].startswith("$.") + assert Path(echoed.path).name == artifact_filename(echoed, echoed.content_type, ArtifactScope.WORKING_MEMORY) assert Path(echoed.path).read_bytes() == _PDF_BYTES def test_resolves_through_the_bulk_route_and_refuses_a_malformed_reference_as_a_value(self) -> None: diff --git a/tests/unit/test_artifacts.py b/tests/unit/test_artifacts.py index f295974..3de0cc9 100644 --- a/tests/unit/test_artifacts.py +++ b/tests/unit/test_artifacts.py @@ -11,17 +11,20 @@ import asyncio import gzip import json +import re from datetime import UTC, datetime, timedelta -from typing import TYPE_CHECKING, Any +from typing import TYPE_CHECKING, Any, cast import httpx import pytest from pipelex_sdk.artifact_models import ( ArtifactItemError, + ArtifactLocation, ArtifactScope, BulkResolvedStorageUrls, DownloadArtifactsOptions, + DownloadedArtifact, FetchArtifactOptions, ResolvedArtifact, ) @@ -31,6 +34,7 @@ download_artifacts, fetch_artifact, is_storage_reference, + locate_artifacts, resolve_artifacts, ) from pipelex_sdk.errors import ( @@ -59,6 +63,9 @@ _RUN_ID = "run-01J" _URI_PNG = "pipelex-storage://org_1/runs/01J/outputs/illustration.png" _URI_PDF = "pipelex-storage://org_1/runs/01J/outputs/report.pdf" +_URI_INPUT = "pipelex-storage://org_1/uploads/brief.pdf" +#: A key whose last segment carries no extension, so the content type decides it. +_URI_BARE = "pipelex-storage://org_1/runs/01J/outputs/report" _STORE = "https://store.example.com" _PDF_BYTES = b"%PDF-1.4 tiny" _PNG_BYTES = b"\x89PNG tiny" @@ -187,13 +194,71 @@ def _results(main_stuff: Any, *, working_memory: Any = _ABSENT_KEY, run_id: str return RunResults.model_validate(body) +def _at(path: str, uri: str = _URI_PNG) -> ArtifactLocation: + """A location at one path, for the naming rule's tests.""" + return ArtifactLocation(uri=uri, found_at=[path]) + + def _content(uri: str) -> dict[str, Any]: """A produced file as the runtime serializes it: the durable reference beside an expiring link.""" return {"url": uri, "public_url": f"{_STORE}/signed-by-the-runtime?sig=stale"} class TestArtifacts: - # ── collect_artifacts ──────────────────────────────────────────── + # ── locate_artifacts / collect_artifacts ───────────────────────── + + def test_locates_every_reference_once_in_discovery_order_with_every_path_it_sits_at(self) -> None: + walked = { + "image": _content(_URI_PNG), + "pages": [{"url": _URI_PNG}, {"deeper": {"url": _URI_PDF}}, _URI_INPUT], + "text": "not a reference", + "count": 3, + "nothing": None, + } + assert locate_artifacts(walked) == [ + ArtifactLocation(uri=_URI_PNG, found_at=["$.image.url", "$.pages[0].url"]), + ArtifactLocation(uri=_URI_PDF, found_at=["$.pages[1].deeper.url"]), + ArtifactLocation(uri=_URI_INPUT, found_at=["$.pages[2]"]), + ] + assert collect_artifacts(walked) == [location.uri for location in locate_artifacts(walked)] + + @pytest.mark.parametrize( + ("walked", "expected"), + [ + (_URI_PNG, "$"), + ({"url": _URI_PNG}, "$.url"), + ({"items": [{"url": _URI_PNG}]}, "$.items[0].url"), + ([[_URI_PNG]], "$[0][0]"), + ], + ) + def test_roots_a_path_at_the_walked_value_itself(self, walked: Any, expected: str) -> None: + assert locate_artifacts(walked) == [ArtifactLocation(uri=_URI_PNG, found_at=[expected])] + + def test_writes_an_identifier_key_dotted_and_any_other_as_a_json_string_in_brackets(self) -> None: + keys = ["_ok", "a key", "2nd", 'say "hi"\\', "", "café 📷", "line\nbreak", "\u2028", "\ud800"] + walked = {key: f"pipelex-storage://org_1/{index_key}.png" for index_key, key in enumerate(keys)} + + # `JSON.stringify`'s escaping, so a path reads the same from the JS twin: the quote, the + # backslash, the control characters and a lone surrogate are escaped, and nothing else is. + assert [location.found_at[0] for location in locate_artifacts(walked)] == [ + "$._ok", + '$["a key"]', + '$["2nd"]', + '$["say \\"hi\\"\\\\"]', + '$[""]', + '$["café 📷"]', + '$["line\\nbreak"]', + '$["\u2028"]', + '$["\\ud800"]', + ] + + def test_walks_a_pydantic_model_by_its_dumped_keys(self) -> None: + assert locate_artifacts(_results({"picture": _content(_URI_PNG)})) == [ + ArtifactLocation(uri=_URI_PNG, found_at=["$.main_stuff.picture.url"]), + ] + + def test_reads_a_key_that_is_not_a_string_as_the_string_it_prints_as(self) -> None: + assert locate_artifacts({7: _URI_PNG}) == [ArtifactLocation(uri=_URI_PNG, found_at=['$["7"]'])] def test_collects_every_reference_once_in_discovery_order(self) -> None: walked = { @@ -214,12 +279,13 @@ def test_collects_every_reference_once_in_discovery_order(self) -> None: ) def test_counts_a_string_only_when_it_is_a_reference(self, value: str, expected: list[str]) -> None: assert collect_artifacts(value) == expected + assert [location.uri for location in locate_artifacts(value)] == expected assert is_storage_reference(value) is (expected != []) def test_ignores_values_that_carry_no_reference(self) -> None: - assert collect_artifacts(None) == [] - assert collect_artifacts(42) == [] - assert collect_artifacts({"a": [1, 2.5, True, None]}) == [] + for value in (None, 42, {"a": [1, 2.5, True, None]}, {"note": f"Saved as {_URI_PDF}."}): + assert collect_artifacts(value) == [] + assert locate_artifacts(value) == [] def test_walks_a_pydantic_model_as_well_as_a_parsed_body(self) -> None: results = _results({"picture": _content(_URI_PNG)}) @@ -227,52 +293,145 @@ def test_walks_a_pydantic_model_as_well_as_a_parsed_body(self) -> None: # ── artifact_filename ──────────────────────────────────────────── + @pytest.mark.parametrize( + ("path", "scope", "expected"), + [ + # The walked value itself, or its `url`, takes the scope's name. + ("$", ArtifactScope.MAIN_STUFF, "main_stuff.png"), + ("$.url", ArtifactScope.MAIN_STUFF, "main_stuff.png"), + ("$.url", ArtifactScope.WORKING_MEMORY, "working_memory.png"), + # A list member is named after the envelope and its index. + ("$.items[0].url", ArtifactScope.MAIN_STUFF, "items-0.png"), + ("$[2].url", ArtifactScope.MAIN_STUFF, "2.png"), + # A nested field is named after its whole path. + ("$.rooms[3].staged_photo.url", ArtifactScope.MAIN_STUFF, "rooms-3-staged_photo.png"), + # Only a final `url` key is dropped: another key, or a `url` key that is not final, is kept. + ("$.photo.src", ArtifactScope.MAIN_STUFF, "photo-src.png"), + ("$.photo.URL", ArtifactScope.MAIN_STUFF, "photo-URL.png"), + ("$.url.original", ArtifactScope.MAIN_STUFF, "url-original.png"), + ("$.links.url[0]", ArtifactScope.MAIN_STUFF, "links-url-0.png"), + ('$["url"]', ArtifactScope.MAIN_STUFF, "main_stuff.png"), + # A key that is not an identifier is reduced to `[A-Za-z0-9_]`, dashes and dots included. + ('$["a key"].url', ArtifactScope.MAIN_STUFF, "a_key.png"), + ('$["staged-photo.v2"].url', ArtifactScope.MAIN_STUFF, "staged_photo_v2.png"), + ('$["café 📷"].url', ArtifactScope.MAIN_STUFF, "caf___.png"), + ('$["2nd"]', ArtifactScope.MAIN_STUFF, "2nd.png"), + ('$[""].url', ArtifactScope.MAIN_STUFF, "main_stuff.png"), + ], + ) + def test_names_a_file_after_the_field_it_fills(self, path: str, scope: ArtifactScope, expected: str) -> None: + assert artifact_filename(_at(path), "image/png", scope) == expected + + @pytest.mark.parametrize( + "path", + ['$["../../etc/passwd"]', '$[".."].url', '$["."]', '$[".env"]', '$["a\\u0000b\\nc"]', '$["C:\\\\Windows"]'], + ) + def test_cannot_name_anything_outside_the_directory_nor_a_hidden_file(self, path: str) -> None: + name = artifact_filename(_at(path, "pipelex-storage://x/.."), None, ArtifactScope.MAIN_STUFF) + assert re.fullmatch(r"[A-Za-z0-9_][A-Za-z0-9_-]*(\.[A-Za-z0-9]+)?", name) + + def test_reduces_a_traversal_key_to_one_plain_name(self) -> None: + assert artifact_filename(_at('$["../../etc/passwd"]', _URI_BARE), None, ArtifactScope.MAIN_STUFF) == "______etc_passwd" + + def test_keeps_the_tail_of_a_long_name_dropping_whole_leading_segments_first(self) -> None: + segments = [f"segment_{index_segment:02d}" for index_segment in range(20)] + name = artifact_filename(_at(f"$.{'.'.join(segments)}.url"), None, ArtifactScope.MAIN_STUFF) + + assert len(name) <= 128 + assert name == f"{'-'.join(segments[9:])}.png" + + def test_cuts_a_single_segment_still_too_long_keeping_the_extension(self) -> None: + long_key = "a" * 300 + assert artifact_filename(_at(f"$.short.{long_key}.url"), None, ArtifactScope.MAIN_STUFF) == f"{'a' * 124}.png" + assert artifact_filename(_at(f"$.{long_key}", _URI_BARE), None, ArtifactScope.MAIN_STUFF) == "a" * 128 + @pytest.mark.parametrize( ("uri", "content_type", "expected"), [ - ("pipelex-storage://org_1/runs/01J/outputs/report.pdf", "application/pdf", "report.pdf"), - # Path separators are the split point, so no traversal and no absolute path survives. - ("pipelex-storage://org_1/../../etc/passwd", None, "passwd"), - # An encoded traversal is one segment, decoded after the split: the separators it hid - # become underscores and the leading dots go, so it still names a file in the directory. - ("pipelex-storage://org_1/x/..%2F..%2Fetc%2Fpasswd", None, "etc_passwd"), - ("pipelex-storage://org_1/x/a\\b\\c.txt", None, "c.txt"), - # A leading dot is stripped, so no hidden file; odd characters are neutralized. - ("pipelex-storage://org_1/.bashrc", None, "bashrc"), - ("pipelex-storage://org_1/my file (1).png", "image/png", "my_file__1_.png"), - # Percent-decoded, with the query and fragment dropped. - ("pipelex-storage://org_1/a%20b.pdf?sig=x#frag", None, "a_b.pdf"), - # The extension comes from the content type only when the key carries none. - ("pipelex-storage://org_1/outputs/report", "application/pdf", "report.pdf"), - ("pipelex-storage://org_1/outputs/report.bin", "application/pdf", "report.bin"), - ("pipelex-storage://org_1/outputs/report", "image/png; charset=binary", "report.png"), - ("pipelex-storage://org_1/outputs/report", "application/x-unknown", "report"), - ("pipelex-storage://org_1/outputs/report", None, "report"), - # Nothing usable in the key: the numbered fallback, one-based. - ("pipelex-storage://", None, "artifact-3"), - ("pipelex-storage://org_1/___", None, "artifact-3"), + # The storage key's extension first, then the content type's, else none. + (_URI_PNG, "application/pdf", "cover.png"), + (_URI_BARE, "application/pdf", "cover.pdf"), + (_URI_BARE, "image/png; charset=binary", "cover.png"), + (_URI_BARE, "application/x-unknown", "cover"), + (_URI_BARE, None, "cover"), + # The key's last segment, percent-decoded, its query and fragment dropped. + ("pipelex-storage://x/hello%20world.pdf?token=1#frag", None, "cover.pdf"), + ("pipelex-storage://x/photo%2Epng", None, "cover.png"), + ("pipelex-storage://a/..\\..\\secret.txt", None, "cover.txt"), + # Reduced to `[A-Za-z0-9]`, never a leading dot, never empty, never long. + ("pipelex-storage://x/photo.P-N_G", None, "cover.PNG"), + ("pipelex-storage://x/.env", None, "cover"), + ("pipelex-storage://x/report.", None, "cover"), + (f"pipelex-storage://x/stem.{'z' * 300}", None, "cover"), + (f"pipelex-storage://x/stem.{'z' * 300}", "text/csv", "cover.csv"), + # Where `decodeURIComponent` would throw — a stray `%`, bytes that are not UTF-8 — the + # segment is read as typed, so both SDKs find the same extension. + ("pipelex-storage://x/bad%zz.pdf", None, "cover.pdf"), + ("pipelex-storage://x/a%2Eb%zz", None, "cover"), + ("pipelex-storage://x/photo.p%FFng", None, "cover.pFFng"), + ], + ) + def test_takes_the_extension_from_the_storage_key_then_the_content_type(self, uri: str, content_type: str | None, expected: str) -> None: + assert artifact_filename(_at("$.cover.url", uri), content_type, ArtifactScope.MAIN_STUFF) == expected + + @pytest.mark.parametrize( + ("path", "uri", "expected"), + [ + ("$.aux.url", _URI_PNG, "aux_.png"), + ("$.NUL", _URI_PNG, "NUL_.png"), + ("$.Com1.url", _URI_PNG, "Com1_.png"), + ("$.lpt9", _URI_BARE, "lpt9_"), + ('$[""].con.url', _URI_PNG, "con_.png"), + # The cap drops every leading segment and leaves the device name alone. + (f"$.{'x' * 130}.prn.url", _URI_PNG, "prn_.png"), + # Only the whole stem is a device name: a join or a longer word is not one. + ("$.a.nul.url", _URI_PNG, "a-nul.png"), + ("$.auxiliary.url", _URI_PNG, "auxiliary.png"), + ("$.com10.url", _URI_PNG, "com10.png"), ], ) - def test_derives_a_filename_that_can_only_name_a_file_in_the_directory(self, uri: str, content_type: str | None, expected: str) -> None: - assert artifact_filename(uri, content_type, 2) == expected + def test_suffixes_a_stem_windows_reserves_for_a_device(self, path: str, uri: str, expected: str) -> None: + assert artifact_filename(_at(path, uri), None, ArtifactScope.MAIN_STUFF) == expected - def test_caps_the_filename_length_keeping_the_extension(self) -> None: - name = artifact_filename(f"pipelex-storage://org_1/{'a' * 400}.pdf", None, 0) - assert len(name) == 128 - assert name.endswith(".pdf") + @pytest.mark.parametrize( + "location", + [ + ArtifactLocation(uri=_URI_PNG, found_at=[]), + _at("rooms.url"), + _at("$.rooms["), + _at("$[-1]"), + _at('$["unterminated]'), + _at('$["\\x"]'), + _at("$.a b"), + # The old signature's bare uri. + cast("ArtifactLocation", _URI_PNG), + ], + ) + def test_refuses_a_location_whose_first_path_is_not_in_the_walks_notation(self, location: ArtifactLocation) -> None: + with pytest.raises(ArtifactOperationError): + artifact_filename(location, None, ArtifactScope.MAIN_STUFF) - def test_drops_an_extension_that_alone_exceeds_the_cap(self) -> None: - name = artifact_filename(f"pipelex-storage://org_1/name.{'z' * 400}", None, 0) - assert len(name) == 128 - assert name.startswith("name.") + @pytest.mark.parametrize("scope", ["other", 0]) + def test_refuses_an_unknown_scope(self, scope: object) -> None: + with pytest.raises(ArtifactOperationError, match='"scope" must be'): + artifact_filename(_at("$.url"), None, cast("ArtifactScope", scope)) - # ── resolve_artifacts ──────────────────────────────────────────── + def test_names_every_location_the_walk_writes_through_the_round_trip_of_its_notation(self) -> None: + walked = { + 'say "hi"\\': {"url": "pipelex-storage://org_1/a.png"}, + "line\nbreak": {"url": "pipelex-storage://org_1/b.png"}, + "\u2028": {"url": "pipelex-storage://org_1/c.png"}, + "[0]": {"url": "pipelex-storage://org_1/d.png"}, + "\ud800": {"url": "pipelex-storage://org_1/e.png"}, + } + names = [artifact_filename(location, None, ArtifactScope.MAIN_STUFF) for location in locate_artifacts(walked)] + assert names == ["say__hi__.png", "line_break.png", "_.png", "_0_.png", "_.png"] - def test_caps_the_length_after_the_guessed_extension(self) -> None: - name = artifact_filename(f"pipelex-storage://org_1/{'a' * 200}", "image/jpeg", 0) + def test_takes_a_verdict_item_as_its_location(self) -> None: + item = DownloadedArtifact(uri=_URI_BARE, found_at=["$.report.url"], content_type="application/pdf", error=None) + assert artifact_filename(item, item.content_type, ArtifactScope.MAIN_STUFF) == "report.pdf" - assert len(name) <= 128 - assert name.endswith(".jpg") + # ── resolve_artifacts ──────────────────────────────────────────── def test_resolves_a_list_within_the_bound_in_one_call_and_keeps_request_order(self) -> None: client = _FakeClient(resolve=_resolver(_resolved(_URI_PNG), _refused(_URI_PDF))) @@ -587,9 +746,10 @@ def _handler(request: httpx.Request) -> httpx.Response: assert [artifact.uri for artifact in verdict.artifacts] == [_URI_PNG, _URI_PDF] assert [artifact.size for artifact in verdict.artifacts] == [len(_PNG_BYTES), len(_PDF_BYTES)] assert [artifact.content_type for artifact in verdict.artifacts] == ["image/png", "application/pdf"] - assert verdict.saved_paths == [str(target / "illustration.png"), str(target / "report.pdf")] - assert (target / "illustration.png").read_bytes() == _PNG_BYTES - assert (target / "report.pdf").read_bytes() == _PDF_BYTES + assert [artifact.found_at for artifact in verdict.artifacts] == [["$.items[0].url"], ["$.items[1].url"]] + assert verdict.saved_paths == [str(target / "items-0.png"), str(target / "items-1.pdf")] + assert (target / "items-0.png").read_bytes() == _PNG_BYTES + assert (target / "items-1.pdf").read_bytes() == _PDF_BYTES def test_takes_results_in_hand_without_re_reading_and_creates_the_directory(self, mocker: MockerFixture, tmp_path: Path) -> None: client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) @@ -614,10 +774,11 @@ def test_walks_working_memory_when_asked_echoed_inputs_included(self, mocker: Mo ) client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF), _resolved(_URI_PNG, content_type="image/png"))) _patch_storage(mocker, _serving(_PDF_BYTES)) + target = tmp_path / "out" verdict = asyncio.run( download_artifacts( client, - dir_path=tmp_path / "out", + dir_path=target, results=results, options=DownloadArtifactsOptions(scope=ArtifactScope.WORKING_MEMORY), ) @@ -625,27 +786,91 @@ def test_walks_working_memory_when_asked_echoed_inputs_included(self, mocker: Mo assert verdict.scope == ArtifactScope.WORKING_MEMORY assert [artifact.uri for artifact in verdict.artifacts] == [_URI_PDF, _URI_PNG] + assert [artifact.found_at for artifact in verdict.artifacts] == [["$.root.doc.content.url"], ["$.root.picture.content.url"]] assert verdict.all_saved is True + assert verdict.saved_paths == [str(target / "root-doc-content.pdf"), str(target / "root-picture-content.png")] def test_never_overwrites_a_name_already_on_disk(self, mocker: MockerFixture, tmp_path: Path) -> None: - other = "pipelex-storage://org_1/runs/01J/second/report.pdf" - client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF), _resolved(other))) + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) _patch_storage(mocker, _serving(_PDF_BYTES)) target = tmp_path / "out" target.mkdir() - (target / "report.pdf").write_bytes(b"do not touch me") + (target / "main_stuff.pdf").write_bytes(b"do not touch me") + (target / "main_stuff-1.pdf").write_bytes(b"nor me") + + verdict = asyncio.run(download_artifacts(client, dir_path=target, results=_results(_content(_URI_PDF)))) + + assert verdict.saved_paths == [str(target / "main_stuff-2.pdf")] + assert (target / "main_stuff.pdf").read_bytes() == b"do not touch me" + assert (target / "main_stuff-1.pdf").read_bytes() == b"nor me" + + def test_names_each_file_after_the_field_it_fills_and_reports_every_path_on_both_arms(self, mocker: MockerFixture, tmp_path: Path) -> None: + client = _FakeClient( + resolve=_resolver( + _resolved(_URI_INPUT), + _resolved(_URI_PNG, content_type="image/png"), + _refused(_URI_PDF, detail="another organization"), + ) + ) + _patch_storage(mocker, _serving(_PNG_BYTES)) + target = tmp_path / "out" + main_stuff = { + "rooms": [ + {"original_photo": _content(_URI_INPUT), "staged_photo": _content(_URI_PNG)}, + {"original_photo": _content(_URI_PDF)}, + ], + "cover": _content(_URI_PNG), + } + + verdict = asyncio.run(download_artifacts(client, dir_path=target, results=_results(main_stuff))) + + assert [(artifact.uri, artifact.found_at, artifact.path) for artifact in verdict.artifacts] == [ + (_URI_INPUT, ["$.rooms[0].original_photo.url"], str(target / "rooms-0-original_photo.pdf")), + (_URI_PNG, ["$.rooms[0].staged_photo.url", "$.cover.url"], str(target / "rooms-0-staged_photo.png")), + (_URI_PDF, ["$.rooms[1].original_photo.url"], None), + ] + assert sorted(entry.name for entry in target.iterdir()) == ["rooms-0-original_photo.pdf", "rooms-0-staged_photo.png"] + + def test_tells_apart_two_paths_that_reduce_to_one_name_with_the_suffix_rule(self, mocker: MockerFixture, tmp_path: Path) -> None: + second = "pipelex-storage://org_1/runs/01J/outputs/second.png" + client = _FakeClient(resolve=_resolver(_resolved(_URI_PNG), _resolved(second))) + _patch_storage(mocker, _serving(_PNG_BYTES)) + target = tmp_path / "out" verdict = asyncio.run( download_artifacts( client, dir_path=target, - results=_results({"items": [_content(_URI_PDF), _content(other)]}), + results=_results({"staged photo": _content(_URI_PNG), "staged-photo": _content(second)}), options=DownloadArtifactsOptions(concurrency=1), ) ) - assert verdict.saved_paths == [str(target / "report-1.pdf"), str(target / "report-2.pdf")] - assert (target / "report.pdf").read_bytes() == b"do not touch me" + assert verdict.saved_paths == [str(target / "staged_photo.png"), str(target / "staged_photo-1.png")] + + def test_saves_every_file_under_the_name_artifact_filename_gives_its_location(self, mocker: MockerFixture, tmp_path: Path) -> None: + uris = ["pipelex-storage://org_1/a.png", "pipelex-storage://org_1/b", "pipelex-storage://org_1/c.pdf"] + client = _FakeClient(resolve=_resolver(*[_resolved(uri, content_type=None) for uri in uris])) + _patch_storage(mocker, _serving(_PNG_BYTES)) + target = tmp_path / "out" + main_stuff = {'say "hi"\\': {"url": uris[0]}, "line\nbreak": [{"url": uris[1]}], "url": uris[2]} + + verdict = asyncio.run(download_artifacts(client, dir_path=target, results=_results(main_stuff))) + + predicted = [artifact_filename(location, None, ArtifactScope.MAIN_STUFF) for location in locate_artifacts(main_stuff)] + assert predicted == ["say__hi__.png", "line_break-0", "main_stuff.pdf"] + assert verdict.saved_paths == [str(target / name) for name in predicted] + # A verdict item is a location too, and names its own file. + assert [artifact_filename(artifact, artifact.content_type, verdict.scope) for artifact in verdict.artifacts] == predicted + + def test_refuses_an_unknown_scope_before_reading_anything(self, tmp_path: Path) -> None: + client = _FakeClient() + # Pydantic refuses the unknown scope at construction; only an unvalidated model reaches here. + options = DownloadArtifactsOptions.model_construct(scope=cast("ArtifactScope", "everything")) + + with pytest.raises(ArtifactOperationError, match='"scope" must be'): + asyncio.run(download_artifacts(client, dir_path=tmp_path, run_id=_RUN_ID, options=options)) + assert client.run_result_calls == [] def test_keeps_a_per_reference_refusal_as_that_items_error_beside_the_saved_ones(self, mocker: MockerFixture, tmp_path: Path) -> None: client = _FakeClient(resolve=_resolver(_refused(_URI_PNG), _resolved(_URI_PDF))) @@ -659,8 +884,9 @@ def test_keeps_a_per_reference_refusal_as_that_items_error_beside_the_saved_ones assert first.error is not None assert first.error.code == "forbidden" assert first.path is None + assert first.found_at == ["$.items[0].url"] assert verdict.artifacts[1].error is None - assert len(verdict.saved_paths) == 1 + assert verdict.saved_paths == [str(tmp_path / "out" / "items-1.pdf")] def test_re_resolves_a_link_that_has_expired_by_the_time_its_task_reaches_it(self, mocker: MockerFixture, tmp_path: Path) -> None: answers = [ @@ -706,6 +932,7 @@ def _refuse(_: list[str]) -> BulkResolvedStorageUrls: assert caught.value.status == 403 assert verdict.saved_paths == [] assert [artifact.error.code for artifact in verdict.artifacts if artifact.error is not None] == ["aborted", "aborted"] + assert [artifact.found_at for artifact in verdict.artifacts] == [["$.items[0].url"], ["$.items[1].url"]] def test_raises_a_credential_failure_part_way_through_carrying_the_verdict_so_far(self, mocker: MockerFixture, tmp_path: Path) -> None: def _resolve(uris: list[str]) -> BulkResolvedStorageUrls: @@ -905,7 +1132,7 @@ def test_unlinks_the_partial_file_when_the_download_is_cancelled(self, mocker: M async def _cancel_midway() -> bool: task = asyncio.create_task(download_artifacts(client, dir_path=target, results=_results({"doc": _content(_URI_PDF)}))) await asyncio.sleep(0.12) - assert (target / "report.pdf").exists() + assert (target / "doc.pdf").exists() task.cancel() try: await task @@ -934,7 +1161,7 @@ def test_the_client_posts_the_bulk_route_and_serves_the_whole_stack( verdict = asyncio.run(api_client.download_artifacts(dir_path=target, results=_results({"doc": _content(_URI_PDF)}))) assert verdict.all_saved is True - assert verdict.saved_paths == [str(target / "report.pdf")] + assert verdict.saved_paths == [str(target / "doc.pdf")] method, url = spy.call_args.args[0], spy.call_args.args[1] assert method == "POST" assert url.endswith("/v1/resolve-storage-url/bulk") From 80e7e398dbf4a8531341a0f670446a8a100b02ad Mon Sep 17 00:00:00 2001 From: Louis Choquel Date: Thu, 24 Sep 2026 13:12:12 +0200 Subject: [PATCH 2/3] =?UTF-8?q?chore/Bump-mthds-0.15.0=20=C2=B7=20L-260922?= =?UTF-8?q?-5f7eeb=20=E2=80=94=20the=20exact=20mthds=20pin=20outlived=20it?= =?UTF-8?q?s=20reason=20(#44)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Moves the exact `mthds` pin from 0.14.0 to 0.15.0. `pipelex` has pinned `mthds==0.15.0` exactly since 0.60.0, so the two packages could not be installed together; this restores that. Nothing in the client changed with the move, and the local method-files codec stays beside `mthds.protocol.method_files` until `PipelineRequestError`'s base class is settled upstream, which `docs/architecture.md` now records. Closes L-260922-5f7eeb 🤖 Generated with [Claude Code](https://claude.com/claude-code) --- ## Summary by cubic Moves the exact `mthds` pin from 0.14.0 to 0.15.0 so `pipelex` and `pipelex-sdk` can be installed together again: `pipelex` has pinned `mthds==0.15.0` exactly since 0.60.0, and two exact pins on different versions left the pair unresolvable. Nothing in the client changed with the move. **Breaking changes in 0.15.0** — none of them affect this SDK: - The one breaking cut (`ConceptAbstract` reduced to a name, `StuffAbstract` dumping its concept as the ref string) sits in protocol models this SDK does not build on. - The base client's run-source guard now rejects a protocol arg smuggled through `extra` before building the body, which this client's own guard already did. - The canonical method-files pair ships as `mthds.protocol.method_files`; this package keeps its own `parse_method_files` / `serialize_method_files` until `PipelineRequestError`'s base class is settled upstream, and `docs/architecture.md` now records the reason. Closes L-260922-5f7eeb. Written for commit 261b2961c448ac1aaab3603ad3150efa27059c88. Summary will update on new commits. Review in cubic Co-authored-by: Claude Opus 5.5 --- CHANGELOG.md | 1 + docs/architecture.md | 2 +- pyproject.toml | 2 +- uv.lock | 8 ++++---- 4 files changed, 7 insertions(+), 6 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 8a871eb..8765652 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,6 +10,7 @@ - **`download_artifacts` names each file after the field it fills, and `artifact_filename` takes a location (Breaking)**: a saved file is named after the first path at which its reference sits — `$.rooms[3].staged_photo.url` is saved as `rooms-3-staged_photo.png`, and an output that is one image as `main_stuff.png` — instead of after the last segment of its storage key, which now supplies only the extension, so the same run saves under the same names from either SDK. `artifact_filename(location, content_type, scope)` replaces `artifact_filename(uri, content_type, index)` and raises `ArtifactOperationError` for anything but an `ArtifactLocation` whose first path is in the walk's notation; a field whose name Windows reserves for a device (`aux`, `nul`, `com1` and the like) is saved with a trailing `_` (`aux_.png`); the full rule is on `docs/artifact-download.md`. - **`DownloadedArtifact` carries a required `found_at` (Breaking)**: every item of a `download_artifacts` verdict carries its reference's `found_at`, on the saved arm and the error arm alike, so a file that was not saved still says which field it would have filled. `DownloadedArtifact` is now an `ArtifactLocation`, so `artifact_filename` takes a verdict item as it is, and code that builds `DownloadedArtifact` values — a test fake standing in for `download_artifacts` — must now supply the field. +- **Requires `mthds` 0.15.0 (Breaking).** The pin moves from 0.14.0, and it moves so that `pipelex` and `pipelex-sdk` can be installed together again: `pipelex` pins `mthds` exactly too and has required 0.15.0 since its 0.60.0, so two exact pins on different versions left the pair unresolvable. Nothing in this client changed with it. The release's one breaking cut — `ConceptAbstract` reduced to a name and `StuffAbstract` dumping its concept as the ref string — sits in protocol models this SDK does not build on, and the base client's run-source guard now refuses a protocol arg smuggled through `extra` before building the body rather than while building it, which this client's own guard already did ahead of both. The release also ships the canonical method-files pair as `mthds.protocol.method_files`; this package keeps its own `parse_method_files` / `serialize_method_files` for now, for the reason `docs/architecture.md` gives. ## [v0.11.0] - 2026-09-23 diff --git a/docs/architecture.md b/docs/architecture.md index 347ad3d..d8ae152 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -292,7 +292,7 @@ Each stays deferred rather than silently missing. Everything else — the protoc **The run-results surface** — `RunResults` tracks `@pipelex/sdk` 0.20.0 field for field, and is at parity with it. Every field is declared, with the two-path semantics [`run-results.md`](./run-results.md) states: `graph_spec` (lifted on the blocking path instead of written `None`), `graph_assembly_error`, the three I/O artifacts typed from `mthds.protocol` with `pipe_io_artifacts_error` beside them, the usage pair, `working_memory` (the hosted artifact relayed as its own key, the blocking one lifted off `pipe_output.working_memory`), and `pipe_output`; and the usage pair's fold, `summarize_usage`, with the summary types it returns. What stays different is idiomatic and deliberate: a key the hosted body did not carry is absent from `model_fields_set` where the JS reads `undefined`, and the page says how to read that. -**Models** — field-for-field across the run-lifecycle types and the product wire models. Deliberate idiomatic ports (not gaps): milliseconds → seconds (`interval_seconds` / `timeout_seconds` / `elapsed_seconds`); the JS `AbortSignal` → Python `asyncio` cancellation (no `signal` field, and no `aborted` flag on the download verdict — a cancelled `download_artifacts` raises `CancelledError` with its partial files already unlinked); the JS artifact request object → `DownloadArtifactsOptions` beside an explicit `dir_path` (`dir` being a Python builtin) and the raw bulk call named `resolve_storage_urls_bulk` after its route rather than `resolveStorageUrls`, so it cannot be misread as the single-reference `resolve_storage_url`; the returned `Response` of `fetchArtifact` → an async context manager, because an httpx stream is only live inside its own block; JS inline string-unions promoted to `StrEnum`s (`OrgRole`, `PipeStatus`, the onboarding fields) with identical wire values; response models are `extra="allow"` for forward-compat. The Pipelex validation narrowing is **owned here** (`pipelex_sdk.validation_models`), narrowing `mthds`'s neutral verdict bases (the resolved follow-up #9); the brand-neutral `Dict*` wire concretes (`DictRunResultExecute`, and the `DictPipeOutputAbstract` / `DictWorkingMemoryAbstract` pair that types `RunResults.pipe_output` and `RunResults.working_memory`) are reused from `mthds` by inheritance or by import — they are a shared wire contract the `pipelex` runtime also builds on — rather than duplicated as `pipelex-sdk-js` does. One addition runs ahead of the JS SDK: `write_codegen_tree` has no `@pipelex/sdk` counterpart, because the JS writer lives inside `pipelex-starter-js`'s own harness, mixed in with project policy; the byte-fidelity contract is small and load-bearing enough to belong to the SDK, where every consumer shares one correct implementation. The drift check goes the other way and takes a different shape on purpose: `@pipelex/sdk`'s `runCodegenCheck` is **pure** — the caller walks its own tree and hands in the text — because that module must stay free of Node builtins for a browser bundle, and its doc consequently loads the caller with obligations (walk the whole tree, do not reformat, decode strictly) whose every breach yields a wrong verdict rather than an error. Python has no such constraint, this package already does filesystem work, and `pipelex`'s own surface takes a root — so `run_codegen_check(root=…)` takes one too. Every caller obligation becomes the library's, the verdict is provable against `pipelex` by calling both with the same directory, and it composes with `write_codegen_tree(report, output_dir=…)` as the same path in and out. Two divergences worth naming: the page envelopes keep the wire's snake_case `next_cursor`, where the JS mirror renamed it `nextCursor` for its own consumers; and the method-files catalog converter (`parse_method_files` / `serialize_method_files`) lives in this package, where the JS pair lives in `mthds-js` because `pipelex-mcp` consumes the same format and wanted one owner. There is no second Python consumer, and the catalog serialization is a Pipelex product concern rather than an MTHDS protocol one, so this SDK is a proper home for it. `mthds` has since grown the canonical pair as `mthds.protocol.method_files`, on its `dev` branch and **not in the `mthds==0.14.0` this package pins exactly** — so adoption waits on a release before anything else. It is a step of its own rather than a rename in any case: that parser raises `PipelineRequestError`, which does not subclass `ValueError`, and pydantic converts only `ValueError` out of a validator — so re-pointing `MethodData`'s validator at it would stop every caller's `except ValidationError` from catching a malformed stored source. The release and the exception's base class are settled first; until then the two implementations are kept in step, and they disagree on blankness (this one is Python's `str.strip`, the canonical one is ECMAScript's), on the serialized bytes (default separators and ASCII escapes here, `JSON.stringify`'s there), and on the JSON constants and integer-literal cap the canonical one closes with `parse_constant=` and `parse_int=float`. +**Models** — field-for-field across the run-lifecycle types and the product wire models. Deliberate idiomatic ports (not gaps): milliseconds → seconds (`interval_seconds` / `timeout_seconds` / `elapsed_seconds`); the JS `AbortSignal` → Python `asyncio` cancellation (no `signal` field, and no `aborted` flag on the download verdict — a cancelled `download_artifacts` raises `CancelledError` with its partial files already unlinked); the JS artifact request object → `DownloadArtifactsOptions` beside an explicit `dir_path` (`dir` being a Python builtin) and the raw bulk call named `resolve_storage_urls_bulk` after its route rather than `resolveStorageUrls`, so it cannot be misread as the single-reference `resolve_storage_url`; the returned `Response` of `fetchArtifact` → an async context manager, because an httpx stream is only live inside its own block; JS inline string-unions promoted to `StrEnum`s (`OrgRole`, `PipeStatus`, the onboarding fields) with identical wire values; response models are `extra="allow"` for forward-compat. The Pipelex validation narrowing is **owned here** (`pipelex_sdk.validation_models`), narrowing `mthds`'s neutral verdict bases (the resolved follow-up #9); the brand-neutral `Dict*` wire concretes (`DictRunResultExecute`, and the `DictPipeOutputAbstract` / `DictWorkingMemoryAbstract` pair that types `RunResults.pipe_output` and `RunResults.working_memory`) are reused from `mthds` by inheritance or by import — they are a shared wire contract the `pipelex` runtime also builds on — rather than duplicated as `pipelex-sdk-js` does. One addition runs ahead of the JS SDK: `write_codegen_tree` has no `@pipelex/sdk` counterpart, because the JS writer lives inside `pipelex-starter-js`'s own harness, mixed in with project policy; the byte-fidelity contract is small and load-bearing enough to belong to the SDK, where every consumer shares one correct implementation. The drift check goes the other way and takes a different shape on purpose: `@pipelex/sdk`'s `runCodegenCheck` is **pure** — the caller walks its own tree and hands in the text — because that module must stay free of Node builtins for a browser bundle, and its doc consequently loads the caller with obligations (walk the whole tree, do not reformat, decode strictly) whose every breach yields a wrong verdict rather than an error. Python has no such constraint, this package already does filesystem work, and `pipelex`'s own surface takes a root — so `run_codegen_check(root=…)` takes one too. Every caller obligation becomes the library's, the verdict is provable against `pipelex` by calling both with the same directory, and it composes with `write_codegen_tree(report, output_dir=…)` as the same path in and out. Two divergences worth naming: the page envelopes keep the wire's snake_case `next_cursor`, where the JS mirror renamed it `nextCursor` for its own consumers; and the method-files catalog converter (`parse_method_files` / `serialize_method_files`) lives in this package, where the JS pair lives in `mthds-js` because `pipelex-mcp` consumes the same format and wanted one owner. There is no second Python consumer, and the catalog serialization is a Pipelex product concern rather than an MTHDS protocol one, so this SDK is a proper home for it. `mthds` has since grown the canonical pair as `mthds.protocol.method_files`, shipped in the `mthds==0.15.0` this package pins exactly, and the local pair is kept beside it on purpose rather than left over. Adoption is a step of its own rather than a rename: that parser raises `PipelineRequestError`, which does not subclass `ValueError`, and pydantic converts only `ValueError` out of a validator — so re-pointing `MethodData`'s validator at it would stop every caller's `except ValidationError` from catching a malformed stored source. The exception's base class is `mthds`'s to settle, and it is settled first; until then the two implementations are kept in step, and they disagree on blankness (this one is Python's `str.strip`, the canonical one is ECMAScript's), on the serialized bytes (default separators and ASCII escapes here, `JSON.stringify`'s there), and on the JSON constants and integer-literal cap the canonical one closes with `parse_constant=` and `parse_int=float`. **Errors** — `ApiResponseError`, `ApiUnreachableError`, `PipelineExecuteTimeoutError`, `RunFailedError`, `RunTimeoutError`, `RunLifecycleUnavailableError`, `PagingNotTerminatingError` are owned here; `RunStillRunningError` is re-exported from `mthds`. `ClientAuthenticationError` is **not** ported: it is a dormant export in the JS barrel (defined and exported but never raised by the client), and in Python it already lives in `mthds.runners.api.exceptions` — importable directly if ever needed, with no barrel here to re-export it through. diff --git a/pyproject.toml b/pyproject.toml index 0a01c2f..57dece0 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -18,7 +18,7 @@ classifiers = [ ] dependencies = [ - "mthds==0.14.0", + "mthds==0.15.0", "pydantic>=2.10.6,<3.0.0", "typing-extensions>=4.0.0", "httpx>=0.23.0,<1.0.0", diff --git a/uv.lock b/uv.lock index f4154b6..0f0fc50 100644 --- a/uv.lock +++ b/uv.lock @@ -212,7 +212,7 @@ wheels = [ [[package]] name = "mthds" -version = "0.14.0" +version = "0.15.0" source = { registry = "https://pypi.org/simple" } dependencies = [ { name = "httpx" }, @@ -221,9 +221,9 @@ dependencies = [ { name = "tomlkit" }, { name = "typing-extensions" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/92/59/4ca9539571a2030f9427aeddb01bc334910953ebdfa4ab7db2051a6b8ae7/mthds-0.14.0.tar.gz", hash = "sha256:d2b4a9cd064004dfd5b802bb71c9894601fba9e598426f96e4bede16d52c3900", size = 221186, upload-time = "2026-09-06T21:44:54.585Z" } +sdist = { url = "https://files.pythonhosted.org/packages/54/16/1aa1219f44bc469213018478f582eb4eb28f69ea43a09fe73afc9cfb4e34/mthds-0.15.0.tar.gz", hash = "sha256:2e612174b089cae799eb236f5b1bcd183b8847a62a06a265f72f0e50b4c89555", size = 250552, upload-time = "2026-09-18T20:23:09.888Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/28/0b/32908eeed33396c5aafd4bcb2a1adc38d0ab8d85c2fde3af0bf4c2f73745/mthds-0.14.0-py3-none-any.whl", hash = "sha256:59a706205b6e6df47345caac01038588d21ab9e9c246afc765ee50b5f27f05a8", size = 87192, upload-time = "2026-09-06T21:44:53.022Z" }, + { url = "https://files.pythonhosted.org/packages/c0/5c/7526815d8bc8cd8690c8cf4de3819e3f322886280af05bf9aed7bdce450c/mthds-0.15.0-py3-none-any.whl", hash = "sha256:98670ca1f97fc2211ace0f404acd416ce5882edb728845d48440d6e73eedc34c", size = 99711, upload-time = "2026-09-18T20:23:08.283Z" }, ] [[package]] @@ -326,7 +326,7 @@ dev = [ [package.metadata] requires-dist = [ { name = "httpx", specifier = ">=0.23.0,<1.0.0" }, - { name = "mthds", specifier = "==0.14.0" }, + { name = "mthds", specifier = "==0.15.0" }, { name = "mypy", marker = "extra == 'dev'", specifier = "==1.19.1" }, { name = "pydantic", specifier = ">=2.10.6,<3.0.0" }, { name = "pylint", marker = "extra == 'dev'", specifier = "==4.0.4" }, From 8c9fcd57f3a96579e36413dfe6d81e4d6e363491 Mon Sep 17 00:00:00 2001 From: Louis Choquel Date: Thu, 24 Sep 2026 13:16:34 +0200 Subject: [PATCH 3/3] Release v0.12.0 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 4 ++-- pyproject.toml | 2 +- uv.lock | 2 +- 3 files changed, 4 insertions(+), 4 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 8765652..0519ca0 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,6 +1,6 @@ # Changelog -## [Unreleased] +## [v0.12.0] - 2026-09-24 ### Added @@ -10,7 +10,7 @@ - **`download_artifacts` names each file after the field it fills, and `artifact_filename` takes a location (Breaking)**: a saved file is named after the first path at which its reference sits — `$.rooms[3].staged_photo.url` is saved as `rooms-3-staged_photo.png`, and an output that is one image as `main_stuff.png` — instead of after the last segment of its storage key, which now supplies only the extension, so the same run saves under the same names from either SDK. `artifact_filename(location, content_type, scope)` replaces `artifact_filename(uri, content_type, index)` and raises `ArtifactOperationError` for anything but an `ArtifactLocation` whose first path is in the walk's notation; a field whose name Windows reserves for a device (`aux`, `nul`, `com1` and the like) is saved with a trailing `_` (`aux_.png`); the full rule is on `docs/artifact-download.md`. - **`DownloadedArtifact` carries a required `found_at` (Breaking)**: every item of a `download_artifacts` verdict carries its reference's `found_at`, on the saved arm and the error arm alike, so a file that was not saved still says which field it would have filled. `DownloadedArtifact` is now an `ArtifactLocation`, so `artifact_filename` takes a verdict item as it is, and code that builds `DownloadedArtifact` values — a test fake standing in for `download_artifacts` — must now supply the field. -- **Requires `mthds` 0.15.0 (Breaking).** The pin moves from 0.14.0, and it moves so that `pipelex` and `pipelex-sdk` can be installed together again: `pipelex` pins `mthds` exactly too and has required 0.15.0 since its 0.60.0, so two exact pins on different versions left the pair unresolvable. Nothing in this client changed with it. The release's one breaking cut — `ConceptAbstract` reduced to a name and `StuffAbstract` dumping its concept as the ref string — sits in protocol models this SDK does not build on, and the base client's run-source guard now refuses a protocol arg smuggled through `extra` before building the body rather than while building it, which this client's own guard already did ahead of both. The release also ships the canonical method-files pair as `mthds.protocol.method_files`; this package keeps its own `parse_method_files` / `serialize_method_files` for now, for the reason `docs/architecture.md` gives. +- **Requires `mthds` 0.15.0 (Breaking)**: the exact pin moves from 0.14.0, so `pipelex-sdk` can again be installed beside `pipelex`, which has pinned `mthds==0.15.0` exactly since its 0.60.0. Nothing in this client's own surface changes: the release's breaking cut to `ConceptAbstract` and `StuffAbstract` sits in protocol models this SDK does not build on, and `parse_method_files` / `serialize_method_files` stay beside the canonical `mthds.protocol.method_files` for the reason `docs/architecture.md` gives. ## [v0.11.0] - 2026-09-23 diff --git a/pyproject.toml b/pyproject.toml index 57dece0..c745cfa 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -1,6 +1,6 @@ [project] name = "pipelex-sdk" -version = "0.11.0" +version = "0.12.0" description = "The Python client for the Pipelex hosted API — the MTHDS Protocol surface plus the durable run lifecycle and the Pipelex product surface, built on the `mthds` protocol base." authors = [{ name = "Evotis S.A.S.", email = "oss@pipelex.com" }] maintainers = [{ name = "Pipelex staff", email = "oss@pipelex.com" }] diff --git a/uv.lock b/uv.lock index 0f0fc50..071ecac 100644 --- a/uv.lock +++ b/uv.lock @@ -303,7 +303,7 @@ wheels = [ [[package]] name = "pipelex-sdk" -version = "0.11.0" +version = "0.12.0" source = { editable = "." } dependencies = [ { name = "httpx" },