diff --git a/CHANGELOG.md b/CHANGELOG.md index ac62b17..8c4e029 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,22 @@ # Changelog +## [v0.10.1] - 2026-09-22 + +### Added + +- **`RunResults.working_memory`, the run's whole memory declared on both paths.** A completed run now hands back every named stuff it held when it finished — the inputs it was given, the intermediates it produced and the main output — on one field that reads the same whichever path ran, typed as the standard's `DictWorkingMemoryAbstract` imported from `mthds` rather than restated: `root` maps a stuff name onto its `concept` and `content`, `aliases` maps a role such as `main_stuff` onto one of those names, and every level stays extension-open so a runner's per-stuff extras survive. On the hosted path it is the `working_memory.json` artifact the platform relays as its own key, no longer an unnamed value riding `model_extra`; on the blocking path it is lifted off `pipe_output.working_memory`, which the standard declares required, so it is always set there. The `None`-versus-absent reading the other optional fields follow applies to it too, and `download_artifacts` walks the declared field for its `working_memory` scope. This closes the last run-results gap against `@pipelex/sdk`. See `docs/run-results.md`. +- **`summarize_usage`, one run-level reading of what a run consumed.** `pipelex_sdk.usage.summarize_usage(results)` folds a completed run's `tokens_usages` / `usage_assembly_error` pair into a single `UsageSummary` — a `state` of `records`, `no_inference` or `unavailable`, the priced total in USD with `cost_partial` flagging a sum that mixes priced and unrated calls, the two additive token totals, the call count, the relayed assembly error and a `by_pipe` rollup ordered most expensive first, with the calls the runtime did not attribute grouped under a `None` pipe code. Every rule `docs/run-usage.md` states is applied in one place instead of re-derived by each consumer: an unrated call (`cost` of `None`) is kept apart from one priced at zero, only `input` and `output` are summed because the other categories are subsets, and an empty record list is a run that did no inference — a cost of `0` and zero tokens, never `None`. It is pure, so it needs no client and does no I/O. A results body that never carried `tokens_usages` raises the new `FieldNotIncludedError` rather than turning a key the read did not deliver into a run with no usage to show; a key relayed as `null` is a value and reads as `unavailable`. `TokensUsageRecord.cost` is now validated strict and finite, so a cost relayed as a string, a bool or a NaN fails the results body's parse instead of reaching a total as a number. This is the Python twin of `@pipelex/sdk`'s `summarizeUsage`, with the same summary shape. +- **`RunResults` describes a run's graph and its data.** Beside `graph_spec` it already carried, the results object now declares `graph_assembly_error` — the graph's twin of `usage_assembly_error`, non-`None` when the runner's graph assembly failed, which is the only thing that tells a broken graph apart from a run that produced none — and the three I/O artifacts that say what the graph's nodes hold: `pipe_io_contracts`, `input_form` and `output_form`, typed as the standard's `PipeIOContracts`, `InputForm` and `OutputForm` imported from `mthds.protocol` exactly as the validate report types them, built over the library the run executed against and keyed by namespaced `pipe_ref`, with `pipe_io_artifacts_error` beside them as their own "why is there none". Every one reads the same on both paths: the hosted results body relays each as an artifact, and the blocking path lifts each off `pipe_output`, unwrapping the runner's `pipe_io_artifacts` envelope onto the three sibling fields. The artifacts are closed shapes, so a member the pinned `mthds` does not define fails the parse of the whole results body — the same ruling the validate report follows. The hosted body relays neither error key yet, so on that path each is absent from `model_fields_set` rather than `None`, and that is the Python reading of the JS `undefined`. This is the Python half of what `@pipelex/sdk` shipped in 0.18.0 and 0.20.0. +- **`docs/run-results.md`** — "Reading a run's results", a page walking every field of `RunResults`: how an absent key differs from a relayed `null` in Python, the run id as the durable handle, `main_stuff` with a worked `get_run_result` example, what `graph_spec` is and how to keep it, the graph and artifact error twins, the three I/O artifacts and the rule that the contracts and the output form are taken together, the usage pair, `pipe_output`, and the produced files whose signed `public_url` must not be stored. Linked from `docs/architecture.md`, whose parity section now names what the run-results surface still lacks against the JS SDK. +- **`method_source_to_contents`, the reader for a stored method's bundle source**: `pipelex_sdk.product_models.method_source_to_contents(method.mthds)` turns the polymorphic string `get_method` hands back — the catalog `[{name, content}]` array the webapp editor writes, or a bare `.mthds` bundle as plain text — into the `list[str]` that `run`, `start` and `validate` take as `mthds_contents`. It never raises, and an empty list means the method carries no source rather than that reading it failed. It reads the array as the catalog form on key presence and then drops an entry whose `content` is not a non-blank string, keeping its siblings — the same reading the platform's own resolver and `@pipelex/sdk` apply to the same stored row, so one stored method reads the same way wherever it is read. +- **The artifact stack — a run's produced files, from its results to bytes on disk.** `pipelex_sdk.artifacts` carries the download twin of `prepare_inputs`, in four layers a consumer can stop at: `collect_artifacts(value)` lists the `pipelex-storage://` references inside any results value without touching the network, `resolve_artifacts(uris)` mints a fresh link for each through the platform's bulk route (`POST /v1/resolve-storage-url/bulk`, chunked at its bound, one verdict per reference with per-reference refusal as a value), `fetch_artifact(uri)` yields one bounded stream (a timeout, redirects refused, the byte cap enforced mid-stream, no credential of ours forwarded to the object store, the store's headers left neutral), and `download_artifacts(run_id | results, dir_path=…)` saves a whole run's files under a directory — by run id days after the run, or from a `RunResults` in hand — and answers a produced verdict naming every reference walked, errors included. Links are always minted fresh and never read off the content's expiring `public_url`, files are never overwritten, a partial file is unlinked on failure or cancellation, and the `working_memory` scope brings down the echoed inputs and intermediates too. The shapes, options and defaults live in `pipelex_sdk.artifact_models`, the typed failures in `pipelex_sdk.errors` (`ArtifactOperationError` and its subclasses, plus `FieldNotIncludedError`), and the whole contract is `docs/artifact-download.md`. This is the Python twin of what `@pipelex/sdk` shipped in 0.18.0. +- **`resolve_storage_urls_bulk`, one request for a whole list of storage references.** The client gained the raw bulk wire call under the artifact stack: `POST /v1/resolve-storage-url/bulk` takes at most the route's bound of references and answers one item per reference in request order, duplicates included, a refused reference carrying its `{code, detail}` inside a `200`. A run producing one file per page no longer costs one round trip per file, and `resolve_storage_url` stays for the caller with a single link to mint. + +### Fixed + +- **The blocking path stops dropping the executed graph.** Against a bare runner, `start_and_wait` now lifts `pipe_output.graph_spec` onto `RunResults.graph_spec` instead of writing `None` — the runner has always returned the graph there — so the field carries the same document whichever path ran. +- **A stored source nested too deeply to decode now fails as a `ValidationError`**: `parse_method_files` converts the JSON decoder's `RecursionError` into the `ValueError` its contract documents, so `MethodData`'s validator surfaces it as a `pydantic.ValidationError` like any other malformed response body instead of letting a bare `RecursionError` escape `get_method` past a caller's `except ValidationError`. + ## [v0.10.0] - 2026-09-13 ### Added diff --git a/Makefile b/Makefile index e4ca82c..2eacf9d 100644 --- a/Makefile +++ b/Makefile @@ -48,6 +48,7 @@ make update - Upgrade dependencies via uv make build - Build the wheels make test - Run unit tests +make e2e-test - Run the live e2e legs (needs PIPELEX_E2E_BASE_URL + PIPELEX_API_KEY) make test-with-prints - Run unit tests with prints make t - Shorthand -> test make tp - Shorthand -> test-with-prints @@ -83,7 +84,7 @@ make li - Shorthand -> lock install endef export HELP -.PHONY: all help env env-verbose check-uv check-uv-verbose lock install update build test test-with-prints t tp gha-tests agent-test format lint pyright mypy pylint merge-check-ruff-format merge-check-ruff-lint merge-check-pyright merge-check-mypy merge-check-pylint check-unused-imports fix-unused-imports check-TODOs cleanderived cleanenv cleanall c cc li agent-check +.PHONY: all help env env-verbose check-uv check-uv-verbose lock install update build test e2e-test test-with-prints t tp gha-tests agent-test format lint pyright mypy pylint merge-check-ruff-format merge-check-ruff-lint merge-check-pyright merge-check-mypy merge-check-pylint check-unused-imports fix-unused-imports check-TODOs cleanderived cleanenv cleanall c cc li agent-check all help: @echo "$$HELP" @@ -193,6 +194,11 @@ test: env $(VENV_PYTEST) -o log_cli=true -o log_level=WARNING $(if $(filter 1,$(VERBOSE)),-v,$(if $(filter 2,$(VERBOSE)),-vv,$(if $(filter 3,$(VERBOSE)),-vvv,))); \ fi +e2e-test: env + $(call PRINT_TITLE,"Live e2e legs") + @echo "These legs run against a live platform; they skip cleanly when PIPELEX_E2E_BASE_URL and PIPELEX_API_KEY are unset." + $(VENV_PYTEST) tests/e2e -o log_cli=true -o log_level=WARNING -v + test-with-prints: env $(call PRINT_TITLE,"Unit testing with prints") @if [ -n "$(TEST)" ]; then \ diff --git a/README.md b/README.md index e57a2b0..1db41ac 100644 --- a/README.md +++ b/README.md @@ -138,7 +138,7 @@ There is no barrel import — package `__init__.py` files stay empty. Import eac - **Client & construction** — `from pipelex_sdk.client import PipelexAPIClient, DEFAULT_API_BASE_URL, MthdsFile` - **Run lifecycle types** — `from pipelex_sdk.runs import RunStatus, RunPublic, RunRead, RunResults, RunResultState, WaitForResultOptions, PollInfo` -- **Product wire models** — `from pipelex_sdk.product_models import UserProfile, MethodData, MethodWriteInput, Membership, MembershipsResponse, SubscriptionResponse, PlanView, InvoiceView, OnboardingSubmission, UploadInput, UploadedFile, PipelineRun, ...` +- **Product wire models** — `from pipelex_sdk.product_models import UserProfile, MethodData, MethodWriteInput, Membership, MembershipsResponse, SubscriptionResponse, PlanView, InvoiceView, OnboardingSubmission, UploadInput, UploadedFile, PipelineRun, ...`, with the catalog-source readers beside them: `method_source_to_contents` turns a fetched `MethodData.mthds` into the `mthds_contents` a run or a validate takes, and `MethodFile` / `parse_method_files` / `serialize_method_files` are the codec for a method's custom PipeFunc `python`. - **Validation verdict types** — `from pipelex_sdk.validation_models import PipelexValidationResult, PipelexValidationReport, PipelexInvalidReport, ValidationErrorItem, SuggestedFix, VALIDATION_VIEW_INPUT_FORM, ...` - **Codegen tree** — `from pipelex_sdk.codegen_writer import write_codegen_tree, CodegenTreeWriteReport` to write one, `from pipelex_sdk.codegen_check import run_codegen_check, CodegenCheckReport, CodegenDrift, DriftCategory` to verify one, with the format primitives in `pipelex_sdk.codegen_lock` (`CodegenLock`, `parse_lock`, `load_lock`, `validate_artifact_path`, ...) and `pipelex_sdk.codegen_stamp` (`STAMPABLE_SUFFIXES`, `is_stampable_artifact_path`, `compute_content_hash`, `parse_stamped`, ...) - **Typed errors** — `from pipelex_sdk.errors import ApiResponseError, ApiUnreachableError, PipelineExecuteTimeoutError, PagingNotTerminatingError, RunFailedError, RunTimeoutError, RunLifecycleUnavailableError, RunStillRunningError, CodegenError, CodegenLockError, ...` @@ -149,7 +149,7 @@ There is no barrel import — package `__init__.py` files stay empty. Import eac ## Development ```bash -make install # create the venv and install all extras (resolves `mthds` from ../mthds-python) +make install # create the venv and install all extras make agent-check # fix-imports + format + lint + pyright + mypy make agent-test # run the test suite quietly (prints only on failure) make check # full gate: agent-check aggregate + unused-imports + pylint diff --git a/docs/architecture.md b/docs/architecture.md index 2768ed8..24203e9 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -108,9 +108,10 @@ The durable run lifecycle (`pipelex_sdk/runs.py` + the client's lifecycle method - `RunStatus` — the hosted status enum, with `is_terminal` / `is_success` predicates (exhaustive `match`). - `RunRead` — a run record read through the self-healing status path (adds `degraded` + `retry_after_seconds`). -- `RunResults` — result artifacts. `main_stuff` (the resolved main output content) is always present for a completed run: on the hosted path it is the `main_stuff.json` S3 artifact; on the bare-runner blocking path the SDK resolves it from the returned working memory via the response's `main_stuff_name`, so both paths deliver the same shape. Consumers read `main_stuff` directly. The full working memory still rides `pipe_output` (blocking path only) for consumers that want it, and `graph_spec` rides the hosted path. A completed run that cannot deliver a main stuff raises `MissingMainStuffError`. Extension-open, so any other server artifact is preserved. +- `RunResults` — result artifacts, every field walked on [`run-results.md`](./run-results.md). `main_stuff` (the resolved main output content) is always present for a completed run: on the hosted path it is the `main_stuff.json` S3 artifact; on the bare-runner blocking path the SDK resolves it from the returned working memory via the response's `main_stuff_name`, so both paths deliver the same shape. Consumers read `main_stuff` directly. The executed graph (`graph_spec`, with `graph_assembly_error` beside it), the three I/O artifacts (`pipe_io_contracts`, `input_form`, `output_form`, typed by import from `mthds.protocol` exactly as on the validate report, with `pipe_io_artifacts_error` beside them) and the usage pair read the same on both paths: the hosted path relays each as an artifact, the blocking path lifts each off `pipe_output` — unwrapping the runner's `pipe_io_artifacts` envelope onto the three sibling fields. `working_memory` reads the same way and is typed by import too, as `mthds`'s `DictWorkingMemoryAbstract` beside the `DictPipeOutputAbstract` that types `pipe_output`: the hosted path relays the `working_memory.json` artifact as its own key, the blocking path lifts `pipe_output.working_memory`, which the standard declares required and which is therefore always set there. A completed run that cannot deliver a main stuff raises `MissingMainStuffError`. Extension-open, so any other server artifact is preserved; and a key the hosted body did not carry is absent from `model_fields_set`, which is how a Python reader tells "not relayed" from "relayed as null" where the JS twin reads `undefined`. - `TokensUsageRecord` — one client-facing usage record per inference call, carried by `RunResults.tokens_usages` on both paths. A mirror of the runtime's own record, not a shape this SDK owns — inference accounting is a Pipelex runtime extension the MTHDS Protocol does not model, so the hosted API is what pins that wire contract: every field is optional and the model is extension-open so pre-contract artifacts (relayed verbatim, never migrated) still parse. See [`run-usage.md`](./run-usage.md) for the field reference, the cost/null semantics, and the old-artifact rules. - `RunResultState` — the single-shot result outcome, a union discriminated on `state` (`running` / `completed` / `failed`). +- `UsageSummary`, `PipeUsageSummary`, `UsageTokenTotals` and `UsageSummaryState` (`pipelex_sdk/usage.py`) — the run-level fold of the usage pair that `summarize_usage(results)` returns, the twin of `@pipelex/sdk`'s `summarizeUsage` down to the field names. Owned here because the fold is a reading of a Pipelex runtime extension the standard does not model, exactly as the record it folds is. The function is pure — no I/O, no client, the input untouched — and the null-aware rules it applies are on [`run-usage.md`](./run-usage.md), including the one divergence from the JS twin: a results body that never carried `tokens_usages` raises `FieldNotIncludedError` instead of reading as `unavailable`, since `model_fields_set` tells an absent key from a relayed `null` where TypeScript cannot. - `WaitForResultOptions` / `PollInfo` — poll-loop tuning and progress info. Async-native cancellation is via `asyncio.CancelledError` (cancel the awaiting task), so there is no `signal` field. ### Polling surface @@ -166,6 +167,8 @@ Several members of the valid arm are the **standard's** artifacts rather than Pi The types are imported and used, never re-exported from `pipelex_sdk`. Re-exporting would put this package's name on a vocabulary it does not own and hand consumers a second import path to drift against; import them from `mthds.protocol` directly. +The same three artifacts ride a completed run's results too — `RunResults.pipe_io_contracts`, `input_form` and `output_form`, built over the library the run executed against and relayed from the three sibling files the worker writes beside `graphspec.json` — and they are typed there by the same imports, under the same ruling and with the same strictness: a member the pinned `mthds` does not define fails the parse of the whole results body, main output included, exactly as it fails the validate report. That is a choice, made once for both surfaces rather than relaxed for one: the artifacts are written by the same `pipelex` the pin tracks, every run older than the artifacts carries them as `None`, and reading a drifted artifact half-way is worse than refusing it. See [`run-results.md`](./run-results.md). + **`bundle_blueprint` and `graph_spec` stay opaque**, for exactly the reason that used to cover all four: no published package declares them, so a type here could only be a copy. When either gets a standard page and a client model, it moves the same way. **Strictness composes; it does not spread.** The imported artifacts are **closed** shapes (`extra="forbid"`): a member this `mthds` version does not define is version drift and fails the parse, where catching it is cheap. The report envelope around them stays **extension-open** — `PipelexValidationReport` inherits `extra="allow"` from `mthds`'s `ValidationReport`, and declaring typed fields on a subclass does not touch that config — so an unrelated field a future server adds to the report still parses and still rides `model_extra`, exactly as before. That is the standard's own arrangement: the report is the envelope and grows, the artifact is the view of one version and does not. A test pins both halves, so a later edit cannot quietly close the envelope. @@ -229,13 +232,19 @@ The wire models are snake_case Pydantic v2. Response models are extension-open ( **`MethodData.python` is a typed `list[MethodFile]`, converted at the boundary.** On the wire it is one string: the JSON text of a `[{name, content}]` array, or `""` for a method with no custom Python. `parse_method_files` / `serialize_method_files` (in `product_models.py`) carry the rules — a blank source or `"[]"` yields `[]`, blank-content entries are dropped in both directions, and an empty list serializes to `""` rather than `"[]"` because `""` is the platform's clear sentinel. `MethodData` applies the parser through a `field_validator(mode="before")` and `MethodWriteInput` the serializer through a `field_serializer`, which is what makes the platform's three-way write contract fall out of `exclude_none=True`: **`None` → key absent → the stored Python is preserved; `[]` → `""` → cleared; a non-empty list → replaced.** `MethodFile` is distinct from `MthdsFile` (the validate input) and `MthdsFileItem` (the build closure) — three shapes for three surfaces. + **`MethodData.mthds` is polymorphic, and `method_source_to_contents` is how you read it.** Where `python` is converted at the boundary, the bundle source is left as the string the platform stored, because it has two at-rest shapes and neither is wrong: the catalog file-array the webapp editor writes, and a bare `.mthds` bundle as plain text. `method_source_to_contents(method.mthds)` resolves it to the `list[str]` that `run` / `start` / `validate` take as `mthds_contents` — the non-blank contents of the file-array, or the whole source as one bundle when it is not that form — and an empty list means the row carries no source at all rather than that reading it failed. It never raises: a `.mthds` file may legally open with a digit or a brace, so a JSON parse failure is a bundle, not an error. It is **not** a delegation to `parse_method_files`, and deliberately so — the two fields have different rules, because `python` is a typed catalog with no second at-rest shape to tell apart while `mthds` has one. What they do share is everything that must not drift: the decoder (`_decode_method_source`, which converts both a `JSONDecodeError` and the `RecursionError` of a source nested past what `json.loads` can descend into one `ValueError`) and blankness (`_is_blank`). + + **Three implementations read this one field, and they agree by ruling rather than by accident.** `@pipelex/sdk`'s `methodSourceToContents` (`src/method-source.ts`) and the platform's own `method_source_to_contents` (`pipelex-server`, `platform/src/pipelex_platform/services/method_resolution.py` — the resolver that expands a `method_id` run, and therefore the one that decides whether a stored method runs) both recognize a catalog entry by the presence of the `name` and `content` keys alone, then drop an entry whose `content` is not a string and keep its siblings. This SDK shipped strict for a moment — the array was the catalog form only when every entry was `{name: str, content: str}` — so `[{"name": "a", "content": "x"}, {"name": "b", "content": 1}]` was `["x"]` in the other two and the whole JSON text here, and a row that runs today by `method_id` would, read through this helper and passed back as `mthds_contents`, have reached the runner as raw JSON and died at the MTHDS parse. **`L-260913-75fdc5` ruled for the lax reading, and this SDK now matches both references exactly.** The reasoning is not that the lax rule is better in the abstract but that a client-side reader of a server-stored field exists so a caller can do locally what the platform does with the same row: a reader that disagrees with the code which actually runs the method is a reader that lies, whatever the merits of its rule. The strict reading's own argument — that a file whose `content` is not a string vanishes without a word — was not dismissed with it. It is a property of the format's decoder and is therefore shared by all three readers, so it is fixed once in the platform rather than three times in its clients, and it is filed as `L-260913-6ec559`. The ruled behaviour is pinned by a test here, so a later change of mind shows up as a failing test rather than as silent drift. + + Two further divergences from the JS twin come from the decoder rather than from this function, and they flip in opposite directions: `json.loads` accepts `NaN` and `Infinity`, which `JSON.parse` refuses, and refuses an integer literal past CPython's digit cap, which `JSON.parse` accepts. Blankness is a third — Python's `str.strip` here, ECMAScript's there — so a source of only U+FEFF is a bundle here and no source there, and one of only U+0085 is the reverse. `mthds.protocol.method_files` closes all three and is the adoption target. + **`delete_method` is asynchronous and its return type says so:** the route answers `202` with a `MethodDeletionAccepted` (`method_id`, `deletion_state`, `deletion_job_id`) the moment the platform has claimed the method and terminated its in-flight workflows — the rest of the cascade (runs, events, S3 objects) is enqueued. A returned value means "accepted", never "gone"; completion is the row disappearing from `list_methods`, not any field of that body. It used to be annotated `-> None` with a docstring promising an empty synchronous delete, which is a misleading contract around a destructive operation. The claim is a conditional write, so a double-clicked delete is a `409 conflict` rather than a second cascade over the same runs. - **Organizations** — `list_memberships()` → `MembershipsResponse` (memberships + active-org feature flags); `create_organization(name)` / `rename_organization(org_id, name)` → `Membership`. Organization *switch* is out of scope (a WorkOS session op, not a `/v1` route). - **Billing** — `get_subscription()`, `list_plans()`, `list_invoices()`, `create_checkout(plan)`. `change_plan(plan)` and `get_billing_portal()` surface a **409 `conflict`** (`ApiResponseError.code`) when there is no subscription yet — start one via `create_checkout` first. - **Pipelex API keys** — `list_pipelex_api_keys()`; `create_pipelex_api_key(label)` and `rotate_pipelex_api_key(id)` return the plaintext `api_key` **once**; `revoke_pipelex_api_key(id)`. Creation surfaces a **409 `pipelex_api_key_limit_reached`** when the per-account limit is hit. Rotation sends no body. - **Gateway (LLM inference) key** — `create_gateway_api_key(promo_code)` **always sends a JSON body** (even with `promo_code=None` → `{"promo_code": null}`); the server 422s an empty body. `get_gateway_api_key()` → status (`gateway_api_key` is `None` until provisioned). - **Onboarding** — `submit_onboarding(OnboardingSubmission)` (`POST /v1/onboarding/submit`, empty 2xx body); absent optional fields are dropped. -- **Storage** — `resolve_storage_url(uri)` → presigned URL; `upload(UploadInput)` → the stored file handle. The higher-level `upload_file` / `prepare_inputs` preparation surface built on top of `upload` is now available — see [input-preparation.md](./input-preparation.md). +- **Storage** — `resolve_storage_url(uri)` → presigned URL for one reference; `resolve_storage_urls_bulk(uris)` → one verdict per reference through `POST /v1/resolve-storage-url/bulk`, the route the artifact stack is built on; `upload(UploadInput)` → the stored file handle. The higher-level `upload_file` / `prepare_inputs` preparation surface built on top of `upload` is now available — see [input-preparation.md](./input-preparation.md) — and its download twin is the artifact stack below. - **Run records** — `list_runs(method_id, …)` / `iterate_runs(method_id, …)` / `get_run_detail(run_id)` (the catalog-style reads, distinct from the lifecycle status/result routes); `update_run(run_id, UpdateRunInput)` (admin/manual status patch, empty 2xx body). **This list is paged too.** `list_runs(method_id, created_from=…, created_to=…, limit=…, cursor=…)` returns a `RunPage` with the same opaque-cursor contract as `MethodPage`. `created_from` / `created_to` are **instants** — ISO-8601 with a UTC offset — and inclusive; they are index key conditions rather than filters, so a bare date or a naive timestamp is a platform `400` surfaced as `ApiResponseError`. Worth knowing and not obvious from the route: every `/v1/runs*` product route sits behind the platform's surface-access gate, which for API-key auth demands the `ff_api_keys` feature flag and fails closed with a `403` — so a `403` here means "flag", not "wrong key". @@ -244,6 +253,14 @@ The wire models are snake_case Pydantic v2. Response models are extension-open ( **`PipelineRun` fields the platform genuinely serves as null are typed nullable.** `method_id` is `None` for an ad-hoc run from an inline bundle, which belongs to no stored method; `pipe_code` is `None` for a run that let the bundle's `main_pipe` decide. `org_id`, `created_by_user_id`, and a narrowed `error: RunErrorReport | None` (`message`, `error_type` — the two fields a consumer may rely on out of the runner's verbose report) join them. `RunDetail`, returned only by `get_run_detail`, adds `mthds_contents` and `inputs`: what the run actually executed, and the only record of it, since a method edited since the run no longer describes what happened. Both are left out of the list and the polled status on purpose — their cost scales with page size and poll rate respectively. +## Artifact stack (`pipelex_sdk/artifacts.py` + `pipelex_sdk/artifact_models.py`) + +The download twin of input preparation, and the Python twin of `@pipelex/sdk`'s `src/artifacts.ts`: `collect_artifacts` walks a value for `pipelex-storage://` references, `resolve_artifacts` mints fresh links for a whole list through the bulk route (chunked at its bound), `fetch_artifact` yields one bounded stream, and `download_artifacts` saves a run's produced files under a directory as a produced verdict. The operations take a `Protocol` rather than the client, so `PipelexAPIClient` satisfies them structurally and a test injects a fake; the client also carries all four as methods, the way it carries `upload_file` and `prepare_inputs`. The whole contract, the options and the error taxonomy are in [artifact-download.md](./artifact-download.md). + +**The shapes are owned by `pipelex_sdk/artifact_models.py`** — the scope enum, the wire items, the options and the verdict, plus the public defaults — beside `product_models` and `crate_models`. That split is not cosmetic: `pipelex_sdk/errors.py` types `ScopeUnavailableError.scope` and `ArtifactAuthenticationError.verdict` with two of them and the operations module imports those errors, so keeping the shapes in the operations module would be an import cycle (`reportImportCycles` is an error here). The JS twin needs no such split, TypeScript tolerating the cycle. + +**The object store is fetched on its own httpx client**, never the API client's: no `Authorization` header of ours can ride along to the store, redirects are refused rather than followed, and the call's timeout is both a budget over the whole exchange and httpx's per-operation timeout. `_new_storage_client` is the one seam the module opens, which is where the unit suite injects an `httpx.MockTransport`. + ## Health probe `health()` → `GET {origin}/health`. The one route served at the **origin**, NOT under the `/v1` prefix — the origin is derived from the base URL (`_origin_of`, exposed as `self.origin_url`), so a base URL of `https://api.pipelex.com/v1/...` still probes `https://api.pipelex.com/health`. It is **out-of-protocol**: the MTHDS Protocol defines no health route, and `/health` is neither a protocol nor a product surface. It rides `_request_json` (the plainer regime), so a non-2xx raises `PipelineRequestError` rather than the product `ApiResponseError` — liveness needs no `code` taxonomy. Transport failures still map to `ApiUnreachableError`. (Checkpoint-5 decision: kept the plainer regime — see "Error regimes" above.) @@ -265,13 +282,15 @@ This SDK is a port of the TypeScript `@pipelex/sdk` (`PipelexApiClient`) and tra - **Tooling routes** — `lint`, `format` (`resolve` and `codegen` shipped in 0.8.0 with the method selectors and are **not** gaps). - **Authoring helpers** — `build_output`, `build_runner`, `concept`, `pipe_spec`. JS still exports its `buildInputs` wrapper; Python's was **deleted** rather than left unused, because `prepare_inputs` was its only caller. A deliberate divergence, not a gap: the JS wrappers retire together as their own step of the same program. -- **Offline helpers** — `get_method_closure` (client-side sugar that parses the polymorphic `mthds` source into a run-ready closure — in the JS SDK it is the documented migration target for the deleted by-id expansion legs; this SDK never had such legs, so the utility stays deferred rather than required). `run_codegen_check` was the other entry here and is no longer a gap: it shipped as `pipelex_sdk.codegen_check`, with a filesystem-shaped signature where the JS export is pure. +- **Offline helpers** — `get_method_closure` (the *client* sugar that fetches a method by id and labels each of its files with that id as provenance, raising when the row carries no source; in the JS SDK it is the documented migration target for the deleted by-id expansion legs, and this SDK never had such legs, so it stays deferred rather than required). Its parsing half is **not** a gap: `method_source_to_contents` is the Python counterpart of `methodSourceToContents`, so a caller composes the closure from `get_method` in two lines. `run_codegen_check` was the other entry here and is no longer a gap: it shipped as `pipelex_sdk.codegen_check`, with a filesystem-shaped signature where the JS export is pure. Each stays deferred rather than silently missing. Everything else — the protocol routes, the durable lifecycle, the whole product surface, and the errors — does have a Python equivalent. **Methods** — everything outside the gap list above has a counterpart: protocol (`execute`, `start`, `validate`, `validate_files`, `models`, `version`), durable lifecycle (`get_run_status`, `get_run_result`, `wait_for_result`, `start_and_wait`, the private `_supports_run_lifecycle` / `_execute_blocking`), the whole product surface (profile, methods CRUD with paged listing and the two iterators, organizations, billing, Pipelex API keys, gateway key, onboarding, storage, run records with `get_run_detail`), the crate routes (`resolve`, `codegen`) with the verbatim `write_codegen_tree` and the offline `run_codegen_check`, the input-preparation surface (`upload_file` / `prepare_inputs`, taking all three method selectors and reading its signature from the input-form descriptor), and `health`. The method selectors (`method_ref` / `method_id`) match the JS v0.16.0 surface across the run and tooling methods, with one signature-shape divergence: JS `validate` takes a `ValidateMethodSelector` object in place of its first argument, while Python takes `method_ref=` / `method_id=` keyword parameters — same wire, same XOR, idiomatic per language. -**Models** — field-for-field across the run-lifecycle types and the product wire models. Deliberate idiomatic ports (not gaps): milliseconds → seconds (`interval_seconds` / `timeout_seconds` / `elapsed_seconds`); the JS `AbortSignal` → Python `asyncio` cancellation (no `signal` field); JS inline string-unions promoted to `StrEnum`s (`OrgRole`, `PipeStatus`, the onboarding fields) with identical wire values; response models are `extra="allow"` for forward-compat. The Pipelex validation narrowing is **owned here** (`pipelex_sdk.validation_models`), narrowing `mthds`'s neutral verdict bases (the resolved follow-up #9); the brand-neutral `Dict*` wire concretes (`DictRunResultExecute`) are reused from `mthds` by inheritance — they are a shared wire contract the `pipelex` runtime also builds on — rather than duplicated as `pipelex-sdk-js` does. One addition runs ahead of the JS SDK: `write_codegen_tree` has no `@pipelex/sdk` counterpart, because the JS writer lives inside `pipelex-starter-js`'s own harness, mixed in with project policy; the byte-fidelity contract is small and load-bearing enough to belong to the SDK, where every consumer shares one correct implementation. The drift check goes the other way and takes a different shape on purpose: `@pipelex/sdk`'s `runCodegenCheck` is **pure** — the caller walks its own tree and hands in the text — because that module must stay free of Node builtins for a browser bundle, and its doc consequently loads the caller with obligations (walk the whole tree, do not reformat, decode strictly) whose every breach yields a wrong verdict rather than an error. Python has no such constraint, this package already does filesystem work, and `pipelex`'s own surface takes a root — so `run_codegen_check(root=…)` takes one too. Every caller obligation becomes the library's, the verdict is provable against `pipelex` by calling both with the same directory, and it composes with `write_codegen_tree(report, output_dir=…)` as the same path in and out. Two divergences worth naming: the page envelopes keep the wire's snake_case `next_cursor`, where the JS mirror renamed it `nextCursor` for its own consumers; and the method-files catalog converter (`parse_method_files` / `serialize_method_files`) lives in this package, where the JS pair lives in `mthds-js` because `pipelex-mcp` consumes the same format and wanted one owner. There is no second Python consumer, and the catalog serialization is a Pipelex product concern rather than an MTHDS protocol one, so this SDK is a proper home for it. If `mthds-python` ever grows an owner for the format, this SDK adopts it then. +**The run-results surface** — `RunResults` tracks `@pipelex/sdk` 0.20.0 field for field, and is at parity with it. Every field is declared, with the two-path semantics [`run-results.md`](./run-results.md) states: `graph_spec` (lifted on the blocking path instead of written `None`), `graph_assembly_error`, the three I/O artifacts typed from `mthds.protocol` with `pipe_io_artifacts_error` beside them, the usage pair, `working_memory` (the hosted artifact relayed as its own key, the blocking one lifted off `pipe_output.working_memory`), and `pipe_output`; and the usage pair's fold, `summarize_usage`, with the summary types it returns. What stays different is idiomatic and deliberate: a key the hosted body did not carry is absent from `model_fields_set` where the JS reads `undefined`, and the page says how to read that. + +**Models** — field-for-field across the run-lifecycle types and the product wire models. Deliberate idiomatic ports (not gaps): milliseconds → seconds (`interval_seconds` / `timeout_seconds` / `elapsed_seconds`); the JS `AbortSignal` → Python `asyncio` cancellation (no `signal` field, and no `aborted` flag on the download verdict — a cancelled `download_artifacts` raises `CancelledError` with its partial files already unlinked); the JS artifact request object → `DownloadArtifactsOptions` beside an explicit `dir_path` (`dir` being a Python builtin) and the raw bulk call named `resolve_storage_urls_bulk` after its route rather than `resolveStorageUrls`, so it cannot be misread as the single-reference `resolve_storage_url`; the returned `Response` of `fetchArtifact` → an async context manager, because an httpx stream is only live inside its own block; JS inline string-unions promoted to `StrEnum`s (`OrgRole`, `PipeStatus`, the onboarding fields) with identical wire values; response models are `extra="allow"` for forward-compat. The Pipelex validation narrowing is **owned here** (`pipelex_sdk.validation_models`), narrowing `mthds`'s neutral verdict bases (the resolved follow-up #9); the brand-neutral `Dict*` wire concretes (`DictRunResultExecute`, and the `DictPipeOutputAbstract` / `DictWorkingMemoryAbstract` pair that types `RunResults.pipe_output` and `RunResults.working_memory`) are reused from `mthds` by inheritance or by import — they are a shared wire contract the `pipelex` runtime also builds on — rather than duplicated as `pipelex-sdk-js` does. One addition runs ahead of the JS SDK: `write_codegen_tree` has no `@pipelex/sdk` counterpart, because the JS writer lives inside `pipelex-starter-js`'s own harness, mixed in with project policy; the byte-fidelity contract is small and load-bearing enough to belong to the SDK, where every consumer shares one correct implementation. The drift check goes the other way and takes a different shape on purpose: `@pipelex/sdk`'s `runCodegenCheck` is **pure** — the caller walks its own tree and hands in the text — because that module must stay free of Node builtins for a browser bundle, and its doc consequently loads the caller with obligations (walk the whole tree, do not reformat, decode strictly) whose every breach yields a wrong verdict rather than an error. Python has no such constraint, this package already does filesystem work, and `pipelex`'s own surface takes a root — so `run_codegen_check(root=…)` takes one too. Every caller obligation becomes the library's, the verdict is provable against `pipelex` by calling both with the same directory, and it composes with `write_codegen_tree(report, output_dir=…)` as the same path in and out. Two divergences worth naming: the page envelopes keep the wire's snake_case `next_cursor`, where the JS mirror renamed it `nextCursor` for its own consumers; and the method-files catalog converter (`parse_method_files` / `serialize_method_files`) lives in this package, where the JS pair lives in `mthds-js` because `pipelex-mcp` consumes the same format and wanted one owner. There is no second Python consumer, and the catalog serialization is a Pipelex product concern rather than an MTHDS protocol one, so this SDK is a proper home for it. `mthds` has since grown the canonical pair as `mthds.protocol.method_files`, on its `dev` branch and **not in the `mthds==0.14.0` this package pins exactly** — so adoption waits on a release before anything else. It is a step of its own rather than a rename in any case: that parser raises `PipelineRequestError`, which does not subclass `ValueError`, and pydantic converts only `ValueError` out of a validator — so re-pointing `MethodData`'s validator at it would stop every caller's `except ValidationError` from catching a malformed stored source. The release and the exception's base class are settled first; until then the two implementations are kept in step, and they disagree on blankness (this one is Python's `str.strip`, the canonical one is ECMAScript's), on the serialized bytes (default separators and ASCII escapes here, `JSON.stringify`'s there), and on the JSON constants and integer-literal cap the canonical one closes with `parse_constant=` and `parse_int=float`. **Errors** — `ApiResponseError`, `ApiUnreachableError`, `PipelineExecuteTimeoutError`, `RunFailedError`, `RunTimeoutError`, `RunLifecycleUnavailableError`, `PagingNotTerminatingError` are owned here; `RunStillRunningError` is re-exported from `mthds`. `ClientAuthenticationError` is **not** ported: it is a dormant export in the JS barrel (defined and exported but never raised by the client), and in Python it already lives in `mthds.runners.api.exceptions` — importable directly if ever needed, with no barrel here to re-export it through. diff --git a/docs/artifact-download.md b/docs/artifact-download.md new file mode 100644 index 0000000..05f0477 --- /dev/null +++ b/docs/artifact-download.md @@ -0,0 +1,155 @@ +# Artifact download (`collect_artifacts` / `resolve_artifacts` / `fetch_artifact` / `download_artifacts`) + +> **Status: implemented** (`pipelex_sdk/artifacts.py`, with its shapes in `pipelex_sdk/artifact_models.py`). This is the download twin of [input preparation](./input-preparation.md): where `prepare_inputs` turns local files into `pipelex-storage://` references before a run, these four operations turn the references a run produced back into bytes on disk, or into a bounded stream, afterwards. They are layered so that each is usable without the next, and they are the Python twins of `@pipelex/sdk`'s `collectArtifacts` / `resolveArtifacts` / `fetchArtifact` / `downloadArtifacts`. +> +> **They need a platform that serves the bulk resolve route.** Everything below `collect_artifacts` mints its links through `POST /v1/resolve-storage-url/bulk`, a hosted-platform route. The public bare runner (`pipelex-api`) has no resolve route at all, single or bulk, and a hosted deployment that predates the route answers a `404`; in both cases the operation raises the existing `ApiResponseError` and nothing is downloaded. `resolve_storage_url`, the single-reference primitive, stays for the callers that have one link to mint. + +## Why this exists + +A run that produces an image, a PDF or a document does not embed the bytes in its results. The content carries the file's durable reference — a `pipelex-storage://` URI in its `url` field — beside a `public_url` the storage provider signed when the run wrote the file. That signed link is short-lived and must not be stored, so every consumer that wanted the file had to walk the result for references, mint a fresh link per reference, and stream each link to disk within sensible bounds. `download_artifacts` makes it one explicit operation, and the layers under it make each step reusable on its own. + +Downloading is **explicit and separate from running**, the input-preparation rule in reverse: `execute` / `start` never silently upload, and `start_and_wait` never silently downloads. There is no `download` option on any run call. A download is its own gesture, inspectable and repeatable — by run id, days after the run. + +Nothing here reads the embedded `public_url`. Every link is minted fresh by the platform, which is what makes a download work long after the embedded link died, and what keeps the tenant boundary where the platform enforces it. + +## The four operations + +### `collect_artifacts(value)` — the pure walk + +```python +from pipelex_sdk.artifacts import collect_artifacts + +uris = collect_artifacts(results.main_stuff) +# ["pipelex-storage://org/runs/01J…/outputs/illustration.png", …] +``` + +Walks any JSON-shaped value and returns every string that **is** a `pipelex-storage://` reference — the whole string, scheme first, with something after the scheme. A string that merely contains a reference does not count; the bare scheme does not count; nothing else is looked at. The result is deduplicated and kept in discovery order. It is a contract rather than a heuristic, because the scheme is unambiguous: the runtime serializes a produced file as content carrying its reference in `url`, and nothing else on the wire starts that way. Mappings, lists and pydantic models are walked alike, so a parsed `main_stuff` and a whole `RunResults` both work. + +It is pure — no network, no key — so a consumer can count or list a result's files without resolving any of them. + +### `resolve_artifacts(uris)` — fresh links for a whole list + +```python +resolved = await client.resolve_artifacts(uris) +for item in resolved: + if item.error is None: + ... # item.url is fetchable now; item.expires_at says until when; item.content_type may be None + else: + ... # item.error is {code, detail} — the route's own per-reference refusal +``` + +One call to the platform's bulk route for the whole list, chunked at the route's bound of `BULK_RESOLVE_MAX_URIS` references per request, answering one `ResolvedArtifact` per reference **in request order, duplicates included**. A resolved item carries `url`, `expires_at` (UTC, ISO 8601) and `content_type` (`str | None` — the platform's guess from the reference's extension, `None` when it has none) with `error` at `None`; a refused item carries `error` as an `ArtifactItemError` (`code`, `detail`) with the three link fields `None`. The codes are the route's: `invalid_storage_uri` for a malformed reference and `forbidden` for one belonging to another organization. A consumer branches on `error`, never on an HTTP status, because the request is a `200` whenever every reference got a verdict. + +Only what is not about a reference raises: the route's whole-request refusals (a caller with no organization, a request over the bound or with an unknown field, a signing failure, a deployment without the route) as `ApiResponseError`, and an unreachable host as `ApiUnreachableError`. An empty list resolves to an empty list with no request made. A link lives about fifteen minutes; resolve close to the moment of use. + +This is where every browser-side or server-rendering consumer stops: it mints links on the server and hands each one to a client that uses it immediately. `resolve_storage_urls_bulk` is the raw wire call underneath it (one request, at most the bound), the way `upload` sits under `upload_file`. + +### `fetch_artifact(uri)` — one bounded stream + +```python +from pipelex_sdk.artifacts import fetch_artifact + +async with fetch_artifact(client, uri) as stream: + # stream.status_code, stream.headers and stream.content_type are the store's own + async for chunk in stream.aiter_bytes(): + ... +``` + +Resolves the reference fresh and hands the object store's response over as a bounded stream. It is an **async context manager**, not a returned response: an httpx stream is only live inside its own block, so the block is where the bytes are read and where the connection is released. `client.fetch_artifact(uri)` is the same thing as a method. The bounds: + +- **A timeout** covering the connection, the headers and the whole body (`timeout_seconds`, default 120 s), applied twice: as a budget over the whole exchange and as httpx's own per-operation timeout, which is the per-stall bound a total budget alone does not give. +- **Redirects refused**: a presigned link has no reason to redirect, and one that does is refused rather than followed. +- **The byte cap enforced mid-stream** (`max_bytes`, default 1 GiB): a declared `Content-Length` over the cap is refused before a byte is read, and a body that crosses the cap while streaming raises out of the iteration — never buffered. +- **No credentials forwarded**: the request carries no headers of ours, and it runs on its own httpx client rather than the API client's. The link's authorization is in its query string, and nothing else may ride along to the store. +- **Plain `http:` refused** unless `allow_http=True`. A general-purpose library does not fetch over plain http silently; the local compose stack's object store hands out such links, and that is what the option opts into. + +It is **header-neutral**: the status and headers are the store's own. The one change is that a `Content-Encoding` the installed httpx actually decoded — `gzip` and `deflate`, plus `br` and `zstd` only when their optional packages are installed, read from httpx's own decoder table — is dropped with the encoded `Content-Length`, since the body handed on is the decoded bytes; any other coding keeps both headers beside its still-encoded body. A proxy relaying it therefore owns the response hygiene — `X-Content-Type-Options: nosniff`, a sandboxing CSP on the asset response, a controlled `Content-Disposition`, private caching — and must set them itself. + +Only a `2xx` is yielded. Anything else raises an `ArtifactFetchError` whose `code` says why, in the same closed vocabulary the download verdict uses per item: the route's `invalid_storage_uri` / `forbidden`, then `unsupported_url` (not a URL, not http(s), or carrying credentials), `plain_http_refused`, `redirect_refused`, `store_refused` (a 401/403 from the store — the link is freshly minted, so this is the store refusing a fresh signature, not an expired link), `not_found` (404/410), `store_error` (any other non-2xx, with `status`), `too_large`, `timeout` and `network`. Cancelling the awaiting task raises `asyncio.CancelledError` as-is. `download_artifacts` and a proxy share this boundary: the download is the same fetch followed by a write. + +### `download_artifacts(run_id | results, dir_path, …)` — the files on disk + +```python +from pipelex_sdk.artifact_models import ArtifactScope, DownloadArtifactsOptions +from pipelex_sdk.artifacts import download_artifacts + +verdict = await download_artifacts( + client, + dir_path="./out/01J…", + run_id="01J…", # or: results= + options=DownloadArtifactsOptions(scope=ArtifactScope.WORKING_MEMORY), # default ArtifactScope.MAIN_STUFF +) + +print(verdict.saved_paths) +if not verdict.all_saved: + for artifact in verdict.artifacts: + if artifact.error is not None: + print(artifact.uri, artifact.error.code, artifact.error.detail) +``` + +`client.download_artifacts(dir_path=…, run_id=…)` is the same thing as a method. + +**Where it reads from.** Exactly one of `run_id` and `results`. A `run_id` re-reads the results through `get_run_result`, so a completed run is downloadable days later from its id alone; a `RunResults` already in hand is read as it is, with no request. `scope` picks the artifact walked for references: `main_stuff` (the default) is the run's output, and `working_memory` is the opt-in that also brings down the echoed inputs and every intermediate stuff — it is read off `RunResults.working_memory`, the declared field both paths deliver ([`run-results.md`](./run-results.md#working_memory--every-named-stuff-of-the-run)). + +**How it downloads.** The whole set is resolved through the bulk route ahead of the work, then an `asyncio.Semaphore` bounds how many references are in flight at once (`concurrency`, default 4) over the whole per-reference pipeline: fetch, create the file exclusively, stream the body in. Resolution is just-in-time where it matters: a link that has expired by the time its task reaches it — a large set downloaded a few at a time can outlive the fifteen-minute link — is resolved again for that reference alone, so no fetch ever runs on a stale signature. + +**Filenames.** Each file is named by `artifact_filename` (exported): the last segment of the storage key, reduced to `[A-Za-z0-9._-]` with leading dots stripped so it can never name anything outside the directory, capped in length with the extension preserved, given an extension from the content type when the key has none, and falling back to `artifact-N`. Files are **never overwritten**: a name already on disk gets a numeric suffix (`report-1.pdf`, `report-2.pdf`), through exclusive creation (`os.O_EXCL`) rather than an exists-check, so two tasks cannot race for one name. The directory is created if missing. + +**Cleanup.** A failed or cancelled download unlinks its partial file; nothing truncated is ever left under a final name. + +**The verdict.** A `DownloadArtifactsResult`, one entry per reference in discovery order: `scope`, `artifacts` (each a `DownloadedArtifact` with `uri`, `path`, `content_type`, `size` and `error`), `saved_paths` (the absolute paths of the saved ones, same order) and `all_saved`. `len(verdict.artifacts)` is the count of references walked, errors included. An empty walk over a present scope — an output that references no stored file — is a verdict with empty lists and `all_saved=True`, not an error, and it touches neither the network nor the disk. `content_type` is the platform's guess from the reference, known before the fetch, on both arms. + +Per-item `error.code` is the fetch vocabulary above plus the download's own: `resolve_failed` (an expired link could not be re-resolved, for a reason that is not the credential), `total_limit_exceeded` (the item that would take the call past `max_total_bytes`; an item refused only because files still in flight hold the room stops nothing else, since one of them may yet fail and give it back), `write_failed` (the file could not be created, written or closed) and `aborted` (not yet started when a credential failure stopped the call). + +**What it raises.** Only conditions with no verdict, all typed: + +- `RunStillRunningError` (with the retry hint) or `RunFailedError` — a `run_id` naming a run that has not completed; +- `FieldNotIncludedError` — the results read never carried the scope's key, so it is absent from `results.model_fields_set`. This is the Python reading of the JS `undefined`: ask for the key and read again; +- `ScopeUnavailableError` — the key WAS relayed and its value is `None`, which is the platform saying it has no such artifact for this run (`scope` and `run_id` on the error). Reading by `run_id`, a null `main_stuff` is already `MissingMainStuffError` from `get_run_result`; +- `ArtifactAuthenticationError` — the resolve route refused the credential (`401` / `403`), on the first resolve or on a re-resolve part-way through. It carries `verdict`, the result as it stood: the refusal stops the remaining references being taken but lets the fetches already running finish, since they are on presigned links that do not carry the credential, so every file saved is real and listed and the rest are marked `aborted` with a detail naming the credential failure; +- `ArtifactOperationError` — an unusable directory, both selectors or neither, or nonsense bounds; +- and the transport and lifecycle errors of the reads it makes, unchanged: `ApiResponseError` for a deployment without the bulk route, `RunLifecycleUnavailableError` for a bare runner asked by `run_id`, `ApiUnreachableError`. + +Everything else that can go wrong with one reference is that reference's `error`. + +**Options and defaults**, all on `DownloadArtifactsOptions` (`FetchArtifactOptions` is its first three): + +| Option | Default | What it bounds | +| ----------------- | ------------------------ | ----------------------------------------------------------------------------------- | +| `scope` | `ArtifactScope.MAIN_STUFF` | the artifact walked for references | +| `concurrency` | `4` | artifacts in flight at once, held by an `asyncio.Semaphore` | +| `max_bytes` | 1 GiB | one file, from `Content-Length` and again mid-stream | +| `timeout_seconds` | 120 s | one file's whole exchange, and httpx's per-operation timeout with it | +| `max_total_bytes` | 4 GiB | the bytes the whole call saves, a file in flight counting its declared length | +| `allow_http` | `False` | whether a plain `http:` link is fetched | + +The caps are accident guards against filling a disk from a runaway output, not judgments about artifact size. + +## What differs from `@pipelex/sdk`, and why + +The contract is the JS one — the same operations, the same defaults, the same verdict shape, the same error taxonomy and the same safety rules. What differs is idiomatic, and is the same set of ports the SDK already makes elsewhere (see the parity section of [`architecture.md`](./architecture.md)): + +- **Seconds, not milliseconds** (`timeout_seconds`), as everywhere else in this package. +- **`asyncio` cancellation, not an `AbortSignal`.** There is no `signal` option and no `aborted` field on the verdict: cancel the awaiting task, and `asyncio.CancelledError` comes out of `download_artifacts` with every partial file already unlinked. The item code `aborted` stays, for the references a credential failure stopped before they were taken. +- **An async context manager, not a returned response**, because an httpx stream is only live inside its own block. +- **`dir_path`, not `dir`**, which is a Python builtin. +- **Options as pydantic models**, so a caller passes `DownloadArtifactsOptions(...)` where the JS twin spreads keys into one request object. +- **The shapes live in `artifact_models.py`**, beside `product_models` and `crate_models`, because `pipelex_sdk.errors` types two of its artifact errors with them and the operations module imports those errors — one home for the shapes keeps that from being an import cycle. + +## Round trip + +The two directions compose. A file uploaded by `prepare_inputs` is echoed in the run's working memory under the same reference, so a pass-through run brings it back byte for byte: + +```python +prepared = await client.prepare_inputs(files=files, inputs={"doc": "./brief.pdf", "note": "hi"}) +results = await client.start_and_wait(pipe_code=pipe_code, mthds_contents=[bundle], inputs=prepared.inputs) +verdict = await download_artifacts( + client, + dir_path="./out", + run_id=results.pipeline_run_id, + options=DownloadArtifactsOptions(scope=ArtifactScope.WORKING_MEMORY), +) +# verdict.artifacts finds prepared.uploads[0].uri among the saved files +``` + +That is also the live e2e leg (`tests/e2e/test_artifacts_e2e.py`, run by `make e2e-test`), which needs a platform carrying the bulk route and skips itself when `PIPELEX_E2E_BASE_URL` and `PIPELEX_API_KEY` are unset. diff --git a/docs/run-results.md b/docs/run-results.md new file mode 100644 index 0000000..4be51ab --- /dev/null +++ b/docs/run-results.md @@ -0,0 +1,178 @@ +# Reading a run's results + +A completed run hands back one object, `RunResults` (`pipelex_sdk/runs.py`), and every field of it is described here. The same object comes back from `start_and_wait`, from `wait_for_result(run_id)`, and from the `completed` arm of `get_run_result(run_id)` — one accessor set whichever path ran, which is the point of the type. It is the Python twin of the page of the same name in `@pipelex/sdk`, and the fields are the same; what differs is idiomatic and is called out where it matters, above all how an absent key is read. + +Two paths produce it. Against the hosted API the SDK starts a durable run and polls `GET /v1/runs/{id}/results`, where the platform relays the run's S3 artifacts verbatim. Against a bare `pipelex-api` runner, which has no run store, the SDK falls back to the blocking `POST /v1/execute` and maps the runner's native `pipe_output` onto the same shape, lifting the artifacts that ride it onto their own fields. `start_and_wait` picks between the two from the `GET /v1/version` handshake, so a consumer does not choose. + +| field | type | hosted (durable) path | bare-runner (blocking) path | +|---|---|---|---| +| `pipeline_run_id` | `str` | the run store's id | the runner's own id for the call | +| `main_stuff` | `Any` | the `main_stuff.json` artifact | resolved out of the returned working memory | +| `graph_spec` | `Any` | the `graphspec.json` artifact | lifted off `pipe_output` | +| `graph_assembly_error` | `str \| None` | absent until the platform relays it | lifted off `pipe_output` | +| `pipe_io_contracts` | `PipeIOContracts \| None` | the `pipe_io_contracts.json` artifact | lifted off `pipe_output` | +| `input_form` | `InputForm \| None` | the `input_form.json` artifact | lifted off `pipe_output` | +| `output_form` | `OutputForm \| None` | the `output_form.json` artifact | lifted off `pipe_output` | +| `pipe_io_artifacts_error` | `str \| None` | absent until the platform relays it | lifted off `pipe_output` | +| `tokens_usages` | `list[TokensUsageRecord] \| None` | the `tokens_usages.json` artifact | lifted off `pipe_output` | +| `usage_assembly_error` | `str \| None` | relayed | lifted off `pipe_output` | +| `working_memory` | `DictWorkingMemoryAbstract \| None` | the `working_memory.json` artifact | lifted off `pipe_output.working_memory` | +| `pipe_output` | `DictPipeOutputAbstract \| None` | absent | the runner's whole native output | + +## `None` versus absent — the reading this page relies on + +Every field but the first two is optional, and two readings of an optional field are distinct on purpose. The JS SDK reads a key the hosted body did not carry as `undefined` and a key relayed as `null` as `null`. Python has one `None`, so the SDK keeps the distinction where pydantic keeps it: in `model_fields_set`. A key the hosted body did not carry is not in the set and reads `None`; a key relayed as `null` is in the set and reads `None` too. + +```python +results = await client.wait_for_result(run_id) + +if "graph_assembly_error" not in results.model_fields_set: + ... # the platform relayed no such key: no information, not "assembly succeeded" +elif results.graph_assembly_error is None: + ... # relayed as null: assembly did not fail +``` + +That distinction matters on the hosted path only. On the blocking path the SDK lifts every field off the runner's output and passes each one explicitly, so every field is set there whether or not the runner carried the key — the blocking path always answers, exactly as the JS twin writes `null` for each. Most consumers never need the set: a check of `is None` is the right branch for "is there a value", and `model_fields_set` is for the one question it answers, whether the wire said anything at all. + +## `pipeline_run_id` — the durable handle + +The run id is what makes a run readable after the process that started it has gone. `start` returns it in its acknowledgement before the run finishes, and every lifecycle read takes it: `get_run_status(run_id)` for the status row, `get_run_result(run_id)` for a single result lookup, `wait_for_result(run_id)` to resume polling a run an earlier session started. It is also what a `RunTimeoutError` leaves you with — the run keeps executing server-side, so the timeout is a reason to re-poll by id, not a reason to run the method again. + +```python +ack = await client.start(pipe_code="my_domain.summarize", inputs={"text": "..."}) +print(ack.pipeline_run_id) # persist this — it outlives the process + +# …later, in another process +results = await client.wait_for_result(ack.pipeline_run_id) +``` + +Against a bare runner the id identifies the call the runner just answered, but there is no run store behind it: the lifecycle routes are absent, so re-reading it raises `RunLifecycleUnavailableError`. Durable resumption is a hosted capability. + +## `main_stuff` — the output + +`main_stuff` is the resolved content of the run's main output and is always present for a completed run. On the hosted path it is the `main_stuff.json` artifact; on the blocking path the SDK resolves it out of the returned working memory through the response's `main_stuff_name`. Both deliver the same content shape, so there is no shape-guessing and no path-dependent branch to write. A completed run that cannot deliver one raises `MissingMainStuffError` rather than handing back a half-filled result. + +It is typed `Any` because the content is polymorphic: a structured output arrives as a dict of the concept's fields, and a multiple output as the envelope `{"items": [...]}` that the runtime's `ListContent` serialises to. Every content type serialises to an object, natives included — a text output is `{"text": "…"}` and a number `{"number": 0}` — so a guard written for a bare `""` or `0` never fires, and an empty multiple output is `{"items": []}` rather than `[]`. Narrow it where you read it, ideally through the types generated for the method rather than a hand-written cast. + +For a multiple output that means reading `items` off the dict and validating each member with the generated per-concept model, because codegen emits a model per concept and no wrapper type for the envelope. **Do not rely on the model to catch the mistake for you.** The generated models are not strict, so a concept with a required field rejects the envelope loudly, while one whose fields are all optional validates it to an empty instance and discards the output in silence. + +```python +from pipelex_sdk.runs import RunResultCompleted, RunResultFailed, RunResultRunning + +state = await client.get_run_result(run_id) + +match state: + case RunResultCompleted(): + # `state.result` is the RunResults; `main_stuff` is the output content. + summary = state.result.main_stuff + print(summary["title"], len(summary["bullets"])) + case RunResultRunning(): + print(f"not finished — poll again in {state.retry_after_seconds or 2}s") + case RunResultFailed(): + print(f"run ended as {state.status}: {state.message}") +``` + +`get_run_result` is the single-shot lookup and returns that discriminated state. `wait_for_result(run_id)` drives the same lookup in a loop, honouring the server's `Retry-After`, and returns the `RunResults` directly — raising `RunFailedError` on a terminal non-completed status and `RunTimeoutError` when the budget runs out. + +## `working_memory` — every named stuff of the run + +`working_memory` is everything the run held when it finished — the inputs it was given, the intermediates it produced and the main output, each under the name the method gave it. It is a declared field on both paths and reads the same on each: on the hosted path the platform relays the `working_memory.json` artifact as its own key, and on the blocking path the SDK lifts it off `pipe_output.working_memory`. The standard declares that member required on the runner's output, so on the blocking path the field always carries a value. + +It is typed as the standard's `DictWorkingMemoryAbstract`, imported from `mthds.runners.api.models` rather than restated here — the same ruling that types `pipe_output`. Two members: `root`, a dict of stuff name to stuff, each stuff carrying a `concept` (the namespaced ref, or the whole concept object when the runner dumps one) and its `content`; and `aliases`, a dict mapping a role onto a root key, which is where `main_stuff` names the entry `results.main_stuff` already resolved for you. Every level is extension-open, so a runner's per-stuff extras — `stuff_code`, `stuff_name` — ride `model_extra` instead of being dropped. + +Reading a stuff by name means going through `root`, and the alias table is what turns a role into that name: + +```python +results = await client.wait_for_result(run_id) + +memory = results.working_memory +if memory is not None: + draft = memory.root["draft"] + print(draft.concept_ref, draft.content) # e.g. "my_domain.Draft" {...} + main_name = memory.aliases.get("main_stuff", "main_stuff") + print(memory.root[main_name].content) # the same content as `results.main_stuff` +``` + +`content` is typed `Any` for the same reason `main_stuff` is: it is the serialized content of whatever concept the stuff holds, so narrow it where you read it, ideally through the types generated for the method. + +**When it is `None`, and when it is absent.** The two readings the page states above apply here: on the hosted path a relayed `null` means the platform has no such artifact for this run (it was not written), and is in `model_fields_set`; a body that carried no `working_memory` key at all leaves the field `None` and out of the set, which is no information rather than "the run held nothing". On the blocking path the field is always set and never `None`. `download_artifacts` makes that distinction an error rather than a branch when its scope is `working_memory` — a never-relayed key raises `FieldNotIncludedError` and a relayed `null` raises `ScopeUnavailableError` ([`artifact-download.md`](./artifact-download.md)). + +## `graph_spec` — the executed graph + +`graph_spec` is the graph the run actually executed: `meta.mode` is `"live"`, and there is one node per pipe with its execution status, its start and end timestamps, its inputs and outputs, and the inference models and cost attributed to it. It is the same document a local `pipelex` run writes as `graphspec.json`, so anything that reads one of those files reads this value unchanged. + +The field is typed `Any` by a standing ruling ([`architecture.md`](./architecture.md#typed-by-import-the-descriptors-and-the-pipe-io-contracts)): no published Python package declares the graph spec, so a type here could only be a copy that drifts from the runtime that emits it. The value is relayed verbatim either way — the typing says where the schema lives, not that the content is uncertain. + +**Rendering it.** The viewer that consumes it is `@pipelex/mthds-ui`'s `GraphViewer`, a React component; a Python service hands the JSON to the front end that renders it. The viewer takes `graph_spec` with the pair `pipe_io_contracts` and `output_form` below to show each node's value rather than its concept's structure, so a service that serves the graph should serve the three artifacts beside it. + +**Keeping it.** The value is plain JSON, so persisting it is a write; there is no SDK helper and none is needed. Keeping it is worth doing for anything you may have to explain later, because it is the only record of what the run did pipe by pipe: + +```python +import json +from pathlib import Path + +Path("graphspec.json").write_text(json.dumps(results.graph_spec, indent=2)) +``` + +**When it is `None`.** On the hosted path, the artifact may not have been written when the results were delivered. On either path, the runner may have assembled no graph at all — which is what the next field is for. + +## `graph_assembly_error` — why there is no graph + +`graph_assembly_error` is the graph's twin of `usage_assembly_error`, and it exists for the same reason: a `None` `graph_spec` alone cannot say whether graph assembly was off, broke, or simply had not finished writing. When the runner's assembly failed, this field carries the runner's message. + +On the blocking path the SDK lifts it off `pipe_output`, beside the graph itself. **On the hosted path the key is absent, so the field reads `None` and is not in `model_fields_set`**: the platform's results body relays no such key and the SDK parses that body as it arrives, so the failure the bare runner reports is not yet observable through the hosted API. The field is declared ahead of that relay so consumers have one accessor to write against and nothing breaks the day the wire gains the key — the value appears on its own, with no SDK change. Until then, treat an unset field on the hosted path as "no information", not as "assembly succeeded". + +```python +if results.graph_assembly_error is not None: + log.warning("graph assembly failed for this run: %s", results.graph_assembly_error) +elif results.graph_spec is None: + ... # no graph: assembly was off, the artifact was not written, or (hosted) the error is not relayed +``` + +## `pipe_io_contracts`, `input_form` and `output_form` — what the graph's data is + +`graph_spec` carries the values a run produced; these three say what those values ARE. They are the validate report's own artifacts — the standard's `PipeIOContracts`, `InputForm` and `OutputForm`, imported from `mthds.protocol` rather than restated here, under the same ruling that governs them on the validate report ([`architecture.md`](./architecture.md#typed-by-import-the-descriptors-and-the-pipe-io-contracts)) — built over the library the run actually executed against and keyed by namespaced `pipe_ref` (`domain.code`) over one shared key set. They are the same documents a local `pipelex` run writes beside its `graphspec.json` as `pipe_io_contracts.json`, `input_form.json` and `output_form.json`, so a consumer reads one thing whether the artifacts came from `/v1/validate`, from a results directory, or from a hosted run. + +The contract names each pipe's inputs and its output — the concept, the multiplicity, the JSON Schema of the payload — and the two form descriptors say what each of those slots IS as a typed field, which is what a renderer needs to lay a value out without inspecting it. The types are the standard's: a contract entry is a `PipeIOContract`, a form entry a `PipeInputFormDescriptor` or `PipeOutputFormDescriptor` whose nodes are the kind-discriminated field union — narrow a node with `match node: case ListField(): …`, importing the per-kind models from `mthds.protocol.input_form`. + +**Read the contracts and the output form together.** `@pipelex/mthds-ui`'s `GraphViewer` gates a data node's value on holding both: given the pair it renders the payload, and given one or neither it falls back to the concept's structure table with no data tab. That is why they arrive as a set rather than one at a time. `input_form` is optional even then — it is what lets the method's own inputs show their values, since no pipe produced them and no output descriptor describes them. + +```python +contracts = results.pipe_io_contracts +output_form = results.output_form +if contracts is not None and output_form is not None: + summarize = contracts["my_domain.summarize"] + print(summarize.output.concept_ref) # e.g. "my_domain.Summary" + print(output_form["my_domain.summarize"].field.kind) # e.g. "object" +``` + +**They are closed shapes, and the parse is strict.** The three artifacts are `extra="forbid"` in `mthds.protocol`: a member the pinned `mthds` does not define is version drift, and it fails the parse of the whole results body — main output included — rather than being read half-way. What is raised is pydantic's own `ValidationError`, out of `get_run_result`, `wait_for_result` and `start_and_wait` alike, and it is neither an `ApiResponseError` nor a `PipelineRequestError`, so a handler written for this SDK's errors does not catch it. On the hosted path nothing is lost by it: the artifacts stay in the run store, and the same results fetch parses from a `pipelex-sdk` release whose `mthds` pin knows the member. On the blocking path there is no store, the runner's response is discarded with the exception, and the completed `main_stuff` is not recoverable from that call; the one way to read such a run is `execute()`, whose extension-open `pipe_output` carries the artifacts raw — a different call, not a recovery of the one that failed. That is the same ruling the validate report follows, for the same reason: one declaration per language, drift refused at the parse. It bites only against a runtime newer than the artifacts this package's `mthds` pin describes, because the artifacts are written by the same `pipelex` the pin tracks, and every run older than the artifacts carries them as `None`, which parses. The envelope around them stays open: an unrelated key the platform adds still parses and rides `model_extra`. + +**When they are `None`.** On the blocking path the SDK unwraps the runner's `pipe_io_artifacts` envelope — the runner carries the three together, since they share a key set and are built in one pass — onto these three fields, so each has one accessor whichever path ran. On the hosted path the platform relays each as its own key, and the key is always in the body, so all three are set there: `None` when the artifact was not written. On either path `None` means the run described no data at all — graph tracing off, a runtime older than the artifacts, or a build that broke, which only `pipe_io_artifacts_error` tells apart, where it is relayed. + +## `pipe_io_artifacts_error` — why there is no description + +`pipe_io_artifacts_error` is the three artifacts' twin of `graph_assembly_error`, and it exists for the same reason: three `None` artifacts alone cannot say whether the run described no data or whether building the description broke. When the runner's build failed, this field carries its message. It is lifted off `pipe_output` on the blocking path and, like `graph_assembly_error`, the hosted results body relays no such key — so treat an unset field there as "no information", not as "the build succeeded". + +## `tokens_usages` and `usage_assembly_error` — what the run consumed + +The usage pair reports what each inference call consumed and cost — one `TokensUsageRecord` per call, in completion order — and reads identically on both paths. The `None`-versus-empty semantics, the cost rules (`None` is unrated, `0` is priced at zero), the non-additive token categories and the pre-contract artifacts that still parse all have their own page: [`run-usage.md`](./run-usage.md). For the run's totals, do not add the records up by hand — `pipelex_sdk.usage.summarize_usage(results)` folds the pair into one null-aware reading with a per-pipe rollup, under every rule that page states. It is the one place the `None`-versus-absent distinction above becomes an error rather than a branch: a results body that never carried `tokens_usages` raises `FieldNotIncludedError`, because a key the read did not carry is not a run that reported no usage. + +## `pipe_output` — the runner's native output + +`pipe_output` is the bare runner's whole native output, and it is present on the blocking path only — the hosted results body carries no such key, so on that path it reads `None`. It is supplementary: `main_stuff`, the graph pair, the three I/O artifacts, the working memory and the usage pair are all lifted out of it onto fields that read the same on both paths, so a consumer that reads those fields keeps working against the hosted API. What `pipe_output` adds is the runner's output exactly as it arrived, typed as the standard's `DictPipeOutputAbstract`, which is extension-open — the runner's Pipelex extension fields, the `pipe_io_artifacts` envelope among them, stay reachable in their raw form through `model_extra`. + +## Produced files + +A run that produces an image, a PDF or a document does not embed the bytes. The content inside `main_stuff` carries the file's durable reference — a `pipelex-storage://` URI, in the content's `url` — beside a `public_url` the storage provider signed when the run wrote the file. **That signed link is short-lived and must not be stored**: it expires on the provider's own schedule, so a link persisted in a database or rendered into a cached page stops working without warning, while the `pipelex-storage://` reference beside it is permanent and is what belongs in your records. + +**The whole download direction is [`artifact-download.md`](./artifact-download.md)**, and it is where a consumer should start: `collect_artifacts(results.main_stuff)` lists the references without touching the network, `resolve_artifacts` mints a fresh link for each through the platform's bulk route, `fetch_artifact` streams one within bounds, and `download_artifacts` saves a whole run's files under a directory and answers a produced verdict. None of them reads the embedded `public_url`. + +For the single reference you already hold, the raw primitive is still there: + +```python +resolved = await client.resolve_storage_url("pipelex-storage://...") +# resolved.url, resolved.expires_at, resolved.content_type — fetch `url` now; re-resolve for the next reader. +``` + +The same rule holds in a browser: resolve on the server, hand the client a link it uses immediately, and never let a presigned URL outlive the request it was minted for. The upload direction — turning local files into `pipelex-storage://` references before a run — is the mirror of this and has its own page, [`input-preparation.md`](./input-preparation.md). diff --git a/docs/run-usage.md b/docs/run-usage.md index e9781b0..eea5340 100644 --- a/docs/run-usage.md +++ b/docs/run-usage.md @@ -1,6 +1,6 @@ # Run usage — reading what a run consumed -A completed run reports what its inference calls consumed as a list of `TokensUsageRecord` objects on `RunResults`, one per inference call, in the order the calls completed. This page covers how to read them, what each field means, and the edge cases the model is deliberately shaped around. +A completed run reports what its inference calls consumed as a list of `TokensUsageRecord` objects on `RunResults`, one per inference call, in the order the calls completed. This page covers how to read them, what each field means, the edge cases the model is deliberately shaped around, and `summarize_usage`, which folds them into one run-level summary under those rules. The wire shape is not this SDK's invention. Inference accounting is a Pipelex runtime extension — the MTHDS Protocol does not model it, and says nothing about usage reporting — so the hosted API is what pins the contract, and `pipelex_sdk.runs.TokensUsageRecord` is a client-side mirror of the runtime's own record. `@pipelex/sdk` carries the same mirror in TypeScript. @@ -10,11 +10,12 @@ The wire shape is not this SDK's invention. Inference accounting is a Pipelex ru result = await client.start_and_wait(pipe_code="my_domain.summarize", inputs={"text": "..."}) if result.tokens_usages is not None: - total_cost = sum(record.cost or 0.0 for record in result.tokens_usages) for record in result.tokens_usages: print(record.pipe_code, record.inference_model_name, record.nb_tokens_by_category, record.cost) ``` +For the run's totals, do not add the records up by hand: [`summarize_usage`](#summarizing-a-run--summarize_usage) does it under the rules below. + The accessor is the same whichever path ran. `start_and_wait` picks a path from the `GET /v1/version` handshake: - **Hosted (durable) path** — the records come from the runner's `tokens_usages.json` artifact, which `GET /v1/runs/{id}/results` unpacks onto the results body as top-level keys and relays verbatim. @@ -52,10 +53,11 @@ Two traps worth naming explicitly: - `cost is None` means the model has **no rate table at all** — an own-GPU model, a mock run, a dry run. - `cost == 0` means a rate table existed and priced the call at zero. +- `cost` is strict and finite on the wire: a value relayed as a string, a bool or a NaN fails the whole results body's parse, so a number that reaches a sum is a number the artifact carried. Those are different facts; `record.cost or 0.0` conflates them, which is fine for a sum but wrong for "was this call priced?". -There is no per-category cost breakdown and no run-level aggregate on the wire. Sum the records for a run total. +There is no per-category cost breakdown and no run-level aggregate on the wire. The run total is the sum of the records, and [`summarize_usage`](#summarizing-a-run--summarize_usage) computes it with the `None` / `0` distinction kept. ## Null and empty semantics @@ -65,7 +67,7 @@ There is no per-category cost breakdown and no run-level aggregate on the wire. - usage assembly **broke** (an event-read failure); - on the hosted path, the run was **delivered before the artifact existed**. -It is `[]` when assembly ran, succeeded, and no inference happened, and non-empty otherwise. +It is `[]` when assembly ran, succeeded, and no inference happened, and non-empty otherwise. An empty list is a run that did no inference, so the run's total cost is `0` and its token totals are zero, never `None`: `None` stays reserved for calls that were not rated. `usage_assembly_error` is the **only** field that distinguishes the broken case from the other two — they are otherwise indistinguishable on the wire. A caller that needs to tell "we have no usage data because something failed" from "there was nothing to report" must branch on `usage_assembly_error`, not on `tokens_usages` alone: @@ -78,6 +80,58 @@ elif not result.tokens_usages: ... # ran, but no inference happened ``` +## Summarizing a run — `summarize_usage` + +`summarize_usage(results)` folds a run's usage pair into one run-level reading and a per-pipe rollup, applying every rule on this page so that no consumer has to re-derive them. It is pure: it does no I/O, needs no client, and leaves its input untouched. It takes the whole `RunResults` a completed run came back with. + +```python +from pipelex_sdk.usage import UsageSummaryState, summarize_usage + +usage = summarize_usage(result) + +match usage.state: + case UsageSummaryState.RECORDS: + cost = "not rated" if usage.total_cost_usd is None else f"${usage.total_cost_usd:.4f}" + partial = " (partial)" if usage.cost_partial else "" + print(f"{usage.calls} calls, {cost}{partial}") + for row in usage.by_pipe: + print(row.pipe_code or "(unattributed)", row.total_cost_usd, row.calls) + case UsageSummaryState.NO_INFERENCE: + print("no inference happened: $0") + case UsageSummaryState.UNAVAILABLE: + print(usage.assembly_error or "usage was not reported for this run") +``` + +| field | type | meaning | +|---|---|---| +| `state` | `UsageSummaryState` | Which reading of `tokens_usages` the summary describes. Read it first. | +| `total_cost_usd` | `float \| None` | Sum of the priced calls' costs, in USD. | +| `cost_partial` | `bool` | True when priced and unrated calls are mixed, so the total covers the priced calls only and is a lower bound. | +| `tokens` | `UsageTokenTotals` | The two additive token totals, `input` and `output`, each summed over the records that reported it. A category no record reported is `None`, which is different from a reported `0`. | +| `calls` | `int` | Number of records summarized, one per inference call. | +| `assembly_error` | `str \| None` | `usage_assembly_error` as relayed. | +| `by_pipe` | `list[PipeUsageSummary]` | One row per `pipe_code`, each carrying `pipe_code`, `total_cost_usd`, `cost_partial`, `tokens` and `calls`, folded over that pipe's calls exactly as the run level is. | + +The three states follow the [null and empty semantics](#null-and-empty-semantics) above, and `UsageSummaryState` is a `StrEnum`, so a state also compares and prints as its wire string: + +| `state` | `tokens_usages` | `total_cost_usd` | `tokens` | `calls` | `by_pipe` | +|---|---|---|---|---|---| +| `records` | a non-empty list | the priced sum, or `None` when no call was priced | the summed totals | the record count | one row per pipe | +| `no_inference` | `[]` | `0` | `input` and `output` both `0` | `0` | `[]` | +| `unavailable` | `None` | `None` | `input` and `output` both `None` | `0` | `[]` | + +A `None` `total_cost_usd` therefore means one of two things, and `state` tells them apart: under `records` no call was rated, and under `unavailable` nothing is known. Within `unavailable`, `assembly_error` is still the only sign that usage assembly broke rather than being off or not yet written. + +`by_pipe` puts the most expensive pipe first. Priced pipes come by cost, descending, and unrated pipes after every priced one; a tie breaks on the call count, descending, and then on the pipe code. The calls the runtime did not attribute to a pipe (`pipe_code` `None`) form one group of their own, which sorts after the named pipes when everything else ties. + +A [pre-contract record](#old-artifacts-parse-too) carries no `cost` and no `pipe_code`, so it counts as unrated and unattributed. Its legacy `job_metadata` and `unit_costs` are never read, so an old artifact shows up as a partial or `None` total rather than as a figure the SDK guessed. + +### A key that was never carried is not a `None` list + +A results body that never carried `tokens_usages` at all raises `FieldNotIncludedError` rather than answering `unavailable`. That is the Python reading of the difference [`run-results.md`](./run-results.md) sets out: a key the platform relayed as `null` is in `results.model_fields_set` and is a value, while a key the body did not carry is absent from the set and says nothing about the run. Answering "nothing is known about this run's usage" for a key the read never delivered would turn a gap in the read into a fact about the run. The error carries the field's name in `field_name`. Today it is dormant: the blocking path lifts the pair off the runner's output and sets it explicitly, and the hosted body relays the key on every read, so the only body without it comes from a runner that relays no usage at all. It becomes reachable the day a results read can leave a key out, which is the include selector's to add, and it then names what the read left out. The TypeScript twin, which cannot tell the two apart, reads both as `unavailable`. + +`usage_assembly_error` is not guarded the same way: both paths relay it beside the list, so its absence carries no such ambiguity and reads as `None`. + ## Old artifacts parse too Durable artifacts written before this contract shipped are relayed verbatim and never migrated. `TokensUsageRecord` therefore keeps **every field optional** and is extension-open (`extra="allow"`) — a pre-contract record parses without raising: diff --git a/pipelex_sdk/artifact_models.py b/pipelex_sdk/artifact_models.py new file mode 100644 index 0000000..9be6a06 --- /dev/null +++ b/pipelex_sdk/artifact_models.py @@ -0,0 +1,165 @@ +"""The artifact stack's wire models, options and verdict — the shapes `pipelex_sdk.artifacts` +produces and consumes, and the defaults its operations apply. + +They live in a module of their own, beside `product_models` and `crate_models`, because +`pipelex_sdk.errors` types two of its artifact errors with them (`ScopeUnavailableError.scope`, +`ArtifactAuthenticationError.verdict`) and the operations module imports those errors — one home for +the shapes keeps that from being an import cycle. Mirrors the types of `pipelex-sdk-js`'s +`src/artifacts.ts`, which needs no such split because TypeScript tolerates the cycle. +""" + +from __future__ import annotations + +from enum import StrEnum + +from pydantic import BaseModel, ConfigDict, Field + +from pipelex_sdk._pydantic_utils import empty_list_factory_of + +# ── Constants ──────────────────────────────────────────────────────── + +#: The scheme of a durable storage reference. +PIPELEX_STORAGE_SCHEME = "pipelex-storage://" + +#: How many references one bulk resolve request takes — the route's bound, fixed by the platform +#: contract rather than configured per deployment. A longer list is a `422`, so `resolve_artifacts` +#: chunks at this size. +BULK_RESOLVE_MAX_URIS = 100 + +#: The per-file byte cap. An accident guard against filling a disk from a runaway output, not a +#: judgment about artifact size: a produced file is server-side, so a caller cannot shrink it the +#: way they can shrink an upload. +DEFAULT_ARTIFACT_MAX_BYTES = 1024 * 1024 * 1024 + +#: The per-file budget for connecting, receiving the headers and reading the body. +DEFAULT_ARTIFACT_TIMEOUT_SECONDS = 120.0 + +#: The cap on the bytes one `download_artifacts` call writes in total. +DEFAULT_DOWNLOAD_MAX_TOTAL_BYTES = 4 * 1024 * 1024 * 1024 + +#: How many artifacts one `download_artifacts` call has in flight at once. +DEFAULT_DOWNLOAD_CONCURRENCY = 4 + + +# ── Types ──────────────────────────────────────────────────────────── + + +class ArtifactScope(StrEnum): + """Which of a run's artifacts `download_artifacts` walks for references.""" + + MAIN_STUFF = "main_stuff" + WORKING_MEMORY = "working_memory" + + @property + def results_field(self) -> str: + """The `RunResults` field this scope walks.""" + match self: + case ArtifactScope.MAIN_STUFF: + return "main_stuff" + case ArtifactScope.WORKING_MEMORY: + return "working_memory" + + +class ArtifactItemError(BaseModel): + """Why one reference failed — a value, never a raised error. + + `code` is the resolve route's own per-reference code (`invalid_storage_uri`, `forbidden`) or one + of the fetch boundary's (see `ArtifactFetchError`), plus the download's own `resolve_failed`, + `total_limit_exceeded`, `write_failed` and `aborted`. `detail` is the sentence a person reads. + """ + + model_config = ConfigDict(extra="allow") + + code: str + detail: str + + +class ResolvedArtifact(BaseModel): + """One reference's resolution — the bulk resolve route's item, verbatim. + + Either the link fields are set and `error` is `None`, or `error` is set and the three link fields + are `None`. A consumer branches on `error`, never on an HTTP status: the request was a `200` + whenever every reference got a verdict. + """ + + model_config = ConfigDict(extra="allow") + + #: The reference exactly as sent. + uri: str + #: A presigned link, fetchable now and for about fifteen minutes. + url: str | None = None + #: UTC expiry of the link, ISO 8601. + expires_at: str | None = None + #: The platform's content-type guess from the reference's extension; `None` when it has none. + content_type: str | None = None + error: ArtifactItemError | None = None + + +class BulkResolvedStorageUrls(BaseModel): + """The wire response of `POST /v1/resolve-storage-url/bulk` — one item per requested reference, + in request order. + """ + + model_config = ConfigDict(extra="allow") + + items: list[ResolvedArtifact] = Field(default_factory=empty_list_factory_of(ResolvedArtifact)) + + +class FetchArtifactOptions(BaseModel): + """The bounds `fetch_artifact` applies. Every one has a safe default.""" + + model_config = ConfigDict(extra="forbid") + + #: Refuse (before a byte is read) and cut (mid-stream) a body over this many bytes. Default 1 GiB. + max_bytes: int = DEFAULT_ARTIFACT_MAX_BYTES + #: Budget for the whole exchange — connect, headers and body. Default 120 s. + timeout_seconds: float = DEFAULT_ARTIFACT_TIMEOUT_SECONDS + #: Accept a plain `http:` link. Off by default: a general-purpose library does not fetch over + #: plain http silently. The local compose stack's object store hands out such links, which is + #: the case this opts into. + allow_http: bool = False + + +class DownloadArtifactsOptions(FetchArtifactOptions): + """The bounds and choices of one `download_artifacts` call, the per-file ones included.""" + + #: `main_stuff` (the default) walks the run's main output. `working_memory` is the opt-in that + #: also brings down the echoed inputs and every intermediate. + scope: ArtifactScope = ArtifactScope.MAIN_STUFF + #: How many artifacts are in flight at once. Default 4. + concurrency: int = DEFAULT_DOWNLOAD_CONCURRENCY + #: Cap on the bytes written by the whole call. Default 4 GiB. The item that would cross it is an + #: item error; every later item is checked against the room the saved files leave. + max_total_bytes: int = DEFAULT_DOWNLOAD_MAX_TOTAL_BYTES + + +class DownloadedArtifact(BaseModel): + """One reference's outcome in a download verdict — one shape with nullable fields, like + `ResolvedArtifact`: either `path` and `size` are set and `error` is `None`, or `error` is set and + both are `None`. `content_type` is the platform's guess from the reference's extension, known + before the fetch, on both arms. + """ + + uri: str + #: Absolute path of the written file. + path: str | None = None + content_type: str | None = None + #: Bytes written. + size: int | None = None + error: ArtifactItemError | None = None + + +class DownloadArtifactsResult(BaseModel): + """The produced verdict of `download_artifacts`. + + `len(artifacts)` is the count of references walked, errors included; an empty list over a present + scope is a verdict ("this output references no stored file"), not an error. + """ + + scope: ArtifactScope + #: One entry per reference, in discovery order. + artifacts: list[DownloadedArtifact] = Field(default_factory=empty_list_factory_of(DownloadedArtifact)) + #: The absolute paths of the files saved, in the same order. + saved_paths: list[str] = Field(default_factory=list) + #: True when every walked reference was saved — vacuously true for an empty walk. + all_saved: bool diff --git a/pipelex_sdk/artifacts.py b/pipelex_sdk/artifacts.py new file mode 100644 index 0000000..cb1e0bf --- /dev/null +++ b/pipelex_sdk/artifacts.py @@ -0,0 +1,870 @@ +"""The artifact stack — the download twin of `prepare_inputs`, in layers so each operation is +usable without the next: + +- `collect_artifacts(value)` — a pure walk of any JSON-shaped value for the strings that ARE + `pipelex-storage://` references. No network, no key. +- `resolve_artifacts(client, uris)` — the platform's bulk resolve route over a whole list, + chunked at the route's bound, one verdict per reference. +- `fetch_artifact(client, uri)` — an async context manager yielding a bounded stream for one + reference: resolved fresh, timed out, redirects refused, the byte cap enforced mid-stream, no + credentials forwarded, headers neutral. +- `download_artifacts(client, ...)` — a run's produced files saved under a directory by a bounded + pool of tasks, as a produced verdict. + +A produced file is never embedded in a run's results: the content carries its durable +`pipelex-storage://` reference beside a signed `public_url` that expires on the provider's +schedule. Nothing here reads that embedded link — every link is minted fresh by the platform, and +re-minted when it has expired by the time a task reaches it. See `docs/artifact-download.md`. + +Python counterpart of `pipelex-sdk-js`'s `src/artifacts.ts`, with the idiomatic ports the SDK +already makes elsewhere: seconds instead of milliseconds, an `asyncio.Semaphore` instead of a +worker pool, `asyncio` cancellation instead of an `AbortSignal`, and an async context manager +instead of a returned `Response` (an httpx stream is only live inside its own block). +""" + +from __future__ import annotations + +import asyncio +import math +import os +import re +from contextlib import asynccontextmanager +from dataclasses import dataclass +from datetime import UTC, datetime +from pathlib import Path +from typing import TYPE_CHECKING, Any, BinaryIO, Protocol, cast +from urllib.parse import unquote, urlsplit + +import httpx +from httpx._decoders import SUPPORTED_DECODERS # ruff: ignore[import-private-name] +from mthds.protocol.exceptions import PipelineRequestError +from pydantic import BaseModel, ValidationError + +from pipelex_sdk.artifact_models import ( + BULK_RESOLVE_MAX_URIS, + PIPELEX_STORAGE_SCHEME, + ArtifactItemError, + ArtifactScope, + BulkResolvedStorageUrls, + DownloadArtifactsOptions, + DownloadArtifactsResult, + DownloadedArtifact, + FetchArtifactOptions, + ResolvedArtifact, +) +from pipelex_sdk.errors import ( + ApiResponseError, + ArtifactAuthenticationError, + ArtifactFetchError, + ArtifactOperationError, + FieldNotIncludedError, + RunFailedError, + RunStillRunningError, + ScopeUnavailableError, +) +from pipelex_sdk.runs import RunResultCompleted, RunResultFailed, RunResultRunning, RunResults + +if TYPE_CHECKING: + from collections.abc import AsyncGenerator, AsyncIterator + from contextlib import AbstractAsyncContextManager + + from pipelex_sdk.runs import RunResultState + +# ── Constants ──────────────────────────────────────────────────────── + +#: A link this close to its `expires_at` is re-resolved rather than fetched: the object store checks +#: the signature when the request arrives, and a few seconds of clock skew between the platform and +#: this process must not turn a link the platform still considers live into a `403`. +_EXPIRY_MARGIN_SECONDS = 10.0 + +#: Longest filename `download_artifacts` writes, extension included. +_MAX_FILENAME_LENGTH = 128 + +#: Ceiling on collision suffixes before the never-overwrite rule gives up. +_MAX_UNIQUE_ATTEMPTS = 10_000 + +#: How much of a body one read takes before it is written. +_STREAM_CHUNK_BYTES = 64 * 1024 + +#: The extension to add when the storage key has none and the resolved content type is one of the +#: artifact types a run produces. Deliberately short: an unknown type simply gets no extension, +#: never a guessed one. +_EXTENSION_BY_CONTENT_TYPE: dict[str, str] = { + "image/png": ".png", + "image/jpeg": ".jpg", + "image/webp": ".webp", + "image/gif": ".gif", + "image/svg+xml": ".svg", + "application/pdf": ".pdf", + "text/plain": ".txt", + "text/markdown": ".md", + "text/html": ".html", + "text/csv": ".csv", + "application/json": ".json", +} + +#: The codings the installed httpx decodes for us, read from its own decoder table (the one place +#: that knows): `br` and `zstd` join it only when their optional packages are installed, and +#: `x-gzip` never does. A coding it did not decode keeps its header, since the bytes handed on are +#: still encoded. +_DECODED_CODINGS = frozenset(SUPPORTED_DECODERS) - {"identity"} + +_SKIPPED_CREDENTIAL = "The download stopped on a credential failure before this artifact was fetched." + + +# ── The client surfaces the operations need ────────────────────────── + + +class BulkResolveClient(Protocol): + """The client surface the reading operations need — the raw bulk resolve call. Typed as + `PipelexAPIClient.resolve_storage_urls_bulk`'s own signature, so the client satisfies it + structurally and a test can inject a fake. + """ + + async def resolve_storage_urls_bulk(self, uris: list[str]) -> BulkResolvedStorageUrls: ... + + +class ArtifactCapableClient(BulkResolveClient, Protocol): + """What `download_artifacts` needs on top: the single-shot result lookup, for the `run_id` arm.""" + + async def get_run_result(self, run_id: str) -> RunResultState: ... + + +# ── collect_artifacts ──────────────────────────────────────────────── + + +def is_storage_reference(value: str) -> bool: + """A string that IS a storage reference: the scheme, then at least one character.""" + return value.startswith(PIPELEX_STORAGE_SCHEME) and len(value) > len(PIPELEX_STORAGE_SCHEME) + + +def collect_artifacts(value: Any) -> list[str]: + """Every `pipelex-storage://` reference inside a JSON-shaped value, deduplicated, in discovery + order. + + A string counts only when it IS a reference — the whole string, scheme first, with something + after the scheme; text that merely contains one does not. The scheme is unambiguous, so this walk + is a contract rather than a heuristic: the runtime serializes a produced image or document as + content carrying its reference in `url`, beside an expiring `public_url` this walk ignores. Pure + — no network, no key — so a consumer can count or list a run's produced files without resolving + any of them. Mappings, sequences and pydantic models are walked alike, so `results.main_stuff` + (a parsed JSON value) and a whole `RunResults` both work. + """ + found: dict[str, None] = {} + _walk_for_references(value, found) + return list(found) + + +def _walk_for_references(value: Any, found: dict[str, None]) -> None: + """Depth-first walk collecting every string that is a storage reference, in discovery order.""" + if isinstance(value, str): + if is_storage_reference(value): + found[value] = None + return + if isinstance(value, BaseModel): + _walk_for_references(value.model_dump(), found) + return + if isinstance(value, dict): + for entry in cast("dict[str, Any]", value).values(): + _walk_for_references(entry, found) + return + if isinstance(value, (list, tuple)): + for item in cast("list[Any]", value): + _walk_for_references(item, found) + + +# ── artifact_filename ──────────────────────────────────────────────── + + +def artifact_filename(uri: str, content_type: str | None, index: int) -> str: + """The bare filename a storage reference is saved under. + + The last segment of the storage key, reduced to a conservative character set so it can never name + anything but a regular file directly inside the target directory. Path separators are the split + point, so no traversal survives; leading dots are stripped, so no hidden file and no `..`; + everything outside `[A-Za-z0-9._-]` becomes `_`; an empty result falls back to a numbered + `artifact-N`. Length is capped with the extension preserved, and an extension is added from the + content type when the key carries none. A collision on disk is not this function's concern: + `download_artifacts` suffixes the stem (`name-1.ext`) on exclusive creation, so a file is never + overwritten. + """ + key = uri.removeprefix(PIPELEX_STORAGE_SCHEME) + key = re.split(r"[?#]", key, maxsplit=1)[0] + segments = [part for part in re.split(r"[\\/]", key) if part] + # `unquote` keeps a malformed escape as typed, where the JS twin's `decodeURIComponent` throws + # and falls back to the same thing; sanitization below handles either. + decoded = unquote(segments[-1]) if segments else "" + + name = re.sub(r"[^A-Za-z0-9._-]", "_", decoded) + name = re.sub(r"^[._-]+", "", name) + name = re.sub(r"[._-]+$", "", name) + if not name: + name = f"artifact-{index + 1}" + + if not _extension_of(name) and content_type is not None: + guessed = _EXTENSION_BY_CONTENT_TYPE.get(content_type.split(";")[0].strip().lower()) + if guessed is not None: + name += guessed + + # The cap comes last, so a guessed extension is inside it like any other. + if len(name) > _MAX_FILENAME_LENGTH: + extension = _extension_of(name) + # The extension is kept only if there is room left for a stem; a pathological extension is + # dropped rather than preserved. + if len(extension) < _MAX_FILENAME_LENGTH: + name = name[: _MAX_FILENAME_LENGTH - len(extension)] + extension + else: + name = name[:_MAX_FILENAME_LENGTH] + return name + + +def _extension_of(name: str) -> str: + """`os.path.splitext` for a bare filename: `.ext`, or `""` (a leading dot is not an extension).""" + dot = name.rfind(".") + return name[dot:] if dot > 0 else "" + + +# ── resolve_artifacts ──────────────────────────────────────────────── + + +async def resolve_artifacts(client: BulkResolveClient, uris: list[str]) -> list[ResolvedArtifact]: + """Resolve a list of references through the bulk route, chunked at `BULK_RESOLVE_MAX_URIS` per + request, and answer one `ResolvedArtifact` per reference in request order, duplicates included. + + Per-reference failure is a value on the item; only what is not about a reference raises — the + route's whole-request refusals as `ApiResponseError` (a caller with no organization, a malformed + request, a signing failure, or a `404` from a deployment that does not serve the route) and an + unreachable host as `ApiUnreachableError`. An empty list resolves to an empty list with no + request made. + """ + items: list[ResolvedArtifact] = [] + for start in range(0, len(uris), BULK_RESOLVE_MAX_URIS): + chunk = list(uris[start : start + BULK_RESOLVE_MAX_URIS]) + answer = await client.resolve_storage_urls_bulk(chunk) + if len(answer.items) != len(chunk): + msg = ( + f"The bulk resolve route answered {len(answer.items)} item(s) for {len(chunk)} reference(s) — " + "a malformed answer, so no reference can be matched to its verdict." + ) + raise ArtifactOperationError(msg) + items.extend(answer.items) + return items + + +# ── fetch_artifact ─────────────────────────────────────────────────── + + +@dataclass(frozen=True) +class _FetchBounds: + """The bounds every fetch runs under, defaults filled in and validated.""" + + max_bytes: int + timeout_seconds: float + allow_http: bool + + +class ArtifactStream: + """A bounded, header-neutral view of the object store's response for one reference. + + `status_code` and `headers` are the store's own — the one change being that a `Content-Encoding` + httpx already decoded is dropped with the encoded `Content-Length`, since the bytes handed on are + the decoded ones. A proxy relaying this response owns its hygiene (`X-Content-Type-Options`, a + sandboxing CSP, a controlled `Content-Disposition`, private caching) and must set those itself, + because this object does not. + + The body is read through `aiter_bytes()` or `read()`, either of which raises `ArtifactFetchError` + (`too_large`) the moment the bytes cross the call's `max_bytes`. + """ + + def __init__(self, uri: str, response: httpx.Response, max_bytes: int) -> None: + self.uri = uri + self.status_code = response.status_code + self.headers = _decoded_headers(response.headers) + self.content_type: str | None = self.headers.get("content-type") + self._response = response + self._max_bytes = max_bytes + + async def aiter_bytes(self, chunk_size: int = _STREAM_CHUNK_BYTES) -> AsyncIterator[bytes]: + """Yield the body in chunks, cutting it the moment it crosses the byte cap.""" + total = 0 + async for chunk in self._response.aiter_bytes(chunk_size): + total += len(chunk) + if total > self._max_bytes: + msg = f"The artifact crossed the {_format_mib(self._max_bytes)} cap mid-stream." + raise ArtifactFetchError(msg, uri=self.uri, code="too_large", status=self.status_code) + yield chunk + + async def read(self) -> bytes: + """The whole body in memory, under the same cap. For a large artifact, iterate instead.""" + parts: list[bytes] = [] + async for chunk in self.aiter_bytes(): + parts.append(chunk) + return b"".join(parts) + + +def _new_storage_client(timeout_seconds: float) -> httpx.AsyncClient: + """The httpx client the object store is fetched with: our own, never the API client's. + + It carries no credential of ours (the link's authorization is in its query string, and nothing + must ride along to the store), refuses redirects rather than following them, and bounds every + stalled connect, read and write at the call's budget — the per-stall bound a total timeout alone + does not give. + """ + return httpx.AsyncClient(timeout=httpx.Timeout(timeout_seconds), follow_redirects=False) + + +@asynccontextmanager +async def _stream_resolved_url( + storage_client: httpx.AsyncClient, + uri: str, + download_url: str, + bounds: _FetchBounds, +) -> AsyncGenerator[ArtifactStream, None]: + """The bounded fetch of an already-resolved link — the half of `fetch_artifact` that + `download_artifacts` shares, its links coming from one bulk resolve ahead of the tasks. + """ + checked_url = _checked_url(uri, download_url, allow_http=bounds.allow_http) + try: + async with asyncio.timeout(bounds.timeout_seconds): + try: + async with storage_client.stream("GET", checked_url) as response: + refusal = _status_refusal(uri, response) + if refusal is not None: + raise refusal + declared = _declared_length(response.headers) + if declared is not None and declared > bounds.max_bytes: + msg = f"The artifact is {_format_mib(declared)}, over the {_format_mib(bounds.max_bytes)} cap." + raise ArtifactFetchError(msg, uri=uri, code="too_large", status=response.status_code) + yield ArtifactStream(uri, response, bounds.max_bytes) + except httpx.InvalidURL as exc: + # Not an `HTTPError`: httpx refuses the link at request build, before any transport. + msg = "The platform resolved the reference to a link that is not a valid absolute URL." + raise ArtifactFetchError(msg, uri=uri, code="unsupported_url") from exc + except httpx.TimeoutException as exc: + msg = f"Fetching the artifact timed out after {bounds.timeout_seconds}s." + raise ArtifactFetchError(msg, uri=uri, code="timeout") from exc + except httpx.HTTPError as exc: + msg = f"The artifact could not be fetched: {exc}." + raise ArtifactFetchError(msg, uri=uri, code="network") from exc + except TimeoutError as exc: + msg = f"Fetching the artifact timed out after {bounds.timeout_seconds}s." + raise ArtifactFetchError(msg, uri=uri, code="timeout") from exc + + +def fetch_artifact( + client: BulkResolveClient, + uri: str, + options: FetchArtifactOptions | None = None, +) -> AbstractAsyncContextManager[ArtifactStream]: + """A bounded stream for one reference, as an async context manager. + + The link is minted fresh through the bulk route, then fetched with redirects refused (a presigned + link has no reason to redirect, and one that does is refused rather than followed), no headers of + ours (the link carries its own authorization in the query string, and nothing must ride along to + the object store), a budget covering the headers and the body, and a byte cap checked against + `Content-Length` before a byte is read and again on every chunk. + + Only a `2xx` is yielded. Anything else raises `ArtifactFetchError` with a `code` (a redirect, a + refused or vanished object, a store fault, a declared oversize, a timeout, a network fault, an + unusable link, or the route's own per-reference refusal); a body that crosses the cap mid-stream + raises the same error type (`too_large`) out of the iteration. Cancelling the awaiting task + propagates `asyncio.CancelledError` as-is. Whole-request failures of the resolve step propagate + unchanged (`ApiResponseError`, `ApiUnreachableError`). + + Usage: + async with fetch_artifact(client, uri) as stream: + async for chunk in stream.aiter_bytes(): + ... + """ + return _fetch_artifact(client, uri, options) + + +@asynccontextmanager +async def _fetch_artifact( + client: BulkResolveClient, + uri: str, + options: FetchArtifactOptions | None = None, +) -> AsyncGenerator[ArtifactStream, None]: + """`fetch_artifact`'s body — resolve one reference fresh, then stream its link under the bounds.""" + bounds = _validated_bounds(options or FetchArtifactOptions()) + resolved = await resolve_artifacts(client, [uri]) + entry = resolved[0] + if entry.error is not None: + raise ArtifactFetchError(entry.error.detail, uri=uri, code=entry.error.code) + if entry.url is None: + msg = "The bulk resolve route answered an item with neither a link nor an error." + raise ArtifactOperationError(msg) + storage_client = _new_storage_client(bounds.timeout_seconds) + try: + async with _stream_resolved_url(storage_client, uri, entry.url, bounds) as stream: + yield stream + finally: + await storage_client.aclose() + + +def _validated_bounds(options: FetchArtifactOptions) -> _FetchBounds: + """The per-file bounds, refusing nonsense before anything is resolved.""" + _require_positive("max_bytes", options.max_bytes) + _require_positive("timeout_seconds", options.timeout_seconds) + return _FetchBounds(max_bytes=options.max_bytes, timeout_seconds=options.timeout_seconds, allow_http=options.allow_http) + + +def _require_positive(name: str, value: float) -> None: + """Refuse a bound that is not a positive, finite number.""" + if not math.isfinite(value) or value <= 0: + msg = f'"{name}" must be a positive number, got {value}.' + raise ArtifactOperationError(msg) + + +def _checked_url(uri: str, download_url: str, *, allow_http: bool) -> str: + """The boundary-approved link, or the typed refusal saying why it is not fetched.""" + try: + parsed = urlsplit(download_url) + except ValueError as exc: + msg = "The platform resolved the reference to a link that is not a valid absolute URL." + raise ArtifactFetchError(msg, uri=uri, code="unsupported_url") from exc + if not parsed.scheme or not parsed.netloc: + msg = "The platform resolved the reference to a link that is not a valid absolute URL." + raise ArtifactFetchError(msg, uri=uri, code="unsupported_url") + scheme = parsed.scheme.lower() + if scheme == "http" and not allow_http: + msg = ( + "The platform resolved the reference to a plain http link, which is refused by default; " + "pass allow_http=True to accept it (the local stack's object store hands out such links)." + ) + raise ArtifactFetchError(msg, uri=uri, code="plain_http_refused") + if scheme not in {"http", "https"}: + msg = f'The platform resolved the reference to a "{scheme}" link, which is not fetched; only http(s) links are.' + raise ArtifactFetchError(msg, uri=uri, code="unsupported_url") + if parsed.username or parsed.password: + msg = "The platform resolved the reference to a link carrying credentials, which is not fetched." + raise ArtifactFetchError(msg, uri=uri, code="unsupported_url") + return download_url + + +def _status_refusal(uri: str, response: httpx.Response) -> ArtifactFetchError | None: + """The typed refusal for a non-2xx store answer, or `None` when the status is a `2xx`.""" + status = response.status_code + if 300 <= status < 400: + msg = f"The resolved link redirected (HTTP {status}); redirects are not followed." + return ArtifactFetchError(msg, uri=uri, code="redirect_refused", status=status) + # The link is minted per call, so a 401/403 is the store refusing a fresh signature (clock skew, + # a signing misconfiguration) rather than an expired link. + if status in {401, 403}: + msg = f"The object store refused the resolved link (HTTP {status})." + return ArtifactFetchError(msg, uri=uri, code="store_refused", status=status) + if status in {404, 410}: + msg = f"The stored file is no longer available (HTTP {status})." + return ArtifactFetchError(msg, uri=uri, code="not_found", status=status) + if status < 200 or status >= 300: + msg = f"The object store answered HTTP {status} for the resolved link." + return ArtifactFetchError(msg, uri=uri, code="store_error", status=status) + return None + + +def _decoded_headers(headers: httpx.Headers) -> httpx.Headers: + """The store's headers as they describe the body we hand on. + + When httpx has decoded every coding the `Content-Encoding` lists, the stream is the decoded bytes: + the encoding and the encoded length are dropped, or a proxy relaying the response would label + plain bytes as compressed and give the wrong length. Any other encoding is passed through with the + still-encoded body it describes. + """ + raw = headers.get("content-encoding", "") + codings = [coding.strip().lower() for coding in raw.split(",")] + codings = [coding for coding in codings if coding and coding != "identity"] + if not codings or not all(coding in _DECODED_CODINGS for coding in codings): + return headers + decoded = httpx.Headers(headers) + del decoded["content-encoding"] + if "content-length" in decoded: + del decoded["content-length"] + return decoded + + +def _declared_length(headers: httpx.Headers) -> int | None: + """The body's declared length, or `None` when it is absent or unreadable.""" + raw = headers.get("content-length") + if raw is None: + return None + try: + return int(raw) + except ValueError: + return None + + +def _format_mib(byte_count: float) -> str: + """A byte count as MiB, for the sentence a person reads.""" + mib = byte_count / (1024 * 1024) + return f"{int(mib)} MiB" if mib.is_integer() else f"{mib:.1f} MiB" + + +# ── download_artifacts ─────────────────────────────────────────────── + + +@dataclass +class _DownloadBudget: + """The bytes the whole call may still write, and the flags that stop the tasks taking more. + + `committed` is the bytes of every file saved or being saved, a file in flight counting as the + larger of its declared length and what it has written. Reserving the declared length up front is + what stops parallel files from each passing the check and then all being cut together; a file + that is unlinked gives its share back. Every item is checked against the room that leaves, on + its own declared length: one refused item never skips a smaller one that still fits. + """ + + max_total_bytes: int + committed: int = 0 + credential_failure: ApiResponseError | None = None + + +async def download_artifacts( + client: ArtifactCapableClient, + *, + dir_path: str | Path, + run_id: str | None = None, + results: RunResults | None = None, + options: DownloadArtifactsOptions | None = None, +) -> DownloadArtifactsResult: + """Save a run's produced files under `dir_path`, and answer a produced verdict. + + Takes exactly one of `run_id` (the results are re-read, so a run is downloadable days later) or + `results` (a `RunResults` in hand); walks the requested scope with `collect_artifacts`; resolves + the whole set through the bulk route ahead of the tasks; then a bounded number of tasks + (`concurrency`, held by an `asyncio.Semaphore`) each fetch, create their file exclusively and + stream the body in, re-resolving any link that has expired by the time a task reaches it. The + embedded `public_url` is never used. Files are named by `artifact_filename` and never + overwritten; a failed or cancelled download unlinks its partial file. + + Returns one entry per reference, errors as values. It raises only when no verdict can be produced: + `RunStillRunningError` or `RunFailedError` for a run that has not completed, `FieldNotIncludedError` + when the results read did not carry the scope's key, `ScopeUnavailableError` when the key was + relayed as `None`, `ArtifactAuthenticationError` (carrying the verdict so far) when the resolve + route refuses the credential, `ArtifactOperationError` for an unusable directory or nonsense + bounds, and the transport and lifecycle errors of the reads it makes (`ApiResponseError` for a + deployment without the bulk route, `RunLifecycleUnavailableError` for a bare runner asked by id, + `ApiUnreachableError`). Cancelling the awaiting task raises `asyncio.CancelledError` out of here, + with every partial file unlinked first. + """ + opts = options or DownloadArtifactsOptions() + bounds = _validated_bounds(opts) + if opts.concurrency < 1: + msg = f'"concurrency" must be a positive integer, got {opts.concurrency}.' + raise ArtifactOperationError(msg) + _require_positive("max_total_bytes", opts.max_total_bytes) + + # An empty `run_id` is no selector at all, and is refused here rather than sent to the results + # read, which would answer a 404 about a run nobody named. + if bool(run_id) == (results is not None): + msg = "download_artifacts takes exactly one of `run_id` (the results are re-read) or `results` (a RunResults in hand)." + raise ArtifactOperationError(msg) + read_results = results if results is not None else await _read_completed_results(client, cast("str", run_id)) + walked = _scope_value(read_results, opts.scope) + + uris = collect_artifacts(walked) + if not uris: + return _assemble_verdict(opts.scope, []) + + target_dir = Path(dir_path).resolve() + try: + await asyncio.to_thread(target_dir.mkdir, parents=True, exist_ok=True) + except OSError as exc: + msg = f'The download directory "{target_dir}" cannot be created or used: {exc}.' + raise ArtifactOperationError(msg) from exc + + budget = _DownloadBudget(max_total_bytes=opts.max_total_bytes) + try: + resolved = await resolve_artifacts(client, uris) + except ApiResponseError as exc: + if not _is_credential_refusal(exc): + raise + verdict = _assemble_verdict(opts.scope, [_item_error(uri, None, "aborted", _SKIPPED_CREDENTIAL) for uri in uris]) + msg = f"The resolve route refused the credential ({exc.status}); no artifact was downloaded." + raise ArtifactAuthenticationError(msg, status=exc.status, verdict=verdict) from exc + + semaphore = asyncio.Semaphore(opts.concurrency) + storage_client = _new_storage_client(bounds.timeout_seconds) + try: + outcomes = await asyncio.gather( + *[ + _process_one( + client=client, + storage_client=storage_client, + semaphore=semaphore, + budget=budget, + bounds=bounds, + target_dir=target_dir, + index=index, + uri=uri, + entry=resolved[index], + ) + for index, uri in enumerate(uris) + ] + ) + finally: + await storage_client.aclose() + + verdict = _assemble_verdict(opts.scope, list(outcomes)) + if budget.credential_failure is not None: + status = budget.credential_failure.status + msg = f"The resolve route refused the credential ({status}) part-way through the download; the verdict so far is on this error." + raise ArtifactAuthenticationError(msg, status=status, verdict=verdict) from budget.credential_failure + return verdict + + +async def _read_completed_results(client: ArtifactCapableClient, run_id: str) -> RunResults: + """Read a run's results by id, turning a run that has not completed into its typed error.""" + state = await client.get_run_result(run_id) + if isinstance(state, RunResultRunning): + retry = state.retry_after_seconds + hint = f" — retry in {retry}s." if retry is not None else "." + msg = f"Run {run_id} is still running, so it has no artifacts to download yet{hint}" + raise RunStillRunningError(msg, run_id=run_id, retry_after_seconds=retry) + if isinstance(state, RunResultFailed): + raise RunFailedError(state.message, run_id=run_id, status=state.status) + completed: RunResultCompleted = state + return completed.result + + +def _scope_value(results: RunResults, scope: ArtifactScope) -> Any: + """The artifact the scope names, or the typed error saying why there is none to walk. + + The two readings of an absent value are distinct, and only one of them is the platform's answer: + a key the results read never carried is not in `model_fields_set` and is `FieldNotIncludedError`, + where a key relayed as `None` is a value and is `ScopeUnavailableError`. + """ + field_name = scope.results_field + if field_name not in results.model_fields_set: + raise FieldNotIncludedError(field_name) + value = getattr(results, field_name, None) + if value is None: + raise ScopeUnavailableError(scope, run_id=results.pipeline_run_id) + return value + + +def _is_credential_refusal(exc: ApiResponseError) -> bool: + """True for the resolve route refusing the caller's credential, which stops the whole download.""" + return exc.status in {401, 403} + + +def _item_error(uri: str, content_type: str | None, code: str, detail: str) -> DownloadedArtifact: + """One reference's failure, as a value on the verdict.""" + return DownloadedArtifact(uri=uri, path=None, content_type=content_type, size=None, error=ArtifactItemError(code=code, detail=detail)) + + +def _is_expired(expires_at: str | None) -> bool: + """True when a link is at or past its expiry, margin included; an unreadable stamp is not.""" + if expires_at is None: + return False + try: + parsed = datetime.fromisoformat(expires_at) + except ValueError: + return False + if parsed.tzinfo is None: + parsed = parsed.replace(tzinfo=UTC) + return (parsed - datetime.now(UTC)).total_seconds() <= _EXPIRY_MARGIN_SECONDS + + +async def _process_one( + *, + client: ArtifactCapableClient, + storage_client: httpx.AsyncClient, + semaphore: asyncio.Semaphore, + budget: _DownloadBudget, + bounds: _FetchBounds, + target_dir: Path, + index: int, + uri: str, + entry: ResolvedArtifact, +) -> DownloadedArtifact: + """One reference's whole pipeline — the semaphore's slot, a re-resolve on expiry, then the save.""" + async with semaphore: + if budget.credential_failure is not None: + return _item_error(uri, entry.content_type, "aborted", _SKIPPED_CREDENTIAL) + if entry.error is None and _is_expired(entry.expires_at): + try: + entry = (await resolve_artifacts(client, [uri]))[0] + except ApiResponseError as exc: + if _is_credential_refusal(exc): + budget.credential_failure = exc + return _item_error(uri, entry.content_type, "aborted", _SKIPPED_CREDENTIAL) + msg = f"The expired link could not be re-resolved: {exc}." + return _item_error(uri, entry.content_type, "resolve_failed", msg) + except (PipelineRequestError, ValidationError, ValueError) as exc: + # Anything else the re-resolve can fail with — an unreachable host, a malformed + # answer, a body that does not parse — is this one reference's error, never the + # whole download's: the other references already have their links. + msg = f"The expired link could not be re-resolved: {exc}." + return _item_error(uri, entry.content_type, "resolve_failed", msg) + if entry.error is not None: + return _item_error(uri, None, entry.error.code, entry.error.detail) + if entry.url is None: + msg = "The bulk resolve route answered an item with neither a link nor an error." + return _item_error(uri, entry.content_type, "resolve_failed", msg) + return await _save_one( + storage_client=storage_client, + budget=budget, + bounds=bounds, + target_dir=target_dir, + index=index, + uri=uri, + download_url=entry.url, + content_type=entry.content_type, + ) + + +class _TargetFile: + """A file created exclusively under the download directory, written in the background thread. + + Exclusive creation is what makes "never overwrite" true rather than merely likely: an + exists-check followed by a write would race a concurrent task. + """ + + def __init__(self, handle: BinaryIO, path: Path) -> None: + self._handle = handle + self.path = path + + async def write(self, chunk: bytes) -> None: + """Write one whole chunk; Python's buffered writer never returns a short write.""" + await asyncio.to_thread(self._handle.write, chunk) + + def close(self) -> None: + """Close the handle, surfacing a failed flush.""" + self._handle.close() + + def remove(self) -> None: + """Close and unlink, so nothing truncated is left under a final name. + + Synchronous on purpose: this runs on the cancellation path too, where another `await` could + be interrupted before the partial file is gone. + """ + try: + self._handle.close() + except OSError: + pass + try: + self.path.unlink(missing_ok=True) + except OSError: + pass + + +def _open_unique_file(target_dir: Path, base_name: str) -> _TargetFile: + """Create `base_name` under `target_dir` exclusively, suffixing the stem until a free name is found.""" + extension = _extension_of(base_name) + stem = base_name[: len(base_name) - len(extension)] + for attempt in range(_MAX_UNIQUE_ATTEMPTS): + candidate = target_dir / (base_name if attempt == 0 else f"{stem}-{attempt}{extension}") + try: + descriptor = os.open(candidate, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o644) + except FileExistsError: + continue + return _TargetFile(os.fdopen(descriptor, "wb"), candidate) + msg = f"Could not find a free filename for {base_name} in {target_dir}." + raise OSError(msg) + + +async def _save_one( + *, + storage_client: httpx.AsyncClient, + budget: _DownloadBudget, + bounds: _FetchBounds, + target_dir: Path, + index: int, + uri: str, + download_url: str, + content_type: str | None, +) -> DownloadedArtifact: + """Fetch one resolved link and write it under the download directory, within the total budget.""" + target: _TargetFile | None = None + written = 0 + reserved = 0 + share_taken = False + + def share() -> int: + return max(written, reserved) + + try: + async with _stream_resolved_url(storage_client, uri, download_url, bounds) as stream: + declared = _declared_length(stream.headers) + reserved = declared if declared is not None else 0 + if budget.committed + reserved > budget.max_total_bytes: + msg = ( + f"Saving this {_format_mib(reserved)} artifact would take the download past its " + f"{_format_mib(budget.max_total_bytes)} total limit." + ) + return _item_error(uri, content_type, "total_limit_exceeded", msg) + budget.committed += reserved + share_taken = True + + try: + # Synchronous on purpose: a cancellation landing inside a worker thread would leave a file + # the handler below cannot see, and an exclusive create is not worth a thread. + target = _open_unique_file(target_dir, artifact_filename(uri, content_type, index)) + except OSError as exc: + budget.committed -= share() + share_taken = False + msg = f"The file could not be created: {exc}." + return _item_error(uri, content_type, "write_failed", msg) + + async for chunk in stream.aiter_bytes(): + # Only the bytes past this file's reservation are new to the total. + growth = max(written + len(chunk), reserved) - share() + if budget.committed + growth > budget.max_total_bytes: + target.remove() + budget.committed -= share() + share_taken = False + msg = f"This artifact took the download past its {_format_mib(budget.max_total_bytes)} total limit." + return _item_error(uri, content_type, "total_limit_exceeded", msg) + budget.committed += growth + written += len(chunk) + try: + await target.write(chunk) + except OSError as exc: + target.remove() + budget.committed -= share() + share_taken = False + msg = f"The file could not be written: {exc}." + return _item_error(uri, content_type, "write_failed", msg) + except ArtifactFetchError as exc: + if target is not None: + target.remove() + if share_taken: + budget.committed -= share() + return _item_error(uri, content_type, exc.code, str(exc)) + except (httpx.HTTPError, OSError) as exc: + if target is not None: + target.remove() + if share_taken: + budget.committed -= share() + msg = f"The artifact could not be read: {exc}." + return _item_error(uri, content_type, "network", msg) + except asyncio.CancelledError: + # A cancelled download leaves nothing truncated behind, then lets the cancellation through. + if target is not None: + target.remove() + if share_taken: + budget.committed -= share() + raise + + try: + target.close() + except OSError as exc: + target.remove() + budget.committed -= share() + msg = f"The file could not be closed: {exc}." + return _item_error(uri, content_type, "write_failed", msg) + + # A body shorter than it declared gives the unused reservation back. + budget.committed -= share() - written + return DownloadedArtifact(uri=uri, path=str(target.path), content_type=content_type, size=written, error=None) + + +def _assemble_verdict(scope: ArtifactScope, artifacts: list[DownloadedArtifact]) -> DownloadArtifactsResult: + """The verdict over one walk: the saved paths in discovery order, and whether every one was saved.""" + saved_paths = [artifact.path for artifact in artifacts if artifact.error is None and artifact.path is not None] + return DownloadArtifactsResult( + scope=scope, + artifacts=artifacts, + saved_paths=saved_paths, + all_saved=len(saved_paths) == len(artifacts), + ) diff --git a/pipelex_sdk/client.py b/pipelex_sdk/client.py index 1395ec8..f076704 100644 --- a/pipelex_sdk/client.py +++ b/pipelex_sdk/client.py @@ -32,6 +32,17 @@ from pydantic_core import to_json from typing_extensions import override +from pipelex_sdk.artifact_models import ( + BulkResolvedStorageUrls, + DownloadArtifactsOptions, + DownloadArtifactsResult, + FetchArtifactOptions, + ResolvedArtifact, +) +from pipelex_sdk.artifacts import ArtifactStream +from pipelex_sdk.artifacts import download_artifacts as _download_artifacts_impl +from pipelex_sdk.artifacts import fetch_artifact as _fetch_artifact_impl +from pipelex_sdk.artifacts import resolve_artifacts as _resolve_artifacts_impl from pipelex_sdk.crate_models import ( CodegenRequest, CodegenResponse, @@ -95,6 +106,8 @@ if TYPE_CHECKING: from collections.abc import AsyncIterator + from contextlib import AbstractAsyncContextManager + from pathlib import Path from mthds.protocol.pipe_output import VariableMultiplicity from mthds.protocol.pipeline_inputs import PipelineInputs @@ -1112,6 +1125,50 @@ async def resolve_storage_url(self, uri: str) -> ResolvedStorageUrl: """Resolve a storage URI to a presigned URL — `POST /v1/resolve-storage-url`.""" return ResolvedStorageUrl.model_validate(await self._request_product("POST", "resolve-storage-url", body={"uri": uri})) + async def resolve_storage_urls_bulk(self, uris: list[str]) -> BulkResolvedStorageUrls: + """Resolve a list of storage URIs in one request — `POST /v1/resolve-storage-url/bulk`. + + The single route applied to a list. One item per reference, in request order, duplicates + included; a refused reference is a value on its item (`error`), and the request is a `200` + whenever every reference got a verdict. At most `BULK_RESOLVE_MAX_URIS` references per call + (a longer list is a `422`) — `resolve_artifacts` chunks a longer set. Served by the hosted + platform only: a deployment without the route answers a `404` `ApiResponseError`. + """ + return BulkResolvedStorageUrls.model_validate(await self._request_product("POST", "resolve-storage-url/bulk", body={"uris": uris})) + + async def resolve_artifacts(self, uris: list[str]) -> list[ResolvedArtifact]: + """Resolve a whole list of `pipelex-storage://` references through the bulk route, chunked at + its bound, answering one `ResolvedArtifact` per reference in request order with per-reference + failure as a value. The reading layer of the artifact stack: pair it with `collect_artifacts` + to mint fresh links for everything a run produced. See `docs/artifact-download.md`. + """ + return await _resolve_artifacts_impl(self, uris) + + def fetch_artifact(self, uri: str, options: FetchArtifactOptions | None = None) -> AbstractAsyncContextManager[ArtifactStream]: + """A bounded stream for one `pipelex-storage://` reference, as an async context manager: + resolved fresh, a timeout, redirects refused, the byte cap enforced mid-stream, no credentials + forwarded, the store's headers neutral. What `download_artifacts` and a same-origin proxy + share. See `docs/artifact-download.md`. + """ + return _fetch_artifact_impl(self, uri, options) + + async def download_artifacts( + self, + *, + dir_path: str | Path, + run_id: str | None = None, + results: RunResults | None = None, + options: DownloadArtifactsOptions | None = None, + ) -> DownloadArtifactsResult: + """Save a run's produced files under a directory — the download twin of `prepare_inputs`. + + Keyed on a `run_id` (the results are re-read, so it works days after the run) or a `RunResults` + in hand; walks the `main_stuff` scope by default, `working_memory` on request; resolves every + link fresh (never the embedded `public_url`); and returns a produced verdict, one entry per + reference, errors as values. See `docs/artifact-download.md`. + """ + return await _download_artifacts_impl(self, dir_path=dir_path, run_id=run_id, results=results, options=options) + async def upload(self, upload_input: UploadInput) -> UploadedFile: """Upload a base64 file — `POST /v1/upload`.""" body = upload_input.model_dump(mode="json", exclude_none=True) @@ -1541,19 +1598,41 @@ def _map_run_result_to_run_results(response: PipelexExecuteResult) -> RunResults `response.main_stuff` resolves the main output out of the returned working memory (and raises `MissingMainStuffError` if the run named no locatable main stuff), so the durable and blocking paths hand back the same `main_stuff` content shape. The already-parsed `pipe_output` model is - carried over as-is — no `.model_dump()` round-trip — so the full working memory stays typed - (blocking only; the hosted path has none). - - The usage pair (`tokens_usages` / `usage_assembly_error`) rides the execute response's - extension-open `pipe_output` as Pipelex extension fields. Lifting it onto the two top-level - fields here is what makes `.tokens_usages` read the same on the blocking and durable paths; - `RunResults` validates the raw records into `TokensUsageRecord`s on the way in. + carried over as-is — no `.model_dump()` round-trip — so the runner's whole envelope stays typed + (blocking only; the hosted path has none), and its `working_memory` is lifted onto the field of + that name, where the hosted path relays the artifact as its own key. The standard declares + `DictPipeOutputAbstract.working_memory` required, so that lift always carries a value here. + + The graph pair (`graph_spec` / `graph_assembly_error`), the usage pair (`tokens_usages` / + `usage_assembly_error`), the `pipe_io_artifacts` envelope and its `pipe_io_artifacts_error` all + ride the execute response's extension-open `pipe_output` as Pipelex extension fields. Lifting + each onto its top-level field here is what makes `.graph_spec`, `.tokens_usages` and + `.pipe_io_contracts` read the same on the blocking and durable paths; `RunResults` validates the + raw records into `TokensUsageRecord`s and the raw artifacts into the standard's models on the + way in. The runner carries the three I/O artifacts in one envelope — they share a key set and + are always built together — where the hosted results body relays them as three sibling keys; + `RunResults` follows the hosted shape and this unwraps the envelope onto it. A null envelope + leaves all three `None`, beside whatever `pipe_io_artifacts_error` says about why. + + Every lifted field is passed explicitly, so on this path each is in `model_fields_set` whether + or not the runner carried the key — the blocking path always answers, as the JS twin writes + `null` there; only the hosted path leaves a field unset when the body did not carry its key. """ pipe_output_extras: dict[str, Any] = response.pipe_output.model_extra or {} + pipe_io_artifacts: dict[str, Any] = {} + raw_pipe_io_artifacts = pipe_output_extras.get("pipe_io_artifacts") + if isinstance(raw_pipe_io_artifacts, dict): + pipe_io_artifacts = cast("dict[str, Any]", raw_pipe_io_artifacts) return RunResults( pipeline_run_id=response.pipeline_run_id, main_stuff=response.main_stuff, - graph_spec=None, + graph_spec=pipe_output_extras.get("graph_spec"), + graph_assembly_error=pipe_output_extras.get("graph_assembly_error"), + pipe_io_contracts=pipe_io_artifacts.get("pipe_io_contracts"), + input_form=pipe_io_artifacts.get("input_form"), + output_form=pipe_io_artifacts.get("output_form"), + pipe_io_artifacts_error=pipe_output_extras.get("pipe_io_artifacts_error"), + working_memory=response.pipe_output.working_memory, pipe_output=response.pipe_output, tokens_usages=pipe_output_extras.get("tokens_usages"), usage_assembly_error=pipe_output_extras.get("usage_assembly_error"), diff --git a/pipelex_sdk/errors.py b/pipelex_sdk/errors.py index b61fd04..3ae8d6a 100644 --- a/pipelex_sdk/errors.py +++ b/pipelex_sdk/errors.py @@ -1,7 +1,7 @@ """Pipelex SDK errors — the transport and response errors raised by `PipelexAPIClient`. These are the error classes the Pipelex hosted client adds on top of the `mthds` -protocol base. Both derive from the protocol base `PipelineRequestError` +protocol base. They derive from the protocol base `PipelineRequestError` (`mthds.protocol.exceptions`), mirroring `pipelex-sdk-js/src/errors.ts`: - `ApiUnreachableError` — the HTTP exchange never produced a response (DNS / connect @@ -21,10 +21,19 @@ stays in `mthds` — it belongs to the protocol `execute()` 202-degrade path, not the lifecycle — and is re-exported here so consumers have a single import home. +The artifact errors (`ArtifactOperationError` and its subclasses, plus `FieldNotIncludedError`) +are the download twin of the input-preparation family: they are raised only where the artifact +operations can produce no verdict at all, per-reference failure being a value on the verdict's +item. See `docs/artifact-download.md`. + The codegen tree errors (`CodegenError`, `CodegenLockError`) are not request errors at all: they are raised by `pipelex_sdk.codegen_writer`, `pipelex_sdk.codegen_check`, `pipelex_sdk.codegen_lock` and `pipelex_sdk.codegen_stamp` over bytes and a directory, so they derive from `Exception` rather than from the protocol base. + +`FieldNotIncludedError` is raised by `pipelex_sdk.usage` over an already-validated `RunResults` +whose body did not carry a key the operation needs. Like `MissingMainStuffError`, it reports a +results read that did not deliver what the caller reads, so it stays under the protocol base. """ from __future__ import annotations @@ -38,6 +47,7 @@ from mthds.runners.api.exceptions import RunStillRunningError as RunStillRunningError # ruff: ignore[useless-import-alias] if TYPE_CHECKING: + from pipelex_sdk.artifact_models import ArtifactScope, DownloadArtifactsResult from pipelex_sdk.runs import RunStatus from pipelex_sdk.validation_models import ValidationErrorItem @@ -262,3 +272,78 @@ class CodegenLockError(CodegenError): plain `CodegenError`: it is a containment violation, not corrupt state a writer may recover from by replacing the lock. """ + + +class ArtifactOperationError(PipelineRequestError): + """Base class for the failures the artifact operations raise on their own + (`fetch_artifact` / `download_artifacts`) — the download twin of `InputPreparationError`. + + Catch this to handle any artifact failure; catch a subclass to branch on the category. A + per-reference failure inside a `download_artifacts` verdict is a **value on the item**, never + one of these: the operation throws only when it can produce no verdict at all. The transport + failures of the resolve route (`ApiResponseError`, `ApiUnreachableError`) and the run-lifecycle + errors propagate unchanged, so they are not subclasses. Mirrors `pipelex-sdk-js`'s + `ArtifactOperationError` family. + """ + + +class ScopeUnavailableError(ArtifactOperationError): + """The scope `download_artifacts` was asked to walk is `None` on the run's results. + + The key WAS relayed — it is in `results.model_fields_set` — and its value is `None`, which is + the platform saying it has no such artifact for this run. Distinct from a key the results read + never carried, which is `FieldNotIncludedError`, and from an empty walk over a present scope, + which is a produced verdict with no artifacts. `scope` names the scope, `run_id` the run. + """ + + def __init__(self, scope: ArtifactScope, run_id: str) -> None: + msg = f'Run "{run_id}" carries no "{scope}" artifact to walk for produced files — the results relayed it as null.' + super().__init__(msg) + self.scope = scope + self.run_id = run_id + + +class ArtifactFetchError(ArtifactOperationError): + """One reference could not be turned into a bounded stream by `fetch_artifact`. + + `code` says why, in a closed vocabulary the download verdict shares for its per-item errors: + the resolve route's own per-reference codes (`invalid_storage_uri`, `forbidden`), then the + fetch boundary's — `unsupported_url`, `plain_http_refused`, `redirect_refused`, `store_refused` + (a 401/403 from the object store), `not_found` (404/410), `store_error` (any other non-2xx), + `too_large`, `timeout`, `network`. `status` is the store's HTTP status when one was received. + `download_artifacts` never lets this escape: it becomes the item's `error`. + """ + + def __init__(self, message: str, uri: str, code: str, status: int | None = None) -> None: + super().__init__(message) + self.uri = uri + self.code = code + self.status = status + + +class ArtifactAuthenticationError(ArtifactOperationError): + """The resolve route refused the caller's credential (`401` / `403`) during a download. + + No further reference can be resolved with it, so the download stops — but the files already + saved are real, and `verdict` carries the result as it stood: every item saved before the + refusal, and the rest marked `aborted`. `status` is the route's status; the wrapped + `ApiResponseError` is reachable through `__cause__`. + """ + + def __init__(self, message: str, status: int, verdict: DownloadArtifactsResult) -> None: + super().__init__(message) + self.status = status + self.verdict = verdict + + +class FieldNotIncludedError(PipelineRequestError): + """A `RunResults` field this operation needs was not carried by the results body it was read from. + + Raised when the field is absent from `results.model_fields_set` — the body did not carry the key — + as opposed to relayed as `None`, which is a value. Carries the field's name in `field_name`. + """ + + def __init__(self, field_name: str) -> None: + self.field_name = field_name + msg = f"RunResults field `{field_name}` was not in the results body: the read did not carry it" + super().__init__(msg) diff --git a/pipelex_sdk/product_models.py b/pipelex_sdk/product_models.py index cbd4339..b53ffc6 100644 --- a/pipelex_sdk/product_models.py +++ b/pipelex_sdk/product_models.py @@ -19,7 +19,7 @@ import json from enum import StrEnum -from typing import Any +from typing import Any, cast from pydantic import BaseModel, ConfigDict, Field, TypeAdapter, ValidationError, field_serializer, field_validator @@ -81,6 +81,39 @@ def _is_blank(content: str) -> bool: return not content.strip() +def _decode_method_source(source: str) -> Any: + """Decode a stored source string, with every way the decoder can refuse one as a `ValueError`. + + The two readers of a stored source apply different shape rules to what comes back — see + `method_source_to_contents` for why — but they must fail identically on text the decoder + cannot take at all, so that rule lives here and in one place. `json.loads` refuses in two + ways: `JSONDecodeError` for malformed text, and `RecursionError` — which is NOT a + `ValueError` — for a source nested past a depth that is an interpreter build constant, + because `json.loads` recurses where `JSON.parse` iterates. Converting both gives each + caller the failure its own contract promises: `MethodData`'s validator turns this + `ValueError` into the `ValidationError` a caller catches (pydantic converts only + `ValueError`), and `method_source_to_contents` reads it as "not the catalog form". + """ + try: + return json.loads(source) + except json.JSONDecodeError as exc: + msg = f"Method file source is not valid JSON; expected {_METHOD_FILES_SHAPE}." + raise ValueError(msg) from exc + except RecursionError as exc: + msg = f"Method file source is nested too deeply to decode; expected {_METHOD_FILES_SHAPE}." + raise ValueError(msg) from exc + + +def _is_catalog_entry(entry: object) -> bool: + """The catalog gate the platform and `@pipelex/sdk` both apply: an object carrying both keys. + + Key presence alone, with no check on either value's type — mirroring + `pipelex_platform.services.method_resolution.method_source_to_contents` and + `@pipelex/sdk`'s `isFileEntry`. Each entry's value types are filtered afterwards, per entry. + """ + return isinstance(entry, dict) and "name" in entry and "content" in entry + + def parse_method_files(source: str | None) -> list[MethodFile]: """Parse the catalog wire string into method files. @@ -90,18 +123,15 @@ def parse_method_files(source: str | None) -> list[MethodFile]: Raises: ValueError: For anything else — a non-array JSON value, an entry that is not a - `{name: str, content: str}` object, or unparseable text. Reached through - `MethodData`'s validator, this surfaces as a `pydantic.ValidationError`, the - same way any other malformed response body fails here. + `{name: str, content: str}` object, unparseable text, or a source nested more + deeply than the decoder can descend. Reached through `MethodData`'s validator, + this surfaces as a `pydantic.ValidationError`, the same way any other malformed + response body fails here. """ if source is None or _is_blank(source): return [] - try: - parsed = json.loads(source) - except json.JSONDecodeError as exc: - msg = f"Method file source is not valid JSON; expected {_METHOD_FILES_SHAPE}." - raise ValueError(msg) from exc + parsed = _decode_method_source(source) try: files = _METHOD_FILES_ADAPTER.validate_python(parsed) @@ -126,6 +156,91 @@ def serialize_method_files(files: list[MethodFile]) -> str: return json.dumps([{"name": file.name, "content": file.content} for file in kept]) +def method_source_to_contents(mthds: str | None) -> list[str]: + """Read a stored method's polymorphic `mthds` source as the bundle contents a call takes. + + `MethodData.mthds` is polymorphic at rest. The webapp editor writes the catalog + file-array — the JSON `[{name, content}]` string the gate below recognizes — while a + row written before that editor, or by hand, holds the `.mthds` source itself as plain + text. A caller holding one of those strings cannot tell which it has, so this resolves + it to the `list[str]` that `run`, `start` and `validate` take as `mthds_contents`. + + An empty list means the method carries no MTHDS source — a row that exists but is not + runnable yet. It is never a failure to read one: this function does not raise. + + **The catalog gate is key presence, and each entry's types are filtered afterwards.** An + array is the catalog form when every entry is an object carrying both a `name` and a + `content` key, whatever those values hold; an entry whose `content` is not a non-blank + string is then dropped and its siblings are kept. So + `[{"name": "a", "content": "x"}, {"name": "b", "content": 1}]` reads as `["x"]`. + + That rule is not this SDK's preference — it reproduces, deliberately and character for + character, what the server does with the same stored row: the platform's own + `method_source_to_contents` (`pipelex_platform/services/method_resolution.py`, the resolver + that expands a `method_id` run) and `@pipelex/sdk`'s `methodSourceToContents`. A client-side + reader of a server-stored field exists so a caller can do locally what the platform does + with that row, and a reader that disagrees with the code which actually runs the method is a + reader that lies, however defensible its own rule. One stored method therefore has one + observable reading across `method_id`, `@pipelex/sdk` and this SDK. This was ruled against a + stricter reading that took a partly malformed array as a bundle; `docs/architecture.md` carries + the ruling and the three readers it compares. + + The cost of the ruled rule is real and is not this function's to fix: a file whose `content` + is not a string vanishes from the bundle without a word. That is a property of the format's + decoder, shared by all three readers, so it is fixed once in the platform rather than three + times in its clients. + + This is consequently NOT a pure delegation to `parse_method_files`. That function reads + `MethodData.python`, whose entries are a typed `[{name: str, content: str}]` catalog with no + second at-rest shape to tell apart, so it stays strict and rejects what this accepts. The two + share what must not drift — the decoder (`_decode_method_source`) and blankness + (`_is_blank`) — and nothing else. + + Lesser divergences from the JS twin follow from the decoder rather than from this function, + and they flip in both directions: `json.loads` accepts `NaN` and `Infinity`, which + `JSON.parse` refuses, and refuses an integer literal past CPython's digit cap, which + `JSON.parse` accepts; blankness here is Python's `str.strip`, not ECMAScript's, so the two + disagree on a source of only U+FEFF and on one of only U+0085. `mthds.protocol.method_files` + closes all of these and is the adoption target once its exception base class is settled. + + Args: + mthds: The stored source, as `MethodData.mthds` carries it. `None` is tolerated + although the model types it `str`, so a contract-violating response body reads + as "no source" rather than raising — the twin's own defensive guard. + + Returns: + One content string per bundle file: the non-blank string contents of the catalog + file-array, or the whole source as a single bundle when it is not that form. + """ + if mthds is None or _is_blank(mthds): + # A blank source is no source, not a bundle of whitespace — and `None` although the + # model types the field `str`, the twin's own defensive guard. `_is_blank` is the + # parser's predicate rather than a second one, so the two readings of one stored + # source cannot drift apart on what blank means. + return [] + + try: + parsed: Any = _decode_method_source(mthds) + except ValueError: + # Text the decoder cannot take — not JSON at all, or nested past what it can descend — + # so the whole source is one legacy bare bundle. A `.mthds` file may legally open with a + # digit or a brace, which is why a decode failure is a bundle rather than an error. + return [mthds] + + if isinstance(parsed, list) and all(_is_catalog_entry(entry) for entry in cast("list[Any]", parsed)): + # Vacuously true for `[]`, which is the webapp editor's "no files" sentinel and yields + # no contents — read as a bundle it would send the two characters `[]` to the runner. + entries = cast("list[dict[str, Any]]", parsed) + contents: list[str] = [] + for entry in entries: + content: Any = entry["content"] + if isinstance(content, str) and not _is_blank(content): + contents.append(content) + return contents + + return [mthds] + + class MethodData(BaseModel): """One saved method record.""" @@ -133,7 +248,11 @@ class MethodData(BaseModel): method_id: str name: str - #: The `.mthds` bundle source. + #: The `.mthds` bundle source, polymorphic at rest and left exactly as the platform stored it: + #: the catalog `[{name, content}]` array the webapp editor writes, or a bare bundle as plain + #: text. Read it with `method_source_to_contents`, which resolves either shape to the + #: `mthds_contents` a run or a validate takes exactly as the platform's own resolver reads + #: the same row; unlike `python`, it is not converted here. mthds: str org_id: str created_by_user_id: str diff --git a/pipelex_sdk/runs.py b/pipelex_sdk/runs.py index aca84f7..818c80a 100644 --- a/pipelex_sdk/runs.py +++ b/pipelex_sdk/runs.py @@ -16,13 +16,17 @@ shapes still exist in `mthds-python`; that duplication is deliberate and is removed from `mthds-python` in Phase 6, leaving these as the single home. -Two things in this module are deliberately NOT owned here, and both reuse rather -than redefine. `RunResults.pipe_output` is typed with the protocol's own -`DictPipeOutputAbstract` wire model from `mthds` — a shared wire contract the -`pipelex` runtime also builds on, not a lifecycle concept. `TokensUsageRecord` -mirrors the runtime's own record: inference accounting is a Pipelex runtime -extension the MTHDS Protocol does not model, so the hosted API is what pins that -wire contract; this SDK follows the shape, it does not define it. +Three things in this module are deliberately NOT owned here, and all reuse rather +than redefine. `RunResults.pipe_output` and `RunResults.working_memory` are typed +with the protocol's own `DictPipeOutputAbstract` and `DictWorkingMemoryAbstract` +wire models from `mthds` — a shared wire contract the `pipelex` runtime also builds +on, not a lifecycle concept. The three I/O artifacts +on `RunResults` (`pipe_io_contracts`, `input_form`, `output_form`) are the standard's +own, typed by importing `mthds.protocol` exactly as the validate report does — one +declaration per language, nothing to drift from. `TokensUsageRecord` mirrors the +runtime's own record: inference accounting is a Pipelex runtime extension the MTHDS +Protocol does not model, so the hosted API is what pins that wire contract; this SDK +follows the shape, it does not define it. Wire contract mirrors `pipelex-platform`: POST /v1/start -> RunResultStart (start, 202) @@ -36,8 +40,11 @@ from enum import StrEnum from typing import TYPE_CHECKING, Annotated, Any, Literal, TypeAlias +from mthds.protocol.input_form import InputForm from mthds.protocol.models import RunResultStart -from mthds.runners.api.models import DictPipeOutputAbstract +from mthds.protocol.output_form import OutputForm +from mthds.protocol.pipe_io_contracts import PipeIOContracts +from mthds.runners.api.models import DictPipeOutputAbstract, DictWorkingMemoryAbstract from pydantic import BaseModel, ConfigDict, Field if TYPE_CHECKING: @@ -195,7 +202,9 @@ class TokensUsageRecord(BaseModel): #: Computed USD cost of this call. `None` when the model has no rate table at all (own-GPU, #: mock, dry run); `0` means a rate table existed and priced the call at zero. The underlying #: rate table never crosses the wire and there is no run-level aggregate — sum the records. - cost: float | None = None + #: Strict and finite: a string, a bool or a NaN on the wire fails the whole results body's + #: parse rather than reaching a sum as a number. + cost: float | None = Field(default=None, strict=True, allow_inf_nan=False) #: ISO 8601 start of the call. started_at: str | None = None #: ISO 8601 end of the call. Duration is derivable from the pair and deliberately not shipped. @@ -212,25 +221,90 @@ class RunResults(BaseModel): run's `main_stuff_name`, so both paths deliver the same content shape. Consumers read `main_stuff` directly — no shape-guessing. A completed run that cannot deliver a main stuff raises `MissingMainStuffError`. Extension-open - (`extra="allow"`): any other server artifact (e.g. the hosted `working_memory`) - is preserved without being named by the SDK. + (`extra="allow"`): any other server artifact the SDK does not name is preserved + on `model_extra` rather than dropped. + + Every field but the first two is optional, and two readings of an optional field + are distinct on purpose. A key the hosted body did not carry is not in + `model_fields_set` and reads `None`; a key relayed as `null` is in the set and + reads `None` too. That is how a reader tells "the platform relayed no such key" + from "the platform relayed null", where the JS twin reads `undefined` against + `null`. The blocking path always answers for every field, so each is set there. + Every field is walked on `docs/run-results.md`. """ model_config = ConfigDict(extra="allow") pipeline_run_id: str #: The resolved main output content — always present for a completed run. Typed `Any` because the - #: content is polymorphic (a list output renders to a top-level array, a structured output to an - #: object) and may be a valid falsy value (empty list, `0`); it is never absent for a completed run. + #: content is polymorphic (a structured output is an object of the concept's fields, a multiple + #: output the `{"items": [...]}` envelope the runtime's `ListContent` serialises to, a native is + #: wrapped too — `{"text": ...}`, `{"number": ...}`) and may be a valid empty value (an empty + #: `items`, an empty `text`); it is never absent for a completed run. main_stuff: Any - #: Method graph spec (`graphspec.json`); `None` if missing mid-write or on the bare-runner path. + #: The executed graph — the same document a local run writes as `graphspec.json`: `meta.mode` + #: `"live"`, one node per pipe with its status, its timings and its own usage. It reaches the + #: client on both paths: the hosted path relays the `graphspec.json` artifact verbatim, and on + #: the blocking path the SDK lifts it off `pipe_output`. `None` when the runner assembled no + #: graph (see `graph_assembly_error`) or, on the hosted path, when the artifact was not yet + #: written. Typed `Any` on purpose — no published Python package declares the graph spec, so a + #: type here could only be a copy that drifts. graph_spec: Any = None + #: Non-`None` when the runner's graph assembly failed for the run — the graph's twin of + #: `usage_assembly_error`, and the only thing that separates "the graph broke" from "this run + #: produced no graph". Lifted off `pipe_output` on the blocking path; the hosted results body + #: carries nothing of the kind yet, so on that path the key is absent (not in `model_fields_set`) + #: until the platform writes and relays it. + graph_assembly_error: str | None = None + #: Per-pipe input/output contracts for the library the run executed against, keyed by namespaced + #: `pipe_ref` (`domain.code`) — the standard's `PipeIOContracts`, the same artifact `POST + #: /v1/validate` reports and the same one a local run writes beside its graph as + #: `pipe_io_contracts.json`. Imported from `mthds.protocol` rather than restated, under the + #: standing ruling that keeps the standard's artifacts declared once per language + #: (`docs/architecture.md`). It is what says what a `graph_spec` node's data IS: the graph + #: carries the values, this carries their concepts and their schemas. Read it together with + #: `output_form` — a renderer takes the pair or neither. A closed shape: a member the pinned + #: `mthds` does not define fails the parse of the whole results body with pydantic's + #: `ValidationError` (`docs/run-results.md` says what that costs on each path). The hosted + #: results body relays it as its own key, so it is set on that path: `None` for a run whose + #: artifact was not written. + pipe_io_contracts: PipeIOContracts | None = None + #: Per-pipe input-form descriptors for that same library — the standard's `InputForm`, keyed + #: over the same `pipe_ref` set as `pipe_io_contracts`, describing each declared input as a + #: typed field rather than a schema. It is what lets a rendered run show its own inputs as + #: values; a renderer treats it as optional even when it has the other two. `None` on the same + #: terms as `pipe_io_contracts`. + input_form: InputForm | None = None + #: Per-pipe OUTPUT-form descriptors for that same library — the standard's `OutputForm`, the twin + #: of `input_form` on the other side of the pipe, keyed over the same `pipe_ref` set. The + #: descriptor says what the result IS and the contract's `output.json_schema` names the property + #: its payload arrives under, which together are everything a renderer needs to lay a run's + #: result out without inspecting the value. `None` on the same terms as `pipe_io_contracts`. + output_form: OutputForm | None = None + #: Non-`None` when the runner's build of the three I/O artifacts failed for the run — their twin + #: of `graph_assembly_error`, and the only thing that separates "describing the data broke" from + #: "this run described none". Lifted off `pipe_output` on the blocking path; absent on the hosted + #: path until the platform writes and relays it. + pipe_io_artifacts_error: str | None = None + #: The run's whole working memory — every named stuff it held when it finished, the inputs it was + #: given and the intermediates it produced as well as the main output, as the standard's + #: `DictWorkingMemoryAbstract` (`root` keyed by stuff name, `aliases` mapping a role such as + #: `main_stuff` onto one of those names). It reaches the client on both paths: the hosted path + #: relays the `working_memory.json` artifact as its own key, and on the blocking path the SDK + #: lifts it off `pipe_output.working_memory`, which the standard declares as a required field — + #: so it is always set there. `None` on the hosted path when the platform relayed `null` (the + #: artifact was not written); absent from `model_fields_set`, and `None` too, when the body did + #: not carry the key at all. Extension-open at every level, so a runner's per-stuff extras + #: (`stuff_code`, `stuff_name`, …) ride `model_extra` rather than being dropped. + working_memory: DictWorkingMemoryAbstract | None = None #: Bare runner's native pipe output — the full working memory, blocking-execute path only; - #: `None` on the hosted path. Supplementary to `main_stuff`, which is already resolved out of - #: it; kept for consumers that need the whole working memory. Extension-open, so the Pipelex - #: extension fields the runner rides on it stay reachable via `model_extra` — including the - #: usage pair, in its **raw** form. Read `tokens_usages` below instead: same data, validated - #: into records, and present on the hosted path too (where `pipe_output` is `None`). + #: `None` on the hosted path. Supplementary: `main_stuff`, the graph pair, the three I/O + #: artifacts, the working memory and the usage pair are all lifted out of it onto fields that + #: read the same on both paths; kept for consumers that want the runner's envelope exactly as it + #: arrived, `pipeline_run_id` and all. Extension-open, so the Pipelex + #: extension fields the runner rides on it stay reachable via `model_extra` in their **raw** + #: form — the usage pair, the graph pair, the `pipe_io_artifacts` envelope. Read the lifted + #: fields instead: same data, validated, and present on the hosted path too. pipe_output: DictPipeOutputAbstract | None = None #: Per-call usage records — token counts by category, computed `cost` in USD, model id — for #: LLM and img-gen/extract/search calls alike. On the hosted path this is the diff --git a/pipelex_sdk/usage.py b/pipelex_sdk/usage.py new file mode 100644 index 0000000..553b386 --- /dev/null +++ b/pipelex_sdk/usage.py @@ -0,0 +1,247 @@ +"""`summarize_usage` — one run-level reading of a run's usage pair. + +A completed run reports usage as `RunResults.tokens_usages` (one record per inference call) +beside `RunResults.usage_assembly_error`, and the wire carries no run-level aggregate. This +module folds the pair into a single summary under the rules `docs/run-usage.md` states, so +every consumer reads the same totals instead of re-deriving them: + +- a `None` cost is unrated (the model has no rate table), a `0` cost is priced at zero; +- `input` and `output` are the only additive token categories (`input_cached` is a subset of + `input`, and summing every category double-counts); +- a `None` list means usage is unavailable, and only `usage_assembly_error` says it broke; +- an empty list is a run that did no inference, which costs `0`, not `None`. + +Pure: no I/O, no client, and the input is never mutated. It is the Python twin of +`@pipelex/sdk`'s `summarizeUsage` and reports the same summary shape, with two deliberate +divergences. Where the JS reads an absent `tokens_usages` as `unavailable`, this raises +`FieldNotIncludedError`, because Python tells an absent key from a relayed `null` through +`model_fields_set` and answering "nothing is known" for a key the read never carried would turn +a gap in the read into a fact about the run. And where the JS guards every read with `typeof`, +because its records are relayed JSON nothing validated, this reads the parsed fields directly: +`cost` is validated strict and finite at the client boundary, so a record whose `cost` is not a +number never reaches the fold — it fails the whole results body's parse. +""" + +from __future__ import annotations + +from dataclasses import dataclass +from enum import StrEnum +from typing import TYPE_CHECKING + +from pydantic import BaseModel, ConfigDict + +from pipelex_sdk.errors import FieldNotIncludedError + +if TYPE_CHECKING: + from collections.abc import Sequence + + from pipelex_sdk.runs import RunResults, TokensUsageRecord + +#: The one `RunResults` field the fold cannot do without, named once so the membership test in +#: `model_fields_set` and the error that reports its absence can never drift apart. +_USAGE_RECORDS_FIELD = "tokens_usages" + + +class UsageSummaryState(StrEnum): + """Which of the three readings of `tokens_usages` the summary describes.""" + + #: A non-empty list: the totals are folded from it. + RECORDS = "records" + #: `[]`: usage assembly ran and no inference happened, so the cost is `0` and the token + #: totals are zero. + NO_INFERENCE = "no_inference" + #: The list was relayed as `None`: usage assembly was off, it broke (then `assembly_error` + #: is non-`None`), or on the hosted path the run was delivered before the usage artifact + #: existed. Nothing is known, so the cost and the token totals are `None`. + UNAVAILABLE = "unavailable" + + +class UsageTokenTotals(BaseModel): + """The two additive token totals. + + Each is the sum of that category over the records that reported it, and `None` when no + record reported it at all — which is different from a reported `0`. + """ + + model_config = ConfigDict(extra="forbid") + + input: int | None + output: int | None + + +class PipeUsageSummary(BaseModel): + """One pipe's share of a run's usage — the same fold as the run level, over its calls only.""" + + model_config = ConfigDict(extra="forbid") + + #: The pipe that made the calls; `None` groups the calls the runtime did not attribute. + pipe_code: str | None + #: Sum of the priced calls' costs in USD; `None` when none of this pipe's calls was priced. + total_cost_usd: float | None + #: True when this pipe mixes priced and unrated calls, so `total_cost_usd` is a lower bound. + cost_partial: bool + tokens: UsageTokenTotals + #: Number of inference calls this pipe made. + calls: int + + +class UsageSummary(BaseModel): + """A run's usage pair folded into one null-aware reading.""" + + model_config = ConfigDict(extra="forbid") + + state: UsageSummaryState + #: Sum of the priced calls' costs in USD. `0` for `no_inference`. `None` for `records` when + #: no call was priced (unrated), and for `unavailable`, where nothing is known — read + #: `state` first to tell the two apart. + total_cost_usd: float | None + #: True when priced and unrated calls are mixed, so `total_cost_usd` covers the priced calls + #: only and is a lower bound. Never true outside `records`. + cost_partial: bool + tokens: UsageTokenTotals + #: Number of usage records summarized — one per inference call; `0` outside `records`. + calls: int + #: `usage_assembly_error` as relayed: non-`None` when the runner's usage assembly failed, + #: which is the only thing that separates a broken assembly from an `unavailable` one that + #: was off. + assembly_error: str | None + #: Per-pipe rollup, most expensive first: priced pipes by cost descending, then unrated + #: pipes, with ties broken by call count (descending) and then by pipe code. Empty outside + #: `records`. + by_pipe: list[PipeUsageSummary] + + +@dataclass(frozen=True) +class _UsageFold: + """The null-aware part of the summary a set of records folds to, at any level.""" + + total_cost_usd: float | None + cost_partial: bool + tokens: UsageTokenTotals + + +def summarize_usage(results: RunResults) -> UsageSummary: + """Fold a run's usage pair into one summary: run totals, the call count, the assembly error + and a per-pipe rollup. See `docs/run-usage.md` for the rules it applies. + + Raises `FieldNotIncludedError` when the results body never carried `tokens_usages` — the key + is absent from `results.model_fields_set` — because a read that did not deliver the key says + nothing about the run, and is not a run with no usage to report. A key relayed as `None` IS a + value and reads as `unavailable`. + + Pre-contract records, relayed verbatim from artifacts written before the usage contract, + carry no `cost` and no `pipe_code`: they count as unrated and unattributed. The legacy + `job_metadata` / `unit_costs` fields are relics and are never read. + """ + if _USAGE_RECORDS_FIELD not in results.model_fields_set: + raise FieldNotIncludedError(_USAGE_RECORDS_FIELD) + + records = results.tokens_usages + assembly_error = results.usage_assembly_error + + if records is None: + return UsageSummary( + state=UsageSummaryState.UNAVAILABLE, + total_cost_usd=None, + cost_partial=False, + tokens=UsageTokenTotals(input=None, output=None), + calls=0, + assembly_error=assembly_error, + by_pipe=[], + ) + + if not records: + return UsageSummary( + state=UsageSummaryState.NO_INFERENCE, + total_cost_usd=0.0, + cost_partial=False, + tokens=UsageTokenTotals(input=0, output=0), + calls=0, + assembly_error=assembly_error, + by_pipe=[], + ) + + run_fold = _fold_records(records) + return UsageSummary( + state=UsageSummaryState.RECORDS, + total_cost_usd=run_fold.total_cost_usd, + cost_partial=run_fold.cost_partial, + tokens=run_fold.tokens, + calls=len(records), + assembly_error=assembly_error, + by_pipe=_roll_up_by_pipe(records), + ) + + +def _fold_records(records: Sequence[TokensUsageRecord]) -> _UsageFold: + """Null-aware totals over a non-empty set of records.""" + priced_sum = 0.0 + any_priced = False + any_unrated = False + input_sum: int | None = None + output_sum: int | None = None + + for record in records: + if record.cost is None: + any_unrated = True + else: + priced_sum += record.cost + any_priced = True + + by_category = record.nb_tokens_by_category + if by_category is not None: + # Only the two joined totals are additive; every other category is a subset of one. + input_count = by_category.get("input") + if input_count is not None: + input_sum = (input_sum or 0) + input_count + output_count = by_category.get("output") + if output_count is not None: + output_sum = (output_sum or 0) + output_count + + total_cost_usd: float | None = None + if any_priced: + total_cost_usd = priced_sum + + return _UsageFold( + total_cost_usd=total_cost_usd, + cost_partial=any_priced and any_unrated, + tokens=UsageTokenTotals(input=input_sum, output=output_sum), + ) + + +def _roll_up_by_pipe(records: Sequence[TokensUsageRecord]) -> list[PipeUsageSummary]: + """Group the records by `pipe_code` (a `None` key gathers the unattributed calls) and sort.""" + groups: dict[str | None, list[TokensUsageRecord]] = {} + for record in records: + groups.setdefault(record.pipe_code, []).append(record) + + rows: list[PipeUsageSummary] = [] + for pipe_code, pipe_records in groups.items(): + pipe_fold = _fold_records(pipe_records) + rows.append( + PipeUsageSummary( + pipe_code=pipe_code, + total_cost_usd=pipe_fold.total_cost_usd, + cost_partial=pipe_fold.cost_partial, + tokens=pipe_fold.tokens, + calls=len(pipe_records), + ), + ) + + rows.sort(key=_pipe_row_sort_key) + return rows + + +def _pipe_row_sort_key(row: PipeUsageSummary) -> tuple[bool, float, int, bool, str]: + """Cost descending with unrated pipes last, then calls descending, then pipe code (`None` last). + + The cost and pipe-code slots are constant within the group the flag beside them selects, so + the placeholder each takes for a `None` never decides an order. + """ + return ( + row.total_cost_usd is None, + -(row.total_cost_usd or 0.0), + -row.calls, + row.pipe_code is None, + row.pipe_code or "", + ) diff --git a/pyproject.toml b/pyproject.toml index 9e6fc3b..c653503 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -1,6 +1,6 @@ [project] name = "pipelex-sdk" -version = "0.10.0" +version = "0.10.1" description = "The Python client for the Pipelex hosted API — the MTHDS Protocol surface plus the durable run lifecycle and the Pipelex product surface, built on the `mthds` protocol base." authors = [{ name = "Evotis S.A.S.", email = "oss@pipelex.com" }] maintainers = [{ name = "Pipelex staff", email = "oss@pipelex.com" }] @@ -45,9 +45,13 @@ Repository = "https://github.com/Pipelex/pipelex-sdk-python" Documentation = "https://github.com/Pipelex/pipelex-sdk-python" Changelog = "https://github.com/Pipelex/pipelex-sdk-python/blob/main/CHANGELOG.md" +# `[tool.pytest]` is the native table pytest 9 reads (`ini_options` is its backwards-compatibility +# form, and the two together are an error). `testpaths` is the unit tree alone, so a bare `pytest` +# (and `make test` / `make agent-test` with it) never collects the live e2e legs — those are +# `make e2e-test`, which names `tests/e2e` explicitly and skips itself without credentials. [tool.pytest] minversion = "9.0" -testpaths = ["tests"] +testpaths = ["tests/unit"] pythonpath = ["."] [tool.mypy] diff --git a/tests/e2e/test_artifacts_e2e.py b/tests/e2e/test_artifacts_e2e.py new file mode 100644 index 0000000..1410939 --- /dev/null +++ b/tests/e2e/test_artifacts_e2e.py @@ -0,0 +1,133 @@ +"""The artifact round trip, exercised against a LIVE hosted platform (no mocks). + +Run it with `make e2e-test` against a platform that serves upload, the durable run lifecycle and the +bulk resolve route (`POST /v1/resolve-storage-url/bulk`), with an API key set for it: + + PIPELEX_E2E_BASE_URL=https://api-dev.pipelex.com PIPELEX_API_KEY=plx_sk_… make e2e-test + +The whole module skips when either variable is unset, so the unit suite's `make agent-test` never +reaches it — and `make agent-test` does not collect this directory at all. + +What the unit suite cannot prove: that the wire shapes the SDK composes — the bulk request, the +per-item answer, the presigned link the store actually honours, the working-memory echo of an +uploaded input — are the ones a real platform produces. Every mock agrees with the client about the +field names; only a live exchange settles whether the platform does. The leg is the JS twin's +(`pipelex-sdk-js/tests/e2e/artifacts.e2e.ts`): prepare a small file, run a pass-through with no +inference, download over `working_memory`, and read the bytes back equal. +""" + +from __future__ import annotations + +import asyncio +import os +import time +from pathlib import Path +from typing import TYPE_CHECKING + +import pytest + +from pipelex_sdk.artifact_models import ArtifactScope, DownloadArtifactsOptions +from pipelex_sdk.artifacts import collect_artifacts, download_artifacts, fetch_artifact +from pipelex_sdk.client import PipelexAPIClient +from pipelex_sdk.crate_models import MthdsFileItem + +if TYPE_CHECKING: + from pipelex_sdk.artifact_models import DownloadArtifactsResult + +_BASE_URL = os.environ.get("PIPELEX_E2E_BASE_URL", "") +_API_KEY = os.environ.get("PIPELEX_API_KEY", "") + +pytestmark = pytest.mark.skipif( + not _BASE_URL or not _API_KEY, + reason="live leg: set PIPELEX_E2E_BASE_URL and PIPELEX_API_KEY to run it", +) + +#: One domain, one main pipe, a Document input beside the Text it echoes — no inference. +_PASS_THROUGH_BUNDLE = """domain = "smoke_artifacts" +main_pipe = "echo_note" + +[pipe.echo_note] +type = "PipeCompose" +description = "Echo the note beside a document, with no inference" +inputs = { doc = "Document", note = "Text" } +output = "Text" +template = "$note" +""" + +#: A minimal PDF with a nonce of its own, so a swapped file could not pass. Nothing in the run reads it. +_PDF_BYTES = ( + b"%PDF-1.4\n1 0 obj<>endobj\n" + b"2 0 obj<>endobj\n" + b"3 0 obj<>endobj\n" + b"%% nonce " + str(time.time()).encode("ascii") + b"\ntrailer<>\n%%EOF\n" +) + + +class TestArtifactRoundTripLive: + def test_brings_an_uploaded_input_back_down_byte_for_byte(self, tmp_path: Path) -> None: + source = tmp_path / "brief.pdf" + source.write_bytes(_PDF_BYTES) + files = [MthdsFileItem(content=_PASS_THROUGH_BUNDLE, source="smoke_artifacts.mthds")] + + async def _round_trip() -> tuple[str, DownloadArtifactsResult]: + async with PipelexAPIClient(api_key=_API_KEY, base_url=_BASE_URL) as client: + # Up: the Document position is uploaded and rewritten to its storage reference. + prepared = await client.prepare_inputs(files=files, inputs={"doc": str(source), "note": "round trip"}) + assert len(prepared.uploads) == 1 + uploaded = prepared.uploads[0].uri + assert collect_artifacts(prepared.inputs) == [uploaded] + + # Across: a pass-through run echoes the input in its working memory. + results = await client.start_and_wait( + pipe_code="smoke_artifacts.echo_note", + mthds_contents=[_PASS_THROUGH_BUNDLE], + inputs=prepared.inputs, + ) + assert results.main_stuff == {"text": "round trip"} + + # Down: by run id, over working_memory, every link minted fresh. + verdict = await download_artifacts( + client, + dir_path=tmp_path / "out", + run_id=results.pipeline_run_id, + options=DownloadArtifactsOptions( + scope=ArtifactScope.WORKING_MEMORY, + # The local stack's object store hands out plain http links; a hosted one never does. + allow_http=_BASE_URL.startswith("http://"), + ), + ) + return uploaded, verdict + + uploaded, verdict = asyncio.run(_round_trip()) + + assert verdict.scope == ArtifactScope.WORKING_MEMORY + assert verdict.all_saved is True + echoed = next(artifact for artifact in verdict.artifacts if artifact.uri == uploaded) + assert echoed.error is None + assert echoed.content_type == "application/pdf" + assert echoed.size == len(_PDF_BYTES) + assert echoed.path in verdict.saved_paths + assert echoed.path is not None + assert Path(echoed.path).read_bytes() == _PDF_BYTES + + def test_resolves_through_the_bulk_route_and_refuses_a_malformed_reference_as_a_value(self) -> None: + async def _resolve_and_fetch() -> None: + async with PipelexAPIClient(api_key=_API_KEY, base_url=_BASE_URL) as client: + record = await client.upload_file(_PDF_BYTES, filename="probe.pdf", content_type="application/pdf") + resolved = await client.resolve_artifacts([record.uri, "pipelex-storage://"]) + + assert len(resolved) == 2 + assert resolved[0].uri == record.uri + assert resolved[0].error is None + assert resolved[0].url is not None + assert resolved[0].url.startswith(("http://", "https://")) + assert resolved[1].error is not None + assert resolved[1].url is None + + # The link the platform minted is honoured by the store: the bytes come back. + options = DownloadArtifactsOptions(allow_http=_BASE_URL.startswith("http://")) + async with fetch_artifact(client, record.uri, options) as stream: + assert stream.status_code == 200 + assert await stream.read() == _PDF_BYTES + + asyncio.run(_resolve_and_fetch()) diff --git a/tests/unit/test_artifacts.py b/tests/unit/test_artifacts.py new file mode 100644 index 0000000..f295974 --- /dev/null +++ b/tests/unit/test_artifacts.py @@ -0,0 +1,960 @@ +"""The artifact stack — the pure walk, the bulk resolve, the bounded fetch and the download. + +Ports `pipelex-sdk-js/tests/artifacts.test.ts` branch for branch, with the Python surface's own +shapes: the client is a fake satisfying the operations' `Protocol`s, and the object store is an +`httpx.MockTransport` injected at the one seam the module opens for it (`_new_storage_client`), so +every case runs at the httpx boundary and no real socket is ever opened. +""" + +from __future__ import annotations + +import asyncio +import gzip +import json +from datetime import UTC, datetime, timedelta +from typing import TYPE_CHECKING, Any + +import httpx +import pytest + +from pipelex_sdk.artifact_models import ( + ArtifactItemError, + ArtifactScope, + BulkResolvedStorageUrls, + DownloadArtifactsOptions, + FetchArtifactOptions, + ResolvedArtifact, +) +from pipelex_sdk.artifacts import ( + artifact_filename, + collect_artifacts, + download_artifacts, + fetch_artifact, + is_storage_reference, + resolve_artifacts, +) +from pipelex_sdk.errors import ( + ApiResponseError, + ApiUnreachableError, + ArtifactAuthenticationError, + ArtifactFetchError, + ArtifactOperationError, + FieldNotIncludedError, + RunFailedError, + RunStillRunningError, + ScopeUnavailableError, +) +from pipelex_sdk.runs import RunResultCompleted, RunResultFailed, RunResultRunning, RunResults, RunStatus + +if TYPE_CHECKING: + from collections.abc import AsyncIterator, Callable + from pathlib import Path + + from pytest_mock import MockerFixture + + from pipelex_sdk.client import PipelexAPIClient + from pipelex_sdk.runs import RunResultState + from tests.unit.conftest import ResponseBuilder, SendPatcher + +_RUN_ID = "run-01J" +_URI_PNG = "pipelex-storage://org_1/runs/01J/outputs/illustration.png" +_URI_PDF = "pipelex-storage://org_1/runs/01J/outputs/report.pdf" +_STORE = "https://store.example.com" +_PDF_BYTES = b"%PDF-1.4 tiny" +_PNG_BYTES = b"\x89PNG tiny" + + +# ── Fakes and builders ─────────────────────────────────────────────── + + +class _FakeClient: + """The two client methods the artifact operations call, scripted per test and recorded.""" + + def __init__( + self, + *, + resolve: Callable[[list[str]], BulkResolvedStorageUrls] | None = None, + run_result: RunResultState | None = None, + ) -> None: + self._resolve = resolve + self._run_result = run_result + self.resolve_calls: list[list[str]] = [] + self.run_result_calls: list[str] = [] + + async def resolve_storage_urls_bulk(self, uris: list[str]) -> BulkResolvedStorageUrls: + self.resolve_calls.append(list(uris)) + if self._resolve is None: + msg = "this test did not script a resolve answer" + raise AssertionError(msg) + return self._resolve(list(uris)) + + async def get_run_result(self, run_id: str) -> RunResultState: + self.run_result_calls.append(run_id) + if self._run_result is None: + msg = "this test did not script a run result" + raise AssertionError(msg) + return self._run_result + + +def _resolved( + uri: str, *, url: str | None = None, expires_in_seconds: float = 900.0, content_type: str | None = "application/pdf" +) -> ResolvedArtifact: + """One resolved item, its link live for `expires_in_seconds` from now.""" + expires_at = (datetime.now(UTC) + timedelta(seconds=expires_in_seconds)).isoformat().replace("+00:00", "Z") + return ResolvedArtifact( + uri=uri, url=url or f"{_STORE}/{uri.rsplit('/', maxsplit=1)[-1]}?sig=fresh", expires_at=expires_at, content_type=content_type + ) + + +def _refused(uri: str, *, code: str = "forbidden", detail: str = "Another organization owns this reference.") -> ResolvedArtifact: + """One refused item — the three link fields null, the verdict on `error`.""" + return ResolvedArtifact(uri=uri, url=None, expires_at=None, content_type=None, error=ArtifactItemError(code=code, detail=detail)) + + +def _answer(*items: ResolvedArtifact) -> BulkResolvedStorageUrls: + return BulkResolvedStorageUrls(items=list(items)) + + +def _resolver(*items: ResolvedArtifact) -> Callable[[list[str]], BulkResolvedStorageUrls]: + """A resolve script answering one item per requested reference, matched by uri.""" + by_uri = {item.uri: item for item in items} + + def _resolve(uris: list[str]) -> BulkResolvedStorageUrls: + return _answer(*[by_uri[uri] for uri in uris]) + + return _resolve + + +def _api_error(status: int) -> ApiResponseError: + return ApiResponseError( + f"API POST /v1/resolve-storage-url/bulk failed ({status})", + api_url=_STORE, + status=status, + status_text="Forbidden", + response_body="{}", + ) + + +def _patch_storage(mocker: MockerFixture, handler: Callable[[httpx.Request], httpx.Response]) -> list[httpx.Request]: + """Route every object-store fetch through a mock transport, and record the requests it saw.""" + seen: list[httpx.Request] = [] + + def _recording(request: httpx.Request) -> httpx.Response: + seen.append(request) + return handler(request) + + def _factory(timeout_seconds: float) -> httpx.AsyncClient: + return httpx.AsyncClient(transport=httpx.MockTransport(_recording), timeout=httpx.Timeout(timeout_seconds), follow_redirects=False) + + mocker.patch("pipelex_sdk.artifacts._new_storage_client", _factory) + return seen + + +def _serving(payload: bytes, *, status: int = 200, headers: dict[str, str] | None = None) -> Callable[[httpx.Request], httpx.Response]: + """A store that answers every link with the same body, `Content-Length` declared by httpx.""" + + def _handler(_: httpx.Request) -> httpx.Response: + return httpx.Response(status, content=payload, headers=headers or {"content-type": "application/pdf"}) + + return _handler + + +def _streamed(chunks: list[bytes], *, delay_seconds: float = 0.0, headers: dict[str, str] | None = None) -> Callable[[httpx.Request], httpx.Response]: + """A store that answers with a chunked body and therefore declares no length.""" + + def _handler(_: httpx.Request) -> httpx.Response: + async def _body() -> AsyncIterator[bytes]: + for chunk in chunks: + if delay_seconds: + await asyncio.sleep(delay_seconds) + yield chunk + + return httpx.Response(200, content=_body(), headers=headers or {"content-type": "application/pdf"}) + + return _handler + + +#: Says a test wants the `working_memory` key left out of the body altogether, which is what the +#: platform does for a key the results read did not carry — distinct from relaying it as null. +_ABSENT_KEY = object() + + +def _results(main_stuff: Any, *, working_memory: Any = _ABSENT_KEY, run_id: str = _RUN_ID) -> RunResults: + """A `RunResults` whose `working_memory` key is present only when the test says so.""" + body: dict[str, Any] = {"pipeline_run_id": run_id, "main_stuff": main_stuff} + if working_memory is not _ABSENT_KEY: + body["working_memory"] = working_memory + return RunResults.model_validate(body) + + +def _content(uri: str) -> dict[str, Any]: + """A produced file as the runtime serializes it: the durable reference beside an expiring link.""" + return {"url": uri, "public_url": f"{_STORE}/signed-by-the-runtime?sig=stale"} + + +class TestArtifacts: + # ── collect_artifacts ──────────────────────────────────────────── + + def test_collects_every_reference_once_in_discovery_order(self) -> None: + walked = { + "picture": _content(_URI_PNG), + "items": [{"doc": _content(_URI_PDF)}, {"doc": _content(_URI_PNG)}], + "note": "no file here", + } + assert collect_artifacts(walked) == [_URI_PNG, _URI_PDF] + + @pytest.mark.parametrize( + ("value", "expected"), + [ + (f"see {_URI_PNG} for the picture", []), + ("pipelex-storage://", []), + ("s3://org_1/runs/01J/outputs/x.png", []), + (_URI_PNG, [_URI_PNG]), + ], + ) + def test_counts_a_string_only_when_it_is_a_reference(self, value: str, expected: list[str]) -> None: + assert collect_artifacts(value) == expected + assert is_storage_reference(value) is (expected != []) + + def test_ignores_values_that_carry_no_reference(self) -> None: + assert collect_artifacts(None) == [] + assert collect_artifacts(42) == [] + assert collect_artifacts({"a": [1, 2.5, True, None]}) == [] + + def test_walks_a_pydantic_model_as_well_as_a_parsed_body(self) -> None: + results = _results({"picture": _content(_URI_PNG)}) + assert collect_artifacts(results) == [_URI_PNG] + + # ── artifact_filename ──────────────────────────────────────────── + + @pytest.mark.parametrize( + ("uri", "content_type", "expected"), + [ + ("pipelex-storage://org_1/runs/01J/outputs/report.pdf", "application/pdf", "report.pdf"), + # Path separators are the split point, so no traversal and no absolute path survives. + ("pipelex-storage://org_1/../../etc/passwd", None, "passwd"), + # An encoded traversal is one segment, decoded after the split: the separators it hid + # become underscores and the leading dots go, so it still names a file in the directory. + ("pipelex-storage://org_1/x/..%2F..%2Fetc%2Fpasswd", None, "etc_passwd"), + ("pipelex-storage://org_1/x/a\\b\\c.txt", None, "c.txt"), + # A leading dot is stripped, so no hidden file; odd characters are neutralized. + ("pipelex-storage://org_1/.bashrc", None, "bashrc"), + ("pipelex-storage://org_1/my file (1).png", "image/png", "my_file__1_.png"), + # Percent-decoded, with the query and fragment dropped. + ("pipelex-storage://org_1/a%20b.pdf?sig=x#frag", None, "a_b.pdf"), + # The extension comes from the content type only when the key carries none. + ("pipelex-storage://org_1/outputs/report", "application/pdf", "report.pdf"), + ("pipelex-storage://org_1/outputs/report.bin", "application/pdf", "report.bin"), + ("pipelex-storage://org_1/outputs/report", "image/png; charset=binary", "report.png"), + ("pipelex-storage://org_1/outputs/report", "application/x-unknown", "report"), + ("pipelex-storage://org_1/outputs/report", None, "report"), + # Nothing usable in the key: the numbered fallback, one-based. + ("pipelex-storage://", None, "artifact-3"), + ("pipelex-storage://org_1/___", None, "artifact-3"), + ], + ) + def test_derives_a_filename_that_can_only_name_a_file_in_the_directory(self, uri: str, content_type: str | None, expected: str) -> None: + assert artifact_filename(uri, content_type, 2) == expected + + def test_caps_the_filename_length_keeping_the_extension(self) -> None: + name = artifact_filename(f"pipelex-storage://org_1/{'a' * 400}.pdf", None, 0) + assert len(name) == 128 + assert name.endswith(".pdf") + + def test_drops_an_extension_that_alone_exceeds_the_cap(self) -> None: + name = artifact_filename(f"pipelex-storage://org_1/name.{'z' * 400}", None, 0) + assert len(name) == 128 + assert name.startswith("name.") + + # ── resolve_artifacts ──────────────────────────────────────────── + + def test_caps_the_length_after_the_guessed_extension(self) -> None: + name = artifact_filename(f"pipelex-storage://org_1/{'a' * 200}", "image/jpeg", 0) + + assert len(name) <= 128 + assert name.endswith(".jpg") + + def test_resolves_a_list_within_the_bound_in_one_call_and_keeps_request_order(self) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PNG), _refused(_URI_PDF))) + items = asyncio.run(resolve_artifacts(client, [_URI_PNG, _URI_PDF, _URI_PNG])) + + assert client.resolve_calls == [[_URI_PNG, _URI_PDF, _URI_PNG]] + assert [item.uri for item in items] == [_URI_PNG, _URI_PDF, _URI_PNG] + assert items[0].error is None + assert items[1].error is not None + assert items[1].error.code == "forbidden" + assert items[1].url is None + + def test_chunks_a_longer_list_at_the_routes_bound(self) -> None: + uris = [f"pipelex-storage://org_1/f{index}.pdf" for index in range(250)] + client = _FakeClient(resolve=lambda chunk: _answer(*[_resolved(uri) for uri in chunk])) + items = asyncio.run(resolve_artifacts(client, uris)) + + assert [len(call) for call in client.resolve_calls] == [100, 100, 50] + assert [item.uri for item in items] == uris + + def test_makes_no_request_for_an_empty_list(self) -> None: + client = _FakeClient() + assert asyncio.run(resolve_artifacts(client, [])) == [] + assert client.resolve_calls == [] + + def test_refuses_a_malformed_answer_rather_than_misattributing_verdicts(self) -> None: + client = _FakeClient(resolve=lambda _: _answer(_resolved(_URI_PNG))) + with pytest.raises(ArtifactOperationError, match="1 item"): + asyncio.run(resolve_artifacts(client, [_URI_PNG, _URI_PDF])) + + def test_lets_a_whole_request_refusal_propagate_unchanged(self) -> None: + def _refuse(_: list[str]) -> BulkResolvedStorageUrls: + raise _api_error(404) + + client = _FakeClient(resolve=_refuse) + with pytest.raises(ApiResponseError) as caught: + asyncio.run(resolve_artifacts(client, [_URI_PNG])) + assert caught.value.status == 404 + + # ── fetch_artifact ─────────────────────────────────────────────── + + def test_fetches_a_fresh_link_with_no_credentials_and_relays_the_stores_response(self, mocker: MockerFixture) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) + seen = _patch_storage(mocker, _serving(_PDF_BYTES, headers={"content-type": "application/pdf", "x-store": "yes"})) + + async def _read() -> tuple[int, bytes, str | None]: + async with fetch_artifact(client, _URI_PDF) as stream: + return stream.status_code, await stream.read(), stream.headers.get("x-store") + + status, body, store_header = asyncio.run(_read()) + + assert status == 200 + assert body == _PDF_BYTES + assert store_header == "yes" + assert len(seen) == 1 + assert "authorization" not in seen[0].headers + assert str(seen[0].url).startswith(f"{_STORE}/report.pdf") + # The link came from the route, never from the content's embedded `public_url`. + assert "sig=fresh" in str(seen[0].url) + + def test_raises_the_routes_per_reference_refusal_as_a_typed_fetch_error(self, mocker: MockerFixture) -> None: + client = _FakeClient(resolve=_resolver(_refused(_URI_PDF, code="invalid_storage_uri", detail="Not a reference."))) + _patch_storage(mocker, _serving(_PDF_BYTES)) + + async def _read() -> None: + async with fetch_artifact(client, _URI_PDF): + pass + + with pytest.raises(ArtifactFetchError) as caught: + asyncio.run(_read()) + assert caught.value.code == "invalid_storage_uri" + assert caught.value.uri == _URI_PDF + + def test_refuses_a_plain_http_link_by_default_and_accepts_it_on_request(self, mocker: MockerFixture) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF, url="http://localhost:9000/report.pdf"))) + seen = _patch_storage(mocker, _serving(_PDF_BYTES)) + + async def _read(options: FetchArtifactOptions | None) -> bytes: + async with fetch_artifact(client, _URI_PDF, options) as stream: + return await stream.read() + + with pytest.raises(ArtifactFetchError) as caught: + asyncio.run(_read(None)) + assert caught.value.code == "plain_http_refused" + assert seen == [] + + assert asyncio.run(_read(FetchArtifactOptions(allow_http=True))) == _PDF_BYTES + assert len(seen) == 1 + + @pytest.mark.parametrize("url", ["ftp://store/x.pdf", "not-a-url", "https://user:pass@store/x.pdf", "https://[::1", "https://host:abc/file"]) + def test_refuses_an_unusable_link_before_any_request(self, mocker: MockerFixture, url: str) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF, url=url))) + seen = _patch_storage(mocker, _serving(_PDF_BYTES)) + + async def _read() -> None: + async with fetch_artifact(client, _URI_PDF): + pass + + with pytest.raises(ArtifactFetchError) as caught: + asyncio.run(_read()) + assert caught.value.code == "unsupported_url" + assert seen == [] + + @pytest.mark.parametrize( + ("status", "code"), + [(302, "redirect_refused"), (401, "store_refused"), (403, "store_refused"), (404, "not_found"), (410, "not_found"), (500, "store_error")], + ) + def test_maps_the_stores_statuses_onto_the_fetch_codes(self, mocker: MockerFixture, status: int, code: str) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) + _patch_storage(mocker, _serving(b"", status=status, headers={"location": f"{_STORE}/elsewhere"})) + + async def _read() -> None: + async with fetch_artifact(client, _URI_PDF): + pass + + with pytest.raises(ArtifactFetchError) as caught: + asyncio.run(_read()) + assert caught.value.code == code + assert caught.value.status == status + + def test_refuses_a_declared_oversize_without_reading_the_body(self, mocker: MockerFixture) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) + _patch_storage(mocker, _serving(b"x" * 64)) + + async def _read() -> None: + async with fetch_artifact(client, _URI_PDF, FetchArtifactOptions(max_bytes=32)): + pass + + with pytest.raises(ArtifactFetchError) as caught: + asyncio.run(_read()) + assert caught.value.code == "too_large" + + def test_cuts_a_body_that_crosses_the_cap_mid_stream(self, mocker: MockerFixture) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) + _patch_storage(mocker, _streamed([b"a" * 16, b"b" * 16, b"c" * 16])) + + async def _read() -> list[bytes]: + chunks: list[bytes] = [] + async with fetch_artifact(client, _URI_PDF, FetchArtifactOptions(max_bytes=20)) as stream: + async for chunk in stream.aiter_bytes(chunk_size=16): + chunks.append(chunk) + return chunks + + with pytest.raises(ArtifactFetchError) as caught: + asyncio.run(_read()) + assert caught.value.code == "too_large" + + def test_times_out_an_exchange_that_outlives_its_budget(self, mocker: MockerFixture) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) + _patch_storage(mocker, _streamed([b"a", b"b"], delay_seconds=0.2)) + + async def _read() -> bytes: + async with fetch_artifact(client, _URI_PDF, FetchArtifactOptions(timeout_seconds=0.05)) as stream: + return await stream.read() + + with pytest.raises(ArtifactFetchError) as caught: + asyncio.run(_read()) + assert caught.value.code == "timeout" + + def test_reports_a_transport_failure_as_a_network_fault(self, mocker: MockerFixture) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) + + def _broken(request: httpx.Request) -> httpx.Response: + msg = "connection refused" + raise httpx.ConnectError(msg, request=request) + + _patch_storage(mocker, _broken) + + async def _read() -> None: + async with fetch_artifact(client, _URI_PDF): + pass + + with pytest.raises(ArtifactFetchError) as caught: + asyncio.run(_read()) + assert caught.value.code == "network" + + @pytest.mark.parametrize(("name", "value"), [("max_bytes", 0), ("timeout_seconds", -1)]) + def test_refuses_nonsense_bounds_before_resolving_anything(self, name: str, value: float) -> None: + client = _FakeClient() + + async def _read() -> None: + async with fetch_artifact(client, _URI_PDF, FetchArtifactOptions.model_validate({name: value})): + pass + + with pytest.raises(ArtifactOperationError, match=name): + asyncio.run(_read()) + assert client.resolve_calls == [] + + def test_drops_a_content_encoding_httpx_already_decoded_with_its_length(self, mocker: MockerFixture) -> None: + packed = gzip.compress(_PDF_BYTES) + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) + _patch_storage(mocker, _serving(packed, headers={"content-type": "application/pdf", "content-encoding": "gzip"})) + + async def _read() -> tuple[bytes, httpx.Headers]: + async with fetch_artifact(client, _URI_PDF) as stream: + return await stream.read(), stream.headers + + body, headers = asyncio.run(_read()) + assert body == _PDF_BYTES + assert "content-encoding" not in headers + assert "content-length" not in headers + + @pytest.mark.parametrize("coding", ["exi", "x-gzip"]) + def test_keeps_an_encoding_httpx_did_not_decode_beside_its_body(self, mocker: MockerFixture, coding: str) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) + _patch_storage(mocker, _serving(_PDF_BYTES, headers={"content-type": "application/pdf", "content-encoding": coding})) + + async def _read() -> httpx.Headers: + async with fetch_artifact(client, _URI_PDF) as stream: + return stream.headers + + headers = asyncio.run(_read()) + assert headers["content-encoding"] == coding + assert headers["content-length"] == str(len(_PDF_BYTES)) + + # ── download_artifacts ─────────────────────────────────────────── + + def test_takes_exactly_one_of_run_id_or_results(self, tmp_path: Path) -> None: + client = _FakeClient() + results = _results({"picture": _content(_URI_PNG)}) + + with pytest.raises(ArtifactOperationError, match="exactly one"): + asyncio.run(download_artifacts(client, dir_path=tmp_path)) + with pytest.raises(ArtifactOperationError, match="exactly one"): + asyncio.run(download_artifacts(client, dir_path=tmp_path, run_id=_RUN_ID, results=results)) + # An empty run id names nothing, so it is neither selector. + with pytest.raises(ArtifactOperationError, match="exactly one"): + asyncio.run(download_artifacts(client, dir_path=tmp_path, run_id="")) + + @pytest.mark.parametrize(("name", "value"), [("concurrency", 0), ("max_total_bytes", 0), ("max_bytes", -1), ("timeout_seconds", 0)]) + def test_validates_the_download_bounds(self, tmp_path: Path, name: str, value: float) -> None: + client = _FakeClient() + with pytest.raises(ArtifactOperationError, match=name): + asyncio.run( + download_artifacts( + client, + dir_path=tmp_path, + results=_results({"picture": _content(_URI_PNG)}), + options=DownloadArtifactsOptions.model_validate({name: value}), + ) + ) + + def test_raises_run_still_running_with_the_retry_hint(self, tmp_path: Path) -> None: + client = _FakeClient(run_result=RunResultRunning(pipeline_run_id=_RUN_ID, retry_after_seconds=7)) + with pytest.raises(RunStillRunningError, match="retry in 7s"): + asyncio.run(download_artifacts(client, dir_path=tmp_path, run_id=_RUN_ID)) + + def test_raises_run_failed_for_a_run_that_ended_without_a_result(self, tmp_path: Path) -> None: + client = _FakeClient(run_result=RunResultFailed(pipeline_run_id=_RUN_ID, status=RunStatus.FAILED, message="the run failed")) + with pytest.raises(RunFailedError) as caught: + asyncio.run(download_artifacts(client, dir_path=tmp_path, run_id=_RUN_ID)) + assert caught.value.status == RunStatus.FAILED + assert caught.value.run_id == _RUN_ID + + def test_raises_field_not_included_when_the_scope_key_was_never_relayed(self, tmp_path: Path) -> None: + client = _FakeClient() + with pytest.raises(FieldNotIncludedError) as caught: + asyncio.run( + download_artifacts( + client, + dir_path=tmp_path, + results=_results({"picture": _content(_URI_PNG)}), + options=DownloadArtifactsOptions(scope=ArtifactScope.WORKING_MEMORY), + ) + ) + assert caught.value.field_name == "working_memory" + + def test_raises_scope_unavailable_when_the_key_was_relayed_as_null(self, tmp_path: Path) -> None: + client = _FakeClient() + with pytest.raises(ScopeUnavailableError) as caught: + asyncio.run( + download_artifacts( + client, + dir_path=tmp_path, + results=_results({"picture": _content(_URI_PNG)}, working_memory=None), + options=DownloadArtifactsOptions(scope=ArtifactScope.WORKING_MEMORY), + ) + ) + assert caught.value.scope == ArtifactScope.WORKING_MEMORY + assert caught.value.run_id == _RUN_ID + + def test_answers_an_empty_walk_over_a_present_scope_as_a_verdict(self, tmp_path: Path) -> None: + client = _FakeClient() + target = tmp_path / "out" + verdict = asyncio.run(download_artifacts(client, dir_path=target, results=_results({"text": "no file here"}))) + + assert verdict.scope == ArtifactScope.MAIN_STUFF + assert verdict.artifacts == [] + assert verdict.saved_paths == [] + assert verdict.all_saved is True + assert client.resolve_calls == [] + assert not target.exists() + + def test_reads_by_run_id_resolves_in_one_bulk_call_and_saves_every_file(self, mocker: MockerFixture, tmp_path: Path) -> None: + results = _results({"items": [_content(_URI_PNG), _content(_URI_PDF)]}) + client = _FakeClient( + resolve=_resolver(_resolved(_URI_PNG, content_type="image/png"), _resolved(_URI_PDF)), + run_result=RunResultCompleted(pipeline_run_id=_RUN_ID, result=results), + ) + + def _handler(request: httpx.Request) -> httpx.Response: + payload = _PNG_BYTES if request.url.path.endswith(".png") else _PDF_BYTES + return httpx.Response(200, content=payload) + + _patch_storage(mocker, _handler) + target = tmp_path / "out" + verdict = asyncio.run(download_artifacts(client, dir_path=target, run_id=_RUN_ID)) + + assert client.run_result_calls == [_RUN_ID] + assert client.resolve_calls == [[_URI_PNG, _URI_PDF]] + assert verdict.all_saved is True + assert [artifact.uri for artifact in verdict.artifacts] == [_URI_PNG, _URI_PDF] + assert [artifact.size for artifact in verdict.artifacts] == [len(_PNG_BYTES), len(_PDF_BYTES)] + assert [artifact.content_type for artifact in verdict.artifacts] == ["image/png", "application/pdf"] + assert verdict.saved_paths == [str(target / "illustration.png"), str(target / "report.pdf")] + assert (target / "illustration.png").read_bytes() == _PNG_BYTES + assert (target / "report.pdf").read_bytes() == _PDF_BYTES + + def test_takes_results_in_hand_without_re_reading_and_creates_the_directory(self, mocker: MockerFixture, tmp_path: Path) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) + _patch_storage(mocker, _serving(_PDF_BYTES)) + target = tmp_path / "nested" / "out" + verdict = asyncio.run(download_artifacts(client, dir_path=target, results=_results({"doc": _content(_URI_PDF)}))) + + assert client.run_result_calls == [] + assert verdict.all_saved is True + assert target.is_dir() + + def test_walks_working_memory_when_asked_echoed_inputs_included(self, mocker: MockerFixture, tmp_path: Path) -> None: + results = _results( + {"text": "round trip"}, + working_memory={ + "root": { + "doc": {"concept": "native.PDF", "content": _content(_URI_PDF)}, + "picture": {"concept": "native.Image", "content": _content(_URI_PNG)}, + }, + "aliases": {}, + }, + ) + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF), _resolved(_URI_PNG, content_type="image/png"))) + _patch_storage(mocker, _serving(_PDF_BYTES)) + verdict = asyncio.run( + download_artifacts( + client, + dir_path=tmp_path / "out", + results=results, + options=DownloadArtifactsOptions(scope=ArtifactScope.WORKING_MEMORY), + ) + ) + + assert verdict.scope == ArtifactScope.WORKING_MEMORY + assert [artifact.uri for artifact in verdict.artifacts] == [_URI_PDF, _URI_PNG] + assert verdict.all_saved is True + + def test_never_overwrites_a_name_already_on_disk(self, mocker: MockerFixture, tmp_path: Path) -> None: + other = "pipelex-storage://org_1/runs/01J/second/report.pdf" + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF), _resolved(other))) + _patch_storage(mocker, _serving(_PDF_BYTES)) + target = tmp_path / "out" + target.mkdir() + (target / "report.pdf").write_bytes(b"do not touch me") + + verdict = asyncio.run( + download_artifacts( + client, + dir_path=target, + results=_results({"items": [_content(_URI_PDF), _content(other)]}), + options=DownloadArtifactsOptions(concurrency=1), + ) + ) + + assert verdict.saved_paths == [str(target / "report-1.pdf"), str(target / "report-2.pdf")] + assert (target / "report.pdf").read_bytes() == b"do not touch me" + + def test_keeps_a_per_reference_refusal_as_that_items_error_beside_the_saved_ones(self, mocker: MockerFixture, tmp_path: Path) -> None: + client = _FakeClient(resolve=_resolver(_refused(_URI_PNG), _resolved(_URI_PDF))) + _patch_storage(mocker, _serving(_PDF_BYTES)) + verdict = asyncio.run( + download_artifacts(client, dir_path=tmp_path / "out", results=_results({"items": [_content(_URI_PNG), _content(_URI_PDF)]})) + ) + + assert verdict.all_saved is False + first = verdict.artifacts[0] + assert first.error is not None + assert first.error.code == "forbidden" + assert first.path is None + assert verdict.artifacts[1].error is None + assert len(verdict.saved_paths) == 1 + + def test_re_resolves_a_link_that_has_expired_by_the_time_its_task_reaches_it(self, mocker: MockerFixture, tmp_path: Path) -> None: + answers = [ + _answer(_resolved(_URI_PDF, url=f"{_STORE}/stale.pdf", expires_in_seconds=1.0)), + _answer(_resolved(_URI_PDF, url=f"{_STORE}/fresh.pdf")), + ] + client = _FakeClient(resolve=lambda _: answers.pop(0)) + seen = _patch_storage(mocker, _serving(_PDF_BYTES)) + verdict = asyncio.run(download_artifacts(client, dir_path=tmp_path / "out", results=_results({"doc": _content(_URI_PDF)}))) + + assert client.resolve_calls == [[_URI_PDF], [_URI_PDF]] + assert [str(request.url) for request in seen] == [f"{_STORE}/fresh.pdf"] + assert verdict.all_saved is True + + @pytest.mark.parametrize("failure", [_api_error(500), ApiUnreachableError("host unreachable", api_url=_STORE, code="ENOTFOUND")]) + def test_marks_an_item_whose_expired_link_cannot_be_re_resolved(self, mocker: MockerFixture, tmp_path: Path, failure: Exception) -> None: + calls: list[int] = [] + + def _resolve(uris: list[str]) -> BulkResolvedStorageUrls: + calls.append(len(uris)) + if len(calls) == 1: + return _answer(_resolved(_URI_PDF, expires_in_seconds=1.0)) + raise failure + + client = _FakeClient(resolve=_resolve) + _patch_storage(mocker, _serving(_PDF_BYTES)) + verdict = asyncio.run(download_artifacts(client, dir_path=tmp_path / "out", results=_results({"doc": _content(_URI_PDF)}))) + + assert verdict.all_saved is False + error = verdict.artifacts[0].error + assert error is not None + assert error.code == "resolve_failed" + + def test_raises_a_credential_failure_on_the_first_resolve_with_an_all_aborted_verdict(self, tmp_path: Path) -> None: + def _refuse(_: list[str]) -> BulkResolvedStorageUrls: + raise _api_error(403) + + client = _FakeClient(resolve=_refuse) + with pytest.raises(ArtifactAuthenticationError) as caught: + asyncio.run(download_artifacts(client, dir_path=tmp_path / "out", results=_results({"items": [_content(_URI_PNG), _content(_URI_PDF)]}))) + + verdict = caught.value.verdict + assert caught.value.status == 403 + assert verdict.saved_paths == [] + assert [artifact.error.code for artifact in verdict.artifacts if artifact.error is not None] == ["aborted", "aborted"] + + def test_raises_a_credential_failure_part_way_through_carrying_the_verdict_so_far(self, mocker: MockerFixture, tmp_path: Path) -> None: + def _resolve(uris: list[str]) -> BulkResolvedStorageUrls: + if len(uris) > 1: + return _answer(_resolved(_URI_PDF), _resolved(_URI_PNG, expires_in_seconds=1.0)) + raise _api_error(401) + + client = _FakeClient(resolve=_resolve) + _patch_storage(mocker, _serving(_PDF_BYTES)) + with pytest.raises(ArtifactAuthenticationError) as caught: + asyncio.run( + download_artifacts( + client, + dir_path=tmp_path / "out", + results=_results({"items": [_content(_URI_PDF), _content(_URI_PNG)]}), + options=DownloadArtifactsOptions(concurrency=1), + ) + ) + + verdict = caught.value.verdict + assert caught.value.status == 401 + assert len(verdict.saved_paths) == 1 + assert verdict.artifacts[0].error is None + second = verdict.artifacts[1].error + assert second is not None + assert second.code == "aborted" + assert "credential failure" in second.detail + + def test_refuses_a_declared_oversize_and_leaves_nothing_behind(self, mocker: MockerFixture, tmp_path: Path) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) + _patch_storage(mocker, _serving(b"x" * 64)) + target = tmp_path / "out" + verdict = asyncio.run( + download_artifacts( + client, + dir_path=target, + results=_results({"doc": _content(_URI_PDF)}), + options=DownloadArtifactsOptions(max_bytes=32), + ) + ) + + error = verdict.artifacts[0].error + assert error is not None + assert error.code == "too_large" + assert list(target.iterdir()) == [] + + def test_cuts_an_undeclared_oversize_mid_stream_and_unlinks_the_partial_file(self, mocker: MockerFixture, tmp_path: Path) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) + _patch_storage(mocker, _streamed([b"a" * 16, b"b" * 16, b"c" * 16])) + target = tmp_path / "out" + verdict = asyncio.run( + download_artifacts( + client, + dir_path=target, + results=_results({"doc": _content(_URI_PDF)}), + options=DownloadArtifactsOptions(max_bytes=20), + ) + ) + + error = verdict.artifacts[0].error + assert error is not None + assert error.code == "too_large" + assert list(target.iterdir()) == [] + + def test_enforces_the_total_cap_on_each_later_item_by_its_own_size(self, mocker: MockerFixture, tmp_path: Path) -> None: + third = "pipelex-storage://org_1/runs/01J/outputs/third.pdf" + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF), _resolved(_URI_PNG), _resolved(third))) + _patch_storage(mocker, _serving(b"x" * 8)) + verdict = asyncio.run( + download_artifacts( + client, + dir_path=tmp_path / "out", + results=_results({"items": [_content(_URI_PDF), _content(_URI_PNG), _content(third)]}), + options=DownloadArtifactsOptions(concurrency=1, max_total_bytes=10), + ) + ) + + codes = [artifact.error.code if artifact.error is not None else None for artifact in verdict.artifacts] + assert codes == [None, "total_limit_exceeded", "total_limit_exceeded"] + refused = verdict.artifacts[2].error + assert refused is not None + assert refused.detail.startswith("Saving this") + assert len(verdict.saved_paths) == 1 + + def test_a_refused_item_never_skips_a_smaller_one_that_still_fits(self, mocker: MockerFixture, tmp_path: Path) -> None: + big = "pipelex-storage://org_1/runs/01J/outputs/big.pdf" + small = "pipelex-storage://org_1/runs/01J/outputs/small.pdf" + client = _FakeClient(resolve=_resolver(_resolved(big), _resolved(small))) + + def _sized(request: httpx.Request) -> httpx.Response: + return httpx.Response(200, content=b"x" * (16 if "big" in request.url.path else 2)) + + _patch_storage(mocker, _sized) + verdict = asyncio.run( + download_artifacts( + client, + dir_path=tmp_path / "out", + results=_results({"items": [_content(big), _content(small)]}), + options=DownloadArtifactsOptions(concurrency=1, max_total_bytes=10), + ) + ) + + codes = [artifact.error.code if artifact.error is not None else None for artifact in verdict.artifacts] + assert codes == ["total_limit_exceeded", None] + assert verdict.artifacts[1].size == 2 + assert len(verdict.saved_paths) == 1 + + def test_enforces_the_total_cap_mid_stream_on_an_undeclared_body(self, mocker: MockerFixture, tmp_path: Path) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) + _patch_storage(mocker, _streamed([b"a" * 8, b"b" * 8])) + target = tmp_path / "out" + verdict = asyncio.run( + download_artifacts( + client, + dir_path=target, + results=_results({"doc": _content(_URI_PDF)}), + options=DownloadArtifactsOptions(max_total_bytes=10), + ) + ) + + error = verdict.artifacts[0].error + assert error is not None + assert error.code == "total_limit_exceeded" + assert list(target.iterdir()) == [] + + def test_keeps_a_store_refusal_as_the_items_error(self, mocker: MockerFixture, tmp_path: Path) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) + _patch_storage(mocker, _serving(b"", status=404)) + verdict = asyncio.run(download_artifacts(client, dir_path=tmp_path / "out", results=_results({"doc": _content(_URI_PDF)}))) + + error = verdict.artifacts[0].error + assert error is not None + assert error.code == "not_found" + assert verdict.all_saved is False + + def test_reports_a_file_that_cannot_be_created_as_write_failed(self, mocker: MockerFixture, tmp_path: Path) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) + _patch_storage(mocker, _serving(_PDF_BYTES)) + mocker.patch("pipelex_sdk.artifacts._open_unique_file", side_effect=OSError("no space left on device")) + target = tmp_path / "out" + verdict = asyncio.run(download_artifacts(client, dir_path=target, results=_results({"doc": _content(_URI_PDF)}))) + + error = verdict.artifacts[0].error + assert error is not None + assert error.code == "write_failed" + assert list(target.iterdir()) == [] + + def test_raises_for_a_directory_that_cannot_be_created(self, tmp_path: Path) -> None: + blocked = tmp_path / "a-file" + blocked.write_bytes(b"not a directory") + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) + with pytest.raises(ArtifactOperationError, match="cannot be created or used"): + asyncio.run(download_artifacts(client, dir_path=blocked, results=_results({"doc": _content(_URI_PDF)}))) + + def test_lets_a_deployment_without_the_bulk_route_surface_as_the_transport_error(self, tmp_path: Path) -> None: + def _refuse(_: list[str]) -> BulkResolvedStorageUrls: + raise _api_error(404) + + client = _FakeClient(resolve=_refuse) + with pytest.raises(ApiResponseError) as caught: + asyncio.run(download_artifacts(client, dir_path=tmp_path / "out", results=_results({"doc": _content(_URI_PDF)}))) + assert caught.value.status == 404 + + def test_holds_the_concurrency_bound_across_the_pipeline(self, mocker: MockerFixture, tmp_path: Path) -> None: + uris = [f"pipelex-storage://org_1/runs/01J/outputs/f{index}.pdf" for index in range(6)] + client = _FakeClient(resolve=_resolver(*[_resolved(uri) for uri in uris])) + live = {"now": 0, "peak": 0} + + def _handler(_: httpx.Request) -> httpx.Response: + async def _body() -> AsyncIterator[bytes]: + live["now"] += 1 + live["peak"] = max(live["peak"], live["now"]) + await asyncio.sleep(0.02) + yield _PDF_BYTES + live["now"] -= 1 + + return httpx.Response(200, content=_body()) + + _patch_storage(mocker, _handler) + verdict = asyncio.run( + download_artifacts( + client, + dir_path=tmp_path / "out", + results=_results({"items": [_content(uri) for uri in uris]}), + options=DownloadArtifactsOptions(concurrency=2), + ) + ) + + assert verdict.all_saved is True + assert live["peak"] == 2 + + def test_unlinks_the_partial_file_when_the_download_is_cancelled(self, mocker: MockerFixture, tmp_path: Path) -> None: + client = _FakeClient(resolve=_resolver(_resolved(_URI_PDF))) + _patch_storage(mocker, _streamed([b"a" * 8] * 20, delay_seconds=0.05)) + target = tmp_path / "out" + + async def _cancel_midway() -> bool: + task = asyncio.create_task(download_artifacts(client, dir_path=target, results=_results({"doc": _content(_URI_PDF)}))) + await asyncio.sleep(0.12) + assert (target / "report.pdf").exists() + task.cancel() + try: + await task + except asyncio.CancelledError: + return True + return False + + assert asyncio.run(_cancel_midway()) is True + assert list(target.iterdir()) == [] + + # ── the client's own methods ───────────────────────────────────── + + def test_the_client_posts_the_bulk_route_and_serves_the_whole_stack( + self, + api_client: PipelexAPIClient, + wire_response: ResponseBuilder, + patch_send: SendPatcher, + mocker: MockerFixture, + tmp_path: Path, + ) -> None: + item = _resolved(_URI_PDF).model_dump() + spy = patch_send(api_client, wire_response(200, json_body={"items": [item]})) + seen = _patch_storage(mocker, _serving(_PDF_BYTES)) + target = tmp_path / "out" + + verdict = asyncio.run(api_client.download_artifacts(dir_path=target, results=_results({"doc": _content(_URI_PDF)}))) + + assert verdict.all_saved is True + assert verdict.saved_paths == [str(target / "report.pdf")] + method, url = spy.call_args.args[0], spy.call_args.args[1] + assert method == "POST" + assert url.endswith("/v1/resolve-storage-url/bulk") + assert json.loads(spy.call_args.kwargs["content"]) == {"uris": [_URI_PDF]} + # The store was reached on its own client, carrying none of the API client's credential. + assert "authorization" not in seen[0].headers + + def test_the_clients_fetch_artifact_yields_the_same_bounded_stream( + self, + api_client: PipelexAPIClient, + wire_response: ResponseBuilder, + patch_send: SendPatcher, + mocker: MockerFixture, + ) -> None: + item = _resolved(_URI_PDF).model_dump() + patch_send(api_client, wire_response(200, json_body={"items": [item]})) + _patch_storage(mocker, _serving(_PDF_BYTES)) + + async def _read() -> bytes: + async with api_client.fetch_artifact(_URI_PDF) as stream: + return await stream.read() + + assert asyncio.run(_read()) == _PDF_BYTES diff --git a/tests/unit/test_client_run_fallback.py b/tests/unit/test_client_run_fallback.py index acbe6d0..0689e9f 100644 --- a/tests/unit/test_client_run_fallback.py +++ b/tests/unit/test_client_run_fallback.py @@ -5,10 +5,11 @@ """ import asyncio -from typing import Any +from typing import Any, cast import httpx import pytest +from mthds.protocol.pipe_io_contracts import PipeIOContract from pytest_mock import MockerFixture from pipelex_sdk.client import PipelexAPIClient @@ -38,6 +39,39 @@ }, } +# The executed graph as the runner returns it inside `pipe_output` — the same document a local run +# writes as `graphspec.json`. +_GRAPH_SPEC: dict[str, Any] = { + "meta": {"format": "mthds", "mode": "live"}, + "nodes": [{"id": "pipe_1", "status": "COMPLETED"}], + "edges": [], +} + +# The runner's `pipe_io_artifacts` envelope: the three I/O artifacts together, since they share a +# key set and are built in one pass. The hosted results body relays them as three siblings instead. +_PIPE_IO_ARTIFACTS: dict[str, Any] = { + "pipe_io_contracts": { + "x.greet": { + "inputs": {}, + "output": { + "concept_ref": "native.Text", + "multiplicity": "single", + "item_count": None, + "optional": False, + "json_schema": {"type": "object", "properties": {"text": {"type": "string"}}}, + }, + }, + }, + "input_form": {"x.greet": {"fields": []}}, + "output_form": {"x.greet": {"field": {"name": "text", "kind": "prose", "required": True}}}, +} + + +def _execute_body_with(**pipe_output_extension_fields: object) -> dict[str, object]: + """The completed blocking body with Pipelex extension fields added onto its `pipe_output`.""" + base_pipe_output = cast("dict[str, object]", _EXECUTE_BODY["pipe_output"]) + return {**_EXECUTE_BODY, "pipe_output": {**base_pipe_output, **pipe_output_extension_fields}} + def _response(status_code: int, *, json: object = None, headers: dict[str, str] | None = None) -> httpx.Response: request = httpx.Request("GET", f"{_BASE_URL}/x") @@ -203,6 +237,158 @@ def test_blocking_fallback_without_usage_pair_defaults_to_none(self, mocker: Moc assert result.tokens_usages is None assert result.usage_assembly_error is None + def test_blocking_fallback_lifts_the_working_memory_off_pipe_output(self, mocker: MockerFixture) -> None: + """The standard declares `pipe_output.working_memory` required, so the SDK lifts it onto + `RunResults.working_memory` and the field always carries a value on this path — set in + `model_fields_set` like every other lifted field, where the hosted path may leave it unset. + """ + client = self._client() + mocker.patch.object( + client, + "_send", + mocker.AsyncMock(side_effect=[_response(200, json=_BARE_VERSION), _response(200, json=_EXECUTE_BODY)]), + ) + + result = asyncio.run(client.start_and_wait(pipe_code="p", mthds_contents=["x"])) + assert result.working_memory is not None + assert "working_memory" in result.model_fields_set + assert result.working_memory.root["result"].content == {"text": "hello"} + assert result.working_memory.aliases == {"main_stuff": "result"} + # The same memory the runner's own envelope carries — lifted, not copied or re-parsed. + assert result.pipe_output is not None + assert result.working_memory is result.pipe_output.working_memory + + def test_blocking_fallback_lifts_the_executed_graph_off_pipe_output(self, mocker: MockerFixture) -> None: + """The runner returns the executed graph inside `pipe_output`; the SDK lifts it onto + `RunResults.graph_spec` so the field carries the same document whichever path ran. + Regression: this path used to write `graph_spec=None` and drop the graph the runner had + already returned. + """ + client = self._client() + mocker.patch.object( + client, + "_send", + mocker.AsyncMock( + side_effect=[ + _response(200, json=_BARE_VERSION), + _response(200, json=_execute_body_with(graph_spec=_GRAPH_SPEC, graph_assembly_error=None)), + ] + ), + ) + + result = asyncio.run(client.start_and_wait(pipe_code="p", mthds_contents=["x"])) + assert result.graph_spec == _GRAPH_SPEC + assert result.graph_assembly_error is None + + def test_blocking_fallback_lifts_a_graph_assembly_failure_off_pipe_output(self, mocker: MockerFixture) -> None: + """A broken assembly and a run with no graph both leave `graph_spec` None; the lifted + `graph_assembly_error` is what separates them. + """ + client = self._client() + mocker.patch.object( + client, + "_send", + mocker.AsyncMock( + side_effect=[ + _response(200, json=_BARE_VERSION), + _response(200, json=_execute_body_with(graph_spec=None, graph_assembly_error="failed to assemble the graph for the run")), + ] + ), + ) + + result = asyncio.run(client.start_and_wait(pipe_code="p", mthds_contents=["x"])) + assert result.graph_spec is None + assert result.graph_assembly_error == "failed to assemble the graph for the run" + + def test_blocking_fallback_unwraps_the_io_artifacts_envelope_off_pipe_output(self, mocker: MockerFixture) -> None: + """The runner carries the three I/O artifacts in one `pipe_io_artifacts` envelope; the SDK + unwraps it onto the hosted shape's three sibling fields, typed from the standard, so each + artifact has one accessor whichever path ran. + """ + client = self._client() + mocker.patch.object( + client, + "_send", + mocker.AsyncMock( + side_effect=[ + _response(200, json=_BARE_VERSION), + _response(200, json=_execute_body_with(pipe_io_artifacts=_PIPE_IO_ARTIFACTS, pipe_io_artifacts_error=None)), + ] + ), + ) + + result = asyncio.run(client.start_and_wait(pipe_code="p", mthds_contents=["x"])) + assert result.pipe_io_contracts is not None + assert isinstance(result.pipe_io_contracts["x.greet"], PipeIOContract) + assert result.pipe_io_contracts["x.greet"].output.concept_ref == "native.Text" + assert result.input_form is not None + assert result.input_form["x.greet"].fields == [] + assert result.output_form is not None + assert result.output_form["x.greet"].field.name == "text" + assert result.pipe_io_artifacts_error is None + # The dumped artifacts are the envelope's members, verbatim: unwrapping moved them, nothing else. + dumped = result.model_dump(mode="json") + assert dumped["pipe_io_contracts"] == _PIPE_IO_ARTIFACTS["pipe_io_contracts"] + assert dumped["input_form"] == _PIPE_IO_ARTIFACTS["input_form"] + # The envelope itself is not a field of `RunResults`: it stays on the runner's own output. + assert "pipe_io_artifacts" not in dumped + assert result.pipe_output is not None + assert (result.pipe_output.model_extra or {})["pipe_io_artifacts"] == _PIPE_IO_ARTIFACTS + + def test_blocking_fallback_lifts_an_io_artifacts_build_failure_off_pipe_output(self, mocker: MockerFixture) -> None: + """A null envelope leaves all three artifacts None, and the lifted `pipe_io_artifacts_error` + is what separates a broken build from a run that described nothing. + """ + client = self._client() + mocker.patch.object( + client, + "_send", + mocker.AsyncMock( + side_effect=[ + _response(200, json=_BARE_VERSION), + _response( + 200, + json=_execute_body_with(pipe_io_artifacts=None, pipe_io_artifacts_error="failed to build the I/O artifacts for the run"), + ), + ] + ), + ) + + result = asyncio.run(client.start_and_wait(pipe_code="p", mthds_contents=["x"])) + assert result.pipe_io_contracts is None + assert result.input_form is None + assert result.output_form is None + assert result.pipe_io_artifacts_error == "failed to build the I/O artifacts for the run" + + def test_blocking_fallback_without_graph_or_artifacts_sets_every_lifted_field_to_none(self, mocker: MockerFixture) -> None: + """A blocking response whose `pipe_output` carries neither the graph pair nor the artifacts + (tracing off, or an older runner) maps every lifted field to None — never a validation + error — and, unlike a hosted body missing the keys, marks each as set: on the blocking path + the SDK always answers, as the JS twin writes `null` there. + """ + client = self._client() + mocker.patch.object( + client, + "_send", + mocker.AsyncMock(side_effect=[_response(200, json=_BARE_VERSION), _response(200, json=_EXECUTE_BODY)]), + ) + + result = asyncio.run(client.start_and_wait(pipe_code="p", mthds_contents=["x"])) + assert result.graph_spec is None + assert result.graph_assembly_error is None + assert result.pipe_io_contracts is None + assert result.input_form is None + assert result.output_form is None + assert result.pipe_io_artifacts_error is None + assert { + "graph_spec", + "graph_assembly_error", + "pipe_io_contracts", + "input_form", + "output_form", + "pipe_io_artifacts_error", + } <= result.model_fields_set + def test_blocking_fallback_raises_when_main_stuff_unlocatable(self, mocker: MockerFixture) -> None: """A completed blocking response whose `main_stuff_name` names no root stuff is a hard fail.""" client = self._client() diff --git a/tests/unit/test_method_files.py b/tests/unit/test_method_files.py index eb8528c..f4020d4 100644 --- a/tests/unit/test_method_files.py +++ b/tests/unit/test_method_files.py @@ -8,10 +8,16 @@ from __future__ import annotations import json +from typing import TYPE_CHECKING import pytest +from pydantic import ValidationError -from pipelex_sdk.product_models import MethodFile, parse_method_files, serialize_method_files +from pipelex_sdk import product_models +from pipelex_sdk.product_models import MethodData, MethodFile, parse_method_files, serialize_method_files + +if TYPE_CHECKING: + from pytest_mock import MockerFixture class TestMethodFiles: @@ -53,6 +59,40 @@ def test_malformed_source_raises_naming_the_expected_shape(self, bad_source: str with pytest.raises(ValueError, match=r"\{name, content\}"): parse_method_files(bad_source) + def test_a_source_too_deep_for_the_decoder_raises_value_error(self, mocker: MockerFixture) -> None: + """`json.loads` recurses where `JSON.parse` iterates, and `RecursionError` is not a `ValueError`. + + The depth that trips the real decoder is an interpreter build constant — it moved by an + order of magnitude in CPython 3.14 — so the decoder is made to raise rather than a nesting + literal being pinned: the conversion is what is under test, not the threshold. + """ + mocker.patch.object(product_models.json, "loads", side_effect=RecursionError) + + with pytest.raises(ValueError, match=r"nested too deeply"): + parse_method_files('[{"name": "a.py", "content": "x = 1"}]') + + def test_a_source_too_deep_for_the_decoder_fails_method_data_as_a_validation_error(self, mocker: MockerFixture) -> None: + """The parser's two call sites must fail the same way; pydantic converts only `ValueError`. + + Unconverted, the decoder's `RecursionError` escapes `MethodData` untouched and a caller's + `except ValidationError` around `get_method` misses a malformed stored source entirely. + """ + mocker.patch.object(product_models.json, "loads", side_effect=RecursionError) + + with pytest.raises(ValidationError): + MethodData.model_validate( + { + "method_id": "mt_x", + "name": "deep", + "mthds": 'domain = "demo"', + "org_id": "org_x", + "created_by_user_id": "user_x", + "python": '[{"name": "a.py", "content": "x = 1"}]', + "created_at": "2026-01-01T00:00:00Z", + "updated_at": "2026-01-01T00:00:00Z", + } + ) + def test_empty_list_serializes_to_the_clear_sentinel(self) -> None: """`""` is the platform's clear signal; the literal `"[]"` would not clear anything.""" assert serialize_method_files([]) == "" diff --git a/tests/unit/test_method_source.py b/tests/unit/test_method_source.py new file mode 100644 index 0000000..1b63a37 --- /dev/null +++ b/tests/unit/test_method_source.py @@ -0,0 +1,142 @@ +"""Tests for `method_source_to_contents` — the adapter over a stored method's polymorphic source. + +Mirrors `pipelex-sdk-js/tests/method-source.test.ts` case for case, so the two SDKs read one +stored method the same way, and adds the cases the JS twin leaves untested: an array entry whose +`content` is not a string — where the reading is the platform's and is pinned here because it was +ruled rather than inherited — and a source the Python decoder cannot take, which has no JS twin +at all because `JSON.parse` iterates where `json.loads` recurses. + +`MethodData.mthds` is either the catalog file-array or a bare `.mthds` bundle, and the reader +cannot ask which. Every case here is therefore about telling the two apart — above all the +degenerate ones, where reading a sentinel as a bundle yields a method that runs the string +`"[]"` as MTHDS source. +""" + +from __future__ import annotations + +import json +from typing import TYPE_CHECKING + +import pytest + +from pipelex_sdk import product_models +from pipelex_sdk.product_models import method_source_to_contents + +if TYPE_CHECKING: + from pytest_mock import MockerFixture + + +class TestMethodSourceToContents: + def test_a_raw_bundle_is_one_content(self) -> None: + """A `.mthds` source is not JSON, and the whole of it is the bundle.""" + source = 'domain = "demo"\nmain_pipe = "main"' + + assert method_source_to_contents(source) == [source] + + def test_a_catalog_array_yields_each_content_in_order(self) -> None: + source = json.dumps( + [ + {"name": "bundle.mthds", "content": 'domain = "demo"'}, + {"name": "pipes.mthds", "content": 'main_pipe = "main"'}, + ] + ) + + assert method_source_to_contents(source) == ['domain = "demo"', 'main_pipe = "main"'] + + def test_blank_contents_are_dropped_from_a_catalog_array(self) -> None: + """A zero-source file is not a bundle file — it would fail the MTHDS parse downstream.""" + source = json.dumps( + [ + {"name": "empty.mthds", "content": ""}, + {"name": "blank.mthds", "content": " \n\t"}, + {"name": "bundle.mthds", "content": 'domain = "demo"'}, + ] + ) + + assert method_source_to_contents(source) == ['domain = "demo"'] + + def test_an_all_blank_catalog_array_is_no_source(self) -> None: + source = json.dumps([{"name": "empty.mthds", "content": ""}]) + + assert method_source_to_contents(source) == [] + + def test_the_empty_catalog_array_is_no_source_not_a_bundle(self) -> None: + """`"[]"` is what the webapp editor writes for a method with no files. + + Read as a bundle it would send the two characters `[]` to the runner as MTHDS source. + """ + assert method_source_to_contents("[]") == [] + + @pytest.mark.parametrize("blank_source", ["", " \n", "\t "]) + def test_a_blank_source_is_no_source(self, blank_source: str) -> None: + assert method_source_to_contents(blank_source) == [] + + def test_a_contract_violating_none_is_no_source_not_a_crash(self) -> None: + """`MethodData` types the field `str`; a server that sends `null` must not raise here.""" + assert method_source_to_contents(None) == [] + + @pytest.mark.parametrize( + "source", + [ + "42", + '{"name": "x", "content": "y"}', + '"just a string"', + "null", + ], + ) + def test_non_array_json_is_a_raw_bundle(self, source: str) -> None: + """Valid JSON that is not the catalog form IS the source — a bundle may open with a digit or a brace.""" + assert method_source_to_contents(source) == [source] + + def test_a_json_array_of_non_entries_is_a_raw_bundle(self) -> None: + assert method_source_to_contents('["a", "b"]') == ['["a", "b"]'] + + @pytest.mark.parametrize( + "source", + [ + '[{"name": "a.mthds"}]', + '[{"content": "domain = \\"demo\\""}]', + ], + ) + def test_an_array_missing_a_catalog_key_is_a_raw_bundle(self, source: str) -> None: + """The catalog gate is key presence, so an entry lacking either key fails the whole array.""" + assert method_source_to_contents(source) == [source] + + def test_a_non_string_name_does_not_disqualify_the_catalog_form(self) -> None: + """Neither reference checks the `name`'s type, only that the key is there.""" + source = '[{"name": 1, "content": "domain = \\"demo\\""}]' + + assert method_source_to_contents(source) == ['domain = "demo"'] + + def test_a_partly_malformed_catalog_array_keeps_its_valid_siblings(self) -> None: + """The ruled reading: reproduce the server's reading of its own field. + + `@pipelex/sdk` and the platform's own resolver — the one that expands a `method_id` run, + and therefore the one that decides whether a stored method runs — both recognize an entry + by key presence alone, keep `"x = 1"` and drop the sibling whose `content` is a number. + This SDK now does the same, so one stored method has one observable reading wherever it is + read. An earlier strict reading took the whole array as a bundle instead; it refused to + lose a file silently, but it disagreed with the code that actually runs the method, which + is the property the ruling chose. The silent drop is a defect of the shared format and is + fixed in the platform, not worked around here. + """ + source = '[{"name": "a.mthds", "content": "x = 1"}, {"name": "b.mthds", "content": 123}]' + + assert method_source_to_contents(source) == ["x = 1"] + + def test_a_source_too_deep_for_the_decoder_is_a_raw_bundle_not_an_escape(self, mocker: MockerFixture) -> None: + """Python's JSON decoder recurses where `JSON.parse` iterates, and `RecursionError` is not a `ValueError`. + + Unconverted it escapes a function whose whole contract is that it never raises. The rule has + one owner — `parse_method_files` turns the decoder's `RecursionError` into the `ValueError` + its contract promises — so this pins the through-path rather than a mock of the delegate. The + depth that trips the real decoder is an interpreter build constant that moved by an order of + magnitude in CPython 3.14, so the decoder is made to raise instead of a nesting literal being + pinned: the conversion is what is under test, not the threshold. + """ + # A source that parses cleanly unpatched, so the assertion below fails if the patch or the + # conversion is absent: without them it reads as the catalog form and yields `["x = 1"]`. + source = '[{"name": "a.mthds", "content": "x = 1"}]' + mocker.patch.object(product_models.json, "loads", side_effect=RecursionError) + + assert method_source_to_contents(source) == [source] diff --git a/tests/unit/test_runs.py b/tests/unit/test_runs.py index 35536fa..0086b09 100644 --- a/tests/unit/test_runs.py +++ b/tests/unit/test_runs.py @@ -3,7 +3,11 @@ from typing import Any import pytest -from pydantic import TypeAdapter +from mthds.protocol.input_form import PipeInputFormDescriptor, ProseField +from mthds.protocol.output_form import PipeOutputFormDescriptor +from mthds.protocol.pipe_io_contracts import IOMultiplicity, PipeIOContract +from mthds.runners.api.models import DictWorkingMemoryAbstract +from pydantic import TypeAdapter, ValidationError from pipelex_sdk.runs import RunResults, RunStatus, TokensUsageRecord @@ -41,6 +45,60 @@ }, } +# The executed graph as the runner writes it to `graphspec.json`: opaque to this SDK, relayed as is. +_GRAPH_SPEC: dict[str, Any] = { + "meta": {"format": "mthds", "mode": "live"}, + "nodes": [{"id": "pipe_1", "status": "COMPLETED"}], + "edges": [], +} + +# The three I/O artifacts for a one-pipe library, in the standard's own shapes and keyed over the +# one shared `pipe_ref` set — the same fixture the JS SDK's tests carry, so the two mirrors are +# exercised on one document. +_PIPE_IO_CONTRACTS: dict[str, Any] = { + "x.greet": { + "inputs": {}, + "output": { + "concept_ref": "native.Text", + "multiplicity": "single", + "item_count": None, + "optional": False, + "json_schema": {"type": "object", "properties": {"text": {"type": "string"}}}, + }, + }, +} +_INPUT_FORM: dict[str, Any] = {"x.greet": {"fields": []}} +_OUTPUT_FORM: dict[str, Any] = {"x.greet": {"field": {"name": "text", "kind": "prose", "required": True}}} + +_GRAPH_ASSEMBLY_ERROR = "failed to assemble the graph for the run" +_PIPE_IO_ARTIFACTS_ERROR = "failed to build the I/O artifacts for the run" + +# The run's whole working memory as the platform relays the `working_memory.json` artifact: one +# root entry per named stuff, and the alias table that names which of them the main stuff is. The +# runner's per-stuff extras (`stuff_code`, `stuff_name`) are carried here too — they are not the +# standard's fields and must ride `model_extra` rather than fail the parse. +_WORKING_MEMORY: dict[str, Any] = { + "root": { + "topic": {"concept": "native.Text", "content": {"text": "kites"}, "stuff_name": "topic"}, + "greeting": {"concept": "x.Greeting", "content": {"text": "hello, kites"}, "stuff_name": "greeting"}, + }, + "aliases": {"main_stuff": "greeting"}, +} + +# Every field the hosted results body may carry beside the ones always present. Their absence, +# their explicit null and their value are three different readings, and the tests below pin each. +_OPTIONAL_RESULT_KEYS = ( + "graph_spec", + "graph_assembly_error", + "pipe_io_contracts", + "input_form", + "output_form", + "pipe_io_artifacts_error", + "tokens_usages", + "usage_assembly_error", + "working_memory", +) + class TestRuns: @pytest.mark.parametrize( @@ -116,6 +174,18 @@ def test_tokens_usage_record_keeps_unrated_cost_null(self) -> None: assert priced_at_zero.cost == 0.0 assert priced_at_zero.cost is not None + @pytest.mark.parametrize("cost", ["0.5", True, "NaN", float("nan"), float("inf")]) + def test_tokens_usage_record_rejects_a_cost_that_is_not_a_finite_number(self, cost: Any) -> None: + """`cost` is strict and finite: a coerced or non-finite value fails the record rather than reaching a sum.""" + with pytest.raises(ValidationError): + TokensUsageRecord.model_validate({**_RATED_RECORD, "cost": cost}) + + def test_tokens_usage_record_accepts_an_integer_cost(self) -> None: + """A JSON integer is a number, and strict mode keeps reading it as one.""" + priced = TokensUsageRecord.model_validate({**_RATED_RECORD, "cost": 1}) + + assert priced.cost == 1.0 + def test_run_results_validates_usage_records(self) -> None: """A results body's raw records become typed records; the null branch stays None.""" results = RunResults.model_validate( @@ -163,3 +233,251 @@ def test_run_results_defaults_usage_pair_to_none(self) -> None: assert results.tokens_usages is None assert results.usage_assembly_error is None assert results.pipe_output is None + + # ── The graph pair and the I/O artifacts ───────────────────── + + def test_run_results_declares_every_optional_field_absent_as_none_and_unset(self) -> None: + """A hosted body carrying none of the optional keys (an older platform, or a lean relay) + leaves each declared field `None` and out of `model_fields_set` — which is how a Python + reader tells "the platform relayed no such key" from "the platform relayed null", where the + JS twin reads `undefined`. + """ + results = RunResults.model_validate({"pipeline_run_id": "run_1", "main_stuff": {"answer": "42"}}) + + assert results.graph_spec is None + assert results.graph_assembly_error is None + assert results.pipe_io_contracts is None + assert results.input_form is None + assert results.output_form is None + assert results.pipe_io_artifacts_error is None + assert results.working_memory is None + assert results.model_fields_set == {"pipeline_run_id", "main_stuff"} + # Nothing rode `model_extra` either: every key of the body is declared. + assert results.model_extra == {} + + def test_run_results_reads_a_relayed_null_as_set(self) -> None: + """A key the platform relays as `null` (the artifact was not written, or the run described no + data) reads `None` too, but is IN `model_fields_set`: relayed-as-null and not-relayed are + distinguishable, which is what the JS `null` / `undefined` split carries. + """ + body: dict[str, Any] = {"pipeline_run_id": "run_1", "main_stuff": {"answer": "42"}} + for key in _OPTIONAL_RESULT_KEYS: + body[key] = None + results = RunResults.model_validate(body) + + for key in _OPTIONAL_RESULT_KEYS: + assert getattr(results, key) is None, key + assert set(_OPTIONAL_RESULT_KEYS) <= results.model_fields_set + + def test_run_results_types_the_io_artifacts_from_the_standard(self) -> None: + """The three artifacts parse into the standard's own models, imported from `mthds.protocol` + rather than restated: a contract entry is a `PipeIOContract`, a form entry a descriptor whose + field is the kind-discriminated node union. + """ + results = RunResults.model_validate( + { + "pipeline_run_id": "run_1", + "main_stuff": {"text": "hello"}, + "graph_spec": _GRAPH_SPEC, + "graph_assembly_error": None, + "pipe_io_contracts": _PIPE_IO_CONTRACTS, + "input_form": _INPUT_FORM, + "output_form": _OUTPUT_FORM, + "pipe_io_artifacts_error": None, + } + ) + + # The graph stays opaque and rides through unchanged. + assert results.graph_spec == _GRAPH_SPEC + assert results.graph_assembly_error is None + assert results.pipe_io_contracts is not None + contract = results.pipe_io_contracts["x.greet"] + assert isinstance(contract, PipeIOContract) + assert contract.inputs == {} + assert contract.output.concept_ref == "native.Text" + assert contract.output.multiplicity == IOMultiplicity.SINGLE + assert contract.output.item_count is None + assert contract.output.optional is False + assert contract.output.json_schema == {"type": "object", "properties": {"text": {"type": "string"}}} + assert results.input_form is not None + input_descriptor = results.input_form["x.greet"] + assert isinstance(input_descriptor, PipeInputFormDescriptor) + assert input_descriptor.fields == [] + assert results.output_form is not None + output_descriptor = results.output_form["x.greet"] + assert isinstance(output_descriptor, PipeOutputFormDescriptor) + assert isinstance(output_descriptor.field, ProseField) + assert output_descriptor.field.name == "text" + assert output_descriptor.field.required is True + assert results.pipe_io_artifacts_error is None + # The three share one key set — the reader's rule that they are taken together. + assert set(results.pipe_io_contracts) == set(results.input_form) == set(results.output_form) == {"x.greet"} + + def test_run_results_round_trips_the_io_artifacts_verbatim(self) -> None: + """Dumping the parsed artifacts in JSON mode gives back the relayed documents: typing them + adds nothing and drops nothing. + """ + results = RunResults.model_validate( + { + "pipeline_run_id": "run_1", + "main_stuff": {"text": "hello"}, + "pipe_io_contracts": _PIPE_IO_CONTRACTS, + "input_form": _INPUT_FORM, + "output_form": _OUTPUT_FORM, + } + ) + dumped = results.model_dump(mode="json", exclude_unset=True) + + assert dumped["pipe_io_contracts"] == _PIPE_IO_CONTRACTS + assert dumped["input_form"] == _INPUT_FORM + assert dumped["output_form"] == _OUTPUT_FORM + + @pytest.mark.parametrize( + ("key", "drifted_value"), + [ + pytest.param( + "pipe_io_contracts", + {"x.greet": {**_PIPE_IO_CONTRACTS["x.greet"], "not_in_this_standard": 1}}, + id="contract-member", + ), + pytest.param("input_form", {"x.greet": {"fields": [], "not_in_this_standard": 1}}, id="input-form-member"), + pytest.param( + "output_form", + {"x.greet": {"field": {"name": "text", "kind": "prose", "required": True}, "not_in_this_standard": 1}}, + id="output-form-member", + ), + ], + ) + def test_run_results_refuses_an_artifact_member_the_pinned_standard_does_not_define(self, key: str, drifted_value: dict[str, Any]) -> None: + """The artifacts are closed shapes: a member the pinned `mthds` does not define is version + drift, refused at the parse of the whole results body rather than read half-way — the same + ruling the validate report follows. The envelope around them stays open (see the extras test + below), so strictness composes rather than spreads. + """ + with pytest.raises(ValidationError): + RunResults.model_validate({"pipeline_run_id": "run_1", "main_stuff": {}, key: drifted_value}) + + def test_run_results_stays_extension_open_around_the_typed_artifacts(self) -> None: + """Declaring typed fields does not close the envelope: a key the SDK does not name still + parses and rides `model_extra`, so the platform can add an artifact without a client release. + """ + results = RunResults.model_validate( + { + "pipeline_run_id": "run_1", + "main_stuff": {}, + "pipe_io_contracts": _PIPE_IO_CONTRACTS, + "some_future_artifact": {"k": "v"}, + } + ) + + assert results.model_extra == {"some_future_artifact": {"k": "v"}} + + @pytest.mark.parametrize( + ("graph_spec", "graph_assembly_error"), + [ + pytest.param(None, None, id="no-graph-or-not-written"), + pytest.param(None, _GRAPH_ASSEMBLY_ERROR, id="assembly-broke"), + pytest.param(_GRAPH_SPEC, None, id="assembled"), + ], + ) + def test_run_results_keeps_the_graph_null_semantics_distinct(self, graph_spec: dict[str, Any] | None, graph_assembly_error: str | None) -> None: + """A run with no graph and a run whose assembly broke both carry a null `graph_spec`; + `graph_assembly_error` is the only field that tells them apart, as `usage_assembly_error` + does for the usage pair. + """ + results = RunResults.model_validate( + { + "pipeline_run_id": "run_1", + "main_stuff": {}, + "graph_spec": graph_spec, + "graph_assembly_error": graph_assembly_error, + } + ) + + assert results.graph_spec == graph_spec + assert results.graph_assembly_error == graph_assembly_error + + def test_run_results_keeps_the_artifact_null_semantics_distinct(self) -> None: + """Three null artifacts alone cannot say whether the run described no data or the build + broke; `pipe_io_artifacts_error` is the only field that tells the two apart. + """ + described_nothing = RunResults.model_validate( + { + "pipeline_run_id": "run_1", + "main_stuff": {}, + "pipe_io_contracts": None, + "input_form": None, + "output_form": None, + "pipe_io_artifacts_error": None, + } + ) + build_broke = RunResults.model_validate( + { + "pipeline_run_id": "run_1", + "main_stuff": {}, + "pipe_io_contracts": None, + "input_form": None, + "output_form": None, + "pipe_io_artifacts_error": _PIPE_IO_ARTIFACTS_ERROR, + } + ) + + assert described_nothing.pipe_io_contracts is None + assert described_nothing.pipe_io_artifacts_error is None + assert build_broke.pipe_io_contracts is None + assert build_broke.input_form is None + assert build_broke.output_form is None + assert build_broke.pipe_io_artifacts_error == _PIPE_IO_ARTIFACTS_ERROR + + # ── The working memory ─────────────────────────────────────── + + def test_run_results_types_the_relayed_working_memory_from_the_standard(self) -> None: + """The hosted `working_memory.json` artifact parses into the standard's own + `DictWorkingMemoryAbstract`, imported from `mthds` rather than restated: a root of named + stuffs, each with its concept ref and its content, and the alias table that says which of + them the main stuff is. The runner's per-stuff extras ride `model_extra`. + """ + results = RunResults.model_validate( + { + "pipeline_run_id": "run_1", + "main_stuff": {"text": "hello, kites"}, + "working_memory": _WORKING_MEMORY, + } + ) + + memory = results.working_memory + assert isinstance(memory, DictWorkingMemoryAbstract) + assert "working_memory" in results.model_fields_set + assert set(memory.root) == {"topic", "greeting"} + greeting = memory.root["greeting"] + assert greeting.concept_ref == "x.Greeting" + assert greeting.content == {"text": "hello, kites"} + assert greeting.model_extra == {"stuff_name": "greeting"} + assert memory.aliases == {"main_stuff": "greeting"} + # The alias table is what turns the `main_stuff` role into a root key, and the content it + # names is the one `main_stuff` already resolved. + assert memory.root[memory.aliases["main_stuff"]].content == results.main_stuff + + def test_run_results_round_trips_the_working_memory_verbatim(self) -> None: + """Dumping the parsed memory in JSON mode gives back the relayed artifact: typing it adds + nothing and drops nothing, extras included. + """ + results = RunResults.model_validate({"pipeline_run_id": "run_1", "main_stuff": {"text": "hello, kites"}, "working_memory": _WORKING_MEMORY}) + dumped = results.model_dump(mode="json", exclude_unset=True) + + assert dumped["working_memory"] == _WORKING_MEMORY + + def test_run_results_reads_a_relayed_null_working_memory_as_set(self) -> None: + """A `working_memory` relayed as `null` — the platform has no such artifact for this run — + reads `None` and IS in `model_fields_set`, which is what separates it from a body that + carried no such key at all. + """ + relayed_null = RunResults.model_validate({"pipeline_run_id": "run_1", "main_stuff": {}, "working_memory": None}) + never_relayed = RunResults.model_validate({"pipeline_run_id": "run_1", "main_stuff": {}}) + + assert relayed_null.working_memory is None + assert "working_memory" in relayed_null.model_fields_set + assert never_relayed.working_memory is None + assert "working_memory" not in never_relayed.model_fields_set + # It is a declared field either way — an absent key never falls through to `model_extra`. + assert never_relayed.model_extra == {} diff --git a/tests/unit/test_usage.py b/tests/unit/test_usage.py new file mode 100644 index 0000000..60eb6e2 --- /dev/null +++ b/tests/unit/test_usage.py @@ -0,0 +1,304 @@ +"""Tests for pipelex_sdk.usage — the null-aware fold of a run's usage pair into one summary.""" + +from typing import Any + +import pytest + +from pipelex_sdk.errors import FieldNotIncludedError +from pipelex_sdk.runs import RunResults +from pipelex_sdk.usage import UsageSummaryState, summarize_usage + +# A record in the shape the current runtime emits: the full contract key set, a value the runtime +# has none of sent as an explicit null. Overridden per case by `_record`. +_FULL_RECORD: dict[str, Any] = { + "model_type": "llm", + "inference_model_name": "test-model", + "inference_model_id": "test-model-2026-01-01", + "pipe_code": "test_domain.summarize", + "job_category": "llm_job", + "unit_job_id": "llm_gen_text", + "nb_tokens_by_category": {"input": 100, "output": 20}, + "cost": 0.01, + "started_at": "2026-09-22T10:00:01+00:00", + "completed_at": "2026-09-22T10:00:03+00:00", +} + +# A durable artifact written BEFORE the usage contract shipped, relayed verbatim ever since: no +# computed `cost`, no flattened `pipe_code`, and the legacy relics riding `model_extra`. +_PRE_CONTRACT_RECORD: dict[str, Any] = { + "model_type": "llm", + "inference_model_name": "legacy-model", + "nb_tokens_by_category": {"input": 40, "output": 8}, + "unit_costs": {"input": 3.0, "output": 15.0}, + "job_metadata": {"pipe_code": "legacy_domain.summarize", "job_category": "llm_job"}, +} + + +def _record(**overrides: Any) -> dict[str, Any]: + """One wire record, the full key set with the case's own values written over it.""" + return {**_FULL_RECORD, **overrides} + + +def _results(**body: Any) -> RunResults: + """A completed run's results body, validated from the wire so `model_fields_set` is truthful.""" + return RunResults.model_validate({"pipeline_run_id": "run_1", "main_stuff": {"text": "out"}, **body}) + + +class TestSummarizeUsage: + def test_a_body_that_did_not_carry_the_key_raises(self) -> None: + results = _results() + + with pytest.raises(FieldNotIncludedError) as exc_info: + summarize_usage(results) + + assert exc_info.value.field_name == "tokens_usages" + assert "tokens_usages" in str(exc_info.value) + + def test_a_key_relayed_as_null_is_a_value_and_does_not_raise(self) -> None: + summary = summarize_usage(_results(tokens_usages=None, usage_assembly_error=None)) + + assert summary.state == UsageSummaryState.UNAVAILABLE + assert summary.total_cost_usd is None + + @pytest.mark.parametrize( + ("tokens_usages", "expected_state", "expected_cost", "expected_input", "expected_output", "expected_calls"), + [ + pytest.param(None, UsageSummaryState.UNAVAILABLE, None, None, None, 0, id="unavailable"), + pytest.param([], UsageSummaryState.NO_INFERENCE, 0.0, 0, 0, 0, id="no_inference"), + pytest.param([_FULL_RECORD], UsageSummaryState.RECORDS, 0.01, 100, 20, 1, id="records"), + ], + ) + def test_each_state_reports_its_own_totals( + self, + tokens_usages: list[dict[str, Any]] | None, + expected_state: UsageSummaryState, + expected_cost: float | None, + expected_input: int | None, + expected_output: int | None, + expected_calls: int, + ) -> None: + summary = summarize_usage(_results(tokens_usages=tokens_usages, usage_assembly_error=None)) + + assert summary.state == expected_state + assert summary.total_cost_usd == expected_cost + assert summary.cost_partial is False + assert summary.tokens.input == expected_input + assert summary.tokens.output == expected_output + assert summary.calls == expected_calls + assert summary.assembly_error is None + + def test_an_empty_list_is_a_run_that_did_no_inference_not_an_unrated_one(self) -> None: + summary = summarize_usage(_results(tokens_usages=[], usage_assembly_error=None)) + + assert summary.state == UsageSummaryState.NO_INFERENCE + assert summary.total_cost_usd == 0.0 + assert summary.cost_partial is False + assert summary.tokens.input == 0 + assert summary.tokens.output == 0 + assert summary.by_pipe == [] + + def test_an_unavailable_run_knows_nothing_at_all(self) -> None: + summary = summarize_usage(_results(tokens_usages=None, usage_assembly_error=None)) + + assert summary.total_cost_usd is None + assert summary.cost_partial is False + assert summary.tokens.input is None + assert summary.tokens.output is None + assert summary.calls == 0 + assert summary.by_pipe == [] + + def test_it_sums_the_priced_calls(self) -> None: + summary = summarize_usage(_results(tokens_usages=[_record(cost=0.25), _record(cost=0.5)], usage_assembly_error=None)) + + assert summary.state == UsageSummaryState.RECORDS + assert summary.total_cost_usd == 0.75 + assert summary.cost_partial is False + assert summary.calls == 2 + + def test_a_zero_cost_stays_priced_rather_than_unrated(self) -> None: + summary = summarize_usage(_results(tokens_usages=[_record(cost=0), _record(cost=0)], usage_assembly_error=None)) + + assert summary.total_cost_usd == 0.0 + assert summary.cost_partial is False + + def test_a_run_whose_every_call_is_unrated_reports_no_total(self) -> None: + summary = summarize_usage(_results(tokens_usages=[_record(cost=None), _record(cost=None)], usage_assembly_error=None)) + + assert summary.state == UsageSummaryState.RECORDS + assert summary.total_cost_usd is None + assert summary.cost_partial is False + assert summary.calls == 2 + + def test_mixing_priced_and_unrated_calls_flags_the_total_as_partial(self) -> None: + summary = summarize_usage( + _results(tokens_usages=[_record(cost=0.5), _record(cost=None), _record(cost=0)], usage_assembly_error=None), + ) + + assert summary.total_cost_usd == 0.5 + assert summary.cost_partial is True + + def test_only_the_joined_input_and_output_categories_are_summed(self) -> None: + summary = summarize_usage( + _results( + tokens_usages=[ + _record(nb_tokens_by_category={"input": 1000, "input_cached": 800, "output": 50, "output_reasoning": 30}), + _record(nb_tokens_by_category={"input": 200, "output": 10, "some_future_category": 7}), + ], + usage_assembly_error=None, + ), + ) + + assert summary.tokens.input == 1200 + assert summary.tokens.output == 60 + + def test_a_category_no_record_reported_totals_to_none(self) -> None: + summary = summarize_usage( + _results( + tokens_usages=[_record(nb_tokens_by_category={"output": 12}), _record(nb_tokens_by_category=None)], + usage_assembly_error=None, + ), + ) + + assert summary.tokens.input is None + assert summary.tokens.output == 12 + + def test_a_reported_zero_stays_apart_from_an_unreported_count(self) -> None: + summary = summarize_usage(_results(tokens_usages=[_record(nb_tokens_by_category={"input": 0, "output": 0})], usage_assembly_error=None)) + + assert summary.tokens.input == 0 + assert summary.tokens.output == 0 + + @pytest.mark.parametrize("tokens_usages", [None, [], [_FULL_RECORD]], ids=["unavailable", "no_inference", "records"]) + def test_the_assembly_error_is_relayed_verbatim_in_every_state(self, tokens_usages: list[dict[str, Any]] | None) -> None: + summary = summarize_usage(_results(tokens_usages=tokens_usages, usage_assembly_error="failed to read pipeline events")) + + assert summary.assembly_error == "failed to read pipeline events" + + def test_an_absent_assembly_error_reads_as_none(self) -> None: + summary = summarize_usage(_results(tokens_usages=[])) + + assert summary.assembly_error is None + assert summary.state == UsageSummaryState.NO_INFERENCE + + def test_it_groups_calls_per_pipe_and_gathers_the_unattributed_ones_under_none(self) -> None: + summary = summarize_usage( + _results( + tokens_usages=[ + _record(pipe_code="extract", cost=0.1, nb_tokens_by_category={"input": 10, "output": 1}), + _record(pipe_code=None, cost=0.2, nb_tokens_by_category={"input": 20, "output": 2}), + _record(pipe_code="extract", cost=0.3, nb_tokens_by_category={"input": 30, "output": 3}), + _record(pipe_code=None, cost=None, nb_tokens_by_category=None), + ], + usage_assembly_error=None, + ), + ) + + extract_row, unattributed_row = summary.by_pipe + assert extract_row.pipe_code == "extract" + assert extract_row.total_cost_usd == 0.1 + 0.3 + assert extract_row.cost_partial is False + assert extract_row.tokens.input == 40 + assert extract_row.tokens.output == 4 + assert extract_row.calls == 2 + assert unattributed_row.pipe_code is None + assert unattributed_row.total_cost_usd == 0.2 + assert unattributed_row.cost_partial is True + assert unattributed_row.tokens.input == 20 + assert unattributed_row.tokens.output == 2 + assert unattributed_row.calls == 2 + + def test_it_sorts_by_cost_descending_with_unrated_pipes_after_every_priced_one(self) -> None: + summary = summarize_usage( + _results( + tokens_usages=[ + _record(pipe_code="unrated", cost=None), + _record(pipe_code="cheap", cost=0.01), + _record(pipe_code="free", cost=0), + _record(pipe_code="expensive", cost=2), + ], + usage_assembly_error=None, + ), + ) + + assert [row.pipe_code for row in summary.by_pipe] == ["expensive", "cheap", "free", "unrated"] + assert [row.total_cost_usd for row in summary.by_pipe] == [2.0, 0.01, 0.0, None] + + def test_it_breaks_a_cost_tie_on_call_count_then_on_pipe_code_with_the_unattributed_group_last(self) -> None: + summary = summarize_usage( + _results( + tokens_usages=[ + _record(pipe_code=None, cost=0.5), + _record(pipe_code="beta", cost=0.5), + _record(pipe_code="alpha", cost=0.5), + _record(pipe_code="busy", cost=0.25), + _record(pipe_code="busy", cost=0.25), + ], + usage_assembly_error=None, + ), + ) + + assert [(row.pipe_code, row.calls) for row in summary.by_pipe] == [("busy", 2), ("alpha", 1), ("beta", 1), (None, 1)] + + def test_it_orders_unrated_pipes_among_themselves_by_call_count_then_pipe_code(self) -> None: + summary = summarize_usage( + _results( + tokens_usages=[ + _record(pipe_code="bbb", cost=None), + _record(pipe_code="aaa", cost=None), + _record(pipe_code="ccc", cost=None), + _record(pipe_code="ccc", cost=None), + ], + usage_assembly_error=None, + ), + ) + + assert [row.pipe_code for row in summary.by_pipe] == ["ccc", "aaa", "bbb"] + assert summary.total_cost_usd is None + + def test_a_pre_contract_record_counts_as_unrated_and_unattributed(self) -> None: + summary = summarize_usage( + _results( + tokens_usages=[_PRE_CONTRACT_RECORD, _record(pipe_code="summarize", cost=0.02)], + usage_assembly_error=None, + ), + ) + + assert summary.state == UsageSummaryState.RECORDS + assert summary.total_cost_usd == 0.02 + assert summary.cost_partial is True + assert summary.tokens.input == 140 + assert summary.tokens.output == 28 + priced_row, legacy_row = summary.by_pipe + assert priced_row.pipe_code == "summarize" + assert priced_row.total_cost_usd == 0.02 + assert priced_row.calls == 1 + assert legacy_row.pipe_code is None + assert legacy_row.total_cost_usd is None + assert legacy_row.cost_partial is False + assert legacy_row.tokens.input == 40 + assert legacy_row.tokens.output == 8 + + def test_a_record_carrying_no_field_at_all_is_tolerated(self) -> None: + summary = summarize_usage(_results(tokens_usages=[{}], usage_assembly_error=None)) + + assert summary.state == UsageSummaryState.RECORDS + assert summary.total_cost_usd is None + assert summary.cost_partial is False + assert summary.tokens.input is None + assert summary.tokens.output is None + assert summary.calls == 1 + assert len(summary.by_pipe) == 1 + assert summary.by_pipe[0].pipe_code is None + assert summary.by_pipe[0].calls == 1 + + def test_it_takes_a_completed_runs_results_without_mutating_them(self) -> None: + results = _results( + tokens_usages=[_record(pipe_code="cheap", cost=0.01), _record(pipe_code="expensive", cost=1)], + usage_assembly_error=None, + ) + snapshot = results.model_dump() + + summary = summarize_usage(results) + + assert [row.pipe_code for row in summary.by_pipe] == ["expensive", "cheap"] + assert results.model_dump() == snapshot diff --git a/uv.lock b/uv.lock index d0a1725..03afa00 100644 --- a/uv.lock +++ b/uv.lock @@ -303,7 +303,7 @@ wheels = [ [[package]] name = "pipelex-sdk" -version = "0.10.0" +version = "0.10.1" source = { editable = "." } dependencies = [ { name = "httpx" }, diff --git a/wip/method-source-adapter/review-deferrals.md b/wip/method-source-adapter/review-deferrals.md new file mode 100644 index 0000000..cbd0399 --- /dev/null +++ b/wip/method-source-adapter/review-deferrals.md @@ -0,0 +1,34 @@ +--- +status: active +item: L-260907-28adb9 +--- + +# Deferred review findings — `method_source_to_contents` and the catalog decoder + +What review rounds on `pipelex-sdk-python#31` confirmed but did not fix, with enough detail to pick each up cold. Everything here was verified against the code; nothing rests on a reviewer's word alone. Findings owned by another repo are not here — they are ledger items (`L-260913-6ec559`, `L-260913-37ca47`). + +## The `RecursionError` conversion diagnoses the wrong cause when the caller's stack is deep + +`pipelex_sdk/product_models.py`, `_decode_method_source`. + +`json.loads` raises `RecursionError` for two different reasons that are indistinguishable at the point of the catch: the *source* is nested past what the decoder can descend, or the *caller* was already near the recursion limit when it called. The conversion reports both as "Method file source is nested too deeply to decode", which is a false statement about the data in the second case, and `method_source_to_contents` then reads a perfectly valid catalog array as a raw bundle and returns it without an error. + +Reproduced at the default recursion limit of 1000: a recursive caller that reaches depth 995 and then passes the valid catalog string `[{"name": "a.py", "content": "x = 1"}]` gets `ValueError("Method file source is nested too deeply to decode…")` from `parse_method_files`, and through `method_source_to_contents` gets `['[{"name": "a.py", "content": "x = 1"}]']` instead of `['x = 1']`. At depth 990 both behave correctly; at 998 a bare `RecursionError` escapes from the frame setup before `json.loads` is even reached, so the "never raises" contract has a floor no function in Python can lift. + +Deferred because the precondition is a program already within a handful of frames of the recursion limit, where essentially nothing behaves. It is recorded rather than fixed because the two causes genuinely cannot be told apart from inside the handler, so any "fix" is either a guard against an untestable state or a reworded message. If it is ever picked up, the honest change is the message — say what was observed ("the decoder ran out of stack") rather than what was inferred about the payload. + +## The integer digit-cap `ValueError` bypasses both curated handlers + +`pipelex_sdk/product_models.py`, `_decode_method_source`. + +`json.loads` on an integer literal past CPython's 4300-digit conversion cap raises a bare `ValueError` ("Exceeds the limit (4300 digits) for integer string conversion") which is neither a `JSONDecodeError` nor a `RecursionError`, so it passes through both `except` clauses untouched and reaches the caller with CPython's message and no `__cause__` chain instead of one naming the expected shape. + +The type contract still holds — it *is* a `ValueError`, which is what the docstring promises and what pydantic converts — and behaviour matches the platform, whose own `except ValueError` catches it identically, so `method_source_to_contents` reads it as a bundle on both sides. Deferred as message quality rather than correctness. Worth noting that its `RecursionError` sibling got an owner in the same commit while this one did not, which is the only reason it looks like an oversight. + +## The tests monkeypatch stdlib `json.loads` process-wide + +`tests/unit/test_method_files.py` and `tests/unit/test_method_source.py`, the deep-nesting tests. + +`product_models.json` *is* the stdlib `json` module object, so `mocker.patch.object(product_models.json, "loads", …)` replaces the attribute on the shared module rather than on a seam local to the module under test. Anything else running in the same process during those tests would get the fake. It is harmless here — pydantic-core is Rust and never routes through `json.loads`, and pytest-mock restores the attribute at teardown — and it is the only way to make the real decoder fail without pinning a nesting depth that is an interpreter build constant. + +Deferred as test hygiene. A module-local seam is available if it is ever wanted: replace the module reference in the module's own namespace (`mocker.patch.object(product_models, "json", …)` with a stub exposing `loads` and `JSONDecodeError`) rather than the attribute on the shared module.