Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
158 changes: 50 additions & 108 deletions .claude/skills/release/SKILL.md

Large diffs are not rendered by default.

5 changes: 5 additions & 0 deletions .gitattributes
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
# The codegen check's fixture tree is genuine `pipelex codegen types` output, committed byte for byte so
# the suite re-proves the verdict over real engine bytes. Its content hashes cover those bytes exactly, so
# a checkout that rewrote its line endings (`core.autocrlf=true`) would hand the suite a different tree
# than the one the stamp and the lock were computed over. Pin it.
tests/unit/data/real_codegen_tree/** -text
25 changes: 25 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,30 @@
# Changelog

## [v0.10.0] - 2026-09-13

### Added

- **`run_codegen_check`, the offline codegen drift check**: `pipelex_sdk.codegen_check.run_codegen_check(root=…)` verifies a generated tree against its `codegen.lock` by hashing alone — no engine boot, no network, no API key, and no dependency on `pipelex`. That last point is the reason it exists: until now the only offline gate a Python project could use was the `pipelex` CLI, the whole runtime, which is exactly the dependency a hosted-API consumer took this SDK to avoid, so the same integration advice left a TypeScript project with a CI drift gate and a Python one without. It is the counterpart of `write_codegen_tree` and takes the directory that writer wrote into. The verdict rides a structured `CodegenCheckReport` — `lock_found`, `is_current`, and a `drifts` list whose categories are `missing`, `modified`, `hand-edited` and `orphan`, at most one per locked artifact, locked drifts first in ascending path order and then orphans — with the lock header's `crate_fingerprint` and `engine_version` surfaced so a caller can ask the one question the offline check cannot, whether the tree still matches what the method resolves to. A lock that cannot be read, or a tree whose paths are not safe and canonical, raises `CodegenLockError`: the absence of a verdict rather than a drift. It is a mirror of `pipelex`'s own `run_codegen_check` down to the drift sentences, so a consumer reads the same report from either, and the agreement was established by running both over scenario trees derived from real `pipelex codegen types` output, one drift class at a time. Two divergences are deliberate and documented: a relaxation inherited from `@pipelex/sdk`, which accepts a well-formed projection line whose axes are outside this SDK's vocabulary rather than report a tree generated by a newer engine as entirely hand-edited; and a tightening of this reader's own, which refuses a Python artifact declaring a PEP 263 source encoding, because that declaration chooses the codec CPython decodes the file with and so makes the header's comment-prefix gate unsound — a tree can otherwise read as current while importing it executes a statement hidden in the unhashed header. Unlike the JS export, which is pure because it must stay browser-bundleable, this one owns the tree walk and the decoding, so none of that SDK's caller obligations carry over; an unreadable file or directory under the root is therefore a `CodegenLockError` rather than the reference's bare `PermissionError`, leaving a CI caller one class to catch.
- **`pipelex_sdk.codegen_stamp` gains the stamp reader the check needs** — `parse_stamped`, `compute_content_hash`, `is_stampable_artifact_path` and the `ParsedStamp` model, beside the ownership predicate and stampable suffixes it already carried — and `CodegenLock` gains `hash_by_path()`. A `codegen.lock` is now read in text mode, so universal-newline translation folds a CRLF away before the TOML parser sees it, as `pipelex`'s reader does; the writer still compares raw bytes, deliberately, because translation there would make one response two different trees across platforms.
- **`write_codegen_tree`, the verbatim codegen tree writer**: `pipelex_sdk.codegen_writer.write_codegen_tree(report, output_dir=…)` writes a valid `codegen()` response to disk byte for byte — every artifact at its `path`, the lock as `codegen.lock` — so a Python project regenerates a tree identical to a local `pipelex codegen types` run without installing `pipelex`. Like that command it validates every path before writing, raises `CodegenError` rather than overwrite a file codegen does not own, rewrites only what changed, and prunes stamped artifacts the previous lock tracked; unlike it, it also checks the paths it is about to prune and requires the response's lock to track exactly its artifacts before the first write; the lock format and path rules it reads are public in `pipelex_sdk.codegen_lock` and `pipelex_sdk.codegen_stamp`.
- **`prepare_inputs` takes the method three ways.** Beside inline `files`, it accepts a `method_ref` address (resolved by the runner) or a stored `method_id` (resolved by the hosted platform) — exactly one per call, all three server-resolved, nothing expanded client-side. Empty is absent (`files=[]`, a blank `method_ref` / `method_id`) but the wrong type is not: none, several, or a non-string selector raises `InputPreparationError` before any request leaves the process — reading a mistyped selector as absent would let the exclusivity check pass and prepare against the wrong method. A non-string `pipe_ref` is refused on the same boundary, rather than being absorbed by the pipe defaulting. A method addressed by URL that declares a file input now has an input-preparation path; previously it had none, even though the request beneath accepted the address.
- **`PipelexValidationReport.default_pipe_ref`** — the qualified `pipe_ref` a caller gets by omitting the pipe selector, or `null` when the closure declares none or several. Optional and read leniently: a runner that predates the field sends nothing, and `prepare_inputs` falls back to the opaque `bundle_blueprint.main_pipe`.

### Changed

- **Breaking: `prepare_inputs` reads its signature from the input-form descriptor, not the inputs template.** It composes one `POST /v1/validate` with `views: ["input_form"]` and `allow_signatures=True`, and walks the standard's `InputForm` artifact — `document` / `image` mark a file position, `object` recurses through `fields`, `list` through `item`, everything else passes through. Source-compatible for every caller passing `files`; the SDK no longer calls `/v1/build/inputs` at runtime. A valid report carrying no descriptor is an error naming the pipelex-api floor, never a silent degrade to "no uploads".
- **Breaking: `build_inputs` and its models are removed.** `client.build_inputs`, `BuildInputsRequest`, `BuildInputsValidReport`, `BuildInputsResponse`, `BuildInputsResponseAdapter` and `InputsTemplateFormat` are gone — the route wrapper existed only to be the signature source `prepare_inputs` read, and nothing calls it now. This is the Python SDK's step of the workspace program retiring `/v1/build/*`; a caller that still needs a fill-in template projects one from the descriptor with `mthds.protocol.inputs_template` (`render_inputs_template` / `project_inputs_template`, in both the compact and explicit shapes, as JSON or TOML), which is also where `InputsTemplateFormat` now lives. The wrapper is not a capability lost but one relocated to the standard's own package — and projected client-side, so a method reached by `method_ref` or `method_id` gets a template with no server round-trip at all.
- **Breaking: the shared crate envelope moved to `pipelex_sdk.crate_models`.** `MthdsFileItem`, `CrateRequestBase` and `CrateInvalidReport` now live beside the routes that use them (`/v1/resolve`, `/v1/codegen`) and `pipelex_sdk/build_models.py` is deleted — a module named for the build routes could not go on holding the envelope after they left. The models themselves are unchanged; update the import path.
- **Breaking: a canonical file dict nested inside a `Dynamic` input is no longer uploaded.** Such an input is `kind: "unknown"` in the descriptor — the standard's escape hatch — and the walk does not enter it. Uploading on the strength of a `url` key is the value-shape guess this change removes; a caller with a Dynamic input uploads with `upload_file` first and passes the storage URI, which `docs/input-preparation.md` has always prescribed.
- **`prepare_inputs` accepts the explicit `{concept, content}` input envelope**, not only compact values, closing a parity gap with the JS SDK. An agent that fills an explicit template — the shape the hosted console and MCP hand out — can now hand it straight back; previously every file-bearing envelope position raised `InputPreparationError: Unsupported value at a file input … got dict`. The envelope's `content` is interpreted identically and preserved on output, so the concept annotation rides through to the run.
- **The documentation is rewritten around the descriptor.** `docs/input-preparation.md` now describes the three call shapes, the signature call, pipe selection and its manifest-only `main_pipe` gap, the envelope, and why the template was the wrong signature source; `docs/architecture.md` follows the removal. It also catches up with v0.9.0, which shipped `output_form` and the `mthds` 0.13.0 bump without touching the docs: the views list is no longer described as having one token, the report's typed fields now include `output_form` and `default_pipe_ref`, and the pipe I/O contracts no longer claim an output carries no schema — 0.13.0 made `json_schema` required there, reversing the reasoning the doc still quoted.
- **Requires `mthds` 0.14.0 (breaking).** The pin moves from 0.13.0. Nothing in this client changed with it: the release's substantive cut is `MTHDS_STANDARD_VERSION` going from `1.0.0` to `2.0.0`, which this SDK never reads — it stamps no crates and ships no manifest — and the new `is_mthds_version_satisfied` helper and the `parse_constraint` whitespace fix sit in `mthds.package.manifest.schema`, which nothing here imports. `PROTOCOL_VERSION` stays at `0.6.0`, so the routes and the wire contract are untouched. The move still matters to anyone installing this package: `pipelex` pins `mthds` exactly too and has already moved to 0.14.0, so two exact pins on different versions made the pair unresolvable — this is what lets `pipelex` and `pipelex-sdk` co-install again.

### Security

- **An optional nested file field is now uploaded.** The required-only inputs template never rendered one, so its file position was invisible and the caller's local path travelled to the runner as a literal string. The descriptor states `required: false` and the walk enters it.
- **A text field merely *named* `url` is no longer read from disk.** The template marked a file position by rendering a `url`-bearing dict — a side effect of the field's *name*, not of its concept — so a path-shaped text value was uploaded. `kind: "text"` ends that.

## [v0.9.0] - 2026-09-02

### Added
Expand Down
40 changes: 39 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -69,6 +69,43 @@ report = await client.validate(method_ref="github.com/Pipelex/methods/documents@
report = await client.validate(method_id="mt_123")
```

### Generate typed code into your project

`codegen()` projects a method into stamped typed artifacts plus their `codegen.lock`, and `write_codegen_tree` writes that response to disk verbatim, so the tree is byte-identical to a local `pipelex codegen types` run and no `pipelex` install is needed:

```python
from pathlib import Path

from pipelex_sdk.codegen_writer import write_codegen_tree
from pipelex_sdk.crate_models import CodegenRequest, CodegenValidReport

report = await client.codegen(CodegenRequest(method_ref="github.com/Pipelex/methods/documents@v0.1.0", target="python-pydantic"))
if not isinstance(report, CodegenValidReport):
raise SystemExit(report.message)
written = write_codegen_tree(report, output_dir=Path("src/generated/documents"))
print(written.written, written.removed)
```

It never overwrites a file codegen does not own, rewrites only what changed, and prunes stamped artifacts that dropped out of the set. Commit the tree; do not run a formatter over it.

### Gate a committed tree in CI, with no key and no `pipelex`

`run_codegen_check` is the writer's counterpart: pure hashing over the tree and its lock, so it boots no engine, reaches no network and needs no API key. Point it at the directory you generated into:

```python
from pathlib import Path

from pipelex_sdk.codegen_check import run_codegen_check

report = run_codegen_check(root=Path("src/generated/documents"))
if not report.is_current:
for drift in report.drifts:
print(f"{drift.path}: {drift.category} — {drift.detail}")
raise SystemExit(1)
```

The drift categories are `pipelex codegen check`'s, and so are the sentences: an artifact edited below its stamp is `hand-edited`, one off the locked hash is `modified`, one the lock tracks and disk has lost is `missing`, and a stamped file the lock does not track is an `orphan` — the stale-artifact class a per-file stamp cannot catch alone. The two readers reach the same verdict over the same bytes apart from two deliberate divergences, both documented in `docs/architecture.md`: this one accepts a projection line whose axes are outside its own vocabulary, where the CLI calls such a tree hand-edited, and it refuses a Python artifact that declares a PEP 263 source encoding, where the CLI calls that one current. Regeneration stays a developer action, because it needs the engine; the check is the CI action, because it needs only hashes, so an upstream template improvement never reddens your pipeline. Whether the tree still matches what the *method* resolves to is a separate question the engine alone can answer — compare `report.crate_fingerprint` against a live `codegen()` response to close it.

### Long runs: start + poll explicitly

Behind the hosted gateway, a synchronous `execute()` is cut off at ~30s and surfaces a `PipelineExecuteTimeoutError` pointing here. For long methods, drive the durable lifecycle yourself — the run survives client disconnects and is resumable by `pipeline_run_id`:
Expand Down Expand Up @@ -103,7 +140,8 @@ There is no barrel import — package `__init__.py` files stay empty. Import eac
- **Run lifecycle types** — `from pipelex_sdk.runs import RunStatus, RunPublic, RunRead, RunResults, RunResultState, WaitForResultOptions, PollInfo`
- **Product wire models** — `from pipelex_sdk.product_models import UserProfile, MethodData, MethodWriteInput, Membership, MembershipsResponse, SubscriptionResponse, PlanView, InvoiceView, OnboardingSubmission, UploadInput, UploadedFile, PipelineRun, ...`
- **Validation verdict types** — `from pipelex_sdk.validation_models import PipelexValidationResult, PipelexValidationReport, PipelexInvalidReport, ValidationErrorItem, SuggestedFix, VALIDATION_VIEW_INPUT_FORM, ...`
- **Typed errors** — `from pipelex_sdk.errors import ApiResponseError, ApiUnreachableError, PipelineExecuteTimeoutError, PagingNotTerminatingError, RunFailedError, RunTimeoutError, RunLifecycleUnavailableError, RunStillRunningError, ...`
- **Codegen tree** — `from pipelex_sdk.codegen_writer import write_codegen_tree, CodegenTreeWriteReport` to write one, `from pipelex_sdk.codegen_check import run_codegen_check, CodegenCheckReport, CodegenDrift, DriftCategory` to verify one, with the format primitives in `pipelex_sdk.codegen_lock` (`CodegenLock`, `parse_lock`, `load_lock`, `validate_artifact_path`, ...) and `pipelex_sdk.codegen_stamp` (`STAMPABLE_SUFFIXES`, `is_stampable_artifact_path`, `compute_content_hash`, `parse_stamped`, ...)
- **Typed errors** — `from pipelex_sdk.errors import ApiResponseError, ApiUnreachableError, PipelineExecuteTimeoutError, PagingNotTerminatingError, RunFailedError, RunTimeoutError, RunLifecycleUnavailableError, RunStillRunningError, CodegenError, CodegenLockError, ...`
- **Version** — `from pipelex_sdk.version import __version__`
- **Protocol surface** (the MTHDS standard's wire types) comes from the `mthds` dependency — e.g. `from mthds.protocol.exceptions import PipelineRequestError`, `from mthds.protocol.models import ValidationResult` (the neutral verdict union that `PipelexValidationResult` narrows).
- **Input-form descriptors and pipe I/O contracts** come from `mthds` too, because they are the standard's artifacts and this SDK only carries them: `from mthds.protocol.input_form import InputForm, InputFormField, ListField, TextField, ...` and `from mthds.protocol.pipe_io_contracts import PipeIOContracts, PipeInputContract, PresenceMarker, IOMultiplicity, ...`. `PipelexValidationReport.input_form` and `.pipe_io_contracts` are typed with them, so a node narrows on its `kind` and a slot's presence and multiplicity read as enums — but `pipelex_sdk` does not re-export the vocabulary, and importing it from here is the one supported path.
Expand Down
Loading
Loading