This is the implementation tracker for the design in wip/updates.md. The design answers what and why; this file is the how, broken into phases with checkboxes. Tick a box when the item is done and verified, not when it is started. Every design choice that was open has been decided (see wip/updates.md §7) and is treated here as settled: input_form stays opaque, MethodData.python is a typed list[MethodFile] with the converter in this repo, the method_id type guard lands now, and an unknown FixOp.kind raises.
One of those settled choices has since been superseded: input_form is no longer opaque, and neither is pipe_io_contracts — both are typed by importing the standard's client models now that mthds publishes them. The record of that change, and why it honours rather than overrides the ownership argument the opaque ruling rested on, is wip/input-form-typed-narrowing.md. Everything else below stands as written.
Ground rules for every phase, from CLAUDE.md:
- Branch:
feature/Typed-method-id-run-option(already carries the typedmethod_idoption and thedelete_methodcontract fix). The PR targetsdev. - Gate each phase on
make agent-checkandmake agent-test; nothing is "done" before both pass. Runmake checkonce at the end as well, since it adds pylint on top of the agent gate. - Tests use
pytest-mockonly, oneTestClassper module, no__init__.pyundertests/. Mock at the httpx boundary (_send), as the existing modules do. - Everything accumulates under
## [Unreleased]inCHANGELOG.md. No version bump in this work: the/releaseskill cuts the version, and with the breaking items in Phase 3 (plus those already on the branch) that will be a minor bump. - Docs move with the code, in the same commit.
docs/architecture.mdis the main one to keep truthful; its "Parity with@pipelex/sdk" section currently claims surface-completeness and is wrong. - No hardcoded counts in code, docs, or commit messages. No volatile state in tracked files (this file records decisions and what landed, never whether the tree is clean or tests are green right now).
- Line numbers quoted below are as of the design date (2026-08-25) and will drift; they are anchors, not contracts.
Suggested order is the order below. Phase 3 is the most urgent fix (a crash against the deployed platform) but is also the largest breaking change, so the plan puts the purely additive Phase 1 first to keep each commit reviewable; reorder if the crash needs to ship alone.
-
make installand confirm the pinnedmthdsbase is the one the code was written against (uv pip show mthds; the floor is>=0.8.2and the workspace copy is 0.8.2). Nothing in this plan needs a newermthds. Confirmed: 0.8.2. - Run
make agent-checkandmake agent-testbefore touching anything, so a later failure is attributable to this work and not to the starting point. Both green at the starting point. - Re-read
wip/updates.md§6 and §7 once; if anything there contradicts this file, this file is the stale one. No contradiction found.
Purely additive. Every affected response model is extra="allow", so nothing parses differently for a body that lacks the new keys; the work is to type what already arrives and to add the one request knob (views) that has no way through today.
- Add
FixSafety(StrEnum)withSAFE = "safe",UNSAFE = "unsafe", and anis_safeproperty (house rule: never compare enum values inline; amatchinside the property). - Add
FixOpKind(StrEnum)with the kinds the runtime emits:SET_KEY = "set_key",ENSURE_TABLE = "ensure_table",DELETE_KEY = "delete_key",DELETE_TABLE = "delete_table",RENAME_TABLE_KEY = "rename_table_key",MOVE_KEY = "move_key",REMAP_VALUE = "remap_value". Source of truth:pipelex/pipelex/suggested_fix.pyand the OpenAPI artifactpipelex-api/docs/openapi/pipelex-api.openapi.yaml. - Add
TomlScalar: TypeAlias = str | int | float | boolandTomlValue: TypeAlias = TomlScalar | dict[str, TomlScalar]. Deeper nesting is not modelled because the server does not emit it; say so in a comment. - Add a private
_FixOpBase(BaseModel)withmodel_config = ConfigDict(extra="allow")andtable_path: list[str](empty list means the document root), then one subclass per kind, each withkind: Literal[FixOpKind.X]and exactly its own members:SetKeyOp(key: str, value: TomlValue),EnsureTableOp(),DeleteKeyOp(key: str),DeleteTableOp(),RenameTableKeyOp(key: str, new_key: str),MoveKeyOp(key: str, new_table_path: list[str], new_key: str),RemapValueOp(key: str, mapping: dict[str, str]). OnEnsureTableOpandDeleteTableOpdeclaretable_path: list[str] = Field(min_length=1), mirroring the artifact'sminItems: 1. - These are reader models: no
frozen, noextra="forbid", none of the runtime's wildcard-refusing validators. Put the two runtime invariants a type cannot carry in the docstrings (*is the wildcard segment and is refused as akeyon every kind butremap_value;ensure_table/delete_tableneed a non-emptytable_path). - Add
FixOp: TypeAlias = Annotated[SetKeyOp | EnsureTableOp | DeleteKeyOp | DeleteTableOp | RenameTableKeyOp | MoveKeyOp | RemapValueOp, Field(discriminator="kind")]. Narrowing ismatch op: case SetKeyOp(): …, exhaustive, nocase _. - Verify that pydantic accepts the raw wire string (
"set_key") againstLiteral[FixOpKind.SET_KEY]both invalidate_pythonandvalidate_json, and that the discriminator resolves on it. If it does not on the pinned pydantic, useLiteral["set_key"]on the models and keepFixOpKindas the documented vocabulary (with a test that the two sets agree). - Add
SuggestedFix(BaseModel, extra="allow"):fix_code: str,description: str,safety: FixSafety,source: str | None = None,ops: list[FixOp](typed default factory viaempty_list_factory_ofonly if a default is warranted — the runtime always sendsops, so leaving it required is fine). - Add
LiftablePipeEntry(BaseModel, extra="allow"):pipe_ref: str,within_pipe_ref: str,skipped_when_absent: list[str] = Field(default_factory=list),absence_source: str. Mirrorspipelex/pipelex/pipeline/liftable_pipes.py. - Add the view token constant next to the field it gates:
VALIDATION_VIEW_INPUT_FORM: Final[str] = "input_form". A constant, not a closed enum — the request boundary is deliberately open so a stale token never fails a call. - On
ValidationErrorItemaddmissing_pipe_code: str | None = None(symmetrical withmissing_concept_code) andsuggested_fix: SuggestedFix | None = None. Leaveerror_type: str | Noneas an open string; do not enum it. - On
PipelexValidationReportaddwarnings: list[ValidationErrorItem] = Field(default_factory=empty_list_factory_of(ValidationErrorItem)),liftable_pipes: list[LiftablePipeEntry] = Field(default_factory=empty_list_factory_of(LiftablePipeEntry)), andinput_form: dict[str, Any] | None = None, each with a docstring:warningsnever flipsis_valid; the two lists default empty so a pre-0.52 runner's body still parses;input_formis present only when the request named theinput_formview (a 0.17.0 runner emitted it unconditionally, whichNone-by-default also reads correctly) and is opaque on purpose, keyed likepipe_io_contracts. -
PipelexInvalidReportgains nothing; add one sentence to its docstring saying why (the invalid arm never carrieswarningsorinput_form— they derive from a crate that was never assembled). - Update the module docstring's list of neutrally-named supporting types to include the new ones, and keep the brand rule stated there (the
Pipelexprefix stays on the two envelopes only).
-
validate(...): appendviews: list[str] | None = Noneafterrender. Whenviews is not None, setextra["views"] = viewsverbatim — no injection, no de-duplication, an explicit[]is sent as[]. WhenNone, the key is absent from the body. It rides_post_validate'sextraexactly likerenderandmthds_sources; nomthds-pythonchange. -
validate_files(...): appendviews: list[str] | None = Noneafterrenderand thread it through tovalidate. - Docstrings: the sentence "differs from the inherited protocol
validatein two Pipelex-API ways" becomes three (render injection,mthds_sources,views); document theviewssemantics (opt-in structured views;input_formis the only token today, named byVALIDATION_VIEW_INPUT_FORM; unknown tokens are lenient-ignored server-side, never a422; the default response stays byte-identical for consumers that discard views). - The
Returns:section ofvalidateshould mention that a valid report now carrieswarnings,liftable_pipes, and (when asked)input_form.
-
tests/unit/test_client_validate.py:viewssent verbatim when given; theviewskey absent from the body when the parameter is omitted; an explicit[]sent as[];validate_filesthreadsviewsthrough;renderinjection unchanged whenviewsis also passed. -
tests/unit/test_validation_contract.py: a valid body carryingwarnings,liftable_pipes, andinput_formparses into typed fields, withinput_formkeyed likepipe_io_contracts; the pre-0.52VALID_BODYstill parses with both lists empty andinput_formNone; the JS null-bearing warning fixture (pipelex-sdk-js/tests/client.test.ts, "carries advisory warnings on the VALID arm, with the valid arm's explicit nulls") parses with every explicitnullreading asNone— this is the regression guard against a future "tighten to required" edit, and it answers the inbox item../wip/inbox/2026-08-25-workspace-validation-error-item-spec-gaps.mdfor the Python mirror. -
tests/unit/test_validation_contract.py: an invalid body carryingmissing_pipe_codeand asuggested_fixwith at least two ops of different kinds parses, andmatch-narrowing reaches each op's own members; anensure_tableop with an emptytable_pathis rejected; an unknownkindraisespydantic.ValidationError;FixSafetyandFixOpKindvalue sets are asserted as the locked vocabularies (same style astest_category_vocabulary_is_the_locked_set). - Keep the canonical bodies where the module already keeps them (module-level constants next to
VALID_BODY); move to atests/unit/test_data.pyonly if the module becomes unreadable.
-
docs/architecture.md→ "validateoverride": add aviewsbullet beside the render bullet; list the typed valid-arm additions (warnings,liftable_pipes,input_form) and theValidationErrorItemadditions with theSuggestedFix/FixOp/FixSafetyvocabulary; one sentence thatPipeInputContract.optionalbecamepresenceand that thefixedmultiplicity carriesitem_countinside the opaquepipe_io_contracts, so nobody discovers the new spellings by surprise. -
docs/architecture.md→ "Brand boundary": the list of neutrally-named supporting types gains the new ones. -
README.mdquickstart: one line showingviews=[VALIDATION_VIEW_INPUT_FORM](or a comment that the input form is opt-in), so the knob is discoverable. -
CHANGELOG.md[Unreleased]→ Added:viewsonvalidate/validate_files; the typed valid-arm fields;missing_pipe_code/suggested_fixand the fix vocabulary; a note that a body from an older runner still parses (the lists default empty,input_formdefaultsNone).
-
make agent-checkandmake agent-testpass. - Commit (suggested message: "Type the pipelex-api 0.17/0.18 validate contract and add the views opt-in").
- Update this file: tick what landed, record the SHA of the Phase 1 commit, note whether pydantic accepted the enum
Literaltags or the string fallback was needed, and any deviation from §1 of the design with its reason.
Landed in 434b2e3. Notes:
- The enum
Literaltags work as written; the string fallback was not needed. Verified against the pinned pydantic (2.13.4) before writing the models: a raw wire"set_key"validates againstLiteral[FixOpKind.SET_KEY]in bothvalidate_pythonandvalidate_json, the discriminator resolves on it, an unknownkindraises, andField(min_length=1)onEnsureTableOp.table_pathrejects an empty path. RemapValueOp.mappingis left unconstrained, where the runtime and the OpenAPI artifact both declareminProperties: 1. These are reader models: an empty mapping is an advisory no-op, not a parse hazard, and refusing it would fail a whole verdict over a harmless op. TheminItems: 1onensure_table/delete_tablewas kept because there the empty case is genuinely meaningless (the document root always exists, and cannot be deleted).- One Phase-2 item landed early, because
validation_models.pywas rewritten wholesale here: theconformance/conformance/validation_contract.pycitation onValidationErrorCategoryis already replaced with the rule it was citing. Phase 2 covers the rest. - No other deviation from §1 of the design.
No behaviour change. Two fixes @pipelex/sdk 0.14.0 shipped under "Fixed" that apply here for the same reason (pipelex-sdk is a public PyPI package).
-
TokensUsageRecordattribution.pipelex_sdk/runs.py(module docstring near line 23 and the class docstring near line 128),docs/run-usage.md(line 5),docs/architecture.md(theTokensUsageRecordbullet near line 103): the record is a Pipelex runtime extension the MTHDS Protocol does not model, and the hosted API pins the wire contract. Reword aspipelex-sdk-js/src/runs.tsand itsdocs/architecture.mddid.docs/run-usage.mdline 5 already says the right thing in its second sentence; make the first sentence agree with it. - Citations a reader cannot open. Replace each bare workspace-private path with the rule it was citing:
pipelex_sdk/client.pynear line 138 (docs/specs/pipelex-platform-api.md→ "the layered extension policy: a hosted client types its own platform's arguments and guards them per layer");pipelex_sdk/validation_models.pynear line 44 (conformance/conformance/validation_contract.py→ "the locked category vocabulary shared with the conformance corpus");tests/unit/test_validation_contract.pymodule docstring and the docstring oftest_category_vocabulary_is_the_locked_set;tests/unit/test_runs.pynear line 12;tests/unit/test_client_method_id.pymodule docstring;docs/architecture.mdnear line 84. Keep the JS mirror references (pipelex-sdk-js/...) where they explain a port — those are a sibling public repo, not a private path. -
CHANGELOG.md[Unreleased]→ Fixed: two entries mirroring 0.14.0's wording. -
make agent-checkandmake agent-testpass. - Commit (suggested message: "Correct the TokensUsageRecord attribution and drop unopenable citations").
Breaking, and the most urgent fix in this plan: list_methods and list_runs crash against the deployed platform because both routes now answer a {items, next_cursor} envelope, and PipelineRun requires fields the platform serves as nullable. Wire fields stay snake_case (next_cursor), where the JS mirror renamed to nextCursor for its own consumers.
- Add
MethodFile(BaseModel, extra="allow")withname: str,content: str, defined beforeMethodData. Docstring: the at-rest catalog form of one source file ([{name, content}]), distinct fromMthdsFile(client.py, validate input) andMthdsFileItem(build_models.py, build closure) — three shapes for three surfaces, name the difference so nobody merges them. - Add
parse_method_files(source: str | None) -> list[MethodFile]: blank source (None,"", whitespace) and"[]"both yield[]; a JSON array of{name, content}yields those files with blank-content entries dropped; anything else (non-array JSON, a malformed entry, unparseable text) raisesValueErrorwith a message naming the expected shape. Implementation:json.loadsthenTypeAdapter(list[MethodFile])(built once at module level, TypeAdapter construction is expensive), wrappingjson.JSONDecodeError/pydantic.ValidationErrorinto theValueError. - Add
serialize_method_files(files: list[MethodFile]) -> str: drop blank-content entries; an empty result serializes to""(the platform's "no source / clear" sentinel), never"[]"; otherwisejson.dumpsof[{name, content}]only (no extras), stable key order. -
MethodData: addorg_id: str,created_by_user_id: str,description: str | None = None,deletion_state: MethodDeletionState | None = None,python: list[MethodFile] = Field(default_factory=empty_list_factory_of(MethodFile)), plus a@field_validator("python", mode="before")that appliesparse_method_fileswhen the incoming value is astrorNoneand passes a list through unchanged (so programmatic construction in tests still works). AValueErrorfrom the parser surfaces aspydantic.ValidationErrorfrommodel_validate, the same way any malformed response body fails here. -
MethodWriteInput: addpython: list[MethodFile] | None = Nonewith a@field_serializer("python")returningserialize_method_files(value). Docstring the three-way contract:None→ key absent (the write body dumps withexclude_none=True) → the stored Python is preserved onPUT;[]→ sent as""→ clears it; a non-empty list → replaces it. - Add
MethodSummary(BaseModel, extra="allow"):method_id: str,name: str,description: str | None = None,created_at: str,deletion_state: MethodDeletionState | None = None. Docstring: deliberately not aMethodData— nomthds, nopython, noupdated_at— because none is in the index projection and puttingmthdsback is what restored the truncation bug; a method mid-deletion stays in the list so a UI can render "Deleting…" whileget_methodrefuses it with a409. - Add
MethodPage(BaseModel, extra="allow"):items: list[MethodSummary],next_cursor: str | None = None. Docstring: opaque cursor, pass it straight back;Nonemeans last page; no total by design. - Add
RunErrorReport(BaseModel, extra="allow"):message: str | None = None,error_type: str | None = None— the two fields a consumer may rely on out of the runner's verbose report. -
PipelineRun:method_id: str | None = None(an ad-hoc run from an inline bundle belongs to no stored method),pipe_code: str | None = None(resolved from the bundle'smain_pipe); addorg_id: str | None = None,created_by_user_id: str | None = None,error: RunErrorReport | None = None. Leavepipe_statusesas it is. - Add
RunDetail(PipelineRun):mthds_contents: list[str] | None = None,inputs: dict[str, Any] | None = None. Docstring:mthds_contentsis what the run actually executed and the only record of it; both fields are absent from the list and the polled status read on purpose (size × page size, size × poll rate). - Add
RunPage(BaseModel, extra="allow"):items: list[PipelineRun],next_cursor: str | None = None. - Update the section comments in the module (
# ── Methods catalog,# ── Run records) so the new models sit under the right banner.
- Add one error for
iterate_methodsrefusing to keep paging past the ceiling (a name likePagingNotTerminatingError), extending whatever base the module's other product errors extend — check the existing hierarchy there first. Message mirrors the JS one: the iterator did not terminate after the ceiling; this is a server-side fault, not a coverage limit.
- Add a module helper
_product_query(params: dict[str, str | int | None]) -> strthat keeps entries on presence (is not None, never truthiness — an explicit emptyqor cursor is bad input the API should reject, not something to drop silently) and encodes withurllib.parse.urlencode, returning""or?…. The existingtest_list_runs_encodes_query_valueassertion (method_id=m%2F1) must stay green, so keep/percent-encoded. - Add
_MAX_LIST_PAGES: int = 10_000beside the other module constants (the JSMAX_PAGES), with the comment that it is a runaway backstop set far beyond any real catalog, not a coverage cap. -
list_methods(self, *, q: str | None = None, limit: int | None = None, cursor: str | None = None) -> MethodPage—GET /v1/methodswith the query built by the helper; parseMethodPage. Docstring:qis a server-side case-insensitive substring match over name and description across the whole catalog;limitdefaults to and is capped by the API; ordering is by creation, newest first. -
iterate_methods(self, *, q: str | None = None, limit: int | None = None) -> AsyncIterator[MethodSummary]as anasync defgenerator: request a page; before yielding, stop ifcursor is not None and page.next_cursor == cursor(the server did not advance; yielding first would double-count); yield every item; stop whenpage.next_cursor is None; otherwise count the page and raise the new error once the count reaches_MAX_LIST_PAGES; continue through empty pages with a live cursor, becauseqis a post-read filter over a bounded index slice per request and{items: [], next_cursor: "…"}means "keep going". Docstring says why there is nolist_all_methods(): an all-at-once helper needs a cap, and a cap is the silent truncation paging removed. -
list_runs(self, method_id: str, *, created_from: str | None = None, created_to: str | None = None, limit: int | None = None, cursor: str | None = None) -> RunPage—GET /v1/runs?method_id=…plus the presence-kept query; parseRunPage. Docstring:created_from/created_toare instants (ISO-8601 with a UTC offset), inclusive, key conditions rather than filters; a bare date or naive timestamp is a platform400surfaced asApiResponseError. Also document the gate the JS mirror does not spell out: every/v1/runs*product route sits behind the platform'srequire_surface_access(), which for API-key auth demands theff_api_keysfeature flag and fails closed with a403— a403here means "flag", not "wrong key". -
iterate_runs(self, method_id: str, *, created_from: str | None = None, created_to: str | None = None, limit: int | None = None) -> AsyncIterator[PipelineRun]— same loop, except an empty page ends it (date bounds are index key conditions, so a run page is never empty-with-a-cursor; the difference is the server, not the client — say so in the docstring).No page ceiling needed: the empty-page stop already catches a server minting fresh cursors while returning nothing.Deviation, taken on PR review: the ceiling applies here too. The plan's reason was too narrow — the empty-page stop only catches a server returning nothing, so a cursor cycling across two or more values (c1 → c2 → c1) over non-empty pages trips neither it nor the adjacent-cursor check and would loop forever re-yielding the same runs. Greptile and Codex flagged it independently. The fix reuses_MAX_LIST_PAGESandPagingNotTerminatingErrorrather than tracking every cursor seen, which would cost unbounded memory for the same protection. -
get_run_detail(self, run_id: str) -> RunDetail—GET /v1/runs/{id}via_request_productwith the id path-encoded like the other id routes (f"{_RUNS}/{quote(run_id, safe='')}"). Distinct fromget_run_status(/status) andget_run_result(/results). - Update the
PipelexAPIClientclass docstring's product-surface bullet and the import block (MethodPage,MethodSummary,RunDetail,RunPage,AsyncIteratorfromcollections.abcunderTYPE_CHECKINGif only used in annotations — it is used at runtime as a return annotation withfrom __future__ import annotations, soTYPE_CHECKINGis fine).
-
tests/unit/test_client_product.py: replace the bare-array fixtures attest_list_methodsandtest_list_runs_encodes_query_valuewith envelopes and assertMethodPage/RunPagecome back withnext_cursor; add query-encoding cases forq/limit/cursorand forcreated_from/created_to, including that an explicit empty string is forwarded rather than dropped and that an absent parameter leaves no key in the query; a run row withnullpipe_codeandmethod_idparses;get_run_detailhits/v1/runs/{id}with encoding and returnsmthds_contentsandinputs;MethodDataparses the new fields withpythonread from the wire string intoMethodFileentries and""reading as[];create_method/update_methodsendpythonthree ways (Noneabsent,[]as"", a list as the JSON text). - New
tests/unit/test_method_files.py(oneTestClass):parse_method_fileson blank /"[]"/ a valid array / an array with a blank-content entry / a non-array / a malformed entry / unparseable text;serialize_method_fileson empty / blank-only / mixed; a round-trip is stable. - New
tests/unit/test_client_paging.py(oneTestClass):iterate_methodscontinues through an empty page with a live cursor and stops onNone; both iterators stop on an unchanged cursor without re-yielding the page;iterate_runsstops on an empty page;iterate_methodsraises the new error past the ceiling (patch_MAX_LIST_PAGESdown viamocker.patchrather than looping ten thousand times); the cursor sent on page N+1 is thenext_cursorreceived on page N. Usemocker.AsyncMock(side_effect=[...])on_sendto script the page sequence. - The
_response/_Sent/_mock_sendhelpers live as private members oftests/unit/test_client_product.py. Rather than importing private test helpers across modules, promote a response builder and a_sendspy to fixtures in a newtests/unit/conftest.pyfor the new modules to use (house rule: fixtures go inconftest.py). Migratingtest_client_product.pyonto those fixtures is optional and not part of this change.
-
docs/architecture.md→ "Pipelex product surface": rewrite the Methods catalog bullet forMethodPage/MethodSummary/iterate_methodsand thepythonthree-way write contract withMethodFile; rewrite the Run records bullet forRunPage/iterate_runs/get_run_detail, the nullablePipelineRunfields anderror, the instant-only date bounds, and theff_api_keys403. State the two iterator stop rules and why they differ. -
docs/architecture.md→ "Parity with@pipelex/sdk": it must stop claiming "surface-complete, with no silent gaps". Rewrite it to list the conscious exclusions honestly:lint,format,resolve,codegen,build_output/build_runner/concept/pipe_spec,run_codegen_check,get_method_closure— unchanged by the cited releases and deferred. While there, fix the stale "Out of scope for v0.1" bullet that still lists/v1/build/*helpers as deferred even thoughbuild_inputsshipped in 0.5.0. -
README.md: if the quickstart gains a listing example, useiterate_methodsrather than a page loop, so the idiom people copy is the one that cannot truncate. -
CHANGELOG.md[Unreleased]: Changed (breaking) —list_methodsreturnsMethodPage(items areMethodSummary, notMethodData),list_runsreturnsRunPage,PipelineRun.method_id/pipe_codeare nullable,MethodData.pythonislist[MethodFile]; Added —iterate_methods,iterate_runs,get_run_detail,MethodSummary/MethodPage/RunPage/RunDetail/RunErrorReport,MethodFilewithparse_method_files/serialize_method_files, the newMethodDatafields,MethodWriteInput.python, the paging error; Fixed — name the crash plainly (iterating the envelope dict yielded its keys, so the first call wasMethodData.model_validate("items")), and that the unit tests mocked the pre-paging shape.
-
make agent-checkandmake agent-testpass. - Commit (suggested message: "Follow the platform's paged method and run lists and stop requiring nullable run fields").
- Update this file: tick what landed, record the Phase 2 and Phase 3 commit SHAs, and note any place the Python shapes deliberately diverge from the JS mirror beyond snake_case (there should be none besides the
pythonconverter living here).
Phase 2 landed in 2a8c589, Phase 3 in 5e01c1b. Notes:
- Divergences from the JS mirror beyond snake_case: the
pythonconverter lives here rather than inmthds-python(the decision ofwip/updates.md§7.2), andparse_method_filesraisesValueErrorwhere the JS pair raisesPipelineRequestError— because the Python parser is reached through a pydanticfield_validator, where aValueErroris the idiomatic signal and surfaces to the caller as apydantic.ValidationErrorlike any other malformed response body. Nothing else diverges. MethodPage.items/RunPage.itemsare required, not defaulted empty. A page body with noitemsis malformed, and failing loudly is the whole point of this phase — the previous shape failed silently in the tests and loudly in production.- The shared test fixtures landed in a new
tests/unit/conftest.py(api_client,wire_response,patch_send), used by the two new modules. Migratingtest_client_product.pyonto them was left out as the plan allowed; it keeps its own equivalent private helpers.
The 2026-08-25 decision in pipelex-sdk-js/wip/boundary-option-type-validation.md: a published client validates request-option types at its boundary and raises PipelineRequestError rather than dropping or forwarding a wrong-typed value. Its evidence names this repo's bare if method_id: in _merge_hosted_run_extensions, which drops falsy non-strings and forwards truthy ones to a server 422.
-
_merge_hosted_run_extensions(client.pynear line 975): replaceif method_id:with an explicit presence check (if method_id is not None) followed byif not isinstance(method_id, str): raise PipelineRequestError(msg)naming the received type, then the existing empty-string-is-absent normalization.Noneand""still contribute nothing. - Update the docstring's
Raises:and the "An absent or emptymethod_id" paragraph. -
tests/unit/test_client_method_id.py: one parametrized test over wrong-typed values (0,123,[],["mt_1"],{}) assertingPipelineRequestErroronexecuteandstartbefore any request is sent; keeptest_empty_method_id_is_absentgreen. -
docs/architecture.md→ "Hosted run extensions (method_id)": add a bullet that a non-string raises at the boundary, and why (one partition of wrong values across both SDKs). -
CHANGELOG.md[Unreleased]→ Changed: the guard, with the JS decision as the reason. -
make agent-checkandmake agent-testpass. - Commit (suggested message: "Reject a non-string method_id at the client boundary").
-
make check(adds pylint to the agent gate) andmake agent-testpass on the final tree. - Re-read
docs/architecture.mdend to end for any remaining claim the code no longer supports (parity,Out of scope, the validate section, the product section). Three further corrections beyond the ones §3.5 named: the intro line claiming "the full0.1.0surface", the parity section's "Methods — full coverage" (which contradicted the gap list directly above it), and the "Models — full field-for-field match" claim; the Models paragraph now also names the two deliberate divergences (snake_casenext_cursor, the converter's home). - Re-read
CHANGELOG.md[Unreleased]: every breaking item is labelled "breaking", no counts, no WIP-doc mentions, the version line untouched (still0.5.0). -
wip/updates.mdstays where it is, with its §7 decisions; this file stays too. Neither is deleted or emptied as part of finishing the work. - Open the PR against
dev, then follow/review-pr-agentsfor the Greptile / Codex loop (compare SHAs, not notifications; reply and resolve each thread in one pass). PR #14.
- Update this file with the final commit SHAs, the PR number, and anything deferred out of the plan with the reason.
PR #14, against dev. The five implementation commits, in order:
| SHA | What |
|---|---|
434b2e3 |
Phase 1 — the validate surface: the views opt-in and the typed 0.17/0.18 contract |
2a8c589 |
Phase 2 — the TokensUsageRecord attribution and the unopenable citations |
5e01c1b |
Phase 3 — paged method and run lists, nullable run fields, the python converter |
cb70fbe |
Phase 4 — the non-string method_id boundary guard |
a857d02 |
Phase 5 — the remaining untrue parity claims in docs/architecture.md |
Deferred out of the plan, with the reason:
RemapValueOp.mappingis not constrained to be non-empty, where the runtime and the OpenAPI artifact both sayminProperties: 1. Reader models should not fail a whole verdict over an op that is merely a no-op. Recorded at Checkpoint 1.- Migrating
test_client_product.pyonto the newtests/unit/conftest.pyfixtures was explicitly optional in §3.4 and was not done; the module keeps its own equivalent private helpers. pipelex_sdk/runs.py's "Wire contract mirrorspipelex-platform" line was left alone. It names a service, not a file path, so it is not one of the unopenable citations Phase 2 was about, and widening that phase's scope on my own judgement was not warranted.
Cubic's pass on fa8d62a reported findings against files this plan touched. Three were real and are fixed; the rest were judged not worth acting on, and the reasons are recorded here rather than only in the resolved GitHub threads.
Fixed:
- The
TokensUsageRecordattribution indocs/architecture.mdwas missed. Phase 2's first checkbox names that site explicitly, andCHANGELOG.mdasserts it was corrected, but2a8c589only touched themethod_idcitation in that file.pipelex_sdk/runs.pyanddocs/run-usage.mdwere correct. The bullet now carries the same wording as the other two. CHANGELOG.mdstill citeddocs/specs/pipelex-platform-api.md. The changelog was outside Phase 2's enumerated citation sites, so the sweep did not reach it — leaving the[Unreleased]section claiming "no more citations a reader cannot open" a few entries below a citation a reader cannot open. The sentence now states the rule and points at the sibling public JS SDK instead.- The
validateoverride intro indocs/architecture.mdstill said "two Pipelex-API extensions" after Phase 1 added the third. The client docstring was updated at the time; the doc's intro sentence was not.
Not acted on, with the reason:
- Hoisting the
method_idtype guard abovestart_and_wait's lifecycle handshake. The claim that a handshake failure can mask the guard is wrong:_supports_run_lifecycleswallows its errors and every downstream path still reaches_merge_hosted_run_extensions, so no wrong-typed value escapes. The real cost is one probe request, memoized for the client's lifetime. Hoisting would mean either calling the merge helper twice or duplicating the check outside the single documented guard site. - A claimed 194-character line in
tests/unit/test_client_validate.py. Measured at 127, under the configured 150, and no line in that file is longer;ruff checkpasses. The measurement was simply wrong. - Rewriting Checkpoint 2's parenthetical to match the two divergences the note beneath it records. The parenthetical is the expectation the plan set out with, and the note is the finding; editing the prediction to match the outcome removes the only evidence that the plan's expectation was slightly off.
- The
../wip/inbox/…reference in §1.3. Phase 2's rule is about the shipped surface — docstrings, comments and doc pages that travel to PyPI, where a workspace path resolves to nothing and reads as rot. This tracker is not that: it is a working document for the reviewers of this PR, it is written from workspace context throughout (it citeswip/updates.mdand the JS repo's ownwip/in the same way), and../wip/inbox/is the notation the workspace guide itself prescribes for a sub-repo checkout. The path resolves where the document is read, and the sentence already states its rationale inline before citing the item, so a reader who cannot open it loses only the provenance. - Replacing the two
# type: ignore[arg-type]comments intests/unit/test_client_method_id.pywithcast(). A cast would assert to the reader that the value is astr | None, which is the exact falsehood that test exists to disprove at runtime; the ignore states the truth that this is a deliberate type error, which is the last resort the coding standard permits.
Once the three PR bots reported clean on 301d96e, a fresh reviewer with no context from this session ran the gstack review procedure against this plan, with an adversarial pass folded in. It found no correctness defect in the shipped code, and it verified the wire contracts against pipelex-server itself rather than against the docstrings that assert them — including that the two iterators' asymmetric stop rules match the two DynamoDB adapters (the method adapter can mint a cursor for an empty page because q filters after the read, the run adapter over-fetches by one and cannot), and that the pydantic three-way write contract for python holds for absent / null / "" / "[]" alike.
Fixed:
README.mdadvertised an import that raisesImportError. "Public import paths" sent readers tofrom mthds.runners.api.models import PipelexValidationResult; that name is not in the installedmthdsat all. The Pipelex narrowing of the verdict union is owned here, whichdocs/architecture.mdstates in its "Brand boundary" section — so the repo contradicted itself. The section gains apipelex_sdk.validation_modelsbullet and now citesmthds.protocol.models.ValidationResultas the neutral union it narrows. Pre-existing, but this branch edited the quickstart three lines above it.- The two page-ceiling tests could not fail. Their only assertion was
exc_info.value.page_limit, which just echoes back the constant the test itself patched, so five scripted responses against a limit of two passed whether the raise fired on page 1 or page 5. Each now also assertssend.call_count, which was mutation-checked: looseningpages_seen >= _MAX_LIST_PAGESto>fails both tests where it previously passed. - Two hardcoded counts that had already gone stale inside this branch.
pipelex_sdk/client.py'svalidatedocstring anddocs/architecture.mdboth said "three Pipelex-API ways/extensions" — the same rot the entry above records having corrected once already, which is exactly why the workspace guide forbids writing counts down. Both now say neither two nor three. executeandstartunder-documented an exception this branch added. TheirRaises:entries named only theextra-smuggling trigger; Phase 4's guard raisesPipelineRequestErrorfor a non-stringmethod_idtoo. The helper's own docstring had been updated at the time, the two public methods' had not.
Deferred to wip/pr-14-review-notes.md, each verified and none blocking: the sdist ships this tracker and the other internal planning documents (a packaging decision that belongs with a release, not inside a feature branch); the client class docstring still dates two surfaces by build-plan phase number; and start_and_wait's Raises: omits the PipelineRequestError it propagates on both paths.
One finding reached outside this repo and was filed rather than fixed: PipelineRun.pipe_statuses is a field no server has ever filled — it is absent from the platform's RunPublic response model and appears nowhere in pipelex-server — yet it is declared in this SDK, in @pipelex/sdk, and in pipelex-app, where run-history progress dots are gated on it and have therefore never rendered. Retiring it here alone would break the parity invariant this package is built on, so the decision belongs to whoever owns the run wire contract: ../wip/inbox/2026-08-25-workspace-pipe-statuses-dead-field-in-three-clients.md.
Considered and declined: guarding iterate_methods against a next_cursor of "", which would let a page double-yield. The adversarial pass demonstrated the mechanism and then confirmed the case is unreachable, since the platform's cursor is a base64 LastEvaluatedKey. That is the impossible-scenario defensiveness this plan set out to avoid.
Cubic's pass on cb391c9 reported four more. Two were real:
docs/architecture.mddocumented an API that does not exist. The hosted-run-extensions section contrastedmethod_idwith "build_inputs(method_id=…)andprepare_inputs, which are client-side sugar that resolves an id to inline files". Neither helper accepts amethod_id:build_inputstakes aBuildInputsRequestoffiles/pipe_ref/format/explicit,prepare_inputstakesfiles/pipe_ref/inputs, and catalog-id resolution for input preparation is recorded as deferred in the 0.5.0 changelog. A reader following that sentence gets aTypeError. The contrast it wanted to draw is real, so the bullet now draws it truthfully. This is exactly the class of untrue claim Phase 5 set out to remove from this file, in a section Phase 4 added.- This tracker recorded the wrong
mthdsfloor. Phase 0 wrote>=0.8.1;pyproject.tomlhas said>=0.8.2sincebc17c07, which is an ancestor of this branch's base — so the floor was already 0.8.2 when the preflight ran, and the number was wrong the day it was written.
Two were not worth a commit:
- Adding this branch's own review-notes file to the sdist inventory in
wip/pr-14-review-notes.md. That listing istar -tzfoutput captured from a build made before the notes file existed. Extending it changes nothing about the finding it supports — that the sdist ships internal planning documents — or about what someone picking the item up would do. - The counts in the completion narrative above ("two page-ceiling tests", "two hardcoded counts"). The no-hardcoded-counts rule exists because a live inventory drifts: it "creates diff churn, goes stale silently, and adds no value". None of that applies to a closed record of a finished review round, which gains no members and cannot go stale. The rule targets counting things that are still moving.
- No e2e suite exists in this repo (
tests/holds onlyunit/), so the liveviewsgate and the live paging envelope are pinned only by mocked bodies here; the JS suite pins both live. Adding an e2e suite is separate work. - The remaining
@pipelex/sdksurfaces without a Python counterpart (lint,format,resolve,codegen,build_output/build_runner/concept/pipe_spec,run_codegen_check,get_method_closure) are untouched by the cited releases and stay deferred; Phase 3 makesdocs/architecture.mdsay so. - The protocol-argument type guards (
pipe_code,mthds_contents) belong inmthds-pythonand arrive here with themthdsfloor bump once that package ships its Phase 1.