diff --git a/docs/rfcs/0012-netsukefile-property-testing.md b/docs/rfcs/0012-netsukefile-property-testing.md index 2045c00e8..6c801cfc0 100644 --- a/docs/rfcs/0012-netsukefile-property-testing.md +++ b/docs/rfcs/0012-netsukefile-property-testing.md @@ -98,8 +98,8 @@ adopts: - Give assertions a structured, host-independent view of every build edge's command invocation and constructed environment. - Let a test case quantify over a bounded, declaratively specified - family of inputs, with deterministic default execution and seed-based - replay of failures. + family of inputs, with deterministic default execution and exact replay + of a failure from its reported seed and tuple inputs. - Provide built-in metamorphic relations for the generator's structural invariants: determinism, order invariance, locality, and rename isomorphism. @@ -149,6 +149,24 @@ quoting belong to the escaping seam's own tests, not to Netsukefile authors. The view carries helpers mirroring the graph view's surface: `action(target)`, `actions_for_rule(name)`, and `has_action(target)`. +The projections use a canonical order. `result.actions` sorts first by the +primary key `target`, then by `rule`, and finally by the RFC 8785 canonical +JSON form of `argv`, `env`, `cwd`, `inputs`, `outputs`, `pool`, `depfile`, and +`dyndep` as tie-breakers. `actions_for_rule(name)` applies the same ordering to +its filtered entries, with `target` as its primary key and the same canonical +JSON form as its tie-breaker. Identical manifests therefore produce identical +ordered projections regardless of declaration order or IR iteration order, and +quantified-action results use that same order. + +That canonical form is the canonical value contract of +[RFC 0006 §6.7](0006-ansible-inspired-template-standard-library.md#67-canonical-value-equality), +already produced in Netsuke by the `serde_json_canonicalizer` dependency. +Encoding it settles the representation of every field kind: optional fields are +omitted when absent, scalars and paths take their RFC 8785 lexical form, arrays +compare order-sensitively, and map keys sort lexicographically. Two projections +are byte-identical exactly when their encodings are, so the tie-break order is +total and unambiguous. + Because the result views are additive-only, this projection introduces no compatibility burden on existing tests, and internal IR types remain unexposed. @@ -165,6 +183,20 @@ A failing quantified assertion reports the binding that falsified it, alongside the substituted actual values, using the FAIL versus ERROR taxonomy of UX design §11.3 unchanged. +Environment data has two representations, separated by a redaction boundary. +Assertion evaluation reads the constructed environment map carried by +`result.actions.env` in full: its keys and values remain available to +MiniJinja, so a property can name a variable and compare two environment values +semantically. Every copy that crosses an external boundary is redacted on the +way out, replacing environment keys with stable opaque key tokens and +environment values with the fixed `` marker. The redacted form is +what reaches failure reports, diagnostics, rendered action views, and persisted +regression artefacts. Redaction happens only as data leaves the runner, never +in the projection that assertions read, so it cannot change whether a case is a +FAIL or an ERROR. This is +[ADR-009](../adr-009-bounded-redacted-manifest-telemetry.md)'s +redact-at-the-boundary rule applied to the actions projection. + ### 3. The `forall` block A test case may declare a `forall` mapping from binding names to domain @@ -177,6 +209,8 @@ draws a target name and a flag ordering, builds the graph, and asserts environment closure and source uniqueness over every generated action. ```yaml +netsuke_test_version: "1.1" + test_command_env_is_closed_over_declared_vars: description: Generated invocations never leak ambient environment. forall: @@ -185,6 +219,7 @@ test_command_env_is_closed_over_declared_vars: steps: - given: let: + declared_env: [] name: "{{ target_name }}" flags: "{{ flag_order }}" - when: build_graph @@ -211,10 +246,12 @@ _Table 2: Initial `forall` domain constructors._ Expansion is deterministic. When the Cartesian product of the declared domains is at most the expansion ceiling, the family is enumerated exhaustively. Above the ceiling, the runner samples with a pseudo-random number generator seeded -from a fixed default; the seed appears verbatim in every failure report, and -`netsuke test --seed ` replays it. This mirrors the `derandomize=True` and -pinned-`@example` convention of the workflow contract suites and the -`proptest-regressions/` replay convention. +with the fixed default seed `0`; the seed appears verbatim in every failure +report. `netsuke test --seed ` selects the reported seed, but replays the +reported case only when the generated tuple inputs persisted with that report +are also available. A seed without the tuple inputs is insufficient for replay. +This mirrors the `derandomize=True` and pinned-`@example` convention of the +workflow contract suites and the `proptest-regressions/` replay convention. On failure, the runner delta-reduces the drawn tuple towards domain minima and reports the smallest falsifying binding set it finds, then persists the seed @@ -247,6 +284,8 @@ generated build script is invariant under declaration reordering and locally unaffected by unrelated additions. ```yaml +netsuke_test_version: "1.1" + test_generation_is_order_invariant: description: Declaration order never changes the generated script. subject: @@ -296,8 +335,12 @@ drifting past it. The lint is advisory by default and promoted to a failure with per entry and report the falsifying binding on failure. - `forall` accepts the closed domain vocabulary of Table 2, expands exhaustively at or below the ceiling, and samples deterministically above it. -- Failure reports name the seed and the minimized drawn tuple; - `netsuke test --seed` replays a reported failure exactly. +- Failure reports name both the seed and the minimized drawn tuple. +- `netsuke test --seed ` replays a reported failure exactly only when the + generated tuple inputs persisted with that report are also available; the + seed selects the reported case, and the tuple inputs constrain what is drawn. + A seed whose tuple inputs are missing does not replay: the run samples the + case freshly from that seed and is not a replay of the reported failure. - Regression tuples persist under the test tree and replay before fresh generation. - The `mutations` vocabulary of Table 3 derives the mutated manifest, @@ -384,9 +427,6 @@ is tracked separately by [RFC 0008](0008-code-health.md). - Should the `result.actions` view land inside RFC 0007's phase 7.5.4 result-view task rather than waiting for this extension? Landing it early would let example-based tests use it immediately. -- Does the environment projection need redaction alignment with - [ADR-009](../adr-009-bounded-redacted-manifest-telemetry.md) when failure - reports print constructed environments? - Is delta reduction over the drawn tuple sufficient minimization in practice, or do sampled list domains need element-wise shrinking? diff --git a/docs/roadmap.md b/docs/roadmap.md index 058fe80a8..cdcbd5324 100644 --- a/docs/roadmap.md +++ b/docs/roadmap.md @@ -2123,11 +2123,21 @@ and [§2](rfcs/0012-netsukefile-property-testing.md#2-quantified-assertions). RFC 0012 Table 1. - Add the `action`, `actions_for_rule`, and `has_action` helpers, keeping the view additive-only and decoupled from internal IR types. + - Success: every Table 1 field and helper is observable in plan mode, and + identical manifests produce byte-identical, canonically ordered views under + the RFC 8785 canonical JSON contract of + [RFC 0006 §6.7](rfcs/0006-ansible-inspired-template-standard-library.md#67-canonical-value-equality), + regardless of declaration or IR iteration order. - [ ] 10.1.3. Implement quantified assertions. - Requires 7.5.5 and 10.1.2. - Add `for_all_actions` and `for_all_targets`, evaluated by the existing MiniJinja engine, reporting the falsifying binding with substituted actual values under the established FAIL and ERROR taxonomy. + - Success: quantified action results follow the canonical projection order, + keep the constructed environment available to assertion evaluation, expose + only redacted environment keys and values in diagnostics, rendered action + views, and persisted regression artefacts, and preserve the established + FAIL versus ERROR outcomes for passing, failing, and erroneous assertions. ### 10.2. Declarative bounded generation @@ -2143,15 +2153,19 @@ whether further constructors are warranted before the dialect stabilizes. See nearest-known-key diagnostics, matching the mock matcher convention. - Raise `netsuke_test_version` to 1.1 and gate the new keys on it, so 1.0 runners fail closed under the RFC 0007 version contract. + - Success: version 1.1 files parse every Table 2 constructor, version 1.0 + runners reject the new keys, and malformed constructors receive the + nearest-known-key diagnostic. - [ ] 10.2.2. Implement deterministic expansion, sampling, and replay. - Requires 10.2.1 and 7.5.2. - Enumerate exhaustively at or below the expansion ceiling; above it, - sample from the seeded `proptest` generator; name the seed in every - failure report and honour `netsuke test --seed`. + sample from the `proptest` generator with fixed default seed `0`; name the + seed in every failure report and honour `netsuke test --seed`. - Persist regression tuples under the test tree, replay them before fresh generation, and delta-reduce failing tuples towards domain minima. - - Success: repeated default runs are byte-identical, and a reported - failure replays from its seed alone. + - Success: repeated default runs are byte-identical, and a reported failure + replays only when both its reported seed and generated tuple inputs are + supplied through the persisted regression artefact. ### 10.3. Metamorphic relations, coverage, and adoption @@ -2174,16 +2188,33 @@ and - Report rules that interpolate commands or construct environments without property coverage; keep the report advisory by default and promote it to a failure under `--strict-coverage`. -- [ ] 10.3.3. Dogfood and document the property dialect. + - Success: an uncovered command or environment produces an advisory report + by default, the same fixture fails under `--strict-coverage`, and covered + fixtures produce neither report. +- [ ] 10.3.3. Dogfood the property dialect on example manifests. - Requires 10.3.2. - Add property suites over the repository's example manifests covering the order-invariance, locality, and environment-closure relations end to end. - - Extend the users' guide testing chapter, `contents.md`, and the RFC 0012 - open questions with the measured expansion-ceiling evidence. + - Scope boundary: change only property suites and their example-manifest + fixtures; do not add constructors, alter the dialect, or update docs. + - Success: the example-manifest suites pass for order invariance, locality, + and environment closure, and record the measured expansion ceiling and + case count used by each relation. +- [ ] 10.3.4. Document the adopted property dialect and measurements. + - Requires 10.3.3. + - Update the users' guide testing chapter, `docs/contents.md`, and the RFC + 0012 open questions with the measured expansion-ceiling evidence from + 10.3.3, including links to the example-manifest relations. + - Scope boundary: change documentation only; do not change the property + implementation, suites, dialect, or measured results. + - Success: all three documents state the same measured expansion ceiling, + identify the order-invariance, locality, and environment-closure example + suites, and link readers to the adopted property dialect. **Success criterion:** the worked examples in RFC 0012 §§3-4 run to a green result deterministically on a machine with no compiler and no network, and a -seeded property failure replays identically from the values in its report. +seeded property failure replays identically from its reported seed and tuple +inputs. ## 11. Allow-listed structured-command shell selection diff --git a/typos.toml b/typos.toml index 0a8cdbc28..f72c0ebc2 100644 --- a/typos.toml +++ b/typos.toml @@ -55,17 +55,31 @@ extend-ignore-re = [ "(?m)^ should_upload_workflow_artifacts: \\$\\{\\{ steps\\.release_modes\\.outputs\\['should-upload-workflow-artifacts'\\] \\}\\}$", "(?m)^ should_upload_workflow_artifacts:$", "(?s)```.*?```", + "--[a-zA-Z0-9-]*color[a-zA-Z0-9*-]*", "Azure\\s+Architecture\\s+Center\\s*\\|\\s*Microsoft\\s+Learn\\b", "GitHub\\s+Flavored\\s+Markdown\\b(?:\\s\\(GFM\\))?", + "HashiCorp", + "How to organise your Rust tests", "NETSUKE_COLOR", + "\\b(?:background|theme)_color\\b", + "\\b(?:carousel|dropdown|footer|indicator|items|navbar|toast)-center\\b", + "\\bAppFactory