Skip to content

specs vt: built-in visual snapshot shoot and diff over emitted stories #588

Description

@nathanacurtis

Problem

Visual snapshot testing over the emitted Storybook has to be invented by every customer, yet the CLI already knows every story id it emitted. The run hand-built the whole loop in ~70 lines (Playwright shoot of every *--default story via iframe.html?id=… against a static build, pixelmatch diff) and it worked remarkably well: 142/142 stories, 1m55s per full run, and zero pixel flake across three full runs (unchanged stories diffed at exactly 0 pixels — the emitted output + static build is fully deterministic; the only hygiene needed was an animation/transition freeze and caret hide).

Since that run, a fuller harness has been built and exercised against four library fixtures. It is not a sketch any more — it measures Figma variant renders against Storybook renders of emitted output, scores pass/fail per variant, ranks by dependency depth, and publishes a status page. This issue is about moving that into the product, not inventing it.

Evidence it works

Four fixtures, one run each, same day:

Fixture Components Pairs Pass Fail Pass rate
A (large commercial library) 69 2,513 513 33 94.0%
B (mid-size library) 32 377 334 12 96.5%
C (mid-size library) 29 420 323 0 100%
D (adversarial fixture set) 133 421 72 59 55.0%

Fixture D's low rate is the point — those fixtures are deliberately hostile, and a failure there is usually the fixture doing its job.

The hard parts are solved and written down, not waiting to be discovered:

  • Width pinning. A render is pinned to the variant's authored width, because Figma constrains a FILL root by its parent frame and Storybook has no equivalent. Narrowing the pin to FILL roots only reads as the more principled rule and was measured to be worse.
  • White-union flatten before diff. Figma exports carry alpha, Playwright shots are transparent; raw RGBA comparison flags every white-vs-transparent pixel. Both sides flatten onto an opaque-white union canvas first.
  • Enum-casing normalization. The spec carries Figma's value casing; the emitters normalize it. Values are matched case-insensitively against the emitted contract's string-literal union and replaced with the contract's spelling, so a manifest stays correct across output from different emitter versions.
  • Interaction classification. A prop declared in the spec but absent from the emitted contract does not exist at render time (hover/pressed are CSS pseudo-classes). Variants driving such a prop are marked deferred — not shot, not failed.
  • The noise valve. A per-library ignore file carries passPct, dimTolerancePx, threshold, skip, skipVariants, sampleVariants, pinWidth. Every entry requires a note: saying why, so a permanent measurement fact stays distinguishable from a temporary allowance.

Two baselines, not one

The original framing here and the harness as built are different products. Both diff the same shots against a different left-hand image, so they share everything except where the baseline comes from.

Mode Baseline Answers Needs
Fidelity the Figma export "does the code match the design?" FIGMA_TOKEN, the variant↔story join, the manifest
Regression the last accepted render "did my change break anything?" nothing but a running Storybook

Fidelity is the differentiated thing and carries all the complexity. Regression is cheap, offline, and is what the zero-flake evidence above actually measured.

One consequence worth stating: in fidelity mode there is no "accept the new render." Percy and Chromatic are built around approving a changed snapshot; here the baseline is the design, and a render that disagrees is wrong by definition. accept exists only in regression mode, and is named so nobody reaches for it expecting the other thing.

On-disk layout

<workspace>/testing/visual/
  figma/      <kind>/<key>/<nodeId>.png     design truth, captured overtly
  accepted/   <kind>/<key>/<nodeId>.png     accepted renders (regression mode)
  render/     <kind>/<key>/<nodeId>.png     this run
  diff/       triptychs and masks
  report/     visual-report.{json,md}
  manifest.json
  visual-ignore.yaml

Not under storybook/. That directory is generated wholesale by specs storybook and rewritten by init --force — a hand-tuned ignore file kept there is a file waiting to be overwritten.

Commands

specs testing visual, with the testing namespace left open for the other suites that already exist as scripts (parity, emitted-tree typecheck, render round-trip, schema compliance).

The subcommand split is not cosmetic. Each stage has a different cost, and the point of splitting is to pay only the one you need — editing a tolerance and re-scoring must cost seconds, not minutes:

Subcommand Cost Network Browser Flags
init instant — — --force
status instant — — --components, --json
manifest seconds — — --components, --check
baseline slow Figma REST — --components, --force
shoot minutes — Playwright --components, --workers N, --port, --target react|webcomponents
diff seconds — — --components, --against figma|accepted
report instant — — --format json|md
accept instant — — --components, --all
(bare) minutes — yes shoot → diff → report

Universal: --config <path>, --kinds components,compositions, --verbose.

  • baseline is never implicit. Not in the bare command, not on a watch, not on a schedule. It is slow, it is rate-limited, and it is the one stage that costs someone else's API quota. The bare command prints how many shootable variants have no baseline and stops there.
  • init scaffolds, the customer installs. The same contract as specs storybook init: it writes testing/visual/package.json declaring Playwright, pixelmatch and pngjs, then prints the one install command to run. We never ship, vendor, or install them — shoot resolves Playwright from that directory at runtime and fails with a pointer to init when it is absent. Everything except shoot works with nothing installed.
  • manifest --check is a dry run reporting unmapped Figma props. Drive it to zero before trusting a library's numbers — an unmapped prop defers every variant that uses it.
  • diff is the only stage that reads visual-ignore.yaml. sampleVariants and pinWidth are manifest-time keys and need a rebuild plus a re-shoot; everything else re-scores in seconds with no browser.
  • Scoring stays in the ignore file, not in flags. A tolerance is a durable fact about a library that needs a note: explaining it. A flag is something passed once to make red go away, and nobody can tell afterwards whether it hid a bug.

Staleness is provenance, not scheduling

An earlier draft keyed baseline staleness off the Figma file's lastModified. That is the wrong signal: the timestamp moves on any edit anywhere in the file, so a run recaptures the whole catalogue to discover nothing changed.

With capture overt and nothing automatic, there is no trigger to compute, so v1 computes none. What it does instead:

  • status reports present / missing per shootable variant. Missing is knowable with no model at all.
  • The report records each baseline's capturedAt beside the payload's own fetch time, so a diff cannot silently blame the emitter for a baseline that predates the design.

Making baseline self-scoping — so an overt run recaptures the three components that changed rather than all eighty — is a later optimization, tracked separately. The cheap signal for it is a hash of the component's Figma node subtree out of the already-fetched payload: offline, free, exact per component, and unmoved by edits elsewhere in the file. Deliberately not the spec version from the ledger: the spec is lossy by design, so that signal goes blind exactly where visual testing earns its keep.

Not in build or run

Out of scope, and not deferred-pending-design — ruled out for v1 on the merits.

Every candidate trigger fires far more often than the thing it is meant to detect:

Trigger Why it fails
Figma file changed moves constantly; few or no components actually change
Spec regenerated a spec-first authoring session rewrites specs every few seconds
Emitted code changed every save, and a shoot is minutes with a browser

Until there is a signal that identifies what changed at component granularity, any automatic trigger is a browser launch tax on an unrelated edit. There is also a mechanical mismatch: steps.ts steps are in-process function calls over files, and shooting needs an external HTTP server listening on a port — unlike every existing step.

A component-scoped recapture driven by a narrow watch is imaginable later. It is not this issue.

Why this cannot stay outside the CLI

The harness discovers specs with a flat directory read, keeping any child that holds an api.yaml. Under the current layout (ADR-096) the children of specs/ are components/ and compositions/, neither of which holds an api.yaml — so discovery returns nothing. The harness still runs against exactly one fixture, and only because that fixture is the last one in the pre-ADR-096 flat layout.

Contract resolution has the matching problem: it hard-codes react/src/components/<Pascal>, which no composition will ever be under.

Both are one fix — discovery and contract resolution go through the layout seam — and that seam exists only in the CLI, deliberately: it is documented as the only thing that knows the layout, after two divergent filters preceded it and a third kind had nowhere to go. Repairing the harness in place means reimplementing it, which is what the seam exists to prevent.

Compositions

Visual testing targets both kinds. A composition's spec settles what that means: no props, a plain FRAME source node, and exactly one emitted story — the emitter writes the reason into the file ("A composition has no variant axes — no sticker sheet").

Concern Component Composition
Figma node COMPONENT_SET → enumerate variant children single FRAME → one node
Variant name parsing Prop=Value, …, can fail none
Prop mapping the main source of deferred pairs nothing to map
Variant↔story join exact / overlay / no-story heuristic exact by construction
Pairs per spec 1–150 exactly 1

The manifest side is nearly free — the builder already takes a non-COMPONENT_SET branch yielding one variant with an empty config. The cost is downstream:

  1. Baseline keying becomes kind-aware. A component and a composition sharing a name is legal and is handled that way in the version assembler. figma/<key>/ therefore collides; it needs figma/<kind>/<key>/, and the same split in the render and diff trees, the report, and the ignore file.
  2. Crop rule — decided. Compositions are frames placed directly on a page node, so every one of them is fixed-width. Pin width unconditionally, and pin height only when the frame's vertical sizing is FIXED; a hugging frame keeps its natural height. This extends the existing width pin rather than introducing a second rule, and turns vertical overflow into visible pixels instead of a dimension mismatch. The height half is the only part that needs measuring, and it can be measured when the first hugging composition appears rather than up front.
  3. Scoring is advisory — decided. Compositions are captured, diffed, and reported with pixel counts, but their pass/fail does not decide the run's exit status. A composition's diff is dominated by its children, so one leaf defect lights up every composition containing it; the existing depth-ascending ranking already says to fix leaves first, and this makes the report agree with it.
  4. Pro gating. Composition output is Pro (ADR-097). Composition visual testing filters on a free run the same way transform does, reporting the count once.
  5. Images and fonts. Canonical page compositions are where image sources and missing licensed fonts concentrate. The noise valve needs the kind dimension from point 1 before it can absorb that separately from component noise.

What ports, and where it lands

~3,000 lines across 15 files in the existing harness.

Harness module LOC Destination
workspace resolution 128 Delete. storybook/workspace.ts already resolves root, specs, data sources, emitted trees, scaffold state, and the dev-server port
spec index 114 Port, shrinks — the CLI knows its own layout
variant↔story join 270 Port as-is. The hardest and most valuable code
PNG utilities 144 Port as-is
manifest builder 450 Port, shrinks — reads payloads through SectionedFile instead of hand-walking JSON
baseline status 118 Port, shrinks — no staleness model
Figma baseline capture 118 Port as-is
shoot 276 Port as-is, plus the composition pin rule
diff + report 435 Port as-is, plus kind-aware keying and advisory scoring
orchestrator 148 Becomes the bare command
status page 298 Becomes a visualtesting Storybook concern

The concern registry (storybook/concerns/registry.ts) is exactly what the report page needs: detect() = does a report directory exist, build() = emit the page. The three hand edits to .storybook/main.ts that currently mount it disappear — and have to, because init --force now owns that file. The scaffolded host already reserves the mount point, serving a visual-testing baselines directory at /baselines when one exists; that line is currently the only mention of visual testing anywhere in the CLI.

No new CLI dependencies. Playwright, pixelmatch and pngjs are declared in the scaffold init writes and installed by the customer in their own testing/visual/ directory — the same contract as the Storybook host, where we never ship or vendor Storybook (ADR A). A customer who never does visual testing never downloads a browser, and the published CLI stays the nine small dependencies it is today.

Build order

In sequence. Each step is independently verifiable. Fixture A is the muse throughout steps 1–7; the others widen coverage in step 8, once the suite is complete.

# Step Validated by Size
#711 Layout-seam discovery and kind-aware baseline keying both kinds enumerate where zero do today m
#712 Command surface, manifest, status unmapped props at zero; no browser, no network l
#713 Overt Figma baseline capture scoped capture writes only what was named m
#714 Shoot, variant↔story join, pin rules three full shoots diff render-against-render at 0 pixels l
#715 Diff, scoring, ignore resolution, report a tolerance change re-scores in seconds with no browser l
#716 Report page as a Storybook concern survives storybook init --force m
#717 Regression mode against accepted renders runs end to end with no Figma token m
#718 Broaden to the remaining fixtures, and document all three complete a full cycle l

Follow-on, not required to ship:

Steps 1–5 are a straight line. #716 and #717 both depend only on #715 and are independent of each other.

Acceptance criteria

  • A customer can capture a baseline, diff a regeneration against it, and read a report using shipped commands.
  • Changed variants are listed with per-variant pixel counts.
  • Both components and compositions are captured, diffed and reported, keyed separately so a shared name cannot collide.
  • Composition results are reported but advisory — they do not decide the run's exit status.
  • Spec discovery and contract resolution go through the layout seam, so every supported layout works.
  • Baseline capture happens only when asked for by name.
  • The report states each baseline's capture time and the payload's fetch time.
  • The status page is published as a Storybook concern, requiring no hand edits to the scaffolded host.
  • The ignore file's scoring keys take effect on a re-diff alone, with no re-shoot.

Non-goals

  • Any automatic trigger, and any integration with build or run.
  • A staleness model. status reports present/missing only.
  • Drop-shadow bounds. Figma exports include them; a tight render crop does not. Needs a shadow-aware rule before elevated variants are scoreable.
  • Hover and pressed states. Driving them with Playwright is deliberately out of scope.
  • Figma variable modes — everything assumes mode defaults.
  • Fixing what the diff finds. This measures; the fixing workflow lives elsewhere.

Case data

  • Workspace: atlassian — Atlassian Design System (ADS) community files; sealed first-customer onboarding simulation, 2026-09-27
  • Territory: cli
  • Size: l

Notes

Priority: P3. Related: #489, #490 (VT content policies), #405, #406 (existing VT concerns). Determinism evidence above is the strongest argument this belongs in the product: the hard part (stability) is already solved by the output's nature.


Implementation details are tracked internally.

Activity

  1. added
    clispecs-cli commands
    testingspecs-testing parity validation
    docsDocumentation that appears on specsplugin.com
    on Sep 27, 2026
  2. moved this to Backlog in Specson Sep 27, 2026
  3. nathanacurtis commented on Sep 28, 2026

    @nathanacurtis
    MemberAuthor

    Handing one item here from #574, which briefly scoped it before being converted to an umbrella: per-workspace diff scoring tolerances.

    The maintained internal harness carries a per-workspace ignore file holding passPct, dimension tolerance, per-component skips, variant skips, and variant sampling caps — with a convention that every entry states why, so a permanent measurement fact stays distinguishable from a temporary allowance. Without that layer every diff scores against one absolute default, and ordinary rasterization noise between a browser and a design tool reads as failure.

    That belongs with the harness, not with the viewer. #574 now covers only the surface that displays results; how they are scored is yours.

    Two related notes from a second sealed onboarding run (2026-09-27, an unrelated 69-component library), both relevant to scope here:

    • The run built its own Playwright and pixel-diff harness from scratch, not knowing a richer one exists, and diffed render against render — cross-framework and self-baseline. That proves the two frameworks agree but never that either matches the design. A customer reinventing this will likely make the same substitution, since the design-source comparison is the part that needs a manifest, node mapping, and baseline capture.
    • That library declares Light and Dark modes. All 138 of its snapshots were taken in Light, and the run reported complete coverage. Any harness over a multi-mode library needs modes in its coverage model or it overstates what it verified. The toolbar half of that is Color page, generated #609; the coverage half is here.

    Given both, this issue's scope may be understated in the same way #574's was — it was written before this comparison existed.

  4. 1 remaining item

  5. nathanacurtis commented on Oct 8, 2026

    @nathanacurtis
    MemberAuthor

    All ten subissues closed; the suite, its Storybook pages, the docs, the fixture sweep, self-scoping capture, and the workspace cleanup are done. Everything lands via PR #721 (base: feature/compositions-cli).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

clispecs-cli commandsdocsDocumentation that appears on specsplugin.comtestingspecs-testing parity validation

Fields

Priority

Moderate

Projects

  • Status
    Done

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions