You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Problem — amicode main currently carries pre-existing suite failures that every campaign must baseline-diff around: the amico-run suite fails 10 tests (unreachable cloud endpoints, curlSotaFetch, the gh shim, a loadRegistry seed, parseDoctorArgs, a reviewSpec guard) and the extension suite fails 16 (service auth/bootstrap, contract, engine-proxy, fleet-data-plane, runner/wiring, staged-bins, one OPENCODE_DB-dependent open-threads case). They are environment-dependent — green on some hosts, red on others — which forces every campaign to prove "identical to baseline" by hand and makes CI verdicts ambiguous. Approach — Make each failing test honest about its environment requirement: tests that need external endpoints or credentials skip cleanly (explicit skip reason, not a bare fail) when the requirement is absent; tests that are genuinely broken on main get fixed or explicitly marked with a tracked reason. No test is deleted; no assertion is weakened — a skipped test states what it needs and how to run it. Scope — in: the 10 amico-run failures + the 16 extension failures on current main; out: any feature work, the scores/skill-drift real-checkout portability issue (#1531), the re-land PRs (#1526–#1530).
Acceptance Criteria
From a fresh worktree off main, the amico-run suite reports 0 failures (skips allowed with explicit reasons)
From a fresh worktree off main, the extension suite reports 0 failures (skips allowed with explicit reasons)
Every skip carries a stated environment requirement (endpoint/credential/binary) — no silent skips
Tests that were genuinely broken (not env-dependent) are fixed, with the fix described in the PR
Testing Decisions
The existing suites themselves are the surface; no new test files except where a skip-reason contract needs pinning.
The session-curation campaign's S3 baseline-diff proof (the 16-extension-fail set was first proven pre-existing there)
Notes
Found repeatedly during the 2026-09-24 re-land campaign; every implementer cast spent effort proving these failures pre-existing. Label afk: implement + merge unattended per house AFK semantics, merges still sequential-green only.
Important
Problem — amicode
maincurrently carries pre-existing suite failures that every campaign must baseline-diff around: the amico-run suite fails 10 tests (unreachable cloud endpoints,curlSotaFetch, the gh shim, a loadRegistry seed,parseDoctorArgs, a reviewSpec guard) and the extension suite fails 16 (service auth/bootstrap, contract, engine-proxy, fleet-data-plane, runner/wiring, staged-bins, oneOPENCODE_DB-dependent open-threads case). They are environment-dependent — green on some hosts, red on others — which forces every campaign to prove "identical to baseline" by hand and makes CI verdicts ambiguous.Approach — Make each failing test honest about its environment requirement: tests that need external endpoints or credentials skip cleanly (explicit skip reason, not a bare fail) when the requirement is absent; tests that are genuinely broken on main get fixed or explicitly marked with a tracked reason. No test is deleted; no assertion is weakened — a skipped test states what it needs and how to run it.
Scope — in: the 10 amico-run failures + the 16 extension failures on current main; out: any feature work, the scores/skill-drift real-checkout portability issue (#1531), the re-land PRs (#1526–#1530).
Acceptance Criteria
Testing Decisions
The existing suites themselves are the surface; no new test files except where a skip-reason contract needs pinning.
Prior Art
Notes
Found repeatedly during the 2026-09-24 re-land campaign; every implementer cast spent effort proving these failures pre-existing. Label afk: implement + merge unattended per house AFK semantics, merges still sequential-green only.