Skip to content

test: prune excess tests in Node and Python suites - #272

Open
Rome-1 wants to merge 2 commits into
mainfrom
chore/test-purge-rebased
Open

Rome-1 wants to merge 2 commits into
mainfrom
chore/test-purge-rebased

Conversation

@Rome-1

@Rome-1 Rome-1 commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

Prune excess tests in the Node and Python suites

The suites had grown well past what they catch. Every test runs on every gate, so dead weight costs CI time and reviewer attention without guarding anything. This removes tests that duplicate another test's coverage, restate the implementation, pin incidental behavior (formatting, ordering, exact messages nothing relies on), test the framework or a mock of the thing itself, or can never fail — while keeping every regression test for a real fixed bug, every documented-contract test, the Node/Python parity tests, and the security gates.

What changed

  • Node: 2320 → 2046 tests.
  • Python: 1758 → 1664 tests.
  • 66 files touched, net ~5500 lines removed.

Removals, by category:

  • Exact duplicates (same input and assertion twice) across the secret-pattern, risk-rule, policy, audit, and CLI suites.
  • Tests of a hand-copied mirror of production code whose real behavior is already covered elsewhere (e.g. a file that asserted on a locally re-implemented function that does not exist in src/).
  • Vacuous assertions (if-guarded checks, always-true disjunctions, try/except that accepts anything, "contains X" guaranteed by an already-checked fixture).
  • Many narrow table rows consolidated into one table-driven test with the same coverage.
  • Golden snapshot fixtures read only by removed snapshot tests.

A second pass restored five groups the first pass removed but nothing else covered: the MCP server version-from-package.json regression test, the audit-skill unknown---engine exit-code and alias tests, the deep-engine benign-skill false-positive guard, the directory-scan golden snapshots (whose byte-identical Node/Python copies were an implicit parity check), and the only test that appends after a corrupt audit-log line.

Verification

  • Rebased onto current main; the regression tests added by the recent fixes all remain present and pass.
  • Full Node suite: no failures before or after. Full Python suite: one failure (test_version_matches_pyproject) that is a pre-existing local-environment artifact (an editable install frozen at an older metadata version), identical before and after and not touched here.
  • No test was weakened and no CI gate was changed.

Ready for review; intended to land after the 0.10.6 release.

Rome-1 added 2 commits October 6, 2026 03:33
Node: 2329 -> 2046 tests (99 -> 97 files), 178s -> 159s wall time.
Python: 1758 -> 1662 tests (75 files), 60s -> 39s wall time.
Same pre-existing failures before and after in both suites (none
introduced by this change) — see PR/commit description for detail.

Removed, by category:

- Exact duplicates: same input/assertion tested twice under different
  names, within one language's suite (secret-patterns, custom-patterns,
  pattern-engine, risk-rules/risk-rules-comprehensive/bypass,
  command-interceptor/command-interceptor-policy, policy-loader/
  policy-validation, audit-logger/audit-logger-lifecycle, issues,
  e2e-cli, scan-remote, api-scope, test_handle_403, agent-init/
  agent-components).

- Tests of a hand-copied mirror of production logic rather than the
  real code, where real-code coverage already exists elsewhere:
  claude-code-integration.test.ts (whole file — tests a function with
  a signature that doesn't exist in source), most of
  agent-compatibility.test.ts's reimplemented hook-dispatch blocks
  (Codex PreToolUse, both PostToolUse blocks, Cross-Agent Consistency,
  the fs-tautology "Environment Detection" block), the basic
  mcp-server.test.ts handler tests (superseded by mcp-server-
  integration/stdio's real-server tests), and the fake install-hook /
  write-secret-content duplicates in Python's test_hook_integration.py
  (superseded by test_agent_install_hook.py and test_hook.py, which
  import the real functions).

- Vacuous or tautological assertions: if-guarded checks that pass
  whether or not the real behavior works, OR-conditions that always
  hold, try/except accepting any outcome, assertions on a literal the
  test just constructed, "contains X" checks guaranteed true by an
  already-asserted fixture, ANSI/emoji-absence checks implied by an
  adjacent exact-equality assertion.

- Consolidation of near-duplicate per-variant tests into table-driven
  tests, same coverage at lower cost: platform-integration.test.ts's
  8-platform MCP install/idempotency/preserve/corrupted-JSON/
  not-detected matrices, risk-rules tier tables, policy-export/
  command-interceptor-policy tier wiring, pattern-engine generic-key
  false-positive tables, formatter mode-difference tests, codex/
  openclaw skill frontmatter checks, audit-skill high-risk pattern
  detection.

- Dead golden-fixture files only read by the removed snapshot tests
  (directory-scan.json, positions-multi-pattern.json, both languages).

- Cleaned up unused imports left behind by the above (Python, via
  ruff --fix --select F401).

One non-deletion fix: python/tests/test_issues.py's TestParseNaturalText
and its 8 dependent classes called a local mirror of
_parse_natural_text because the import was simply missing, not because
of any technical barrier (other private functions in the same file were
already imported directly) — switched to importing and calling the real
function, turning ~28 previously-fake tests into real coverage.

Kept despite resembling the above: Node's policy-export.test.ts JSON/
TOML-shape tests, posttool.test.ts, baseline.test.ts's applyBaseline
block, hook-formats.test.ts, and hook-integration.test.ts's install-hook
tests — each is the only (if indirect) coverage of real dispatch logic
that isn't exported for direct import, and Python has real coverage of
the equivalent behavior that would otherwise create a parity asymmetry.
An adversarial review of the prune found these had no other coverage:
- mcp-server-integration: version-from-package.json regression (41c0119)
- skill-scanner: audit-skill unknown --engine exits 2 (CLI_SPEC contract,
  Node path separate from agent scan), --engine alias, and the deep-engine
  benign-skill false-positive guard
- snapshot-scan (both runtimes) + directory-scan/positions goldens: the
  byte-identical goldens were an implicit Node/Python parity check
- audit-logger-lifecycle: the only test that appends after a corrupt line,
  which exercises readLastLineHash on the hash chain
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant