Repository navigation
Conversation
Node: 2329 -> 2046 tests (99 -> 97 files), 178s -> 159s wall time. Python: 1758 -> 1662 tests (75 files), 60s -> 39s wall time. Same pre-existing failures before and after in both suites (none introduced by this change) — see PR/commit description for detail. Removed, by category: - Exact duplicates: same input/assertion tested twice under different names, within one language's suite (secret-patterns, custom-patterns, pattern-engine, risk-rules/risk-rules-comprehensive/bypass, command-interceptor/command-interceptor-policy, policy-loader/ policy-validation, audit-logger/audit-logger-lifecycle, issues, e2e-cli, scan-remote, api-scope, test_handle_403, agent-init/ agent-components). - Tests of a hand-copied mirror of production logic rather than the real code, where real-code coverage already exists elsewhere: claude-code-integration.test.ts (whole file — tests a function with a signature that doesn't exist in source), most of agent-compatibility.test.ts's reimplemented hook-dispatch blocks (Codex PreToolUse, both PostToolUse blocks, Cross-Agent Consistency, the fs-tautology "Environment Detection" block), the basic mcp-server.test.ts handler tests (superseded by mcp-server- integration/stdio's real-server tests), and the fake install-hook / write-secret-content duplicates in Python's test_hook_integration.py (superseded by test_agent_install_hook.py and test_hook.py, which import the real functions). - Vacuous or tautological assertions: if-guarded checks that pass whether or not the real behavior works, OR-conditions that always hold, try/except accepting any outcome, assertions on a literal the test just constructed, "contains X" checks guaranteed true by an already-asserted fixture, ANSI/emoji-absence checks implied by an adjacent exact-equality assertion. - Consolidation of near-duplicate per-variant tests into table-driven tests, same coverage at lower cost: platform-integration.test.ts's 8-platform MCP install/idempotency/preserve/corrupted-JSON/ not-detected matrices, risk-rules tier tables, policy-export/ command-interceptor-policy tier wiring, pattern-engine generic-key false-positive tables, formatter mode-difference tests, codex/ openclaw skill frontmatter checks, audit-skill high-risk pattern detection. - Dead golden-fixture files only read by the removed snapshot tests (directory-scan.json, positions-multi-pattern.json, both languages). - Cleaned up unused imports left behind by the above (Python, via ruff --fix --select F401). One non-deletion fix: python/tests/test_issues.py's TestParseNaturalText and its 8 dependent classes called a local mirror of _parse_natural_text because the import was simply missing, not because of any technical barrier (other private functions in the same file were already imported directly) — switched to importing and calling the real function, turning ~28 previously-fake tests into real coverage. Kept despite resembling the above: Node's policy-export.test.ts JSON/ TOML-shape tests, posttool.test.ts, baseline.test.ts's applyBaseline block, hook-formats.test.ts, and hook-integration.test.ts's install-hook tests — each is the only (if indirect) coverage of real dispatch logic that isn't exported for direct import, and Python has real coverage of the equivalent behavior that would otherwise create a parity asymmetry.
An adversarial review of the prune found these had no other coverage: - mcp-server-integration: version-from-package.json regression (41c0119) - skill-scanner: audit-skill unknown --engine exits 2 (CLI_SPEC contract, Node path separate from agent scan), --engine alias, and the deep-engine benign-skill false-positive guard - snapshot-scan (both runtimes) + directory-scan/positions goldens: the byte-identical goldens were an implicit Node/Python parity check - audit-logger-lifecycle: the only test that appends after a corrupt line, which exercises readLastLineHash on the hash chain
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Prune excess tests in the Node and Python suites
The suites had grown well past what they catch. Every test runs on every gate, so dead weight costs CI time and reviewer attention without guarding anything. This removes tests that duplicate another test's coverage, restate the implementation, pin incidental behavior (formatting, ordering, exact messages nothing relies on), test the framework or a mock of the thing itself, or can never fail — while keeping every regression test for a real fixed bug, every documented-contract test, the Node/Python parity tests, and the security gates.
What changed
Removals, by category:
src/).A second pass restored five groups the first pass removed but nothing else covered: the MCP server version-from-
package.jsonregression test, theaudit-skillunknown---engineexit-code and alias tests, the deep-engine benign-skill false-positive guard, the directory-scan golden snapshots (whose byte-identical Node/Python copies were an implicit parity check), and the only test that appends after a corrupt audit-log line.Verification
main; the regression tests added by the recent fixes all remain present and pass.test_version_matches_pyproject) that is a pre-existing local-environment artifact (an editable install frozen at an older metadata version), identical before and after and not touched here.Ready for review; intended to land after the 0.10.6 release.