Repository navigation
Conversation
Coverage Summary
|
Sourcify sweep (Slang frontend vs solx 0.1.8)
candidate wall time over Outcome by corpus trait
Candidate-only failures (744)kinds:
Failures shared with the baseline (baseline diagnostics) (189)kinds:
|
|
Run 1 (d35b135 = #671 + harness, run 33651331011): 66,142 / 112,587 ok (58.75%), 46,320 Slang-only failures, 125 shared with 0.1.8 (the known stack-too-deep gates), 0 timeouts. Single-file contracts pass at 96%, multi-file at 10%. Seven signatures cover 99.4% of the failures (buckets are first-failure only, so fixing the top one will unmask more of the lower rows):
Long tail (≤ 27 each): LLVM worker assertion The first report's signatures hid most of this ( |
2d34148 to
a98f2e6
Compare
c237244 to
6d41503
Compare
6d41503 to
96cebda
Compare
ed6c4df to
31ee803
Compare
69d842a to
b123248
Compare
88b9e22 to
3dc23b2
Compare
175a65e to
d4479be
Compare
3dc23b2 to
fa0180e
Compare
Compiles the pinned Sourcify corpus (112,587 verified solc 0.8.34 contracts, evmVersion >= cancun) through `solx --standard-json` and reports the outcome table plus a ranked failure census. Each corpus record becomes one standard-JSON input with sources inline and evmVersion, libraries, remappings and viaIR passed through verbatim; candidate failures are re-run with the released solc-frontend solx so they split into candidate-only and both-fail. Failure kinds (error, panic, abort, no-bytecode, timeout) are recorded separately because a frontend crash ends the whole standard-JSON call. CI: `.github/workflows/sourcify-sweep.yaml`, label-gated on `ci:sourcify-sweep` or workflow_dispatch, 24 hosted shards that each extract only their slice of the corpus, summarize job posts the report as a sticky PR comment. Report-only. Ten corpus records are committed as fixtures so the harness runs offline against the workspace build.
The first full run showed the recorded first-line signature is too coarse: `Sol pass pipeline failed` hid the MLIR verifier message in stderr, `__datasize__$<hash>$__` link errors split one bug into hundreds of buckets, and the shared-failure census showed the candidate's kind instead of the baseline diagnostic that explains the contract. Bucketing now lives in report.py (`refine`) over the raw material run.py records, so old results re-render: MLIR verifier and worker-crash lines are pulled from stderr, import failures split into non-relative / URL-style / relative, panics carry their source location, hashes and literal `\n` are normalised. run.py records which contracts did come out on a no-bytecode result.
The Slang frontend checks pragmas against Slang's latest version (0.8.36) while the baseline is a 0.8.34 compiler, so any exact pin fails one leg before parsing and the outcome classification is lost. With --rewrite-pragmas each leg compiles the source with every version pragma replaced by the version its own --version reports; the corpus stays the verbatim Sourcify record. Pragmas have no effect on codegen.
Sourcify holds only 9,296 solc 0.8.36 compilations, too few for a census on their own, so the corpus is the 0.8.34 release (byte-identical half) plus the 6,530 unique 0.8.36 contracts, extracted with the same query and layout. The frontend's 0.8.36 language version is handled by the per-leg pragma rewrite, so the records stay verbatim.
#751 removed the build-slang alias and build-toolchain's solc input; the default build is the Slang pipeline now.
fa0180e to
052ead9
Compare
A full sweep is cheap enough to run per merge, so main gets a census for every push instead of only when someone labels a PR. On a PR the label now keeps the sweep running on each push rather than needing a re-label. Main runs are not cancelled by the next merge so each one is attributable.
cbaa7bd to
97f479e
Compare
The pattern matched any text up to the next `;`, so a comment mentioning "pragma solidity" lost the code after it, often an import. Both legs saw the damaged source and 17 verified contracts landed in both-fail with `Undeclared identifier`, hiding a frontend gap. Matching only version-expression characters leaves those comments alone; real pragmas, including ones split across lines, are still rewritten.
A backslash inside an f-string expression is a syntax error before PEP 701, so report.py did not load on older local interpreters.
GitHub keeps one pending run per group, so with a shared main group a third push replaced the queued run and that commit got no census. The shard timing comment now matches measured runs.
97f479e to
3a1e869
Compare
Sourcify test for 0.8.34 and
>= cancunused to previously test solx. Using this PR to get an idea of how much we can successfully compile at this point, and to find bugs.Adds a CI sweep that compiles every verified Sourcify contract for solc 0.8.34 and 0.8.36 with evmVersion >= cancun (119,117 contracts) through solx and reports what fails and how. Each failure is re-run with solx 0.1.8, so a frontend gap (
solx-fail) is told apart from input that no solx compiles (both-fail). The report ranks failures by signature with example contracts, which makes it a progress meter for the frontend now and a regression net later, oncesolx-failis low enough to gate on.main, and on PR pushes while the PR carriesci:sourcify-sweep. On a PR the report is a sticky comment; onmainit is the run summary.tests/sourcify-sweep/README.mdcovers local runs and the corpus pin.