Skip to content

Latest commit

 

History

533 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

webgpu-native-cts

A WebGPU Conformance Test Suite (CTS) written in C++, targeting the WebGPU C API (webgpu.h) directly — no JavaScript engine, no language bindings.

The tests are C++ that call the WebGPU C API. C++ is used because the upstream CTS is object-oriented, fluent, and closure-based; a C++ port maps to it almost 1:1, which is the whole point of a faithful port.

The goal is to port the upstream TypeScript WebGPU CTS as faithfully as practical, so that native WebGPU implementations can be checked for conformance by linking a small native test binary instead of embedding a full JS runtime.

Why

The canonical WebGPU CTS is written in TypeScript and runs in a browser or against native implementations through a JS engine bound to native code:

  • wgpu runs the CTS via Deno (deno_webgpu ops → wgpu_core).
  • Dawn runs the CTS via Node.js NAPI bindings (dawn_node → dawn::native).

Both approaches require shipping and maintaining a JavaScript runtime and a binding layer. This project takes a different path: the tests are written directly against the C API that implementations such as wgpu-native, yawgpu, and Dawn export, so the test binary links straight against the implementation under test.

Target implementations

Tests are written against the canonical webgpu.h from webgpu-native/webgpu-headers, which all of the target backends implement. The suite is link-agnostic: a build-time option (CTS_BACKEND) selects which implementation to link against. Three backends are targeted, each playing a distinct role in the cross-implementation comparison:

  • yawgpu (github.com/infosia/yawgpu) — a from-scratch Rust implementation of webgpu.h (Metal/Vulkan backends). Its WGSL frontend is Tint — the same shader compiler Dawn uses — so its shader-compile, const-eval, and validation behaviour is byte-equivalent to the oracle. The primary conformance subject: this suite exists to validate it. On native Metal it runs the suite crash-free with shader/execution fully green, its fail profile byte-identical to the Dawn oracle, and its coverage at Dawn's level (see Test results). Its vendor-extension surface (the companion yawgpu.h) is not yet exercised by the suite — only the canonical webgpu.h is tested today.
  • Dawn — Google's C++ reference implementation. It passes the ported suite, so it serves as the conformance oracle: the ground-truth behaviour against which any backend disagreement is judged. Since yawgpu and Dawn now share the Tint frontend, they agree on shader compilation and validation.
  • wgpu-native — the mature Rust implementation (wgpu_core), still on the naga WGSL frontend; the backend the harness was first brought up against, and a third independent data point. Being the one remaining naga-based backend, it is now where the naga-lineage shader findings (const-eval, frontend-validation, discard-derivative) still manifest after yawgpu moved to Tint.

Implementation-specific differences (native feature enums, instance-creation extras like yawgpu's YaWGPUInstanceBackendSelect backend selector) are isolated behind a thin backend shim.

Language

  • The entire suite — tests and harness — is C++20.
  • Tests call the WebGPU C API (webgpu.h) directly; they do not depend on any C++ wrapper for WebGPU itself.
  • The harness is a custom C++ framework that mirrors the upstream CTS framework 1:1 (MakeTestGroup / g.test().desc().params().fn(), fluent parameter builders — combine/combineWithParams/filter/expand/beginSubcases with per-case subcase expansion — lambda test bodies, expectValidationError([&]{ ... }, shouldError), the suite:file:test:params query system, and a case/subcase tree). It is not built on GoogleTest, Catch2, doctest, or Criterion — those impose a test model that conflicts with the CTS case/subcase split and query-string identity. The harness's own logic (params expansion, query parsing, format tables, expectation matching) is checked by a small self-test binary, cts_unittests, with no third-party test framework.

Repository layout

webgpu-native-cts/
├── README.md  CLAUDE.md  LICENSE  CMakeLists.txt
├── docs/                  # Design docs + UPSTREAM / COVERAGE / FINDINGS
├── specs/                 # Per-slice task specs + reference/ (workflow, templates)
├── include/cts/           # Public C++ test-author API (gpu.h, test.h, webgpu.h)
├── src/
│   ├── common/            # Harness: registry, params, query, runner, runtime, webgpu/ (async→sync, backend shim)
│   ├── webgpu/            # Ported tests (.spec.cpp) + capability_info / texture_format tables + listing.json
│   └── unittests/         # Harness self-tests (cts_unittests)
├── expectations/          # Per-backend known-failure lists ({wgpu-native,yawgpu,dawn}.txt)
└── tools/gen_listings/    # Listing generator → src/webgpu/listing.json

The build directories (build-*/) are git-ignored. The canonical webgpu.h is not vendored — each backend supplies its own webgpu-headers/webgpu.h (Dawn its generated header), selected by CTS_BACKEND at configure time.

Status

Active. The harness is complete and all three backends build link-agnostically and run on real GPUs (macOS / Apple Metal and Windows 11 / Vulkan, NVIDIA RTX 5060 Ti). Every portable upstream area is ported: the entire api surface (api/validation + api/operation), all of shader/execution (the FP-interval f32/f16/abstract math framework, the subgroup*/quad* execution builtins on a ported subgroup_util engine, and the texture_utils meta-test), and all of shader/validation (parse/statement/expression + the 114 builtin signature/type/const-overflow specs, on a ported ShaderValidationTest enabler). The only unported upstream is compat (todo) and web_platform/idl (N/A — no C-API surface). See coverage below and COVERAGE.

Conformance (current). yawgpu's WGSL frontend is Tint (Dawn's shader compiler), so its shader-compile behaviour is byte-equivalent to the Dawn oracle: the whole suite runs crash-free on every real-hardware path, and shader/execution is fail=0 crash=0 on native Metal. Its Metal coverage is at Dawn's level and its Metal fail profile is byte-identical to Dawn; on Vulkan the residual fails are known non-defects carried as xfail. The naga-lineage findings do not manifest on yawgpu — they appear only on wgpu-native, the one naga-based backend and a panic-heavy bring-up reference. Per-backend numbers: Test results; per-finding detail: FINDINGS.

Port coverage

683 upstream .spec.ts files at the pinned revision:

pie showData
    title Upstream .spec.ts files (683)
    "Ported — complete" : 586
    "Ported — partial" : 56
    "Not portable (N/A)" : 21
    "Todo" : 20
Loading

Every portable area is done — all 201 api/* files and all 446 shader/* files (execution + validation) are accounted for (complete, partial, or classified N/A), with zero remaining todo in those areas. "Addressed" below = complete + partial + N/A (every upstream file resolved); the only remainder is compat (todo) and web_platform/idl (N/A).

Area Addressed* Note
api/validation 129 / 129 ✅ fully ported — every file complete (112), partial (14), or N/A (3); no todo (Y-6 V1–V10: capability_checks/features + all 35 limits)
api/operation 72 / 72 ✅ fully ported — every file complete (28), partial (42), or N/A (2); no todo. Partials leave some native-portable breadth deferred (vertical-first)
shader/execution 239 / 239 ✅ fully ported — structural files + the entire expression/call/builtin family: atomics, the texture built-in family, sync/derivatives, integer/bit/pack, the P4 math/trig builtins on a ported FP-interval acceptance framework (f32 / f16 / abstract-float), all binary/unary operators, conversions, constructors, expression/access/*, the subgroup*/quadBroadcast/quadSwap execution builtins (on a ported subgroup_util compute/fragment/accuracy engine), and the texture_utils meta-test. Dawn-oracle green (fail=0; subgroup-size gates honored — Dawn-Metal subgroupMaxSize=32); yawgpu is fail=0 here too (Tint frontend, Dawn-equivalent)
shader/validation 207 / 207 ✅ fully ported — extension/shader_io/decl/functions/types/const_assert/uniformity, parse, statement, and expression (incl. all 114 builtin signature/type/const-overflow specs) on a ported ShaderValidationTest enabler (expectCompileResult/expectPipelineResult). Dawn-oracle green (fail=0); yawgpu is fail=0 here too — byte-identical to Dawn, including the subgroups-gated cases (parse,requires:wgsl_matches_api, uniformity:uniform_subgroup_ops), since it shares the Tint frontend. The naga-lineage validation gaps appear only on wgpu-native
Total 663 / 683 + web_platform/idl N/A (16); compat + misc todo

* addressed = complete + partial + N/A (every upstream file resolved). Per-file detail and what each batch added: COVERAGE.

What works

  • Harness, 1:1 with upstream — registry, fluent params, the suite:file:test:params query system, fixtures, error scopes, async→sync helpers, listing generator, cts_unittests self-tests.
  • Crash containment — --isolate runs each case in a child process (POSIX + Windows), so a backend abort becomes a contained crash result instead of killing the run.
  • Per-backend expectations (--expectations) — runs with known divergences still exit 0, with nothing silently masked; --workers N shards a full sweep ~10× faster.

Test results

Per-area pass / skip / fail / crash from a full sweep of all 642 ported files: on macOS / Apple Metal, on Windows and Linux / native Vulkan (NVIDIA RTX 5060 Ti), and on GLES (Tier-2 experimental — bring-up snapshots, not conformance results). Every backend runs whole-suite per-subcase (--workers, no --isolate). All tables are raw (no --expectations), so documented non-defects show in fail rather than being masked; each table carries its own sweep date, run mode and backend revision. Per-finding detail is in FINDINGS.

yawgpu — native Metal (Tint frontend), per-subcase

area pass skip fail crash
api/validation (126) 293,557 61,433 2† 0
api/operation (70) 228,852 741 0 0
shader/execution (239) 822,636 21,950 0 0
shader/validation (207) 646,773 20,369 0 0
total 1,991,818 104,493 2† 0

Dawn — native Metal (conformance oracle, Tint frontend), per-subcase

area pass skip fail crash
api/validation (126) 293,547 61,443 2† 0
api/operation (70) 228,849 744 0 0
shader/execution (239) 822,636 21,950 0 0
shader/validation (207) 646,773 20,369 0 0
total 1,991,805 104,506 2 0

† Documented non-defects, carried as xfail.

yawgpu matches Dawn on Metal. Same fail profile (only those 2 cases), both shader areas subcase-identical, and the skip counts within 13 subcases of each other — a capability-exposure difference, not conformance divergence.

wgpu-native — native Metal (bring-up reference, naga frontend), per-subcase

area pass skip fail crash
api/validation (126) 178,872 81,235 9,642 7,116
api/operation (70) 149,347 76,270 3,681 171
shader/execution (239) 635,636 125,436 4,280 32,096
shader/validation (207) 277,329 316,065 73,748 0
total 1,241,184 599,006 91,351 39,383

Swept 2026-09-26 raw — four per-area --workers 6 runs on wgpu-native v29.0.1.1 (the latest release; wgpu-core / naga 29.0.3) / CTS 242cd3e, macOS / Apple M2: 47 minutes (api/validation 2m08s, api/operation 2m00s, shader/execution 42m36s, shader/validation 43s). All-features tests run with every optional feature wgpu-native can grant (it also advertises two experimental features it cannot, which the harness excludes — F-154). Known release defects include the immediates API divergences (F-154) and mapAsync rejecting WGPU_WHOLE_MAP_SIZE (F-155, fixed upstream after the release).

wgpu-native is on the naga WGSL frontend and is a panic-heavy bring-up reference. Because many cases share a worker process, one abort contaminates the rest of that process: crash is inflated versus a per-case isolate run and pass is correspondingly deflated, and an aborted case is counted once rather than per subcase, so the row totals are below the suite's subcase count. Read these numbers as run-mode-sensitive, not a like-for-like comparison to yawgpu/Dawn.

yawgpu — native Vulkan (Windows 11 / NVIDIA RTX 5060 Ti, Tint frontend), per-subcase

area pass skip fail crash
api/validation (126) 244,673 110,315 4‡ 0
api/operation (70) 209,360 20,233 0 0
shader/execution (239) 531,309 313,164 113‡ 0
shader/validation (207) 646,860 20,282 0 0
total 1,632,202 463,994 117‡ 0

Swept 2026-09-26 raw — four per-area --workers 8 runs on yawgpu 53ca8bb (adds subgroup-size-control on Vulkan) / this CTS revision, NVIDIA driver 610.62, crash=0. Versus the 2026-09-25 sweep (yawgpu 2c6ea6f): +94 pass / −94 skip, all from the now-runnable subgroup-size-control cases (compute_builtins:subgroup_size_attribute 6, shader,validation,extension,subgroup_size_control 87, capability_checks,features,subgroup_size_control 1); the fail set is unchanged.

‡ Documented non-defects, carried as xfail in expectations/yawgpu-vulkan.txt; the suite exits fail=0 once expectations are applied.

yawgpu — native Vulkan (Linux / NVIDIA RTX 5060 Ti, Tint frontend), per-subcase

area pass skip fail crash
api/validation (126) 244,676 110,316 4§ 0
api/operation (70) 209,360 20,233 0 0
shader/execution (239) 531,304 313,170 113§ 0
shader/validation (207) 646,773 20,369 0 0
total 1,632,113 464,088 117§ 0

Swept 2026-09-21 raw — four per-area --workers 4 runs on yawgpu 80219df / CTS df58708, NVIDIA driver 595.91, Vulkan 1.4.329: 53 minutes, crash=0 across 2,096,318 subcases. Same GPU as the Windows table on a different OS: skip and fail counts are identical in every area, and pass counts agree to within 5 subcases.

§ Documented non-defects, carried as xfail in expectations/yawgpu-vulkan.txt; the suite exits fail=0 once expectations are applied.

wgpu-native — native Vulkan (Linux / NVIDIA RTX 5060 Ti, naga frontend), per-subcase

area pass skip fail crash
api/validation (126) 132,556 132,375 8,694 7,686
api/operation (70) 150,337 76,267 2,695 170
shader/execution (239) 419,010 416,648 2,990 3,167
shader/validation (207) 277,329 316,065 73,748 0
total 979,232 941,355 88,127 11,023

Swept 2026-09-26 raw — four per-area --workers 6 runs on wgpu-native v29.0.1.1 / CTS 60af4b1, NVIDIA driver 610.57.04: 2h45m (api/validation 21m36s, api/operation 33m51s, shader/execution 1h48m56s, shader/validation 31s). Much slower than Metal because every abort restarts a shard worker (device re-creation, cold pipeline caches). shader/validation is subcase-identical to the Metal table (the naga frontend is platform-independent). The skip gap versus Metal is hardware/feature exposure, not conformance: ~342k skips are ASTC / ETC2 / EAC formats (no such compression on desktop NVIDIA — yawgpu's Linux table skips them too) and ~50k are subgroups not granted. The three api,operation,limits,max_combined_limits: max_storage_buffer_texture_frag_outputs cases, which abort on Metal (F-088), hang on Vulkan (100% CPU, RSS growing ~40 MB/min); they were killed after ~32 minutes and are counted as crash. The same run-mode caveats as the Metal table apply.

yawgpu — GLES / Tier 2 experimental (Linux / NVIDIA RTX 5060 Ti, Tint frontend), per-subcase

This is not a conformance table — see the Tier-2 note under the Haswell table below.

area pass skip fail crash
api/validation (126) 203,412 151,323 257 0
api/operation (70) 141,076 76,704 11,813 0
shader/execution (239) 309,535 517,406 17,645 0
shader/validation (207) 369,753 297,389 0 0
total 1,023,776 1,042,822 29,715 0

Swept 2026-09-21 raw — --workers 2 on yawgpu 80219df built --features gles, OpenGL ES 3.2 via EGL_PLATFORM_DEVICE_EXT (headless), driver 595.91.07; 10 minutes. shader/validation is byte-identical to the Haswell table below — the WGSL→GLSL-ES path is Tint, so it is driver-independent. Nothing is excluded here: the 2 files quarantined on the Haswell host run clean on this driver.

yawgpu — GLES / Tier 2 experimental (Linux / Mesa crocus on Intel Haswell, Tint frontend), per-subcase

This is not a conformance table. GLES is yawgpu's Tier 2 / experimental backend (opt-in gles cargo feature), and this row is a bring-up progress snapshot, not a fail≈0 conformance result like the Metal/Vulkan tables above. Run raw (no --expectations) on crocus — Mesa's native GLES driver for Haswell (native ES, closer to an Android device than to ANGLE's ES→D3D/Vulkan translation).

area pass skip fail crash
api/validation (124§) 194,827 157,163 325 0
api/operation (67) 149,727 76,698 3,036 0
shader/execution (239) 315,602 516,424 2,900 0
shader/validation (207) 369,753 297,389 0 0
total 1,029,909 1,047,674 6,261 0

--workers 2, one process at a time. shader/validation is fully clean (the WGSL→GLSL-ES path is Tint, the same compiler as the Dawn oracle). The shader/execution residual is dominated by catalogued Tier-2 boundaries — see yawgpu specs/blocks/67-gles-backend.md for the per-cluster disposition.

§ 2 api/validation files are quarantined (excluded from the run), not failing: a zero-dimension indirect dispatch hard-wedges this Haswell GPU machine-wide. Real-GPU verification here is Linux/Mesa (the spec'd Tier-2 target is Windows ANGLE); numbers and feature set may change without SemVer guarantees.

Design and roadmap live in docs/ — start with docs/00-overview.md and docs/07-roadmap.md.

Documentation map

Document Contents
00-overview Goals, non-goals, scope, key decisions
01-architecture Component architecture; TS-CTS → C mapping
02-harness Registry, fixtures, params, query, tree, runner, listing
03-webgpu-c-abstraction Async→sync helpers, error scopes, backend shim
04-authoring-tests Test-author API (C++) and worked example
05-porting-guide How to port a .spec.ts to a .spec.cpp
06-build-and-run CMake build, backend selection, running, filtering
07-roadmap Roadmap and remaining work
UPSTREAM Pinned upstream CTS / header / backend revisions
COVERAGE Per-area / per-file port status
FINDINGS Per-backend conformance defects the suite surfaced

License

BSD-3-Clause (see LICENSE), matching the upstream CTS so that ported test logic — a derivative work — stays license-compatible. Each ported file must preserve upstream attribution; see 05-porting-guide §0.

About

WebGPU Native Conformance Test Suite

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages