Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
31 commits
Select commit Hold shift + click to select a range
f155676
simx: resume RTU traversal after any-hit and intersection verdicts
blaisetine Sep 22, 2026
991a8a5
kernel: pin WGATHER to the warp-gather registers; add diverge_loop test
blaisetine Sep 22, 2026
47a98b7
rtu: resume RTL traversal after any-hit and intersection verdicts
blaisetine Sep 22, 2026
dd1188c
simx: WGATHER writes every lane of the warp
blaisetine Sep 22, 2026
43ec716
raytrace.h: value-initialize the BVH builder's child array
blaisetine Sep 22, 2026
2391b06
rtu: require a full-warp ALU for the WGATHER'd trace config
blaisetine Sep 22, 2026
8ba2724
rtu: reject only degenerate triangles, not small ones
blaisetine Sep 22, 2026
3118776
rtu: stage instance ids and the object ray for instanced candidates
blaisetine Sep 22, 2026
3da2d0c
rtu: match lavapipe's ray/primitive intersection exactly
blaisetine Oct 3, 2026
69299a2
rtu: report hit facing and pick lavapipe's hit among near-equal ones
blaisetine Oct 3, 2026
148baa2
rtu: SimX ray log and an RTL-vs-SimX ray replay harness
blaisetine Oct 3, 2026
36e11c0
rtu: RTL tri/box PEs match SimX's intersection tests bit for bit
blaisetine Oct 3, 2026
ce14729
rt_replay: flush mismatches per batch; document the rtlsim stall-time…
blaisetine Oct 3, 2026
a886131
rtu: RTL near-tie oracle picks the hit lavapipe keeps, as SimX does
blaisetine Oct 3, 2026
9c9c6c3
rtu: pipeline the near-tie window bound and the oracle for 200 MHz
blaisetine Oct 3, 2026
85c24f6
rtu: RTL short-stack overflow never loses a hit
blaisetine Oct 3, 2026
a2fa139
rtu: instances carry lavapipe's world->object matrix; no inverse anyw…
blaisetine Oct 3, 2026
92925ba
rtu: a SimX lane without a hit reports t_max and its world ray, as th…
blaisetine Oct 3, 2026
e7b6f1f
simx: ray log snapshots the whole scene image, not just the lines Sim…
blaisetine Oct 3, 2026
dacdb9b
rtu: remove the near-tie visit-order oracle
blaisetine Oct 3, 2026
7f4edff
rtu: commit the nearest hit; on equal t the first one found stays
blaisetine Oct 3, 2026
d76bee1
rtu: fp32 watertight triangle test
blaisetine Oct 3, 2026
05ac466
rtu: cull boxes against the ray interval; one-add box decode
blaisetine Oct 3, 2026
6a7d4d8
rtu: instance transform as FMA chains
blaisetine Oct 3, 2026
7c6e037
fpu: vendor FMA maps FMUL to a*b + -0, not a*b + +0
blaisetine Oct 3, 2026
98058d3
aved: clock the kernel at the rate its image was timed for
blaisetine Oct 3, 2026
642420e
runtime: stage device transfers through a bounded host buffer
blaisetine Oct 4, 2026
6d81dad
runtime: a failed queued command fails the commands behind it
blaisetine Oct 4, 2026
31a3832
aved: keep the CP staging aperture out of device memory
blaisetine Oct 4, 2026
f70cff0
tex: weight bilinear taps by f/256, rounded, per Vulkan texel filtering
blaisetine Oct 4, 2026
833a108
vxbin: read kernel entries from readelf columns, not a decimal-size p…
blaisetine Oct 4, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 9 additions & 1 deletion VX_types.toml
Original file line number Diff line number Diff line change
Expand Up @@ -429,7 +429,7 @@ VX_RT_HIT_BARY_U = 11
VX_RT_HIT_BARY_V = 12
VX_RT_HIT_PRIMITIVE_ID = 13
VX_RT_HIT_INSTANCE_ID = 14
VX_RT_HIT_GEOMETRY_INDEX = 15
VX_RT_HIT_GEOMETRY_INDEX = 15 # VX_RT_HIT_GEOMETRY_MASK bits: gl_GeometryIndexEXT; bit 31: VX_RT_HIT_BACK_FACING
VX_RT_HIT_INSTANCE_CUSTOM = 16
VX_RT_OBJECT_RAY_ORIGIN = 17 # object_ray.origin.{x,y,z} (17..19)
VX_RT_OBJECT_RAY_DIRECTION = 20 # object_ray.direction.{x,y,z} (20..22)
Expand Down Expand Up @@ -477,6 +477,14 @@ VX_RT_FLAG_SKIP_AABBS = 0x200
VX_RT_FLAG_ENABLE_CHS = 0x400
VX_RT_FLAG_ENABLE_MISS = 0x800

[rtu_hit_bits]
# The HIT_GEOMETRY_INDEX word: the leaf's geometry index in the low 28 bits (the
# Vulkan BVH layout keeps flags above them, which the RTU drops), and
# BACK_FACING set when a triangle hit lies on its back face after the instance's
# FLIP_FACING, i.e. gl_HitKindEXT is BACK_FACING.
VX_RT_HIT_GEOMETRY_MASK = 0x0fffffff
VX_RT_HIT_BACK_FACING = 0x80000000

[rtu_cb_actions]
# vx_rt_cb_ret action codes
VX_RT_CB_ACCEPT = 1
Expand Down
7 changes: 7 additions & 0 deletions ci/testcases/unittest.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -31,3 +31,10 @@ tests:
via: script
run: "make -C hw/unittest run-tcu-dsp"
touches: [hw/rtl/tcu, hw/unittest/tcu_fedp]
# Texel blend (VX_tex_lerp, both forms the sampler uses) against the C model
# the SW sampler and SimX share, and the exact Vulkan f/256 blend, over every
# input.
- id: hw-tex-lerp
via: script
run: "make -C hw/unittest/tex_lerp run"
touches: [hw/rtl/tex, hw/unittest/tex_lerp, sw/common/vx_gfx_abi.h, sw/common/gfx_frag_tex.h]
43 changes: 26 additions & 17 deletions docs/designs/ray_tracing_architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -299,9 +299,15 @@ Two work products leave the scheduler:
returned `hitAttribute`).

Robustness details worth naming: a short-stack of depth `RTU_STACK_DEPTH` bounds
per-context node stack RAM; on overflow the walker sets an `ovf` flag and, at
pop-time, **re-descends** the subtree pruned by the tightened `best_t` (bounded by
`RTU_RESTART_CAP = 8` restarts) — a full traversal on a finite stack. A 16-entry
per-context node stack RAM, and an overflow never loses a hit. The walk visits
nodes in rank-path order (each level's child rank in the t-sorted list, an
instance's index in its leaf); a child that does not fit is dropped and the walk
records it, then only descends until it would pop, and instead **restarts** from
the root along the dropped child's rank path, held in a small per-context path
RAM. Everything before that child has been visited, a tightened `best_t` only
culls a suffix of a node's sorted children (so ranks stay valid), and each
restart starts strictly further along, so the walk is exact and terminates on
any tree up to 63 levels deep. A deep tree costs restarts, not hits. A 16-entry
box collector insertion-sorts a node's child hits t-ascending so descent is
nearest-first. The insertion slot is decoded from the **admit thermometer**: the
collected list is sorted and its count mask is a prefix, so the "entries at or
Expand All @@ -321,20 +327,23 @@ subnormals flushed either way), and `VX_CFG_FMA_LATENCY` /
whatever depth results:

- **`VX_rtu_box_pe`** — pipelined ray/AABB slab test, one child box per cycle,
emitting `{hit, t_near}`. Dequantizes the node's int8 child corners
(`origin + q·2^exp`), does the slab test with `VX_fma_unit` + `VX_fncp_unit`,
and subtracts the ray origin *before* multiplying by `inv_d` so axis-aligned
rays (`inv_d = ±inf`) stay NaN-free. Also handles raw/procedural boxes.
- **`VX_rtu_tri_pe`** — pipelined Möller–Trumbore triangle test, one triangle per
cycle, emitting `{hit, t, u, v, back_facing}`; reuses `VX_fma_unit`,
`VX_fdiv_unit` (1/det), `VX_fncp_unit`, and `VX_rtu_fdot3`/`fcross3`. The
dot/cross helpers pipeline their 24×24 mantissa products into DSP multipliers
(`LATENCY_IMUL` deep) fed the **raw** mantissas: a flushed (subnormal/zero)
term is discarded downstream in the `VX_rtu_fmac3` accumulator by its zero
product-exponent, so no subnormal-flush select sits in front of the multiplier
inputs and the DSPs launch straight from the source flops.
- **`VX_rtu_xform`** — TLAS world→object transform, `obj = Rᵀ·(ro−t)` — FMA-only
(an orthonormal TLAS rotation needs no determinant or divide). Always built:
emitting `{hit, t_near}`. Mirrors SimX `box_rel` + `ray_box` bit for bit:
corners relative to the ray, `q·2^exp + (origin − ro)` (the product exact),
slabs `rel·inv_d`, culled against the ray interval `[t_min, t_max]` with
`t_max` the committed hit. `inv_d` is `FLT_MAX` for a zero direction
component, so no slab is ever `0·inf`. Also handles raw/procedural boxes
(`origin = +0`).
- **`VX_rtu_tri_pe`** — pipelined watertight triangle test (Woop, Benthin, Wald,
JCGT 2013), all F32: shear, edge functions as rounded products and a rounded
difference (so a shared edge evaluates to exactly negated weights in its two
triangles), `det = (w0 + w1) + w2`, `T` as an FMA chain, one `1/det` scaling
`t`, `u`, `v`. Accepts `t_min < t < t_max` (Vulkan's open triangle interval),
with `t_max` the committed hit, one triangle per cycle, emitting
`{hit, t, u, v, back_facing}`; bit-exact against SimX `ray_triangle`.
- **`VX_rtu_xform`** — TLAS world→object transform. The instance record holds
the world→object matrix, so no inverse is taken; each object-ray component is
three dependent FMAs (`fma(z, m2, fma(y, m1, fma(x, m0, t)))`, the direction
seeded with `x·m0`), 18 FMA units, `3·FMA` deep. Always built:
the CW-BVH walker descends `LEAF_INST` natively; only the flat walker's
(`WIDTH = 0`) instancing loop is gated by `VX_CFG_RTU_TLAS_ENABLE`.
- **`VX_rtu_recip`** — F32 reciprocal for `inv_d`, either a portable LUT+Newton
Expand Down
12 changes: 6 additions & 6 deletions docs/designs/texture_sampler_architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -256,15 +256,15 @@ stays a plain 4-byte-word interface.
colour here instead (§3.1).
2. **U lerps**: per lane and level, 8 `VX_tex_lerp` instances (4 channels ×
{low, high} texel pairs), each a 3-cycle fixed-point datapath computing
`(s + (s >> 8)) >> 8` with `s = a·(255−f) + b·f + 0x80` — the exact
divide-by-255 rounding, not a plain shift.
`(a·(256−f) + b·f + 0x80) >> 8` — the weight is `f/256`, the 8-bit
subtexel fraction Vulkan's texel filtering defines
(`subTexelPrecisionBits = 8`), rounded to nearest.
3. **V lerp**: 4 more lerps per level blend the two U results, another 3
cycles.
4. **Level lerp**: 4 final lerps blend the two levels' texels by the
request's lod fraction — a fraction of **256** that truncates, the form
the software sampler blends levels in, where a tap weight is a fraction
of 255 (§2.3). A single-level sample carries weight 0 and passes level 0
through.
request's lod fraction — also a fraction of 256, but truncated rather
than rounded, the form the software sampler blends levels in (§2.3). A
single-level sample carries weight 0 and passes level 0 through.

The whole sampler is ~10 cycles fixed latency, one request per cycle
throughput, per-channel 8-bit arithmetic — no floating-point anywhere (the
Expand Down
4 changes: 2 additions & 2 deletions hw/rtl/fpu/VX_fma_unit.sv
Original file line number Diff line number Diff line change
Expand Up @@ -71,7 +71,7 @@ module VX_fma_unit import VX_gpu_pkg::*, VX_fpu_pkg::*; #(

if (USE_VENDOR_IP) begin : g_vendor
// xil_fma / acl_fmadd compute a*b+c, so the FMA-core opcodes are remapped:
// MUL : a*b + 0
// MUL : a*b + -0 (-0 is the additive identity: a -0 product stays -0)
// ADD/SUB : a*1.0 (+/-) b
// MADD/NMADD : (+/-)a*b (+/-) c
// The vendor IP rounds round-to-nearest-even only (frm is ignored).
Expand All @@ -89,7 +89,7 @@ module VX_fma_unit import VX_gpu_pkg::*, VX_fpu_pkg::*; #(
if (is_neg) begin // MUL
a32 = dataa[31:0];
b32 = datab[31:0];
c32 = '0;
c32 = 32'h80000000;
end else begin // ADD/SUB
a32 = dataa[31:0];
b32 = 32'h3f800000; // 1.0f
Expand Down
Loading
Loading