Conversation
Traversal was single-yield per lane: after the first candidate (a non-opaque triangle or a procedural AABB) was decided, the lane was resolved, so a rejected candidate hid everything behind it. Vulkan requires every candidate along the ray to be offered. A verdict that does not end the ray now resumes its walk above the decided candidate. Candidates are offered in ascending (t, key) order, key = (instance_id << 32) | record offset; the resumed walk is seeded with the committed hit and skips everything at or below the floor. An intersection shader's reported t commits only if it is nearer than the committed hit and inside the ray's interval. Also report a procedural leaf's gl_PrimitiveID as prim_base + index, as the RTL and the triangle path already do. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
WGATHER writes every non-source lane of its destination, active or not, so a unit reading the packed operand across the warp sees all of it. In an allocatable register that clobbers inactive lanes' live values. The gather is now pinned to x31 (x30 for a second live operand), which the compiler reserves in any function naming them, and the RTU trace and DXA issue read the pinned registers directly. diverge_loop covers a divergent continue out of a loop body whose other path carries its own loop, both feeding one accumulator. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Mirror the SimX multi-candidate model in the RTL scheduler. A verdict that
does not end the ray re-walks that lane from the root, seeded with the
committed hit as its t_max and the decided candidate as a floor, so every
candidate along the ray is offered in ascending (t, key) order,
key = {instance id, record offset}.
- context word: the staged candidate's key and the resume floor (t, key);
the key compares are precomputed at ALIGN off the store row, so EXEC only
adds the t compares to the staging condition
- barrier walker: a RES job also reads the committed-t row and drops an
intersection shader's accept that is not nearer than the committed hit,
then relaunches the non-terminating deciding lanes (re-walk template keeps
the world-ray reciprocals and skips the setup) and re-arms finalise for
them only (live mask)
- core: T_RWAIT takes the next candidate batch as well as the terminal
record; the scheduler's yield is gated off the cycle a resume job retires,
when the flags it clears still read old
- flat TLAS walk scans every instance instead of stopping at the first one
that staged a candidate
New test rt_smoke_ahs_multi: six stacked triangles (a t tie, an opaque one
behind) with per-lane accept targets; checks the offer order and the final
hit per lane. Passes on simx and rtlsim.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The gather's lane loop started at the warp's first active lane, so a warp whose low lanes were masked never wrote those lanes' gathered values. A consumer reading fixed lanes then saw stale data: the RTU takes a trace's scene, payload and flags|cull from lanes 1-3 of the config register, so a trace issued with lane 0 inactive ran on whatever an earlier trace had left there (in LumiBench, shadow rays traced against a stale scene with cull mask 0 and missed every occluder). Visit every lane, falling back to the warp's last active lane for a masked source, as the RTL does. New test rt_smoke_partial_mask: an all-lane trace of an empty scene, then a trace from the last lane alone; fails without the fix, passes on simx and rtlsim. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
GCC 13 flags ch[] as maybe-uninitialized under -Werror, which broke the host build of seven tests/raytracing apps. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A trace's scene/payload/flags ride lanes 1-3 of a WGATHER'd register. The ALU writes those lanes under a partial warp only when it runs the whole warp in one packet; a narrower ALU skips all-masked packets and picks its fallback source lane per packet, so the RTU would trace a stale config. Fail the build for such a config instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The ray-triangle test dropped any triangle with |det| < 1e-6. det scales with the triangle's area, so small (or finely tessellated) geometry vanished: in LumiBench's Spring scene the character's eyes and mouth were missing and their shadows with them. Reject only |det| < FLT_MIN (edge-on or zero-area, where 1/det overflows) in both SimX and the RTL tri PE. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two RTL-only defects in the candidate record the RTU hands a warp, both invisible on SimX: 1. A procedural (IS) candidate never staged instance_id or the instance custom index: the commit engine skipped those rows for CK_YLDP, so the intersection shader -- and, after an accept, the committed hit -- read whatever an earlier record left there. Every sphere of LumiBench WKND is a procedural instance whose shaders index per-sphere data by instance id, so WKND rendered only sky (RTLsim) or sky and ground (V80). Both candidate kinds now stage the full t/u/v/prim/inst/geom/custom record. 2. The object-ray slots (gl_ObjectRayOrigin/DirectionEXT) of any candidate were filled from the world-ray staging, so an any-hit or intersection shader inside a translated instance tested the wrong ray. A candidate from inside a BLAS now carries the walker's transformed ray through the commit queue into six window-store rows, flagged per context (obj_vld). The core picks those rows per lane. A candidate outside any BLAS keeps using the staging, where object space is world space, so it costs no extra cycles, and its commit-queue walk still ends at the SBT row. New rt_smoke_proc_inst: two lanes of one warp hit two instances (non-zero ids) of one procedural BLAS; it checks the ids the IS sees, the committed ids and t. It fails on the old RTL (garbage ids, IS rejects) and passes on SimX and the fixed RTL. RT suite: SimX 37/37, RTLsim 35/35. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
LumiBench images on SimX now have to be bit-identical to lavapipe, so the RTU intersection tests follow lavapipe's arithmetic rather than a generic formulation: - Triangle: watertight test in lavapipe's op order -- fp32 shear, fp64 edge functions and t, det = w0+(w1+w2), t = f32(T/det); accept tmin < t < tmax (strict on both sides, as lvp_build_triangle_case). SimX rtu_isect.cpp and a new pipelined RTL VX_rtu_tri_pe.sv; hw/unittest/rtu_tri_pe checks the RTL against SimX bit-for-bit (1M random cases, 0 mismatches). - Box (SimX): cull against [0, tmax], not [tmin, tmax] -- tmin belongs to the primitive test only. A huge flat triangle whose slab exit lies below tmin was culled although its rounded t passes tmin (CRNVL_PT). Zero direction components use FLT_MAX as reciprocal, as lavapipe. The RTL box PE and the strict tmin compare in the RTL tri PE still need the same change (tracked for the RTLsim step). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- gl_HitKindEXT: the back-facing flag rides in bit 31 of the hit geometry index (VX_RT_HIT_BACK_FACING); the geometry index is masked to 28 bits (VX_RT_HIT_GEOMETRY_MASK). SimX walker + RTL scheduler. - Hit choice: lavapipe commits only t < tmax and walks its binary BVH nearest-child-first, so among equal-t hits the first visited wins, and its fp32 box cull can hide a hit a few ulps nearer. Coincident/near-coincident geometry (SHIP sails, PARK body, BATH, PARTY, SPNZA) depends on this. The driver now appends lavapipe's tree as visit-order tables (TLAS in the scene; compact 32 B/node BLAS tables in a separate buffer). For an opaque hit within t*2^-19 of the committed one, the SimX walker climbs both leaves to their LCA, replays lavapipe's box test/order and its cull, and keeps the hit lavapipe would. Scenes without tables keep the static (instance, geometry, prim) key, which the RTL scheduler still uses. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
VX_RTU_RAYLOG=<path> makes the SimX RTU append every traced ray (scene root, flags, cull mask, origin/dir/tmin/tmax bits) with its terminal result (status, t/u/v bits, primitive, geometry index incl. facing bit, instance id/custom), plus every scene line a completed walk read, at its device address. Lines are re-emitted under a new epoch if the device rewrites them. Rays whose result was decided by an any-hit or intersection shader are flagged non-replayable. VX_RTU_RAYLOG_MAX / VX_RTU_RAYLOG_EVERY bound the log. tests/raytracing/rt_replay restores the logged lines at their original addresses (vx_buffer_reserve before any other allocation), re-traces the rays one per thread in warp-uniform (scene, flags, cull) groups, and compares each result bit-exactly with the log, classifying mismatches (hit/miss, near tie, same-primitive precision, terminate-on-first-hit, ...). Ray selection by index range, list file (its own mismatch output is accepted) or stride. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Tri PE: accept tmin < t < tmax (strict on both sides, as SimX ray_triangle). The FP units now run with EXCEPT_ENABLE=1: with it off the soft FMA/FDIV treat NaN/inf operands as finite, so a NaN vertex (an inactive triangle) reported a hit. - Box PE: rebuilt to mirror SimX reconstruct_child_aabb + ray_aabb_intersect: mn = origin + q*2^e (exact product, one add; the old q*2^e + (origin - ro) FMA rounds differently in ~3% of slabs), then (mn - ro) * inv_d; lo/hi reduced with fmin/fmax NaN semantics seeded with -/+inf; accept hi >= max(0, lo) && lo <= t_max (culled against [0, t_max], not t_min); t_near = max(t_min, lo), +0 for a zero so the collector's unsigned order holds. Same latency; IEEE specials on. - Recip: a zero (or flushed subnormal) direction component returns +FLT_MAX, as SimX; the divider handles inf/NaN. - hw/unittest/rtu_box_pe (recip + box PE vs SimX: inv_d, accept, t_near and per-node child order) and directed cases in rtu_tri_pe (t on tmin/tmax, origin on the triangle, NaN/inf/-0/subnormal operands). Both report 0 mismatches; the only differences left are subnormal flushes (the PEs are FTZ/DAZ by design), each confirmed equal to SimX run under FTZ/DAZ. The scheduler must now pass the ray's t_max (not the committed t) to the tri PE, as SimX does, or equal-t twins never reach the tie-break. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…out define A model that hangs or aborts mid-run now keeps the mismatches it already reported. On rtlsim, one hard ray can hold every warp in vx_rt_wait past the scheduler's all-warps-stalled watchdog, so the replay must be built with -DVX_DBG_STALL_TIMEOUT=2000000000. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Ports the SimX visit-order oracle (6804d1b0c) to the RTL RTU, so that among equal / near-equal-t opaque hits the RTL commits exactly the hit SimX (and lavapipe) commits. - Scheduler: an opaque hit within t*2^-19 of the hit committed in this walk, with visit-order tables for both and a different triangle, hands both hits to the oracle and parks the context (CS_ORC_REQ / CS_ORC_WAIT); otherwise the static (instance, geometry, prim) key still settles exact ties, so scenes without tables are unchanged. The context word carries the committed hit's near bound, TLAS rank, parent/side, BLAS table and vertices, and the walk's TLAS table / BLAS table / rank / parent/side from the leaf headers. near_t(t) is computed off the tri PE result into its result RAM and the window compares at ALIGN, so EXEC gains no logic depth and the common path no cycles. - Culling as SimX: node children are culled against near_t(best_t) once an opaque hit is committed (procedural AABBs still against best_t), and the tri PE tests the ray's own [tmin, tmax], so a hit at or past the committed t reaches the tie-break. - VX_rtu_oracle: climbs both leaves to their lowest common ancestor in the BLAS table (same instance) or the TLAS table with the world ray (across instances), testing every box on the way in lavapipe's F32 slab test against the other hit's t, orders the two LCA children by entry distance (child 0 on ties), and on a lost-by-cull verdict climbs the nearer hit's BLAS path with lavapipe's object ray rebuilt from the table's world->object matrix in its op order. One VX_fma_unit (IEEE, subnormals), three serial correctly-rounded reciprocals (1/0 -> FLT_MAX), a 64-word LUTRAM register file, and a fetch unit that overlaps a climb step's parent fetch with its box test; table lines come over the scheduler's memory port under the context's tag (oracle first; a context's fetch retries). A climb cap keeps a malformed (cyclic) table from hanging a context. - VX_rtu_near_t: t + |t|*2^-19 with both F32 roundings, exact for all 2^32 inputs; VX_rtu_f32_round: the shared integer-to-F32 rounder. - rt_smoke_tie: coincident and near-coincident triangles within a BLAS and across instances, with lavapipe-style visit-order tables (TLAS in the scene, BLAS tables in a separate buffer), traced with and without the tables; every hit field is compared bit-exactly against SimX's results (golden.h, regenerated with -g), and the tables must change some verdicts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The V80 build of a886131 missed 5 ns by 5.6 ns: near_t(t) was formed combinationally from the tri PE result into the result RAM (53 logic levels), and the oracle had -0.5..-2 ns paths through its register file reads. - VX_rtu_near_t / VX_rtu_f32_round: pipelined (6 / 4 stages; leading one, alignment, round increment and result each registered). Still exact for all 2^32 inputs. - Scheduler: the tri result event (result RAM write and context wake) is delayed through the near_t pipeline, so the bound lands with the result and no PE output reaches a RAM through more than a few levels. The tri PE's SimX cost model gains the same 6 cycles (rtu_isect.cpp). - Oracle: F32 unit operands registered off the LUTRAM reads; the box test's min/max fold split over registered reads; reciprocal operands registered before decode and the rounding pipelined; fetched scalar words captured into a register before they feed address arithmetic. - Oracle: the ray it set up last (world or object ray and reciprocals) is kept, keyed by world ray, TLAS table and instance, so a run of verdicts on one ray skips the setup; it is forgotten whenever the RTU is idle, as it always is between dependent launches. - rt_smoke_tie: thinner coincident stack (12 surfaces instead of 20), so its densest warp trace stays well inside the core's stall watchdog on the one-oracle RTL; the tables now change 117 of 256 verdicts (was 51). Out-of-context Vivado (VX_rtu_scheduler, NUM_CTX=16, V80 part, 5 ns): post-place WNS +0.148 ns, the worst paths all inside the tri PE. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The RTL dropped the children that did not fit the short stack and, once the stack emptied, re-descended the BLAS or the TLAS pruned by best_t, at most 8 times. A re-descent repeats the same nearest-first walk, so it overflows and drops the same children again unless a hit happened to cull them: past the cap the walk ended with those subtrees unvisited. On PARK_PT (deep BLASes, 16-entry stack) a ray from a surface missed its hit at t=0.005 and reported one at t=0.85 (replay ray 657); a 64-entry stack passed, which pinned it. The walk visits nodes in rank-path order: each level's rank in the t-sorted child list, an instance's index in its leaf. A child that does not fit is dropped and recorded (level, rank); the walk then only descends (the stack stays full) and, when it would pop, restarts from the root following the dropped child's rank path, read back from a per-context path RAM that the walk writes on every descent and pop. Every node before that child has been visited; a tightened best_t only culls a suffix of a node's sorted children, so recorded ranks stay valid; each restart starts strictly further along, so the walk is exact and terminates. Stack entries carry their level and rank; the restart cap and the BLAS-only re-descent are gone. rt_smoke_deep_stack gains a decoy scene (boxes walked first whose triangles miss; the one hit sits in dropped subtrees): the old RTL returns a miss, the new one the hit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…here The instance record held the object->world transform and each model inverted it on its own: SimX by cofactor inverse, the RTL as R^T (exact only for an orthonormal R). Neither reproduces the object ray the Vulkan reference traces, which applies its own world->object matrix as rounded products added in the order translation + x + y + z (lvp_mul_vec3_mat; llvmpipe lowers ffma, so nothing is fused). The record now holds that world->object matrix (the driver copies lavapipe's wto, the host builder inverts in double), and SimX (world_to_object_ray, also used by the visit-order oracle) and VX_rtu_xform transform the ray in that order, so the object ray matches lavapipe bit for bit for any affine instance, scale and shear included. The RTL keeps its 4*FMA latency: 18 products, then three dependent adds, every stage registered; IEEE specials handled. The unused VX_rtu_fdot3/fcross3/fmac3 go. hw/unittest/rtu_xform: 1M random affine (rotation, non-uniform scale, shear) + 38K directed special cases vs SimX, 0 mismatches. Direct record writers in tests/raytracing store the inverse translation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e RTL The RTL returns the ray's own t_max as the hit t of a lane with no committed hit (what a ray query reports as the committed t with none: lavapipe initialises it to t_max) and the world ray as its object ray; SimX returned zeros. Both now report the same. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…X read A model that culls differently descends into nodes the SimX walk never fetched; replayed against only the lines SimX read, those come back as zeros and show up as false mismatches. The first trace against a scene now also snapshots its image from device RAM: the header's scene_bytes after the root and the instance table packed below it (all-zero lines skipped). Still opt-in via VX_RTU_RAYLOG; the processor hands the logger its RAM at attach time. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The oracle replayed the Vulkan reference's (lavapipe's) own BVH -- appended by the driver as "visit-order tables" -- to decide which of two near-equal-t opaque hits lavapipe would keep, and culled child boxes against a widened bound near_t(best_t) = best_t + |best_t|*2^-19 so every such hit stayed reachable. Which hit a traversal reports among (near-)coincident ones is implementation-defined (Vulkan: "If t < t_max ... the candidate is set as the current closest hit"); mimicking one implementation's tree order is not a hardware feature, it is parity emulation of the CPU reference. - RTL: VX_rtu_oracle, VX_rtu_near_t and VX_rtu_f32_round are gone, with the scheduler's CS_ORC_REQ/CS_ORC_WAIT states, the oracle's memory-port and line-buffer sharing, the table pointers / ranks / vertices in the context word, the near-window precompute at ALIGN, and the 6-cycle delay of the tri result event that carried near_t. Node children are culled at best_t. - SimX: the LCA climb (src_keeps_a and friends), near_t / box_cull_t and every visit-order-table read are gone from rtu_walker.cpp; the TriPe cost model drops the near_t stages. - Leaf headers: LeafTri/LeafInst flags and LeafInst geometry_index/prim_base no longer carry table words; they are reserved and ignored. - tests/raytracing/rt_smoke_tie only exercised the tables; removed. - Lint: mark the box PE's early tag offset bits unused (pre-existing warning). The exact-t static (instance, geometry, primitive) key is still in place; it goes in the next change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Drop the static (instance, geometry, primitive) key that settled exact-t ties between opaque hits. It existed so two traversal orders would agree on which of two coincident triangles is reported, which the Vulkan spec leaves to the implementation: "If t < t_max, t_max is set to t and the candidate is set as the current closest hit. If t > t_max, the candidate is dropped." (Ray Closest Hit Determination). What real RT units do -- and what this change does -- is shrink the ray interval to the committed hit. - The tri PE (RTL) and ray_triangle (SimX) now test against [t_min, best_t) instead of the ray's own t_max, so a hit is only ever reported strictly nearer than the committed one and the first hit found at a given t stays. A context has at most one triangle in flight and nothing else writes its best_t meanwhile, so the RTL's interval is the one SimX's sequential walk uses. - RTL: best_kv/best_ki/best_kg/best_kp, tri_tie and the separate t < best_t compare go; the PE verdict is the commit condition. - SimX: commit_takes goes; the flat walker also tests against best_t (an opaque hit already occludes every farther candidate, so nothing that could be staged is lost). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace the F64 edge-function/t/divide datapath with the standard F32 watertight test (Woop, Benthin, Wald, "Watertight Ray/Triangle Intersection", JCGT 2013). The F64 path existed only so t would round like the Vulkan reference's (lavapipe's) op order; an F32 test is what RT hardware builds. kz = argmax|dir|, kx/ky follow (swapped when dir[kz] < 0) sz = 1/dir[kz], sx = dir[kx]*sz, sy = dir[ky]*sz r = v - o, px = fma(-sx, rz, rx), py = fma(-sy, rz, ry), pz = sz*rz w0 = px2*py1 - py2*px1 (w1, w2 likewise), det = (w0 + w1) + w2 T = fma(w2, pz2, fma(w1, pz1, w0*pz0)), rcp = 1/det t = T*rcp, u = w1*rcp, v = w2*rcp, back_facing = det < 0 The edge functions stay two rounded products and a rounded difference of the per-vertex sheared coordinates, so an edge shared by two triangles yields exactly negated weights in both and no ray leaks between them (the sheared coordinates are per vertex, so fusing the shear is safe). Range check: the Vulkan spec defines the triangle interval as open at both ends -- "For any primitive that has within its bounds a position x_r = 0, y_r = 0, and (for triangles) t_min < -z_r/||d|| < t_max, or (otherwise) t_min <= -z_r/||d|| <= t_max, an intersection candidate exists" (Ray Intersection Candidate Determination) -- so a hit needs t_min < t < t_max, with t_max the committed hit (previous change). - VX_rtu_tri_pe: 9 subs, 2+6+3 shear FMA/MULs, 6 MULs + 3 SUBs for the edges, 2 adds + 3 FMA/MULs for det and T, 2 F32 dividers (1/dir[kz], 1/det), 3 MULs for t/u/v -- no F64 unit and no F64 divider. Latency 2 + 2*FDIV + 7*FMA (was 3 + FDIV + 3*FMA + 5*FMA64 + FDIV64), every stage registered. - RTU_LATENCY_FMA64 / RTU_FDIV64_LAT / kRtuLatencyFma64 / kRtuFdiv64Lat go; TriPe::pipe_depth follows the new pipeline. - hw/unittest/rtu_tri_pe: 1M random + 109K directed cases, 0 mismatches vs SimX (46 subnormal-flush cases, the PEs being FTZ/DAZ); new shared-edge family: rays through the shared diagonal of 20K quads, none passes between the two triangles. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The box test culled against [0, t_max] instead of [t_min, t_max]. That was added only so a box whose slab exit lies just below t_min would still be entered when the Vulkan reference's (lavapipe's) rounding put a triangle's t inside the interval. The standard slab test culls against the ray's own interval, with t_max shrunk to the committed hit: lo = max(t_min, min(t0, t1)[x,y,z]), hi = min(t_max, max(t0, t1)[x,y,z]) hit = lo <= hi, t_near = lo Kept: a zero (or subnormal) direction component has reciprocal FLT_MAX, so no slab is ever 0 * inf = NaN; fmin/fmax drop a NaN slab instead of poisoning the fold. Box decode: the corner relative to the ray origin is now formed as q*2^exp + (origin - ro) -- one subtract per axis shared by both corners, then one add per corner (q*2^exp is exact) -- instead of (origin + q*2^exp) - ro. Same 3-FMA-deep pipeline, 15 F32 units instead of 18. A raw (procedural) box uses origin = +0. - VX_rtu_box_pe: t_min/t_max fold into the min/max reduction tree, so the verdict stage is a single compare. - SimX: ray_recip / quant_corner / box_rel / ray_box replace reconstruct_child_aabb + ray_aabb_intersect and mirror the PE op for op (quant_corner flushes a subnormal corner as the PE does; the walker sets up the reciprocals once per ray). - hw/unittest/rtu_box_pe: 1.1M boxes (103.5K directed), 0 mismatches in inv_d / verdict / t_near and 0 in child order; the 1904 subnormal-flush cases all come from the PEs' FTZ (e.g. 1/FLT_MAX). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The instance record keeps carrying the world->object matrix (as real RT hardware does; it also handles scaled and sheared instances with no inverse anywhere). What goes is the emulation of the Vulkan reference's (lavapipe's) op order -- every product rounded, then t + x + y + z, nothing fused -- which cost 18 multipliers and 15 adders over 4 FMA stages only to round like lavapipe. Each object-ray component is now three dependent FMAs: obj_ro[i] = fma(ro.z, m[i][2], fma(ro.y, m[i][1], fma(ro.x, m[i][0], m[i][3]))) obj_rd[i] = fma(rd.z, m[i][2], fma(rd.y, m[i][1], rd.x * m[i][0])) - VX_rtu_xform: 18 FMA units, 3*FMA latency (was 33 units, 4*FMA), every stage registered; the direction's first step adds -0 so the product keeps its own zero sign. - SimX world_to_object_ray mirrors it with std::fma; kRtuXformLatency 36 -> 27. - hw/unittest/rtu_xform: 1M random affine (rotation, non-uniform scale, shear) + 38K directed special cases, 0 mismatches vs SimX (1 subnormal flush). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
xil_fma / acl_fmadd only compute a*b + c, so VX_fma_unit maps a multiply to a*b + c with c = +0. Under round-to-nearest -0 + +0 = +0, so every product that is -0 (a negative operand times a zero, or a negative product flushed to zero) came out +0 -- an IEEE 754 sign-of-zero error in fmul.s on the Xilinx/Altera FP IP path, and in every RTU PE multiply on the V80. -0 is the additive identity (x + -0 == x for every x, signed zeros included), so the remap adds -0 instead. The soft core (rtlsim, ASIC) was already correct. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A VRT design write does not touch the user clock, so a freshly programmed image ran at whatever rate the previous image had set: a 200 MHz Vortex image came up at 60 MHz left behind by a FireSim image (silently 3.3x slow), and the reverse case overclocks a design past its timing closure. On open, set the clock to the vbin's timed rate (getMaxFrequency) as XRT does from the xclbin on load, read it back, and refuse to run if it is above that rate. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cp_submit_mem_write/read allocated one CP-visible host staging buffer the
size of the whole transfer. On the V80 a large upload (a 1.9 GB
acceleration structure) asks VRT for a buffer its allocator rejects
("Invalid argument"), so the scene never reached device memory and every
ray missed. Move the data through a 64 MB staging buffer chunk by chunk,
as pageable copies do on other GPU runtimes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An error from an asynchronous command only reached that command's own event, so a vx_enqueue_write without an event lost it, vx_queue_finish returned success, and the kernel launched behind it ran on device memory that was never written. As in-order queues do elsewhere (OpenCL's error-for-wait-list, CUDA's error at the next synchronize), the first failure is kept: later commands complete with it without executing, and finish() returns it, then clears it so the queue can be used again. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
HOST_TAG=HBM1 put the CP's m_axi_host on the HBM_AXI port whose 512 MB slice starts at +512 MB from the HBM base -- inside the 4 GB device-memory window that MEM_TAG=MEM exposes from that same base. VRT allocated the CP's DMA staging buffers there while Vortex's own allocator handed out the same addresses, so any workload whose device buffers grew past 512 MB was overwritten by its own uploads: LumiBench CAR/ROBOT rendered wrong pixels and PARK's corrupted BVH never finished traversing (hit the 30 min timeout). With the window kept clear (runtime experiment) all three render bit-identical to SimX and PARK_SH completes in 46 s. Use HBM8, which starts at +4 GB, past the last device byte, and refuse HBM0..HBM7 for HOST_TAG at parse time when MEM_TAG is MEM or HBM0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The bilinear tap fraction is the low 8 bits of the scaled texel coordinate (x0s & 0xff), i.e. frac*256. Vulkan's texel filtering with subTexelPrecisionBits = 8 defines the tap weight as f/256, but the 8-bit blend (Lerp8888 in the shared sampler and VX_tex_lerp's FRAC_SCALE=255 path in the RTL) normalised it as f/255. A full-weight tap was over-weighted and the bilinear result drifted up to 2.4 LSB from the Vulkan formula (>1 LSB in ~2.9% of random cases). The float filter and the trilinear level blend already used /256, so the unit was also internally inconsistent. Both implementations now compute (a*(256-f) + b*f + 128) >> 8: exact weights, round to nearest, within 1/2 LSB per lerp. The packed lane peaks below 2^16, so Lerp8888 stays carry-free. VX_tex_lerp keeps one shift-only datapath (no /255 correction stage, same 3-cycle latency); the level blend instance keeps its existing truncation (ROUND=0), matching TexLodLerp. The OM unorm blends, which are legitimately /255, are untouched. New hw/unittest/tex_lerp checks VX_tex_lerp exhaustively (all 2^24 inputs, random stalls) against Lerp8888 / TexLodLerp and the exact Vulkan blend, wired into hw/unittest and a hw-tex-lerp CI lane. The box/evilskull gfx_draw3d goldens (box_ref_128, evilskull_ref_32, evilskull_ref_128) move by 1-2 LSB on 2-3% of pixels; FF (SimX, RTL) and the SW sampler agree bit-for-bit on the new images. Goldens are left for human regeneration. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…attern readelf -s -W prints a symbol Size of 100000 or more in hex (0x18794). The symbol regexes required a decimal size, so a kernel_main of 100 KB or more matched nothing: the .vxbin got no VXSYMTAB footer, Module::load_bytes fell back to "main" at min_vma, which is __vx_cta_entry itself, and every warp re-entered its own startup code forever -- a silent hang on SimX. Split each readelf row into its columns instead and take Value and Name, for both the kernel entries and the _edata/_end lookups. Verified: a 100,244-byte probe kernel that hung now runs to completion on SimX; 41 ordinary kernels (regression tests and vortexpipe shaders) produce byte-identical .vxbin files; demo passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tinebp
added a commit
that referenced
this pull request
Oct 4, 2026
rtu_raylog.cpp (from #424) guarded its process-wide logger with a std::mutex, which the SimX threading-boundary rule forbids in sim/simx; ci/check_simx_mt_boundary.sh failed every host cell (dtm, cupbop, gem5, sst) on 4db71f3. The log is a debug trace shared by every RTU core, and SimX debug tracing already requires a serial build. Follow that policy: drop the lock, and when VX_RTU_RAYLOG is set in a parallel build (SIMX_MT > 1) print a notice and leave the log disabled. Serial behaviour is unchanged (rt_smoke with the log on writes the same stream; with SIMX_MT=2 it passes, log off). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tinebp
added a commit
that referenced
this pull request
Oct 4, 2026
#424 changed the bilinear filter weights and VX_types.toml, which left the draw3d golden images and every perf_gate baseline stale. - gfx_draw3d box_ref_128, evilskull_ref_32 and evilskull_ref_128 are re-rendered on SimX. Each one is reproduced bit-exactly by SimX at 4 cores with 2 raster cores, with caches disabled, with software TEX and with an all-software pipeline, and by rtlsim. - The perf baselines are regenerated with --update-baselines. core, tensor and graphics change only in config_hash (cycles unchanged, except om at -0.1%); dxa moves by -1.8% to +0.07% and rt_raycast by -0.3% to +1.3%. These cycle counts match the counts CI measured on 4db71f3 exactly, so the movement comes from #424, not from the dcache fix that follows it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.