Skip to content

afu: route the command processor through the bank adapter on multi-port shells - #422

Merged
tinebp merged 1 commit into
masterfrom
fix/afu-multibank-cp
Oct 3, 2026
Merged

tinebp merged 1 commit into
masterfrom
fix/afu-multibank-cp

Conversation

@tinebp

@tinebp tinebp commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • On the XRT shell the command processor's device master was merged with Vortex at AXI port 0, downstream of VX_mem_to_axi, so every CP transfer left on m_axi_mem_0 whichever bank owned the line. This went unnoticed because xrtsim backed all ports with one RAM and the HBM platforms build the merged single port; with one memory per port (a per-bank Questa testbench, or a per-channel HBM build) only bank 0 receives the kernel image.
  • With several ports the CP now enters VX_mem_to_axi through VX_membus_from_axi as one more input port, so bank selection applies to the CP and Vortex alike (the structure the OPAE shell already uses). The single-port shell keeps its original logic verbatim.
  • New VX_afu_axi_limit holds port 0 to one read and one write in flight, the policy the 2:1 arbiter used to impose, so cycle counts are unchanged.
  • xrtsim gives each AXI port its own memory, so a request on the wrong port misses instead of finding the data in a shared image (the old RTL now hangs at launch under xrtsim).
  • Design docs updated for both structures.

Validation

  • xrtsim CI, RV32: regression 5/5, config2 11/11, stress 2/2, scope 1/1, debug 2/2. RV64: regression 5/5, config2 11/11.
  • Cycles (vecadd, 4 cores, 4 warps, 8 threads): 1 bank 116,164 → 116,164; 2 banks 64,252 → 64,018; 4/8/16 banks within 0.03%.
  • vecadd + basic pass at 4, 8 and 16 banks; 32 banks untested.
  • ci/fpga_gate.py -b top (AFU top, 2 banks): PASS, Fmax 259.1 MHz (baseline 252.6), LUT +0.5%, FF +1.2%, LUTRAM/BRAM/DSP unchanged. The single-port branch is the original logic and was not re-synthesized.
  • Not validated: opencl/vulkan/graphics/hip on xrt, hardware.

Notes

  • With several ports, CP transfers go out line by line since a burst cannot span interleaved banks: a 256 KB upload plus download is about 1.9× slower in simulation. Unchanged with a single port, which is what the U55C/U50 platforms build.
  • Per-port addresses stay global; bank-local addressing under INTERLEAVE=1 is a prerequisite for a per-channel HBM build and is not part of this change.

🤖 Generated with Claude Code

…rt shells

The XRT shell merged the CP's device master with Vortex at AXI port 0,
downstream of VX_mem_to_axi, so every CP transfer left on m_axi_mem_0
whichever bank owned the line. It went unnoticed because xrtsim backed
all ports with one RAM and the HBM platforms build the merged single port.

With several ports the CP now enters VX_mem_to_axi through
VX_membus_from_axi as one more input port, so bank selection applies to
the CP and Vortex alike; the single-port shell keeps its original
structure. VX_afu_axi_limit holds port 0 to one read and one write in
flight, the policy the 2:1 arbiter used to impose, which keeps cycle
counts unchanged.

xrtsim gives each AXI port its own memory, so a request on the wrong port
misses instead of finding the data in a shared image.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@tinebp
tinebp merged commit f63adcf into master Oct 3, 2026
252 of 260 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant