Tinyvortex is a tinygrad fork with a native VORTEX backend to easily compile and run Tensor programs on the open-source RISC-V based Vortex GPGPU architecture. This platform serves as an experimental framework for running, profiling and analyzing MLSys workloads on Vortex, and guides future GPU microarchitecture design-space exploration.
tinygrad programs still have the standard Tensor API, while generated kernels are rendered as Vortex C++ intrinsics, compiled to .vxbin, and launched through the Vortex C runtime with Python ctypes.
The default target is simx (cycle-approx simulator) to keep simulation times reasonable during development and testing. However, rtlsim (verilator based functional RTL simulation) and WIP XRT/xrtsim (AMD Xilinx FPGA) drivers are also supported for accurate hardware benchmarking.
-
Setup a
/builddirectory in the Vortex repo as per project guidelines -
Clone tinyvortex and enter the directory:
git clone https://github.com/NikhilRout/tinyvortex.git
cd tinyvortex
- Create and activate a Python virtual environment, and install tinyvortex:
python3 -m venv venv
source venv/bin/activate
python3 -m pip install -e .- Point tinyvortex to Vortex runtime:
export VORTEX_HOME=$HOME/vortex
export VORTEX_BUILD=$VORTEX_HOME/build
export LD_LIBRARY_PATH=$VORTEX_BUILD/sw/runtime:$LD_LIBRARY_PATH- Run an example smoke test to verify setup:
DEV=VORTEX python3 - <<'PY'
from tinygrad import Tensor
print((Tensor([1, 2, 3], device="VORTEX") * 2).tolist())
PYTinyvortex supports all Vortex runtime drivers; choose one through DEV:
DEV=VORTEX # default: simx
DEV=SIMX+VORTEX # cycle-approx simulator
DEV=RTLSIM+VORTEX # verilator-based RTL simulator
DEV=XRTSIM+VORTEX # Xilinx FPGA, WIPVortex hardware configs are passed through VORTEX_CONFIGS and forwarded to the Vortex kernel build. Example:
DEV=VORTEX VORTEX_CONFIGS='-DVX_CFG_NUM_CLUSTERS=1 -DVX_CFG_NUM_CORES=1 -DVX_CFG_NUM_WARPS=4 -DVX_CFG_NUM_THREADS=32' VORTEX_BUILD_RUNTIME=1 python3 examples/vortex/vecadd.py- Use
VORTEX_BUILD_RUNTIME=1when changing runtime/simulator-affecting configs - Use
VORTEX_REBUILD=1to clean the tinyvortex kernel build directory before compiling. VORTEX_CONFIGSis included in the tinygrad compile cache key and forwarded to the Vortex kernel build- When running
DEV=RTLSIM+VORTEXwith Verilator builds that treat RTL warnings as fatal, add-Wno-fataltoVORTEX_CONFIGS - Add
BROWSER=1 VIZ=1to launch tinygrad's UOp graph/rewrite visualizer after the run - Use
JITBEAM=4with@TinyJitto spend more capture time searching faster kernels for repeated runs
Try out lightweight MNIST inference on Vortex!
DEV=VORTEX VORTEX_CONFIGS='-DVX_CFG_NUM_CLUSTERS=1 -DVX_CFG_NUM_CORES=1 -DVX_CFG_NUM_WARPS=4 -DVX_CFG_NUM_THREADS=32' python3 examples/vortex/tinymnist.py --infer-samples=4Other useful knobs:
VORTEX_KERNEL_LIB=vortex2
VORTEX_MAKE_ARGS='-j8'
VORTEX_TIMEOUT=600000
VORTEX_DEBUG=1 # forwarded to Vortex make as DEBUG
VORTEX_PERF=1 # mirrors Vortex blackbox --perf=1
VORTEX_SCOPE=1
VORTEX_SAIF=1- Add -DEXT_TCU_ENABLE to
VORTEX_CONFIGSto include the Vortex Tensor Core Unit extension in your build - Vortex TCUs are mixed-precision; low-precision input formats are selected by argpassing
--itypeand the higher-precision accumulator/output format is selected through--otype - Matrix dimensions can be specified with
--m,--n, and--base-k(fp32 dim = 2x fp16 = 4x fp8 = 8x fp4) - Example:
VORTEX_CONFIGS='-DVX_CFG_EXT_TCU_ENABLE -DVX_CFG_NUM_THREADS=32' VORTEX_BUILD_RUNTIME=1 python3 examples/vortex/sgemm_tcu.py --itype=bf16 --otype=fp32 --m=64 --n=64 --base-k=64- Use
TC=0to force scalar/SIMT lowering if you want to compare against non-TCU codegen
Note: 4-bit datatype handling, structured sparsity, microscaling formats, WGMMA, DXA, and async barriers are WIP to be ported soon.
tinygrad/device.py: registersVORTEXas a device.tinygrad/runtime/ops_vortex.py: backend implementation.VortexConfigresolves source-tree vs SDK mode, driver, config strings, paths, and cache keys.VortexRendererrenders UOps to Vortex C++ and owns VORTEX-specific tensor-core descriptors/lowering.VortexCompilerinvokesextra/vortex/Makefile.kernel.VortexRuntime,VortexAllocator, andVortexProgrambindlibvortex.so, manage buffers, upload kernels/args, launch withvx_start_g, and wait withvx_ready_wait.
extra/vortex/Makefile.kernel: kernel-only Vortex build helper.tinygrad/codegen/gpudims.py: keeps local-dimension masks correct for Vortex global stores.examples/vortex/: example tinyvortex tensor programs
- Make sure your shell is not running with
DEBUG=release; tinygrad parsesDEBUGas an integer. Useunset DEBUGorDEBUG=0before running Python LD_LIBRARY_PATHmust include the Vortex runtime directory before Python starts because Vortex's runtime stub dlopens driver libraries by name- Config changes intentionally invalidate relevant tinyvortex compiler cache entries
trace/ramulator.log.*is ignored by git; may be useful for debugging memory-system behavior
- Porting Vortex roofline perf plot
- Vortex MUFU intrinsics and tinygrad Uop mapping for accelerating activations/epilogue
- Trying out JITBEAM kernel search
- Run tinyvortex on AMD Xilinx U55C FPGA