-
Notifications
You must be signed in to change notification settings - Fork 1.8k
Pull requests: antirez/ds4
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Fix spec prefix capture on compressor-less Flash layers
#648
opened Aug 1, 2026 by
gilbert-barajas
Loading…
cuda: implement real per-(layer,expert) LRU for --ssd-streaming-cache-experts
#647
opened Aug 1, 2026 by
nexus-cw
Loading…
fix(quantizer): support preserved MXFP4 in DSpark plans
#645
opened Aug 1, 2026 by
apetersson
Loading…
ssd-streaming: opt-in measure mode — per-token page invalidation + true disk-read accounting (Darwin)
#637
opened Aug 1, 2026 by
FabioMalpezzi
Loading…
Laguna: mixed-precision quant support (APEX IQ4_XS/Q6_K, official BF16) + metadata-driven rope
#633
opened Jul 31, 2026 by
jasontitus
Loading…
fix: keep distributed layer-slice HC on active tier
#631
opened Jul 30, 2026 by
edoardolegnaro
Loading…
Add model-first provider boundary and split kernel implementations
#628
opened Jul 29, 2026 by
devteapot
Loading…
ROCm: restore DeepSeek V4 Flash decode performance and Q4 SSD streaming
#623
opened Jul 28, 2026 by
kyuz0
Loading…
cuda: copy model into VRAM on single-GPU via --gpu-resident
#622
opened Jul 28, 2026 by
pvaccarello
Loading…
6 of 7 tasks
Support AProjQ4 GGUFs: Q4_K dense attention projections
#621
opened Jul 28, 2026 by
GiorgioOppo
Loading…
allow overriding FORCE_HF_DOWNLOAD via environment variable
#619
opened Jul 28, 2026 by
alexlipa91
Loading…
Fix Laguna KV ring commits for oversized prefills
#614
opened Jul 26, 2026 by
christophsturm
Loading…
Server: chat/completions live tool continuation via tool_call_id
#611
opened Jul 26, 2026 by
starforge-labs
Loading…
CUDA: resident expert cache and faster selected-expert uploads for --ssd-streaming decode
#605
opened Jul 25, 2026 by
iCreil
Loading…
Server: fused single-GPU decode for --batched-session (co-decode 2-4 sessions in one graph pass)
#604
opened Jul 25, 2026 by
iCreil
Loading…
rocm: AMD Instinct (CDNA/gfx90a) support with multi-GPU device placement
#602
opened Jul 24, 2026 by
ewindisch
Loading…
rocm: fix build on gfx1201 (RDNA4, R9700) — gate the WMMA v1 kernel
#599
opened Jul 24, 2026 by
cm999club
Loading…
fix(cuda): avoid null output-head dereference in distributed coordinators
#592
opened Jul 22, 2026 by
hy-sde
Loading…
Previous Next
ProTip!
Exclude everything labeled
bug with -label:bug.