Add six measured single-card B70 results (vLLM, OpenVINO GenAI, Cascadia, llama.cpp) - #30
Merged
Merged
Conversation
Single-card Arc Pro B70 cells across three models (Qwen3.8-27B dense GDN, Qwen3.6-27B dense, Qwen3.6-35B-A3B MoE) and four runtimes. Every decode number is a median of n=5 post-warmup samples at C1 (the llama.cpp entry is llama-bench engine-native tg, mean of 5 raw repetitions), with workload coordinates, cache state, speculation policy, sampling class, power, and checkpoint identity in each file's notes. All runs target one card via the component field since the rig is registered as 2x B70.
|
@SergiioB is attempting to deploy a commit to the Community Labs Team on Vercel. A member of the Team first needs to authorize it. |
- results/SergiioB/2026-08-30-qwen3-6-27b-int4-vllm.json: imported (result 42) - results/SergiioB/2026-08-30-qwen3-6-35b-a3b-int4-vllm.json: imported (result 43) - results/SergiioB/2026-09-11-qwen3-8-27b-int4-cascadia.json: imported (result 44) - results/SergiioB/2026-09-14-qwen3-8-27b-int4-vllm.json: imported (result 45) - results/SergiioB/2026-09-14-qwen3-8-27b-int8-openvino-genai.json: imported (result 46) - results/SergiioB/2026-09-14-qwen3-8-27b-q4-k-m-llamacpp.json: imported (result 47) Merge ingestion completed. Rerunning this workflow will not duplicate these results. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this is
Six benchmark results measured on my Arc Pro B70 rig (rig 6, "B70 Workstation"), all targeting one card via the component field since the rig is registered as 2× B70. Coverage: three models — Qwen3.8-27B (dense, GDN hybrid attention), Qwen3.6-27B (dense), Qwen3.6-35B-A3B (MoE, 3B active) — across four runtimes: vLLM XPU, OpenVINO GenAI, Cascadia, and llama.cpp SYCL.
Measurement discipline
Notes on catalog mapping
int4boards, matching how the boards are defined for these models.gptq-4bitboard on the model) and Muse-Glimmer-30B / Ornith-1.5-35B-A3B (need model entries). Happy to send a catalog PR for these if you want them.All numbers were measured on my own hardware and are submitted as self-reported evidence-backed results, per the result-file schema and the ingestion check.