Add sm107 tunings for DeviceHistogram::HistogramEven - #11085
Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
|
update: results came out fantastic. dropping a controversial blob and rebenchning |
88c51fe to
b99e840
Compare
🔬 CUB benchmark SASS comparisonHow to request a benchmark run
Targets with a SASS change
|
😬 CI Workflow Results🟥 Finished in 1h 52m: Pass: 96%/272 | Total: 2d 14h | Max: 57m 27s | Hits: 85%/225823See results here. AI failure analysis1. C Parallel histogram compilation: missing policy_selector::sample_type initializer · 4 jobsExplanation: The change adds `sample_type` to the histogram `policy_selector` aggregate and updates its typed constructor, but the runtime C Parallel initializer in `c/parallel/src/histogram.cu` still supplies only the previous seven fields. GCC promotes the resulting missing-field warning to an error; both Python wheels fail because they build the same C Parallel source. Evidence: Copy this prompt into a coding agentJobs: |
There was a problem hiding this comment.
I think the I32 results on the benchmark comparison are too regressive. We take a lot of 20+% regressions for <10% gains. Can you look whether you have tunings that cause less regressions with some speedup at least?
Those runs are worrying me:
| I32 | 2^16 | 32 | 0.201 | 22.86% | 🔴 SLOW |
| I32 | 2^20 | 32 | 0.201 | 23.37% | 🔴 SLOW |
| I32 | 2^24 | 32 | 0.201 | -2.69% | 🟢 FAST |
| I32 | 2^28 | 32 | 0.201 | -9.15% | 🟢 FAST |
| I32 | 2^16 | 128 | 0.201 | 23.96% | 🔴 SLOW |
| I32 | 2^20 | 128 | 0.201 | 23.50% | 🔴 SLOW |
| I32 | 2^24 | 128 | 0.201 | -0.70% | 🔵 SAME |
| I32 | 2^28 | 128 | 0.201 | -7.17% | 🟢 FAST |
| I32 | 2^16 | 2048 | 0.201 | 29.01% | 🔴 SLOW |
| I32 | 2^20 | 2048 | 0.201 | 26.39% | 🔴 SLOW |
| I32 | 2^24 | 2048 | 0.201 | -1.83% | 🟢 FAST |
| I32 | 2^28 | 2048 | 0.201 | -2.00% | 🟢 FAST |
| I32 | 2^16 | 32 | 1 | 26.24% | 🔴 SLOW |
| I32 | 2^20 | 32 | 1 | 26.98% | 🔴 SLOW |
| I32 | 2^24 | 32 | 1 | -3.23% | 🟢 FAST |
| I32 | 2^28 | 32 | 1 | -9.93% | 🟢 FAST |
| I32 | 2^16 | 128 | 1 | 23.33% | 🔴 SLOW |
| I32 | 2^20 | 128 | 1 | 24.06% | 🔴 SLOW |
| I32 | 2^24 | 128 | 1 | -3.03% | 🟢 FAST |
| I32 | 2^28 | 128 | 1 | -9.53% | 🟢 FAST |
| I32 | 2^16 | 2048 | 1 | 36.90% | 🔴 SLOW |
| I32 | 2^20 | 2048 | 1 | 40.47% | 🔴 SLOW |
| I32 | 2^24 | 2048 | 1 | -3.13% | 🟢 FAST |
| I32 | 2^28 | 2048 | 1 | 1.88% | 🔴 SLOW |
Too little gain for the regression in my opinion.
closes part of https://github.com/NVIDIA-dev/cccl_private/issues/738
performance results
Important Trade-off: F32 with 2048 bins gets ~6% slower at large sizes. No config in the search fixes this without losing the other wins, and we can't dispatch on bin count (runtime value). Kept it because the wins are much bigger: up to −76% on 2M-bin workloads and ~−20% everywhere else.