feat: bucketized futexes - #2469
Conversation
There was a problem hiding this comment.
Benchmark Results
Details
| Benchmark | Current: 2f2342e | Previous: 2e23902 | Performance Ratio |
|---|---|---|---|
| startup_benchmark Build Time | 93.09 s |
80.34 s |
1.16 ❗ |
| startup_benchmark File Size | 0.79 MB |
0.80 MB |
0.99 ❗ |
| Startup Time - 1 core | 0.76 s (±0.02 s) |
0.75 s (±0.02 s) |
1.02 |
| Startup Time - 2 cores | 0.77 s (±0.02 s) |
0.74 s (±0.02 s) |
1.04 ❗ |
| Startup Time - 4 cores | 0.77 s (±0.01 s) |
0.74 s (±0.02 s) |
1.03 ❗ |
| multithreaded_benchmark Build Time | 91.48 s |
82.11 s |
1.11 ❗ |
| multithreaded_benchmark File Size | 0.87 MB |
0.86 MB |
1.02 ❗ |
| Multithreaded Pi Efficiency - 2 Threads | 68.50 % (±6.82 %) |
85.89 % (±6.61 %) |
0.80 ❗ |
| Multithreaded Pi Efficiency - 4 Threads | 40.80 % (±3.75 %) |
43.43 % (±2.56 %) |
0.94 |
| Multithreaded Pi Efficiency - 8 Threads | 20.19 % (±1.85 %) |
25.76 % (±1.53 %) |
0.78 ❗ |
| micro_benchmarks Build Time | 208.68 s |
80.40 s |
2.60 ❗ |
| micro_benchmarks File Size | 0.88 MB |
0.86 MB |
1.02 ❗ |
| Scheduling time - 1 thread | 179.46 ticks (±46.44 ticks) |
62.65 ticks (±4.06 ticks) |
2.86 ❗ |
| Scheduling time - 2 threads | 104.94 ticks (±21.84 ticks) |
34.08 ticks (±4.10 ticks) |
3.08 ❗ |
| Micro - Time for syscall (getpid) | 9.09 ticks (±4.62 ticks) |
3.45 ticks (±0.58 ticks) |
2.64 ❗ |
| Memcpy speed - (built_in) block size 4096 | 56666.56 MByte/s (±40115.71 MByte/s) |
82448.38 MByte/s (±56997.13 MByte/s) |
0.69 |
| Memcpy speed - (built_in) block size 1048576 | 12830.79 MByte/s (±10294.12 MByte/s) |
30585.98 MByte/s (±24707.84 MByte/s) |
0.42 |
| Memcpy speed - (built_in) block size 16777216 | 11468.67 MByte/s (±9455.77 MByte/s) |
26340.06 MByte/s (±21720.96 MByte/s) |
0.44 |
| Memset speed - (built_in) block size 4096 | 56065.06 MByte/s (±40023.08 MByte/s) |
82292.76 MByte/s (±56891.50 MByte/s) |
0.68 |
| Memset speed - (built_in) block size 1048576 | 13112.51 MByte/s (±10452.86 MByte/s) |
31323.85 MByte/s (±25145.86 MByte/s) |
0.42 |
| Memset speed - (built_in) block size 16777216 | 11717.84 MByte/s (±9591.20 MByte/s) |
27104.68 MByte/s (±22209.94 MByte/s) |
0.43 |
| Memcpy speed - (rust) block size 4096 | 55167.71 MByte/s (±39358.35 MByte/s) |
74097.96 MByte/s (±51811.44 MByte/s) |
0.74 |
| Memcpy speed - (rust) block size 1048576 | 14483.14 MByte/s (±12499.07 MByte/s) |
30361.60 MByte/s (±24602.37 MByte/s) |
0.48 |
| Memcpy speed - (rust) block size 16777216 | 12089.45 MByte/s (±9941.70 MByte/s) |
27625.34 MByte/s (±22806.88 MByte/s) |
0.44 |
| Memset speed - (rust) block size 4096 | 55437.24 MByte/s (±39519.26 MByte/s) |
74373.47 MByte/s (±51976.48 MByte/s) |
0.75 |
| Memset speed - (rust) block size 1048576 | 14801.98 MByte/s (±12662.56 MByte/s) |
31110.89 MByte/s (±25033.24 MByte/s) |
0.48 |
| Memset speed - (rust) block size 16777216 | 12324.65 MByte/s (±10049.09 MByte/s) |
28386.93 MByte/s (±23265.03 MByte/s) |
0.43 |
| alloc_benchmarks Build Time | 214.50 s |
74.76 s |
2.87 ❗ |
| alloc_benchmarks File Size | 0.87 MB |
0.87 MB |
0.99 ❗ |
| Allocations - Allocation success | 91.38 % |
91.31 % |
1.00 ❗ |
| Allocations - Deallocation success | 100.00 % |
100.00 % |
1 |
| Allocations - Pre-fail Allocations | 61.60 % |
61.44 % |
1.00 ❗ |
| Allocations - Average Allocation time | 23068.52 Ticks (±1406.62 Ticks) |
5860.58 Ticks (±98.43 Ticks) |
3.94 ❗ |
| Allocations - Average Allocation time (no fail) | 23600.89 Ticks (±1821.12 Ticks) |
6554.81 Ticks (±92.86 Ticks) |
3.60 ❗ |
| Allocations - Average Deallocation time | 5237.83 Ticks (±1501.20 Ticks) |
1805.01 Ticks (±250.35 Ticks) |
2.90 ❗ |
| mutex_benchmark Build Time | 213.60 s |
79.82 s |
2.68 ❗ |
| mutex_benchmark File Size | 0.88 MB |
0.86 MB |
1.02 ❗ |
| Mutex Stress Test Average Time per Iteration - 1 Threads | 35.18 ns (±8.20 ns) |
12.10 ns (±0.41 ns) |
2.91 ❗ |
| Mutex Stress Test Average Time per Iteration - 2 Threads | 30.72 ns (±7.38 ns) |
40.26 ns (±1.68 ns) |
0.76 ❗ |
This comment was automatically generated by workflow using github-action-benchmark.
17f45b2 to
2f2342e
Compare
|
I can indeed make it generic, it's just that AFAIK there is no other usecase in the code for now for such a map. So this may be a case of "premature abstraction" :D But if you insist, I'll gladly do it :) |
| } | ||
|
|
||
| fn hash_key(v: usize) -> usize { | ||
| let v = (v >> 3).to_be_bytes(); |
There was a problem hiding this comment.
If you're trying to remove the zero bits resulting from AtomicU32's alignment, then you have to shift by 2, not 3. Since you're hashing anyway, I don't think this is necessary anyway.
| type Bucket = InterruptSpinMutex<TaskListBucket>; | ||
|
|
||
| #[repr(transparent)] | ||
| struct TaskListBucket(LinkedList<BucketElem>); |
There was a problem hiding this comment.
Using a linked-list here is a bit unfortunate – if there are a lot of tasks, the hashbrown::HashTable as an inner map to avoid having to recompute the hash while keeping hashmap-like lookup performance. Or alternatively, use a BTreeMap – that will also reduce memory usage as addresses become unused.
While tracking #2468, I initially suspected the futex lock to be an issue, so I applied the "Todo" and made a bucket list instead of the single lock.
I have based this off on #2468 so that we get performance results that make sense.