perf: speed up prune pack listing and used-blob IO - #565
Open
BradKollmyer wants to merge 7 commits into
Open
BradKollmyer wants to merge 7 commits into
BradKollmyer wants to merge 7 commits into
Conversation
OpenDAL only sends maxFileCount when ListOptions.limit is set. Without it, B2 defaults to 100 names per page, so prune/check pack listing becomes thousands of round-trips. Request 10000 for scheme b2.
256 KiB MAX_HOLESIZE split prune/restore pack reads into extra HTTP range GETs on high-latency stores. 4 MiB is cheaper than another RTT on a ~100 Mbps link.
Unused gaps in a pack were fetched one after another. Issue those range GETs in parallel; the packer already serializes writes.
stream_list used CPU-count Rayon workers and a zero-capacity channel, so reading index/snapshots was one GET at a time on high-latency backends. Use 16-32 workers with prefetch so GETs overlap decrypt/parse.
TreeStreamerOnce was hardcoded to 4 workers, which left B2 tree-pack GETs idle. Use 2x CPUs (8-32) with a larger out channel, and a HashSet for visited tree ids.
The used-id set was a BTreeMap, so inserts got slower as prune walked more blobs. A HashMap keeps that path O(1).
This was referenced Sep 7, 2026
A single Vec of blob ids doubles on growth. On a large repo that request is hundreds of MiB while the old buffer is still live, and musl aborts (openzwave: memory allocation of 807665664 bytes failed at reading index 906/1551). Store ids and full entries in 1 Mi-entry chunks and binary-search each sorted chunk.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #564.
First PR in the prune warm-cache stack. Speeds up pack listing and the used-blob walk’s IO shape before the CPU-side decode work.
HashMap.Vecgrowth cannot request a multi-hundred-MiB allocation (openzwave backup OOM while reading the index).Unique commits vs
main: main...BradKollmyer:perf/prune-io