Skip to content

perf(verl): reduce peak memory when decoding routed experts - #636

Open
Bozhen Peng (kiteretsu903) wants to merge 1 commit into
microsoft:mainfrom
kiteretsu903:perf-stream-routed-experts
Open

Bozhen Peng (kiteretsu903) wants to merge 1 commit into
microsoft:mainfrom
kiteretsu903:perf-stream-routed-experts

Conversation

@kiteretsu903

Copy link
Copy Markdown
Contributor

I noticed that R3 batch construction keeps every decoded routing array in memory while copying them into the final tensor. For MoE rollouts with long sequences, those temporary arrays can take up a lot of host memory.

This decodes one rollout at a time and releases its array after copying it into the batch. Padding, truncation and the output format stay the same.

In a local CPU benchmark with 64 synthetic rollouts, peak temporary decoding allocations fell from 126.7 MB to 4.7 MB, and process peak RSS fell from 660.3 MB to 543.4 MB. The final 128 MB tensor was the same.

Validation

I added focused tests for releasing each decoded array before loading the next one, padding and truncation with uint8 and int64 inputs, and rejecting routing data that's too short. The retention test fails on the original implementation; the compatibility tests pass on both.

Using Python 3.12 with the locked CPU dependencies:

  • python -m pytest -q tests/verl/test_rollout_adapter.py — 21 passed.
  • python -m pytest -q --durations=10 tests — 142 passed.
  • ruff check ., ruff format --check ., python scripts/check_headers.py, and pre-commit run --all-files --show-diff-on-failure — passed.
  • pyright --venvpath /home/admin1/os-contributions/agent-lightning-research — no errors or warnings.
  • uv build --no-sources --out-dir /mnt/d/Documents/os-contributions/verification/agent-lightning-r3/dist — passed; verified the updated module in both archives.

Copilot AI balanced review requested due to automatic review settings October 4, 2026 03:03

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants