Skip to content

Add format-aware FB conversion with GPU-safe parallelism - #18

Draft
weedmo wants to merge 1 commit into
mainfrom
agent/gpu-aware-fb-conversion
Draft

weedmo wants to merge 1 commit into
mainfrom
agent/gpu-aware-fb-conversion

Conversation

@weedmo

@weedmo weedmo commented Jul 16, 2026

Copy link
Copy Markdown
Owner

What changed

  • Route converter jobs by detected or requested input format (auto, mcap, or fb) and report skipped mixed-format recordings through job progress.
  • Add a dedicated mcap_to_fb_convert job that rewrites legacy MCAP recordings into the Data Foundry FB layout using staging and atomic promotion.
  • Run FB datasets through the data_foundry-conversion-worker GPU image via the Docker API while preserving cancellation and destination rollback behavior.
  • Select FB converter concurrency from explicit GPU profiles:
    • RTX 5090: parallel=3, nvenc_parallel=3
    • RTX 4090: parallel=3, nvenc_parallel=2
    • Other GPUs, including RTX 5060-class cards: 1/1
  • Keep FB_CONVERTER_PARALLEL, FB_CONVERTER_NVENC_PARALLEL, and FB_CONVERTER_GPU_MODEL as operational overrides.
  • Extend the job enum, API types, Docker runtime dependencies, documentation, and regression coverage for the new paths.

Why

The previous FB wrapper always passed 2/2, so lower-memory GPUs inherited concurrency chosen for higher-end workers and could fail with GPU OOM. The converter also lacked a first-class distinction between legacy MCAP recordings and the Data Foundry FB input layout.

This change makes format selection explicit and gives lower-tier or unknown GPUs a conservative default while retaining tuned 4090/5090 profiles and manual overrides.

Impact

  • Existing legacy MCAP conversion remains available through the same convert job.
  • FB conversion requires the configured data_foundry-conversion-worker image and Docker socket access from the converter service.
  • Deployments must rebuild/restart the converter image for the new runtime and profile selection to take effect.
  • The database job enum gains mcap_to_fb_convert through idempotent schema compatibility statements.

Validation

  • uv run pytest -q tests/test_converter_input_format.py tests/test_fb_converter.py tests/test_mcap_to_fb.py tests/test_mcap_to_fb_job.py tests/test_converter_queue_adapter.py tests/test_jobs_router.py tests/test_converter_dockerfile.py tests/test_ui_service_docker_ops.py tests/test_converter_mem_limit.py — 67 passed.
  • npm run build — TypeScript and Vite production build passed.
  • git diff --check — passed.
  • RTX 5060 Ti profile smoke check selected parallel=1, nvenc_parallel=1.

Known validation gaps

  • A real end-to-end FB conversion against the GPU worker image was not run as part of this publish step; the Docker lifecycle is covered with mocked regression tests.
  • The full backend suite reaches an existing unrelated failure in tests/test_config.py: the test expects dataset_sources == ["lerobot", "lerobot_test"], while current settings also contain "raw". The first-failure run reported 37 passed and 19 skipped before that assertion.

Route recordings by source format, add staged MCAP-to-FB conversion, and select FB worker concurrency from explicit GPU profiles with a conservative fallback.

Constraint: 5060-class workers must not inherit concurrency tuned for 5090-class hardware

Rejected: Fixed 2/2 FB parallelism | It can exhaust VRAM on lower-memory GPUs

Confidence: high

Scope-risk: moderate

Directive: Keep numeric environment overrides available until cooperative runtime throttling exists

Tested: 67 targeted backend tests; frontend production build; diff whitespace check; RTX 5060 Ti profile smoke selected 1/1

Not-tested: Full suite is blocked by existing tests/test_config.py expectation that omits the current raw dataset source
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant