Skip to content

Evaluate complex manipulation baselines and paired Astra guidance - #706

Draft
rpuns wants to merge 12 commits into
astra/reasoning-policy-learning-20261003from
astra/complex-manipulation-20261006
Draft

rpuns wants to merge 12 commits into
astra/reasoning-policy-learning-20261003from
astra/complex-manipulation-20261006

Conversation

@rpuns

@rpuns rpuns commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Released checkpoints for mobile parallel grippers and bimanual dexterous hands require different controls and observations from LIBERO. This adds pinned releases, native evaluation runners, a matched Astra language-guidance screen and self-contained reports with videos and traceable outcomes.

Eight native episodes are complete: LoadPreparedFood 0/3, PackIdenticalLunches 0/3, Jigsaw Puzzle Assembly 0/1 and Fridge Wine Interhand Pour 0/1. RoboCasa uses pretrain scenes and seeds 0/1/2; each dexterous task uses one recorded reset anchor. Jigsaw reached two of four stages at some point and retained one at termination; fridge pouring reached two of five and retained one. These are development measurements, not held-out OOD results.

All six paired RoboCasa guidance episodes are now complete. Astra phase prompts scored 0/3 on each task, versus native 0/3 on each task. GPT-6 Astra was configured at medium effort through the authenticated Codex relay. It saw three live cameras, 16D robot state, previous live feedback and four frames from the paired native failure. Up to 16 reviews appended short phase instructions while the frozen policy generated all motor commands. The six episodes used 81 reviews and 1,592,255 tokens (777,728 cached input, 803,452 uncached input and 11,075 output). CPU teacher preflights consumed another 38,215 tokens, including the rejected probe. No TEI/TLI, vision interpolation, flow reversal or policy update was evaluated in this screen.

Guidance executed 16,200 controls in 3,240 prefixes, of which 3,180 were assisted. Two extra resets stopped before any controls; both remain in the ledger. Protocol amendments preserve exact simulator state and proprioception while recording bounded image rounding and equivalent OBJ-format declarations. The final report includes all 14 native/guided videos, paired outcomes, exact prompts, every accepted intervention, token/cache breakdowns, timings and losslessly compressed evidence. Its post-hoc tokenizer audit found 21 clipping warnings; all 81 reconstructed review inputs retained the instruction and state, and the one clipped review lost only a trailing space. Full numerical parity with the upstream RoboCasa policy factory remains unmeasured.

The Codex parser change permits only the identified pre-turn feature-compatibility notices; real tool use and other error events remain rejected. No private reasoning events or signed artifact URLs are included in the portable report.

Validation: 1,815 unit tests passed, 6 skipped; Ruff, compile and source whitespace checks passed. Original XML diff evidence retains its two blank filename headers. Chrome at 1440 px and 390 px verified the report and all paired videos. The portable manifest, original archived evidence, local links and ZIP CRC were checked.

All GPU allocations in this study ledger are closed. Charged usage including earlier studies is 24.462870 of 48 authorized L40S-hours; the three guidance allocations used 1.176271 hours including initialization and rejected resets. The larger guidance/learning comparison at 1/2/4/8/16 collection episodes remains pending.

Latest report: astra_reversal/reports/complex_manipulation_guidance/index.html. Registered method: astra_reversal/complex_manipulation/language_protocol.json. Complete experiment status: astra_reversal/complex_manipulation/protocol.json.

Stacked directly on #705 (astra/reasoning-policy-learning-20261003), preserving the linear PR chain.

@rpuns rpuns changed the title Add native pi0.5 pilots for complex manipulation benchmarks Validate native RoboCasa pilots and prepare dexterous baselines Oct 7, 2026
@rpuns rpuns changed the title Validate native RoboCasa pilots and prepare dexterous baselines Evaluate native RoboCasa baselines and add dexterous pilots Oct 7, 2026
@rpuns rpuns changed the title Evaluate native RoboCasa baselines and add dexterous pilots Evaluate complex manipulation baselines and add matched Astra guidance Oct 7, 2026
@rpuns rpuns changed the title Evaluate complex manipulation baselines and add matched Astra guidance Evaluate complex manipulation baselines and paired Astra guidance Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant