Repository navigation
Conversation
added 4 commits
October 6, 2026 14:09
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Released checkpoints for mobile parallel grippers and bimanual dexterous hands require different controls and observations from LIBERO. This adds pinned releases, native evaluation runners, a matched Astra language-guidance screen and self-contained reports with videos and traceable outcomes.
Eight native episodes are complete: LoadPreparedFood 0/3, PackIdenticalLunches 0/3, Jigsaw Puzzle Assembly 0/1 and Fridge Wine Interhand Pour 0/1. RoboCasa uses pretrain scenes and seeds 0/1/2; each dexterous task uses one recorded reset anchor. Jigsaw reached two of four stages at some point and retained one at termination; fridge pouring reached two of five and retained one. These are development measurements, not held-out OOD results.
All six paired RoboCasa guidance episodes are now complete. Astra phase prompts scored 0/3 on each task, versus native 0/3 on each task. GPT-6 Astra was configured at medium effort through the authenticated Codex relay. It saw three live cameras, 16D robot state, previous live feedback and four frames from the paired native failure. Up to 16 reviews appended short phase instructions while the frozen policy generated all motor commands. The six episodes used 81 reviews and 1,592,255 tokens (777,728 cached input, 803,452 uncached input and 11,075 output). CPU teacher preflights consumed another 38,215 tokens, including the rejected probe. No TEI/TLI, vision interpolation, flow reversal or policy update was evaluated in this screen.
Guidance executed 16,200 controls in 3,240 prefixes, of which 3,180 were assisted. Two extra resets stopped before any controls; both remain in the ledger. Protocol amendments preserve exact simulator state and proprioception while recording bounded image rounding and equivalent OBJ-format declarations. The final report includes all 14 native/guided videos, paired outcomes, exact prompts, every accepted intervention, token/cache breakdowns, timings and losslessly compressed evidence. Its post-hoc tokenizer audit found 21 clipping warnings; all 81 reconstructed review inputs retained the instruction and state, and the one clipped review lost only a trailing space. Full numerical parity with the upstream RoboCasa policy factory remains unmeasured.
The Codex parser change permits only the identified pre-turn feature-compatibility notices; real tool use and other error events remain rejected. No private reasoning events or signed artifact URLs are included in the portable report.
Validation: 1,815 unit tests passed, 6 skipped; Ruff, compile and source whitespace checks passed. Original XML diff evidence retains its two blank filename headers. Chrome at 1440 px and 390 px verified the report and all paired videos. The portable manifest, original archived evidence, local links and ZIP CRC were checked.
All GPU allocations in this study ledger are closed. Charged usage including earlier studies is 24.462870 of 48 authorized L40S-hours; the three guidance allocations used 1.176271 hours including initialization and rejected resets. The larger guidance/learning comparison at 1/2/4/8/16 collection episodes remains pending.
Latest report: astra_reversal/reports/complex_manipulation_guidance/index.html. Registered method: astra_reversal/complex_manipulation/language_protocol.json. Complete experiment status: astra_reversal/complex_manipulation/protocol.json.
Stacked directly on #705 (astra/reasoning-policy-learning-20261003), preserving the linear PR chain.