Repository navigation
Conversation
added 30 commits
October 3, 2026 14:05
added 7 commits
October 4, 2026 01:38
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This study tests whether Astra can teach autonomous OOD behavior more sample-efficiently than RL from the same π0.5 checkpoint. Astra observes real cameras, state and execution history, compares computational corrections with one fixed native proposal, executes one selected prefix, and updates policy LoRA parameters from useful real action windows. FRS action steering, physical candidate retries, privileged object poses and synthetic training targets are excluded.
The research objective remains unmet. No completed checkpoint reaches the preregistered 8/10 autonomous target, and no fresh-task or untouched-reset confirmation was performed. Measured scope is Goal OOD6 / seed173 and Spatial OOD2 / seeds173 and179. These development runs are separate from the earlier 20-task intervention campaign. Each SR uses ten declared reset states with simulator success and verified reset identities.
These endpoints have different budgets; the dashboard retains all 64 completed scheduled curve points with their actual interaction and token costs. The first Spatial teacher gain has not repeated. Its later policy3 was saved but remained unevaluated after external quota preemption.
The final paired control freezes one unassisted native success before either student outcome. Fourteen locally useful windows yield 3/10, versus 2/10 from all eighteen successful-episode windows. Both fresh students receive 40 uniform seven-channel updates, with identical initial adapter hashes and sampled flow-time sequences. Only the local-label student gains reset20; neither loses a native win. This is one extra successful reset, not confirmed superiority. Both branches retain the original 106 controls and 455,499 monitoring tokens; they add no collection or teacher calls and 5,168 evaluation controls. This is not the original online learning trajectory or a measured teacher-free collection.
The intervention audit exposes a bottleneck: the seed179 teacher executes ten locally useful assisted commands but none enters its complete-window training data. Relaxing selection is insufficient. On the same Goal data, local labels / hindsight / whole-success BC score 5/10 / 4/10 / 5/10 at 100 updates. Strict replay scores 5/10, 5/10, 4/10 at 40/100/200 updates; evidence-masked replay scores 5/10, 3/10, 0/10 despite improving training fit. Versioned revisions include RTC weighting, gripper targets, prompt/TEI/TLI candidates, bounded controller targets and per-rollout visual grounding; recorded decisions distinguish offered methods from actual execution.
The frozen V8 evaluation restores the exact saved policy4 and repeats all ten original resets without new training or teacher calls. It preserves the native binary outcomes and adds 1,951 evaluation controls. All seven earlier completed outcomes reproduce; six command sequences match exactly. One failed trajectory diverges after initially identical commands. Earlier retained evaluation costs remain charged; totals remain lower bounds wherever an unarchived tail is unknown. Repeated cases do not become a seventeen-trial denominator.
The study closes at 22.055960 of 24 authorized L40S GPU-hours, including initialization and failed workers. All 26 GPU allocations have ended, with peak concurrency two. Retained totals are at least 18,101 collection controls, 115,777 evaluation controls and 38,616,918 reported teacher tokens including cache and offline diagnostics. Reused source costs count once globally. Tokens are reported usage, not an API bill.
Reproducibility includes pinned full weights and RLinf components, strict native/converted-weight checks, immutable OSMO bundles, exact Astra prompts, saved checkpoints, per-reset outcomes, checksummed execution/judgment/admission/optimizer records, actual camera examples, rollout videos and a standalone HTML dashboard with CSV/JSON/PNG/PDF exports. The serial RL integrations have documented overrides and limited tuning; this is not a stock SOTA benchmark reproduction. DSRL's initial noise distribution differs from the native Gaussian despite sharing decoder weights.
Validation: full
pytest tests/unit— 1,794 passed, six optional skips. Changed Python passes Ruff. Chrome-rendered tables, JavaScript syntax, archive integrity, record checksums and privacy scans pass. Interrupted evaluation accounting has a regression test; private provider event streams and signed catalogs are excluded from public record export.Linear stack: based on
astra/demo-skill-library-20260930(PR #704), with no merges or rewritten history.