Repository navigation
Conversation
(cherry picked from commit ba4509b)
(cherry picked from commit 9d2d649)
(cherry picked from commit 0f1af60)
(cherry picked from commit 271f027)
(cherry picked from commit bb9567f)
(cherry picked from commit 76ec11b)
(cherry picked from commit eaf180d)
(cherry picked from commit 067b250)
…ine eval (cherry picked from commit 3050d4e)
(cherry picked from commit 9965f82)
(cherry picked from commit 9751782)
(cherry picked from commit 44cddce)
(cherry picked from commit ce8543d)
(cherry picked from commit b86931a)
…val for E1 (M,16) tokens Aidan, 2026-09-15: fold the rl2 YAM stationery uploads into the ABC stationery set, hold the middle Step 0 tempo tertile out of training, and train two rollout-ready models -- the lab time-indexed recipe and the arcdur arc row -- scored with Aniketh's open-loop segment evaluator (open_loop_sim, 4481e3d). EVALUATOR open_loop_sim only knew the lab (M+1, 14) / (2M, 14) ARC layouts. Add token_layout (lab | e1_dur | e1_logdur | e1_profile) and truncate_e1_wide_token so an E1 wide (M, 16) prediction is truncated by waypoint rows (each row carries its own timing) and detokenized with TokenizeBimanualArcLengthE1 built exactly as robot/arc_decoder.BimanualArcDecoder builds that layout -- the scored decode is the deployed decode. The scoring loop is unchanged; lab-layout behaviour is unchanged (token_layout defaults to lab). Tests: truncation keeps waypoint/timing rows, and a truncated e1_dur decode equals the robot decoder's first 25 steps. CONFIGS (generated from the split manifest, not hand-edited) data/abc_arc/stationery_midtempo_{time,arcdur}.yaml and eval-only configs for val / test_mid / test_in, plus a *_smoke_* set on a 4-episode subset. Episode-hash frozensets; the task clause admits 'sort the stationery into containers' (ABC) and 'organize_stationary' (rl2 YAM station) -- no SQL relabel. time = Ryan's shorts_extreme_bc data block (Yam lab keymap, raw 100-frame eef_frame chunk); arcdur = shorts_extreme_arcdur (bimanual_arc, D 0.40 m, M 100, (100, 16)). experiment/abc_arc/stationery_midtempo_{time,arcdur}.yaml: the ported shorts-extreme recipe (e1/hpt_flow_wrists, 30k steps, batch 32, 1 GPU, checkpoint every 5k) with evaluator eval_open_loop_sim (10 val episodes every 5k in training, whole episodes). SCRIPTS (scripts/e1/) build_stationery_midtempo.py split v2: train = slow + fast tertiles minus val and test_in; val / test_mid identical to v1 make_stationery_midtempo_configs.py, make_smoke_manifest.py stationery_midtempo_{train,eval}.sbatch, stationery_midtempo_smoke.sh, stationery_midtempo_tests.sbatch Phoenix gpu-h200 / cpu-small launchers harvest_stationery_midtempo.py select on val open-loop paired MSE, report selected + final on test_in / test_mid package_stationery_midtempo.py rollout artefact: checkpoint (sha256), norm stats with normalizer_state, Hydra config, filled yam_rollout.yaml (base_T_model from the rl2 calibration, e1_dur decoder for arcdur), manifest replay_check_stationery_midtempo.py load the package via load_graph_policy and replay a recorded rl2 episode offline Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit ed120944eab5a485e934f104b1047143a19f2e62)
(cherry picked from commit b2f356359784fe68640d2edde31ca88f91aae7cd)
(cherry picked from commit ec32507b3eba2bb02c2680e4d32a8fc39f28d0d7)
(cherry picked from commit 902de75b813f792040239f1ed71b04d82272de45)
…C arc pretrain configs Loader workers are re-forked at every epoch boundary by default and this stack deadlocks there: job 13307279 wedged at step 99 of a 100-step epoch, workers idle in do_poll, main thread in futex_wait, after limping at ~19 s/step. persistent_workers=true crosses three boundaries cleanly at 4.2 it/s (job 13308744). Also adds the ABC-only arc pretrain (Elmo ckpt is variant=time, so it has no arc pretraining) and a per-job hydra run dir. (cherry picked from commit f189f60fc73499078c1bf3fe2ffe17203f58e0d8)
…ent base (ABC+RL2)
Fills the {time,arc} x {scratch,base} matrix. timepre_abc is the ABC-only time twin of arcpre_abc (Elmo ckpt is not a substitute: it trained on ABC + old RL2, 2264 eps). arcpre_abcrl2 mirrors the DATA POOL of Elmo checkpoint run (ABC + RL2 organize_stationary_updated, 2287 eps) with the arc-duration target, so a later RL2 fine-tune from it parallels the time arm.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit d810ec9861a089447705becafdee7c39d1c6327d)
…re all landing in a shared None/checkpoints) Copied from Elmo resolved config where output_dir had already resolved to null; dirpath then rendered as literal None/checkpoints relative to the repo root, shared by every run. Lightning versioned the collisions (-v1/-v2) so nothing was lost, but ownership was only recoverable by mtime. Each job now checkpoints under its own hydra run dir. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit a07d692d3e4d097a02c83209ce5413727bf1e194)
… land in None/checkpoints Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit 9bfee8d17c32aaf784ae6180c545a73da8af1b9e)
… cells, ABC-only towel bases; towel sync
Aidan 09-18: never mix ABC and RL2 unless told to. From-scratch = RL2 YAM data of the task only; pretraining bases = ABC only. Adds scratch_rl2_{stationery,towels}_{time,arcdur}, base_abc_towels_{time,arcdur}, towel data configs, and sync_towels (rl2 task is fold_towels after the 09-18 relabel). arcpre_abcrl2 (ABC+RL2) violated the rule and was cancelled before it ran.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 19f18b9151fb25eab9a10d37a63ce3ec543a6aba)
…one submit helper Untyped gres lets PACE expand the partition list (a job once landed on a V100, no bf16), so grid_submit.sh excludes every non-H100/H200 GPU node and the job exits 42 if the GPU is not H100/H200. norm_stats.save_cache_dir was null (inherited from a resolved config), so no run was saving the stats a later fine-tune must reuse; now saved under the run dir. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit 1dd56aecc1a8bc48e222294c5d84bbe3f24a5c8a)
… tasks; pinned-pool norm-stats recompute Aidan 09-18 asked for two from-scratch runs on BOTH abc and rl2 data with the velocity arc tokenizer (variant arcvel -> profile, (100,16) = 14 canonical + per-arm path speed). This is the explicit exception to the never-mix rule and applies to these two configs only. Also commits the pinned 2089-episode data config + sbatch used to recompute arcpre_abc norm stats exactly. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit 3d97d5ef827430ca797f0109711291c682197889)
…licitly authorised), both tasks; exclude hung H200 node Same pool filter and valid_ratio as the *_mix_arcvel configs; only the action target differs (14-D time chunk). atl1-1-03-018-14-0 hung job 13325462 in its first backward for 8.5 h (GPU 100 %, zero steps) -> added to the exclude list. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit ca5d2f1ff052bf64aa381ebdb198cabdb81b4e3c)
…h runs; arc evaluator on group_balance_sampler: WeightedRandomSampler over MultiDataset.index_map, each named group gets a fixed share of draws, uniform over frames within a group; membership from explicit episode lists (198 RL2 episodes per task). Hooked into MultiDataModuleWrapper.train_dataloader next to anchor_sampler. Four scratch_mix5050_* experiments supersede scratch_mix_* (cancelled); the arcvel pair now runs BimanualTempoEval so Valid/* is logged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit 5af693db2487361b010c77135fe223372e38ce51)
…aws; end-to-end test against SQL ground truth Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit 88e2e5a345383eb41e68fb6eb5ee96d5bff3bd57)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit 6ddf673eadd9e3103b7d1cd7c4e559922295d9f9)
This was referenced Oct 2, 2026
Collaborator
This was referenced Oct 2, 2026
added 4 commits
October 1, 2026 22:42
ElmoPA
changed the base branch from
codex/graph-consolidation-stack-20261001/13-weighted-e1
to
graphite-base/198
October 2, 2026 18:32
ElmoPA
force-pushed
the
graphite-base/198
branch
from
October 2, 2026 18:32
96df46f to
47b9429
Compare
ElmoPA
changed the base branch from
graphite-base/198
to
codex/graph-consolidation-stack-20261001/12-canonical-arc
October 2, 2026 18:32
ElmoPA
approved these changes
Oct 2, 2026
ElmoPA
approved these changes
Oct 2, 2026
ElmoPA
marked this pull request as ready for review
October 2, 2026 19:05
This was referenced Oct 7, 2026
ElmoPA
added a commit
that referenced
this pull request
Oct 7, 2026
…#199) Replace five identical model copies with thin aliases and express the two Qwen 180M variants through explicit flow-head overrides. Existing callers retain their resolved graph behavior. This extends the existing graph-contract stack ending at #159. Review the config deduplication and compatibility layer independently of the ARC and weighted-training integration above it. The branch points at the original consolidation commit `67357575`; no commit or code was rewritten to create this stack. Validation: the committed deduplication evidence records unchanged composition for all 254 contexts in this layer. Stack construction verified exact commit identity and parent ancestry. The complete assembled CPU validation and source dispositions are attached to #198; those results apply to the assembled code commit, not separately to every intermediate layer. Keep this PR in draft pending the assembled real-weight/data OSMO gate and applicable fresh CI/review. Stack order: #150–#159 → #199 → #200 → #201 → #198.
ElmoPA
changed the base branch from
codex/graph-consolidation-stack-20261001/12-canonical-arc
to
graphite-base/198
October 7, 2026 21:28
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Integrate the ARC, graph-runtime and weighted-training work behind #197, #159 and #160 into a consistent runtime. All 34 source PR heads remain ancestors. The result preserves independent per-arm ARC clocks, explicit wide/stacked/E1 layouts, YAM 100/400-frame windows, homogeneous batching and checkpoint-bound normalization. Ten duplicated model configurations become aliases or explicit variants; the 128 source model entries and 64 grid recipes have recorded dispositions.
Review this layer against #200 for the combined net changes to weighted/E1 integration and the lightweight runtime/companion-validation tree.
Keep the runtime tree lightweight: tests, fixtures, reports, design docs, notebooks, documentation images and standalone validation tooling live on
codex/graph-validation-20261001. The runtime branch retains code, configs, launch tools, AGENTS.md guidance, licenses and runtime-required package metadata/support. This separates 364 files / 16.9 MB from the runtime checkout without discarding them or rewriting history.CI loads only the companion files enumerated at commit
34c1d761409cdf74e48b11a57d52421c184748de, pinned in.github/validation-ref, and tests the runtime checkout under review. The restore helper verifies file hashes and refuses to overwrite tracked code or local edits. The isolated wheel gate excludes restored companion artifacts from its build input. ARC parity checks retain exact byte comparisons against frozen #197 source executed on the same platform, replacing macOS-specific output hashes that failed on Linux.Review entry points: consolidation and source/config dispositions, ARC contract, and companion validation workflow.
Validation on runtime commit
7ccb609626c62fafce5c1181bc79f576e0829f2bwith companion pin34c1d761409cdf74e48b11a57d52421c184748de: Linux CI passed all 2,001 tests, zero failures/skips, all 439 YAML contexts and 439 constructor contexts, static checks and the isolated installed-wheel gates. All 1,993 existing test cases remain, with eight added restore-helper regressions; all 23 ARC parity tests pass. The local suite also passed 2,001 tests. Passing CI and permanent receipt / compressed JUnit and audit artifacts record the exact revisions. This CI fix and separation do not change runtime model/tokenizer source.Merge gate remains open: real-weight/data OSMO L40/L40S validation requires separate authorization under the handoff. No GPU job, deployment or checkpoint conversion has been performed. This draft must not merge until that gate and applicable CI/review pass. Main and the source PR branches remain unchanged.
The initial wiki-tracker failure is a missing
OBSIDIAN_VAULT_REPOconfiguration (a request to/repos//contents/...returns 404). It is separate from the ARC parity failure and has not been disabled or represented as passing.