Skip to content

About

Auditable CompactionRL reproduction with synthetic long-horizon experiments and a live THUDM/slime E2E

Resources

Stars

3 stars

Watchers

0 watching

Forks

Repository files navigation

CompactionRL reproduction

An auditable, framework-neutral reproduction of the publicly specified algorithm in CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents (Li et al., 2026), with a thin integration path for THUDM/slime.

This repository reproduces the algorithmic core. It does not claim to reproduce the paper's 30B/106B benchmark scores: the exact prompts, resume template, training subset/order, reward implementation, PPO settings, and several infrastructure details are not public.

For the full experiment narrative—including the 96-step long-horizon result, the value-function design, all Phase 1–17 outcomes, failure analyses, and cost decisions—read Reproducing CompactRL: What Worked, What Failed, and Why We Did Not Scale. A Chinese version is also available.

What is implemented

  • Fixed-budget compaction trigger: C - |history| < T_comp.
  • Atomic assistant-action + environment-observation history steps.
  • Same-policy summary prompt and system + resume(summary) + recent tail reconstruction, with the paper's default k=2 tail behavior.
  • Shared terminal task reward for execution and summary segments.
  • Token-normalized clipped PPO objective.
  • Local token-level GAE plus the paper's cross-trajectory correction: A_hat[s,i] = (gamma * lambda_s) ** N_after[s] * A_local[s,i].
  • VAPO-style length-adaptive lambda = 1 - 1 / (1.5 * response_length).
  • Separate actor/critic summary masks for a clean direct-summary-loss ablation.
  • Paper-reported 50-step critic warm-up and 2 critic : 1 actor update support through a minimal patch against the pinned public slime commit.
  • Optional tokenwise-lambda and stitched-bootstrap value-target ablations, plus compaction-boundary value-gap and explained-variance diagnostics.
  • Exact zero-critic full-trajectory reference used to test reward-distance correction.
  • A dependency-free toy actor-critic experiment comparing joint summary training with a frozen-summary ablation.
  • A 96-step memory-chain experiment with three context resets, multi-seed learning curves, a frozen-summary control, and a readable successful trace.
  • A deterministic 360-task synthetic memory environment with exact rewards, balanced train/dev/test splits, and pinned regeneration hashes.
  • Optional common-random-number SGLang sampling with stable per-task and per-action/summary seeds, auditable seed metadata, and exact saved-trajectory comparison gates for paired training arms.
  • Optional actor-only occurrence credit, bounded post-coverage stopping pressure, and marker-gated causal-prefix redistribution. These preserve task GAE/critic targets and include exact sampled-token mass audits.
  • Versioned reported-paper, eight-GPU debug, and causal-ablation settings.
  • An out-of-tree integration for pinned public slime: native compact list[Sample] rollout generation, exact token provenance, verifier rewards, chronological segment recovery, a custom advantage hook, and an 8-GPU PPO launcher. Source/API compatibility, the official slime container runtime, a four-A10 one-update actor/critic/rollout job, and a bounded five-update Full-vs-summary-loss-masked pilot are validated live.
  • A publication-blocking live E2E audit: immutable run identity, independent advantage/return recomputation, actor/critic gradient artifacts, checkpoint inventory, lightweight evidence download, and a final report that cannot be generated before every real-run gate passes.

Quick start

No third-party package is required for the core tests:

PYTHONPATH=src python3 -m unittest discover -s tests -v
PYTHONPATH=src python3 examples/toy_joint_training.py
PYTHONPATH=src python3 experiments/long_horizon_memory_chain.py

Or run both:

bash scripts/run_checks.sh

Generate or validate the simulation data:

PYTHONPATH=src python3 scripts/generate_sim_data.py
PYTHONPATH=src python3 scripts/summarize_sim_data.py
PYTHONPATH=src:. python3 scripts/build_plan_grammar_sft.py \
  --source data/synthetic/train.jsonl \
  --sft-output data/synthetic/plan_grammar_sft.jsonl \
  --rl-output data/synthetic/ordered_plan_train.jsonl

Validate against a pinned slime checkout:

python3 scripts/check_slime_compat.py --slime-root /path/to/slime

After an authorized Modal GPU run, download only its lightweight evidence and build the gated final report:

bash scripts/download_modal_evidence.sh

The report builder intentionally fails while the live e2e_audit.json is missing or unsuccessful.

The committed final E2E report passed both the 96-step causal-learning gate and the live pinned-slime integration gate. The newer Modal scale-pilot report records nonzero post-compaction rewards and the single-seed arm comparison. Its task-type audit is also the current LLM blocker: both arms scored 0/23 on reward-critical early_fact memory, so the pilot is not presented as proof of learned LLM summary retention. Its historical masked arm also omitted the summary from critic training; the value-function fix below supersedes that ablation contract. The value-function fix report records the corrected mask/critic wiring and a passing bounded 4×A10 runtime audit. The subsequent three-arm scale comparison completed 60 updates per arm with 50 critic-only warmup updates. It verifies the corrected value path at scale, but its saved-checkpoint capability result is superseded: actor and critic inherited one save directory, and the later critic save replaced the actor shards. The deterministic checkpoint evaluation documents the diagnosis, role-specific checkpoint fix, passing fixed-run audit, and a paired dev-60 result. After two actor updates the fixed model scored 12/60 versus 14/60 for the base model, so there is not yet evidence of held-out LLM capability improvement. The follow-up seed replication and checkpoint-resume report records a complete 60-rollout continuation and a second deterministic dev-60 comparison. Seed 17 scored 20/60 while seed 23 scored 12/60 versus the same 14/60 base, so the positive seed-17 capability result did not replicate. The implementation is operational, but stable LLM improvement remains an open gate. The controlled critic learning-rate ablation shows that reducing the resumed critic LR from 3e-6 to 1e-6 halves mean value loss and materially reduces actor/critic gradient excursions. Its dev-60 score changes only from 12/60 to 13/60 (p=1.0) while invalid commands and truncation regress, so the stability gain is not presented as capability gain. The follow-on actor reference-KL ablation verifies slime's actor-only k3 KL path on the same continuation. KL 0.01 partially recovers invalid-command and truncation rates and ties the base model at 14/60, but early-fact memory falls to 1/20. It remains an engineering stability control, not a report-reported setting or a capability result. The KL dose-response follow-up records a positive KL 0.05 candidate: 21/60 overall and 12/20 early-fact versus 13/60 and 2/20 for the matched no-KL continuation (+10/-2, p=0.0386). It is the first positive long-horizon memory signal in the scaled LLM experiment, but it remains provisional until common-random-number and second-seed replication because stochastic training trajectories differ. The subsequent deterministic sampling validation uses two independent 4×A10 runs to verify exact equality of all eight saved temperature-1 token trajectories and all 183 derived request seeds. This enables common-random-number KL replication while retaining an explicit boundary: GPU optimizer updates are not bitwise deterministic. The controlled KL replication then rejects the provisional KL 0.05 advantage: under common-random-number training, no KL scores 26/60 while KL 0.05 scores 19/60 (+3/-10 from the KL arm, p=0.0923). The no-KL checkpoint itself scores 26/60 versus 14/60 for the frozen base (+15/-3, p=0.00754), including 10/20 versus 2/20 early-fact tasks. This is a positive single-run candidate, not yet a replicated capability claim; an identical independent rerun is the next gate. That identical no-KL rerun scores 25/60, including 16/20 early-fact, versus the original 26/60 and frozen base 14/60. The original and repeat each beat base in paired tests (p=0.00754 and p=0.0192), so the aggregate gain is robust to one independent same-seed GPU rerun. Their early/latest composition and value diagnostics still vary materially; cross-model-seed replication was the remaining gate at that stage. The follow-up deterministic training validation adds the pinned slime Megatron reproducibility recipe and runs two independent 2-update 4×A10 jobs. Both updates' 16 token trajectories, actor gradients, four critic updates, and value diagnostics match exactly. This validates controlled optimizer experiments through two updates; it is infrastructure evidence, not a model-quality result. The full deterministic second-seed KL replication extends that control to eight updates from the seed-17 checkpoint. Update 60 matches all sampled tokens, value/advantage diagnostics, and both critic updates exactly before the KL objective separates the arms. No KL scores 29/60 versus frozen-base 14/60 (+17/-2, p=0.000729), giving positive no-KL results on both seed 17 and seed 23. KL 0.05 scores 32/60 on seed 17 but its direct advantage over no KL is not significant and reverses on seed 23, so KL 0 remains the default. Every arm remains 0/20 on ordered-plan; the next gate is a 4096-token no-compaction learnability control. That full-context learnability control scores 40/60: early-fact and latest-update are both 20/20, with zero compactions and zero truncations, but ordered-plan remains 0/20. Trace analysis shows a grammar failure: all plan responses invent an audit_records array instead of the required {"plan":"a>b>c"} string. The repository therefore keeps the default prompt unchanged and adds a versioned, non-leaking plan-schema hint as the next calibration gate before grammar SFT or more sparse-reward RL. The first schema-hint calibration keeps ordered-plan at 0/20. It eliminates invalid commands, but the model emits placeholder ordinals or distractor records instead of the three milestone tokens. A fixed fake-value demonstration is the final prompt-only gate; failure there routes to a small grammar SFT rather than a compacted RL comparison. That fixed-example gate reaches only 1/20 under full context. The repository now includes a pinned-slime, one-GPU grammar SFT warm-start. The first terminal-only recipe is sharply non-monotonic: update 10 scores 0/20, update 20 scores 20/20, and update 30 returns to 0/20 as the action protocol collapses. The passing update-20 checkpoint is a valid gated candidate, while a lower-LR corrected recipe behavior-clones legal next actions and terminal JSON to seek a wider stable window. Compaction summaries remain entirely outside SFT targets. See the grammar-SFT diagnosis. The corrected run passes the full-context gate at updates 10, 20, and 30 (20/20, zero invalid commands). Its selected earliest checkpoint, update 10, scores 8/20 with compaction versus 20/20 without compaction on the exact same task IDs. This clean 12-task paired regression (p=0.000488) establishes the phase-2 summary-state-transfer target while retaining nonzero exploration reward for CompactRL. The first live phase-2 smoke then exposed an important warm-start value bug: slime's random missing critic head generated nonzero returns and a policy gradient on an all-zero-reward batch. The integration now zero-initializes only that missing scalar head. A same-task, byte-identical trajectory replay reduces values, returns, advantages, both critic updates, and the actor gradient to exactly zero; the automated safety audit passes 19/19 checks. See the phase-2 warm-start report. The first 12-rollout phase-2 pilot then passes the full slime E2E audit and receives seven real sparse rewards, but held-out compacted accuracy moves from 8/20 to 7/20. Paired traces show three gains and four regressions; regressed summaries repeat the first milestone and grow by 28.5 tokens on average. Because the nominal four-step critic warm-up contained no reward, the next controlled setting uses 12 critic-only batches and halves actor LR before making any improvement claim. That 20-rollout reward-effective warm-up is the first positive long-horizon result: the critic sees three positive tasks before the actor is enabled, the final actor scores 16/20 versus the frozen baseline's 8/20, and paired outcomes are nine improvements to one regression (p=0.0215). Summary milestone recall rises from 56.10% to 93.90% with zero invented plan tokens. The subsequent independent-seed replication and locked test initially passes the two remaining gates. Seed 29 scores 15/20 on dev, while the fresh test20 comparison is 6/20 for the SFT baseline, 16/20 for RL seed 23, and 13/20 for RL seed 29. Those checkpoint measurements remain valid, but the later seed-sensitivity and reward-gate experiment changes the conclusion: an identical seed-37 run scores only 6/20 on dev, 2/20 on ordered-plan test, and 30/60 on full test. Its test summary recall falls to 21.43%, with increased invention and repetition. The fixed ungated recipe succeeds on two of three observed training seeds and is not robust enough to scale as a default.

The opt-in raw-reward gate fixes the concrete zero-positive-batch failure mode. It continues critic training but skips an actor update unless the batch has at least one positive verifier reward. Under the exact seed-37 counterfactual all eight actor-phase batches remain zero-positive, so actor/critic update counts are 0/40 and deterministic dev20 is exactly unchanged from SFT at 8/20. The remote skip audit passes all 11 applicable checks. This is a verified safe no-op, not learning: the next experiment must create a new dev/test protocol and make reward-bearing actor batches sufficiently likely before model scale-up. Across all runs, the evidence is small-model synthetic long-horizon wiring and capability evidence through pinned slime, not a GLM-scale benchmark claim or a universal Pareto improvement.

That next experiment is now complete. The phase-3 protocol uses new phase3 task IDs, 480 train-only ordered-plan rows, and fresh 60-task ordered-plan dev/test slices. It selects rollout/global batch 24 from the frozen-actor reward estimate (12/288) and requires at least four reward-bearing actor batches plus a 10-point dev gain on two independent seeds before opening the new locked test. All three seeds improve fresh dev from the SFT baseline's 22/60: seed 41 scores 31/60, seed 53 41/60, and diagnostic seed 67 36/60. Seeds 41 and 67 pass all 12 per-seed training/dev checks; seed 53 fails only the preregistered invention guardrail because summary invention incidence rises by 6.45 points versus a five-point maximum. Under the strict first-two-seed rule, locked test remains unopened. The exact evidence and interpretation are in the phase-3 result.

The next stability experiment is preregistered on another fresh namespace. Phase 4 compares matched no-KL and actor reference-KL 0.01 arms while leaving rewards, GAE, returns, critic targets, and the raw-reward actor gate unchanged. It adds an explicit repetition guardrail and requires the KL arm to reduce invention or repetition without losing more than five capability points in two fixed seed pairs before any new locked test can open.

Phase 4 is now complete and the KL control fails that gate. In pair 1, no-KL and KL score 28/60 and 25/60; in pair 2 they score 52/60 and 39/60. KL consistently reduces invention, repetition, and summary length, but loses 5.0 and 21.7 exact-match points plus 6.7 and 23.4 milestone-recall points against its matched no-KL arm. All GAE/value/KL objective audits pass, so this is an actor-objective result rather than a changed critic target. The phase-4 locked test remains unopened. See the phase-4 result.

Across the four phase-4 arms, zero-reward trajectories still account for 26.0%--41.0% of absolute actor advantage mass once a batch-level gate has admitted an update. Phase 5 therefore isolates an opt-in, actor-only control: compute the same GAE/returns for every trajectory and keep critic training unchanged, but zero policy advantages for zero raw-reward logical trajectories inside a reward-bearing batch. This engineering ablation is not part of the report-faithful default. The fixed data, two seed pairs, and open-test rule are in the phase-5 preregistration.

Phase 5 is now complete and does not pass its two-pair gate. Pair 1 passes and favors positive-trajectory advantages: exact match rises from 42/60 to 44/60, milestone recall from 86.9% to 92.3%, invention falls from 6% to 2%, and repetition from 13% to 10%. Pair 2 has zero reward-bearing actor batches after the 12-step critic warmup, so both arms correctly perform zero actor updates, continue all 40 critic updates, and exactly reproduce the SFT dev result (19/60). The second pair therefore cannot exercise the mask; the phase-5 locked test remains unopened. The audits, value-function behavior, and interpretation are in the phase-5 result.

Phase 6 is preregistered on a fresh namespace to replicate the pair-1 effect without the pair-2 opportunity failure. It fixes 36 updates (12 critic-only plus 24 actor-phase), batch size 48, exactly 1728 unique ordered-plan train tasks, a 2048-token dynamic microbatch cap, two new seed pairs, and a minimum of six reward-bearing actor batches per arm. The scale is chosen using only already open phase-5 train outcomes; neither fresh dev nor locked test participates. See the phase-6 protocol.

Phase 6 is now complete. Scaling removes the actor-opportunity failure and all four arms learn the long-horizon task: final dev exact match is 55/60 and 50/60 in pair 1, and 52/60 and 48/60 in pair 2, against the fresh SFT baseline's 19/60. But the positive-only actor mask loses to the batch-gate control in both pairs, and every RL arm violates the preregistered summary invention/repetition guardrails. All critic, GAE, return, mask, and optimizer audits pass. The phase-6 locked test remains unopened; see the phase-6 result.

Phase 7 addresses the exposed objective misspecification without changing the critic architecture. Its candidate keeps one shared scalar return but sets it to terminal exact match x minimum observed-summary plan-token F1. Missing milestones lower recall; invented, future, or repeated milestone tokens lower precision. The scalar critic, local and cross-trajectory GAE, masks, failed trajectory critic training, and optimizer schedule remain intact. It uses a fresh 1728-task train namespace and two fixed seed pairs; a remote auditor must reconstruct every shaped reward before any dev claim or locked-test opening. See the phase-7 preregistration.

Phase 7 is now complete and the multiplicative fidelity reward fails its two-pair gate. It retains the long-horizon capability gain: the candidates score 51/60 and 52/60 against the fresh SFT baseline's 24/60. But neither candidate reduces invention or repetition by the required ten points versus its matched terminal-reward control; both also exceed the SFT-relative quality limits and make summaries 12.7--13.2 tokens longer than control. All scalar reward reconstruction, critic, GAE, return, mask, optimizer, and checkpoint audits pass. The failure is therefore attributed to sparse multiplicative credit (R=0 for every failed task), not to the scalar critic architecture. The phase-7 locked test remains unopened. See the phase-7 result.

Phase 8 is preregistered on another fresh namespace. It keeps the successful phase-7 return unchanged but gives failed trajectories a dense quality cost: R = task_exact + minimum_summary_F1 - 1. Thus faithful failures receive zero and degenerate failures receive a negative scalar. Negative trajectories stay in both actor and critic training whenever a positive trajectory admits the batch; the scalar head, GAE, cross-trajectory correction, masks, schedule, and control remain unchanged. The fixed hashes, seed pairs, negative-return audit, dev gates, and locked-test rule are in the phase-8 preregistration.

Phase 8 is now complete and fails in the same way in both seed pairs. The dense candidates score 0/60 and 2/60, versus 53/60 and 55/60 for the matched terminal controls; their summaries collapse to 2.9 and 5.6 mean tokens. All reward reconstruction, critic, value-mask, GAE, return, gradient, schedule, and checkpoint audits pass, and the dense critics track the supplied negative returns. The failure is therefore actor-side: admitted batches retain predominantly negative trajectories, whose signed policy gradients suppress almost every sampled summary. The locked test remains closed. See the phase-8 result.

Phase 9 freezes the dense critic target and changes only actor routing. Both arms use R = task_exact + minimum_summary_F1 - 1 and train the critic on every trajectory. After GAE, only the candidate zeroes actor advantages for logical returns that are not positive. Fresh data, two fixed seed pairs, strong SFT-relative quality gates, matched rescue gates, and the locked-test policy are frozen in the phase-9 protocol.

Phase 9 is complete and replicates the mechanism in both fixed seed pairs. The signed controls collapse to 5/60 and 0/60 dev, while actor-only any-success routing reaches 51/60 and 55/60. Every reward, value, GAE, return, mask, schedule, checkpoint, matched-stream, and cross-arm tensor audit passes. The two candidates nevertheless reach 18.5%/40.2% invention and 98.9%/94.6% repetition, failing the same two frozen quality guardrails. The locked test remains closed. See the phase-9 result.

Phase 10 is preregistered on another fresh namespace to isolate the remaining actor-admission issue. Both arms keep Phase 9's dense critic and post-GAE actor routing; the control admits any success (R > 0) while the candidate admits only a task success whose minimum sampled-summary F1 is exactly one (R > 0.999). The threshold is shared by the batch gate and trajectory mask, and every below-threshold trajectory remains critic data. See the phase-10 protocol.

Phase 10 is complete and the locked test remains closed. Pair 1's perfect-fidelity arm keeps 54/60 dev exact versus 57/60 for any-success while cutting repetition from 93.75% to 37.89%, but it still fails the frozen SFT-relative invention and repetition limits. Pair 2 supplies the missing strict causal opportunity: at an identical update-12 tensor batch, the candidate masks exactly one additional imperfect-success actor trajectory while retaining 5,474 critic tokens and all return targets. Both Pair 2 arms complete 72 critic updates and reach final value explained variance above 0.73. Pair 2 dev was not run after a low-Modal-balance warning because Pair 1 had already made the locked-test decision irreversible. See the phase-10 result.

Phase 11 passed its bounded preflight and both training gates. Stage 2a completed in 1,556.24 seconds and used at most 103.75 A10 GPU-minutes. Stage 2b resumed that exact checkpoint at update 13, completed through update 35 in 2,673.45 seconds, and used at most 178.23 A10 GPU-minutes (about $3.27 at the recorded A10 rate). Across the two stages the candidate made 19 actor and all 72 critic optimizer updates. Runtime, fidelity, matched task-stream, strict-update-12, summary-KL coverage, and objective-parity audits all passed. The training authorization is consumed and closed. The candidate-only dev screen was fail-closed before GPU allocation and capped at 600 seconds on 2xA10 (20 A10 GPU-minutes per attempt). It is now complete: the candidate kept 54/60 exact and 99.19% summary recall, but failed invention (7.53% > 5%) and repetition (65.59% > 20%). Fresh dev and locked test remain closed. It nevertheless improved the exact matched no-KL control from 50 to 54 exact, 26.88% to 7.53% invention, and 88.17% to 65.59% repetition. The direction works; coefficient 0.01 with the full actor-token denominator is not strong enough. It applies k3 reference KL only to sampled summary tokens while retaining the original full actor-token denominator; PPO advantages, returns, actor admission, and the dense critic path are unchanged. To conserve quota, the completed Phase-10 Pair-2 perfect-fidelity run is the matched no-KL control, so only one candidate would be trained. To avoid paying the full candidate cost before the first strictly comparable actor update is known-good, candidate training is split: Stage 2a runs updates 0--12 from the frozen SFT warm start and performs the strict update-12 tensor audit; only a separately authorized Stage 2b may resume that exact checkpoint for updates 13--35. Both stages emit a persistent, fail-closed gate artifact. The preflight completed one actor update and four critic updates with every built-in audit passing; cumulative training usage across the failed device-placement attempt, passing retry, Stage 2a, and Stage 2b was at most 340.99 A10 GPU-minutes (about $6.26 at the recorded A10 rate). See the phase-11 protocol.

The toy is a behavioral smoke test, not evidence of LLM improvement. It checks the causal idea: a summary decision and a post-compaction execution decision share one terminal reward, so updating the summary policy can beat freezing it.

Algorithm in one screen

For history h_t = (system, user, (action_1, observation_1), ...):

  1. If remaining context is below T_comp, sample a summary S_t from the same trainable policy using a fixed summary instruction.
  2. Continue from system + resume(S_t) + last_k_atomic_steps, reducing k from two if needed to fit the peak context.
  3. Export every execution and summary response as a chronological segment. Optimize only tokens actually sampled by the policy; mask prompts, observations, copied summaries, and copied tail turns. In the summary-loss-masked ablation, keep the summary segment and its critic mask but set only its actor mask to zero.
  4. Give every segment from one logical rollout the same verifier reward.
  5. Compute local GAE inside each segment. If N_after[s] optimized tokens occur in later segments, multiply all local advantages in segment s by (gamma * lambda_s) ** N_after[s].
  6. Optimize PPO by summing over all enabled tokens and dividing once by the total enabled-token count in the global batch.

The crucial invariant is token provenance: summary tokens get loss only at the turn where the policy sampled them. Their copied appearance in the resumed prompt must be masked.

Why this is separate from slime

The public slime main branch audited for this project (commit fb42ae456fac8166afb604f13b30d22bb3c75053, 2026-07-15) provides the required infrastructure primitives:

  • PPO with a critic;
  • --calculate-per-token-loss;
  • --custom-advantage-function-path;
  • list[Sample] fan-out with shared rollout_id;
  • loss masks and token-in/token-out agent adapters;
  • compacted/sub-agent trajectory grouping.

It does not expose a built-in CompactionRL estimator or the paper's rollout, so this repository supplies both through slime's supported customization interfaces. Keeping the research logic here makes the equations testable without a large framework deployment.

See the slime integration guide for the concrete contract and the technical analysis for report claims, experimental results, ambiguities, and a staged reproduction plan.

Repository map

src/compactionrl/advantages.py       paper equations and exact reference
src/compactionrl/compaction.py       trigger and context reconstruction
src/compactionrl/losses.py           token-level PPO reducer and ablation
src/compactionrl/adapters/slime.py   optional public-slime advantage hook
integrations/slime/rollout.py        slime custom compact rollout
integrations/slime/run_*.sh          pinned configurable-GPU PPO launcher
integrations/slime/modal_e2e.py      official-image Modal preflight and E2E
integrations/slime/patches/           minimal pinned-slime critic-loop patch
scripts/audit_slime_e2e.py           independent live-artifact E2E gate
scripts/audit_reward_gated_skip.py   zero-positive actor-skip audit
scripts/summarize_training_curve.py  logical reward curve from slime dumps
scripts/summarize_value_diagnostics.py critic and boundary-value log summary
scripts/summarize_eval_rollout.py    fixed held-out logical-task metrics
scripts/compare_eval_runs.py         paired checkpoint comparison and exact test
scripts/render_megatron_roles.py     isolated actor/critic checkpoint paths
scripts/build_final_e2e_report.py     publication-blocking evidence report
scripts/download_modal_evidence.sh    lightweight Modal evidence retrieval
examples/toy_joint_training.py       joint-vs-frozen summary smoke test
experiments/long_horizon_memory_chain.py  96-step multi-compaction experiment
src/compactionrl/simulation.py       deterministic environment and verifier
data/synthetic/*.jsonl               240/60/60 generated task splits
data/synthetic_phase3/*.jsonl        fresh 1440/180/180 preregistered splits
configs/                             reported settings and debug presets
results/long_horizon_memory_chain.*  measured multi-seed result and path trace
results/modal/image_inspection.json  measured official-image runtime preflight
tests/                               deterministic algorithm tests
docs/TECHNICAL_ANALYSIS.md           deep report analysis and disclosure gaps
docs/SLIME_INTEGRATION.md            integration contract and launch sketch
docs/SIMULATION.md                    data design and first experiment settings
blog/compactionrl-reproduction.zh-CN.md complete Chinese experiment narrative
blog/compactionrl-reproduction.md    complete English experiment narrative
reports/modal_scale_pilot_20260718.md bounded live comparison and claim boundary
reports/modal_scale_value_comparison_20260718.md corrected three-arm scale report
reports/modal_checkpoint_evaluation_20260718.md checkpoint diagnosis and dev eval
reports/modal_kl_dose_response_20260719.md KL 0/0.01/0.05 candidate result
reports/modal_deterministic_sampling_20260719.md exact SGLang trajectory gate
reports/modal_controlled_kl_replication_20260719.md common-random-number KL result
reports/modal_deterministic_no_kl_replication_20260719.md same-seed robustness result
reports/modal_deterministic_training_validation_20260719.md exact optimizer repeat gate
reports/modal_deterministic_seed17_kl_replication_20260719.md eight-update second-seed gate
reports/modal_full_context_learnability_20260719.md no-compaction grammar gate
reports/modal_phase2_seed_sensitivity_reward_gate_20260719.md current phase-2 conclusion
reports/modal_phase3_preregistered_protocol_20260719.md fresh-data scale-up gate
reports/modal_plan_schema_hint_20260719.md non-leaking plan-schema calibration
reports/modal_plan_example_gate_20260719.md fixed-example gate and SFT decision
reports/phase16_offline_credit_redesign_20260724.md coverage-stop live result
reports/phase17_causal_prefix_postmortem_20260724.md causal-credit mechanism result and rejected pairing
results/modal/scale_20260718/        manifests, audits, curves, value diagnostics
results/modal/checkpoint_eval_20260718/ fixed checkpoint/evaluation summary

Claim boundary

There are three separate evidence levels:

  1. Verified here: equations, trigger/reconstruction state transition, token-loss weighting, deterministic tests, the controlled 96-step learning experiment, and live Qwen/slime actor+critic updates with independently audited gradients, GAE, masks, compact paths, checkpoint artifacts, and nonzero exact-match rewards after compaction.
  2. Specified by the paper but not run here: SWE-Dev training through slime, GLM-4.7-Flash / GLM-4.5-Air-SFT, Harbor, Terminus-KIRA, and paper-scale ablations.
  3. Not publicly reproducible exactly: the internal GLM-5.2 training run and undisclosed recipe details.

Sources

About

Auditable CompactionRL reproduction with synthetic long-horizon experiments and a live THUDM/slime E2E

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages