Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .batten/asked.jsonl
Original file line number Diff line number Diff line change
Expand Up @@ -6,3 +6,5 @@
{"question":"Is the family above the set you meant, and what should happen next?","options":[{"label":"Fix the board state","description":"Check CLOUD-2037, CLOUD-2017 and CLOUD-122 against the tree, and move any with no PR behind it back to Backlog with a comment."},{"label":"Pick up the open rows","description":"Plan a branch that works through the open error and advisory prose rows, starting with CLOUD-2037, CLOUD-2017 and CLOUD-1327."},{"label":"Different batch","description":"I meant specific tickets filed together. I'll give a key or a phrase from one of them."}],"answer":"Fix the board state","at":1791005849,"head":"4ed344a3941cf1e9bcc37918eedf555f6c119986"}
{"question":"Dispatch two sessions, one PR each: Bundle A brief:43e89098ea7db7ba6b3d7acf3326c250067b7faf6dd3d1f1fe39baf4b4a6a666 (hook.rs/refusal.rs/lib.rs: CLOUD-1728, 1826, 1996, 1806, 1893, 1470, 2075, 2078, 1583, 2079) and Bundle B brief:79ee5e255f39abeaf60752e5f347a6fde1e7863e1ce0aed48c0a4d2a74feed41 (config caps and commit-msg classes: CLOUD-1960, 1642, 1643)?","options":[{"label":"Approve dispatch","description":"Start both sessions with these exact briefs. Each claims its rows in order, commits once per row, and lands one PR."},{"label":"Revise briefs","description":"Hold the dispatch; tell me what to change."}],"answer":"Approve dispatch","at":1791029074,"head":"4ed344a3941cf1e9bcc37918eedf555f6c119986"}
{"question":"Which permission mode should the two child sessions start in?","options":[{"label":"Auto (Recommended)","description":"They build and land unattended; the PRs and board are the record."},{"label":"Plan","description":"Each child parks for your approval in the web UI before building each row."}],"answer":"Auto (Recommended)","at":1791029074,"head":"4ed344a3941cf1e9bcc37918eedf555f6c119986"}
{"question":"Admit the `[refusal]` ceiling raise? Gate `config-lint` refuses `batten.toml:refusal.max_tokens refusal-ceiling-raised` and `batten.toml:refusal.first_sighting_max_tokens refusal-ceiling-raised`: raising a committed per-line ceiling is a weakening, admitted only by your recorded answer. Why it should not stand: CLOUD-2075's decided mechanism puts labels, every route and the override request on every firing, and no renderer reads the ceiling any more — it only reports. The values (24→96, 64→176) are the measured widest lines, rounded up to a multiple of 16, as the row specifies. Cost if wrong: each refusal line costs up to ~4x the old context per firing (worst measured pointer 84 tokens, full arm 159). I caused this by changing the renderer.","options":[{"label":"Admit the raise","description":"Record the admission; the ceilings stay at 96 and 176 and land proceeds."},{"label":"Refuse","description":"Do not admit; CLOUD-2075 does not land until the ceiling is redesigned."}],"answer":"Admit the raise","at":1791043777,"head":"2f23e7fd7e442fdacae28ce805fdcc3993947e8f"}
{"question":"Admit `refusal-ceiling-raised refusal.max_tokens` and `refusal-ceiling-raised refusal.first_sighting_max_tokens`? (Re-asked: my previous question spelled these pairs in the wrong order, so the ledger could not match your answer to them. Nothing else changed.) Gate `config-lint` reports the [refusal] ceilings raised from 24 to 96 and 64 to 176 as weakenings. Why the raise should stand: CLOUD-2075 puts labels, every route and the override request on every firing; no renderer reads these keys any more, they only report; the values are the measured widest lines (pointer 84, full 159 estimated tokens) rounded up to a multiple of 16. Cost if wrong: each refusal line costs up to ~4x the old context per firing.","options":[{"label":"Admit the raise","description":"Record the admission for both pairs; land proceeds with the ceilings at 96 and 176."},{"label":"Refuse","description":"Do not admit; CLOUD-2075 does not land until the ceilings are redesigned."}],"answer":"Admit the raise","at":1791044811,"head":"2f23e7fd7e442fdacae28ce805fdcc3993947e8f"}
3 changes: 2 additions & 1 deletion .serena/memories/workflow/landing-loop.md
Original file line number Diff line number Diff line change
Expand Up @@ -190,7 +190,8 @@ exist in no other clone, so `history-drop` reads the reset as discarding work an
stops it. Measured 2026-09-19, and the class's own override did not open it: a
`history drop unpushed` admission was requested, answered and spent TWICE — once
against the commit sha, once against the subject its refusal line prints — and
the reset was refused unchanged after each. CLOUD-1871 owns that gap. The rebase
the reset was refused unchanged after each. Fixed by CLOUD-1826: paste the
pointers the refusal prints (`1 <sha>`), not the sha alone. The rebase
went through on the first try.

Take the second anyway, on the rare tree where the guard lets it, and two things
Expand Down
123 changes: 65 additions & 58 deletions batten.toml
Original file line number Diff line number Diff line change
Expand Up @@ -601,7 +601,8 @@ sha256 = "b2c822742e8cbf355ba0cb4cc690c3cd8fdc9ec1916c8148f27bd9098cb7aee4"
#
# `reason` is required on a shape rule, unlike on a file rule: a mediated deny
# reaches a model as the entire explanation, so a refusal that named only its id
# would be un-actionable (CLOUD-122). The crate appends the bypass hatch.
# would be un-actionable (CLOUD-122). It is the row's remedy, which the refusal
# reaches through `batten policy rule '<id>'`.
# THESE FOUR WOULD DECLARE `bypass_env = "BATTEN_GH_GUARD_BYPASS"` AND DO NOT YET
# (CLOUD-437, deferred to CLOUD-1027). They are the ported `gh-guard`, so that
# name is TRUE of them where it was a fossil everywhere else, and declaring it
Expand All @@ -616,12 +617,11 @@ sha256 = "b2c822742e8cbf355ba0cb4cc690c3cd8fdc9ec1916c8148f27bd9098cb7aee4"
# one. Asserting it in the change that performs it is the exact shape §8 refuses.
#
# So until that row is groomed, these four declare no per-row hatch. That is a
# statement about the column, NOT a remedy: the residual global hook hatch still
# technically reaches them only because their class, `call name refused`,
# declares no override route yet (CLOUD-1806), and CLOUD-1357 makes every class
# that DOES declare one unsuppressible by it. A reader looking for the way past a
# refusal reads the class's routes with `batten policy explain "<class>"` and,
# where one is an override, walks `batten override request`/`spend`. This comment
# statement about the column, NOT a remedy: their class, `call name refused`,
# declares an override route (CLOUD-1806). A reader looking for the way past a
# refusal reads the row's remedy with `batten policy rule '<id>'`, and where that
# remedy cannot perform the change, walks `batten override request`/`spend` —
# the recorded way through, bound to the row id at the current commit. This comment
# used to say these rows "take the general" hatch "like every other row", which an
# agent read as advice and proposed for a class it could not have suppressed.
[[rule]]
Expand Down Expand Up @@ -5700,64 +5700,41 @@ files = [".serena/memories/*.md", ".serena/memories/**/*.md"]
max_bytes_per_file = 40960
max_tokens = 110000

# What ONE emitted mediated refusal line may cost (CLOUD-1286).
# What ONE emitted mediated refusal line may cost (CLOUD-1286, re-declared by
# CLOUD-2075).
#
# The neighbouring budget above bounds what loads once per session. This bounds
# what is emitted ~300 times in one, so it is the ceiling that actually
# compounds: measured live on 2026-09-01, a `no-tool-substitution` refusal was 88
# words / ~115 tokens, and ~300 firings of it is ~34,500 tokens against a ~175k
# window — 20%, which is CLOUD-417's headline figure arrived at independently
# from the other direction.
# This bounds the POINTER arm — what every firing after the first in a context
# emits. CLOUD-2075 put both labels, the subjects, every route (override routes
# included, as the ready request that admits) and the `policy rule` hop on that
# arm, because a pointer a ruling needs is never shed. So the number is the
# measured widest pointer the tree emits, rounded up to a multiple of 16, and
# nothing reads it at runtime: `refusal_ceiling.rs`'s corpus case REPORTS an
# over-ceiling line rather than any renderer truncating one. Measured 2026-10-03
# with `every_arm_the_corpus_emits_is_within_its_declared_ceiling` and the
# routed-read and render-bench cases (estimated tokens, `budget.rs`'s len/4):
#
# 24 is chosen against the longest line the tree can actually emit rather than
# against the shortest: a three-word class is ~3 tokens and the pointers are the
# rest, so a deeply-nested `path:line` plus an artifact fits with room, while the
# ~43-token rendered form this row retires does not. A ceiling only the shortest
# class clears would fire on correct output, and the first person it fires on
# switches it off (CLOUD-418).
# 84 pointer branch write unsafe (render bench, two routes + override)
# 78 pointer path read routed (longest committed memory name)
# 49 pointer the Bash corpus (`tool run loose`)
#
# Declared here rather than as a constant in `crates/batten` for the reason
# non-negotiable rule 1 gives: a consumer whose harness renders wider cannot move
# a number compiled into the engine.
[refusal]
max_tokens = 24
# What ONE FIRST SIGHTING may cost (CLOUD-1637).
#
# The key above prices prose a reader has already met. This prices the one firing
# where they have not: the same line plus the class gloss and every non-override
# route, which is what turns a three-word pointer into something actionable. Two
# keys rather than one number doing both jobs, because the quantities differ by
# design — and bounding the long arm by the short arm's 24 would be a ceiling no
# line that arm can compose could ever meet, which `refusal::validate` now refuses
# outright rather than leaving as an arm that silently never renders.
#
# 64 IS READ OFF A MEASUREMENT PLUS THE DECLARED MAXIMA, not chosen. Fired in a
# fixture whose sightings store started empty, 2026-09-08, one command per
# distinct refusing row (estimated tokens, `budget.rs`'s len/4):
#
# 36 verdict read dropped verdict-not-discarded ... read rules/toolchain.md
# 35 tool run loose no-tool-substitution ... read rules/scanning.md
# 34 trunk push forced no-force-push ... run git push --force-with-lease
# 28 call name refused no-bare-cargo ... read batten.toml
#
# So the tree emits 36 at its widest today. The headroom above that is not slack:
# `GLOSS_MAX` admits a 120-character gloss, which is ~30 estimated tokens on its
# own, and a class may declare several routes. 64 clears every line this tree can
# spell while still catching a route list that has run away — the only thing this
# ceiling sheds, since the gloss is undroppable and the first route is the floor.
# A ceiling that fired on correct output would be switched off by the first person
# it fired on (CLOUD-418), and `refusal_ceiling.rs` asserts the corpus stays under
# it in the same direction the repeat arm's anti-vacuity case does.
#
# NOT AN `[epoch]` FLOOR BUMP, and that is a decision rather than an omission.
# `min_batten_version` is compared against the RUNNING build, which in this
# repository is the one built from this tree, so naming the next release makes the
# repository refuse its own config for the whole life of the PR (measured
# 2026-08-11: floor 0.0.62 against a 0.0.61 build, exit 1). The floor can only
# name a version that already exists, so the raise belongs at the release. What
# `[epoch]` actually wants — an older binary REFUSING this key rather than
# ignoring it — `deny_unknown_fields` on `[refusal]` already delivers.
first_sighting_max_tokens = 64
max_tokens = 96
# What ONE FULL ARM may cost (CLOUD-1637, re-declared by CLOUD-2075).
#
# The full arm is the pointer arm plus ` — ` and the definition: the class gloss
# (or an undeclared row's reason), the ROW'S OWN REMEDY, each override's
# precondition and the `policy explain` hop. It is delivered once per context per
# compaction cycle. Measured 2026-10-03 the same way as the key above:
#
# 159 full tool run loose / tool select other (the row's reason is the bulk)
#
# 176 is that rounded up to a multiple of 16. `refusal::validate` still refuses a
# value at or below `max_tokens`. NOT AN `[epoch]` FLOOR BUMP, for the reason the
# history of this key gives: the floor can only name a release that exists.
first_sighting_max_tokens = 176

# What the whole advisory CHANNEL may cost on one boundary (CLOUD-896).
#
Expand Down Expand Up @@ -5851,6 +5828,13 @@ family = "unix"
# `mise run hook-cost`, which is why the figure ships as a command rather than as
# a number in an issue body.
#
# READ PER COMPACTION CYCLE (CLOUD-2075). A finding's full arm is delivered once
# per context per cycle, and a `SessionStart` opens the next cycle — so
# `hookcost::measure` counts repeats per `SessionStart`-bounded segment, judges a
# labelled finding by its key rather than by the bytes around it, and never counts
# a pointer arm: pointing at the first copy is what a repeat is supposed to do.
# Unlabelled output keeps the per-segment `(hook, digest)` reading.
#
# Declared here rather than in `crates/batten` for non-negotiable rule 1's
# reason: how loud a repository's hooks may be is a property of that repository's
# hooks, and a consumer cannot move a number compiled into the engine.
Expand Down Expand Up @@ -8251,6 +8235,19 @@ severity = "warn"
id = "unanswered-human-calls"
count = "unanswered-human-calls"

# A code-host call is handed the memory documenting this host's GitHub access
# (CLOUD-1470). Measured twice: 2026-08-31 an `add_repo` detour the memory says is
# blocked (CLOUD-1259), and 2026-09-05 six turns spent telling the owner the
# lander could not run before `mise run land` drove the loop first try. `warn`,
# so the call is allowed and the pointer rides it on the advisory channel.
[[rule]]
id = "forge read first"
no_fix_reason = "class 3 (a call rewrite): the subject is a tool call refused before it runs, so nothing persists for a command to repair; the refusal's remedy text is the route"
kind = "policy"
scope = "mediated_call"
module = "policy/forge-read-first.rego"
severity = "warn"

# A weaker spelling of a task this project already defines (CLOUD-856, the
# successor shape for `run-shape-guard`'s `cargo-substitutes-for-a-task`).
#
Expand Down Expand Up @@ -13401,6 +13398,16 @@ id = "module read first"
kind = "document"
target = "policy/answer-the-operator.rego"
[[verdict]]
id = "forge read first"
gloss = "a code-host call; read the memory documenting this host's GitHub access first"
class = """
A call reaching the code host was made with the memory documenting this host's GitHub access unread at the moment it mattered: on 2026-08-31 a session took an `add_repo` detour that memory says is blocked, and on 2026-09-05 one spent six turns claiming the lander could not run before `mise run land` drove the loop first try. The pointer rides the allowed call; the remedy is to read the memory before concluding the host is unreachable.
"""
[[verdict.route]]
id = "memory read first"
kind = "document"
target = ".serena/memories/github-access.md"
[[verdict]]
id = "forge check red"
gloss = "the forge judged this commit and its fan-in check did not pass"
class = """
Expand Down
Loading