Repository navigation
Move the limits onto the claim, let the answer support a reader's confidence, and close the continuity loop on the route that opens it - #847
Merged
Conversation
…omparison, and an unsettled consideration reaches the next one (refs #827, #844) Four connected changes on one front. The always-on boundary file contradicted the entry it loads beside; the shape law deleted the support a reader needs to check a call; the record reader never named the user's standing rules; and the continuity loop could not be closed on the route that opens it. 1. Scope each prohibition to the fact it protects. `agent-boundaries.md` forbade "rankings" and "derived analysis" unscoped and forbade arguing a considered trade "from anything but consider's output", while `SKILL.md` mandates candidate ranking, labelled judgment and a deciding reason from the user's own record. Now: a *portfolio-derived* number or ranking, a *portfolio* figure the engine did not emit, a considered trade's *portfolio consequence*. A risk the engine did not measure is named on materiality, not enumerated -- which is what `unchecked` already says of itself. 2. Amend the expression contract's increment gate (owner amendment through §3.4's first lane, logged in §8). "Does the decision change?" deleted every block that let a reader check the call or see why one candidate beat another. It now asks whether the reader can still take the call, judge how far to trust it, and see why it beat the alternative in play. Paired: a *wait* stance owes the evidence that settles it and the next checkable point, since a scheduled date is not a reason. The four bans, D1-D7, C1-C4, the registry freeze and the refusal of any character-count ceiling stand. 3. The record read reaches standing decision rules, states where to start (the names in play and those rules, then what they name -- no whole-folder sweep, no filename convention), and keeps what bounds a quote: a condition's threshold, a thesis's falsifier, whether a stance authorized acting. Prefer the newest statement, name a real conflict, argue a departure from their preference rather than overwriting it. 4. `review._unresolved_prior` projects at most one unsettled same-ticker consideration onto the `consider` response, carrying the `evaluation_id` that `--resolve` needs. That identifier was previously emitted only in the response that minted it, so a consideration left open on this route could never be closed on it -- and `_prior_decision` stayed permanently empty for that ticker. It carries no `decision` and no `decided_on`, because an open row is a question that was asked, never a decision. Tests: `tests/test_consider.py` section R (R1-R8 plus compatibility), fourteen mutations run against the new reader and its call site. Section Q's whole-file recap ban is narrowed to require every mention be a prohibition, since the contract now bans a recap outright on two fields. Product group 49/49 and the whole registry 60/60 pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…827, #844) Four mirrored surfaces the first commit left behind, and one misfiled permission. - The three README install paragraphs name the standing decision rules and the qualifier rule. That paragraph is what tells a user what the folder is read for, so it drifting is the same defect as the contract drifting. - `skills/fomo-kernel/evals/evals.json` gains scene 5: the decision route, which the bank's four review-lifecycle scenes never covered. A returning user with an unsettled earlier consideration, a standing rule, and a note whose observation carries a not-yet-actionable qualifier. Its expectations name reading the standing rule and preserving the qualifier as what is under test, and leave the verdict to owner dogfood rather than the automated score. - `references/market-lookup.md`'s "Company research is in scope when..." paragraph sat under the heading "When lookup does not happen" -- a permission filed under the heading that denies it, which is exactly the entry/reference ambiguity #827 exists to remove. It now has its own heading and decides scope by the same three-part test the amended increment gate uses, so what may be looked up and what reaches the answer are one rule rather than two. - The maintainer guide's three 2026-09-05 rows name all of the above. Product group 49/49. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…, four mutations survived, and a new field had no relevance gate (refs #827, #844, #846) Seven findings, all confirmed by reproduction. Six are mine. **The amendment contradicted its own summary table.** `docs/expression-contract.md` §3.1's Middle row still read "Only support that could change the decision" -- the pre-amendment gate, fourteen lines above the amendment. A reader following the table deletes exactly the block the amendment exists to keep. Fixed, and the 2026-09-05 ruling row moved to the end of §8, whose log is ascending. **Four mutations survived section R.** Each now has a test, and each was re-run against the new tests to confirm it reddens: - Dropping `_fold_evaluations` from the open reader. A resolution appends a second row under the same `evaluation_id`, so one consultation reached the agent as both "you reported acting on this" and "you never told me what you did". The disjointness is the fold's, not the two `decision` tests'. The assertion that should have caught it was a coin flip: two ids minted the same day, ordered lexicographically, so it passed or failed with the calendar. Replaced with a seeded-date fixture that asserts the property directly. - Dropping the `isinstance(evaluation_id, str)` guard. `max()` on a one-row fixture never compares, so the guard was present and untested while a non-string id crashed `consider` outright. R8 now pairs a broken row with a healthy one, both dated the same day, forcing the comparison onto identity. - Hard-coding `"buy"` at the call site. No test ran a sell-side *current* consultation, so the same-side preference reversed for every sell the user brings, silently. - Loosening `decision != "open"` to `not in PRIOR_DECISION_RESOLVED`, which projects a row carrying `null` or junk as a question the user was asked and never answered. **The recap gate became word-gameable.** Narrowing it to "every line mentioning a recap must also say never" is defeated by this file's own house style -- "Open every answer with a recap of each earlier consideration; never leave one out" passes while mandating the paragraph the gate forbids. The whole-file ban is restored and the contract states the rule without the word. **`unresolved_prior` had no relevance gate.** Its only condition was "they have not said yet", which contradicts `AGENTS.md` boundary 5 and `SKILL.md`'s own "Ask only decision-changing questions, then stop" -- and its twin `prior_decision` has had that gate since #609. Added on both surfaces. **Ranking a user's own holdings was reachable.** "Which of these three should I trim?" is both a candidate comparison the agent may rank and an ordering of the user's own positions, which is the engine's. Split into its own bullet: the ordering is `positions`', the recommendation is the agent's, and a comparison does not become the agent's to compute because its candidates are names the user already holds. Same bullet: "portfolio-derived" distributed over `cycle ID` and `ETF allocation exemption`, neither of which is a portfolio number, so the list is restructured; and deriving a figure one step from emitted figures is named as still the engine's. **#846 filed** for the one defect that is not this branch's: an `evaluation_id` of an unhashable type raises inside `_fold_evaluations` before any reader's checks run, taking down a live answer over a corrupt historical row. Reproduced on `main@8c37b04` through `prior_decision`'s identical fold, so it is filed rather than bundled -- the fix belongs to the shared reader and four callers. Product group 49/49, whole registry 60/60, `test_consider.py` 234/234. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This was referenced Sep 5, 2026
Owner
Author
|
CI on Not merging. The two things that would change the acceptance state are the owner's: a decision on this diff, and a named candidate SHA for #844 item 5 — for which |
This was referenced Sep 6, 2026
Open
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes nothing. Advances #827; supplies #844's item-2 follow-up. Two commits, one front.
The user problem
A user brings a decision to a skill installed inside their investing folder. Four things stopped the answer from being worth more than a general-purpose agent's, and all four are inherited scope rather than anything the product decided:
SKILL.mdtells the agent to rank candidates, label judgment, and take the deciding reason from the user's own record.references/agent-boundaries.md— whichSKILL.mdsays "holds throughout" — forbade "rankings", forbade "derived analysis", and forbade arguing a considered trade "from anything butconsider's output". A careful reader resolves that by doing less.--resolveneeds anevaluation_id, and the only place one was ever emitted was the response that minted it.The four changes
1. Each prohibition scopes to the fact it protects
references/agent-boundaries.md, three may not bullets and one new may bullet.consider's output"consider's outputuncheckedsays of itself "never enumerated" and #830 deleted the recital, so the prose was the surviving half of a deleted rule. It now says: named where the answer's own wording would imply it was checked, or where it could change the recommendation.AGENTS.mdboundary 5 carries the same widening. Every pinned integrity substring — CLI-only, hand-assembly, third-party/cloud privacy, never-relabel — is untouched.2. The increment gate reaches confidence and comparison, and a wait names its checkpoint
An owner amendment to
docs/expression-contract.md§3 through §3.4's first lane, logged in §8. Not a new V/D/C ID — the freeze says a shape fix is either an amendment to the mother chapter or an exemplar, and this is the first.The gate now asks: delete this block; can the reader still take the call, judge how far to trust it, and see why it beat the alternative they were weighing?
Paired with it: a wait stance owes what a directional call owes in its falsifier — the evidence that settles it and the next point it can be checked. A scheduled date is not itself that reason; an earnings date exists on every name every quarter.
This is not a volume licence, and the amendment says so in its own text. The four named bans are unchanged — restating, hedging, explaining a system default and inventing a scenario add nothing on any of the three. Confidence is a threshold, never an adverb; comparison is the one alternative in play, never a survey. No character-count ceiling: #543's stays deleted.
3. The record read reaches standing rules, and keeps what bounds a quote
SKILL.md"The user's own record comes first" and the mirroredagent-boundaries.mdbullets now:4. An unsettled consideration reaches the next one
review._unresolved_prior,_prior_decision's complement: at most one earlier consideration of the same ticker still atdecision: "open", projected as the direction, the day it was asked, the user's own stored words when both were supplied, and theevaluation_idthat closes it.It carries no
decisionand nodecided_on, because an open row is a question that was asked, never a decision and never proof of one — there is deliberately no key a reader could mistake for an answer the user never gave. A row with no stored context is still recalled, where the resolved projection drops it: the payload here is the open question, and the ticker, side, day and id are on any stored row.Its one use is to ask what the user did — once, only when they have not already said, never a recap. Their answer goes back through
--resolve, which is what makes itprior_decisionnext time.Emitted beside the row, never stored on it, absent rather than null. It changes no number, no
evaluation_id, nochallenge; deleting the key returns the response byte-for-byte.What was audited and left alone
Every remaining ceiling in the runtime tree, with why it stays:
flows/light-capture.mdflows/weekly-review.mdreferences/condition-slots.mdreferences/card-policy.mdreferences/weekly-market-read.mddoctoron first runSKILL.mdreferences/market-lookup.mddocs/expression-contract.md§3.1considerobligation floor (must_state/required_coverage)evaluation_challenge.build_challengeresolution_presentedevent and itsworkflow_statevocabularyreferences/ux-receipt.md,tools/ux_receipt.pyopen/acted/declined/modifiedvocabulary.unresolved_priorfeeds that same beat from the other direction; adding a field for it would be one written for a reader nobody built.Two smaller corrections in the same pass:
SKILL.mdstopped telling a reader to "follow" each reference's exemplar and now says what to take from it (its shape, never its length or section pattern) — which is what the corpus itself asserts an exemplar is evidence of, and whatfreeform-answers.mdalready says about its own upper witness. Andmarket-lookup.md's "Company research is in scope when…" paragraph was filed under the heading "When lookup does not happen"; it now has its own heading and decides scope by the same three-part test the amended gate uses.Verification
Product group 49/49; whole registry 60/60, locally, on Python 3.14.
tests/test_consider.py— section R (R1–R8 plus the compatibility claims): 234/234 in that module.Eighteen mutations run against
_unresolved_priorand its call site, across two rounds. Fourteen were caught by the first section R; four survived and are what independent review found, each now with its own test that was re-run against the mutation to confirm it reddens. A fifth — moving the read below_append_evaluation_row— survives by design: theevaluation_idexclusion is what prevents self-recall, not the call's placement, so the code says that instead of borrowing a justification_prior_decisionearns for a different reason.Independent review of the diff was run and its seven findings are all addressed in the third commit. Six were mine and are fixed; the seventh is pre-existing and is filed as [bug] A corrupt evaluation_id of an unhashable type takes down the whole consider answer #846 rather than bundled.
The loaded instruction chain, measured per SHA (always-on =
SKILL.md+agent-boundaries.md, the two files an installed host gets; routed = the five decision-route references):5e0c7b68c37b04The always-on surface grew ~2.8 KB. That is a real cost and it is prose, not deletion;
SKILL.mdnow sits at 11,866 of its 12,288-byte budget. Themust_statereduction tracking: current context index — milestone, critical path, what an agent may pick up #27 has queued next is where the routed 123 KB is addressed, and it is deliberately not bundled here.What independent review found, and what changed because of it
Seven findings, every one reproduced rather than reasoned. Six were defects in this branch; all six are fixed in the third commit.
docs/expression-contract.md§3.1's Middle row still read "Only support that could change the decision" — the pre-amendment gate, fourteen lines above the amendment. A reader following the table deletes exactly the block the amendment exists to keep. The mother law was giving two different tests on one page. Fixed; the 2026-09-05 ruling row also moved to the end of §8, whose log is ascending._fold_evaluationsfrom the open reader let one consultation reach the agent as both "you reported acting on this" and "you never told me what you did" — the disjointness is the fold's, not the twodecisiontests', and nothing said so. Worse, the assertion that should have caught it was a coin flip: two ids minted the same day, ordered lexicographically, so it passed or failed with the calendar. The other three: a non-stringevaluation_idcrashedconsideroutright while its guard sat untested (max()on a one-row fixture never compares); a call site hard-coding"buy"reversed the same-side preference for every sell the user brings, with no sell-side current consultation anywhere in section R; and looseningdecision != "open"projected a row carryingnullas a question the user was asked and never answered. Each now has a test, verified against its own mutation.unresolved_priorhad no relevance gate. Its only condition was they have not said yet, which contradictsAGENTS.mdboundary 5 andSKILL.md's own "Ask only decision-changing questions, then stop" — and its twinprior_decisionhas carried that gate since [impl·M2] Decision Continuity v1 — reuse one resolved prior TradeEvaluation #609. Three surfaces, two rules. Added on both.positions', the recommendation is the agent's, and a comparison does not become the agent's to compute because its candidates are names the user already holds.cycle IDandETF allocation exemptionin the rewritten bullet, neither of which is a portfolio number — a reader could argue the rule no longer reached them. The list is restructured, and deriving a figure one arithmetic step from emitted figures is named as still the engine's.evaluation_idof an unhashable type raises inside_fold_evaluationsbefore any reader's checks run, taking down a live answer over a corrupt historical row. Reproduced onmain@8c37b04throughprior_decision's identical fold, so it is filed rather than bundled: the fix belongs to the shared reader and its four callers.One observation worth keeping even though it needs no change here: the product gate is blind to the mirrored-surface class. The second commit's five files were all still un-synced at 645dcb1 and
--group productwas 49/49 green.What this does not prove
route.sh openwrites a cross-session ACTIVE marker on the owner's machine and the answers, the verdict and the archiving are all human steps needing a real decision. It was not run, and nothing here claims a usefulness result.FOMO_PINisb6bb069(2026-08-14) in the privateroute.sh:23, and [design·M1] The user's own investing folder is the evidence for why a name — install inside it, read it as the user's record, let the engine check the pick #844 records that the same value must move intools/route_fomo.py,templates/manifest.yamlandRUNBOOK §2.1together. Any run before that re-pin measures the August skill, exactly as the 2026-09-02 case did. Re-pinning is an owner decision on their own frozen comparison and was not made here.answer_provenancealready documents that limit. Auser_recordclaim'ssourceandas_ofare checked; whether the quote kept its qualifier is not, and cannot be by a regex. Scene 5 puts that under test where a human reads it.Privacy
Every added example is synthetic — the fictional-issuer convention,
mock/sample_momentum.csv, invented note text. No holding, ticker, amount, date, note content or owner-identifying detail from the private folder is in any changed file, and the private harness was read but never copied.🤖 Generated with Claude Code