Skip to content

Move the limits onto the claim, let the answer support a reader's confidence, and close the continuity loop on the route that opens it - #847

Merged
atomchung merged 3 commits into
mainfrom
claude/fomo-kernel-judgment-redesign-7a0073
Sep 5, 2026
Merged

atomchung merged 3 commits into
mainfrom
claude/fomo-kernel-judgment-redesign-7a0073

Conversation

@atomchung

Copy link
Copy Markdown
Owner

Closes nothing. Advances #827; supplies #844's item-2 follow-up. Two commits, one front.

The user problem

A user brings a decision to a skill installed inside their investing folder. Four things stopped the answer from being worth more than a general-purpose agent's, and all four are inherited scope rather than anything the product decided:

  1. The always-on boundary file forbade what the entry mandates. SKILL.md tells the agent to rank candidates, label judgment, and take the deciding reason from the user's own record. references/agent-boundaries.md — which SKILL.md says "holds throughout" — forbade "rankings", forbade "derived analysis", and forbade arguing a considered trade "from anything but consider's output". A careful reader resolves that by doing less.
  2. The shape law deleted the support a reader needs. The increment gate asked: delete this block; does the decision change? A block that lets the reader check a call, or that says why this pick beat the one they came in holding, flips no action — so the literal test deleted it, and the answer kept its stance while losing every reason anyone could weigh.
  3. The record read went to the ticker page and stopped. The reader named the per-name thesis, falsifiers, open questions, prior decisions and stances. It never named the user's standing decision rules, and it said nothing about carrying a passage's qualifier.
  4. The continuity loop could not be closed on the route that opens it. --resolve needs an evaluation_id, and the only place one was ever emitted was the response that minted it.

The four changes

1. Each prohibition scopes to the fact it protects

references/agent-boundaries.md, three may not bullets and one new may bullet.

Was Is The failure it protects
"Calculate or alter numbers, rankings, weights…" a portfolio-derived number, ranking, weight… A book figure the agent computed is not reproducible next week, and reconciliation rests on this week's number and next week's meaning the same thing. The engine ranks the user's positions; a preference among candidates they asked you to compare is judgment, labelled, and enters no canonical state.
"derived analysis is not [allowed]" deriving a portfolio figure the engine did not emit is not Same failure. Read literally the old form forbade every valuation judgment and every thesis — the product's own job.
"Argue a considered trade's case from anything but consider's output" state its portfolio consequence from anything but consider's output A portfolio claim with nothing deterministic behind it. The case is argued from the user's record, sourced facts and labelled judgment — which is what the entry already said.
"the risks the engine does not measure are named rather than passed over" kept, restated as a materiality test Silence about an unchecked risk reads as a clean bill of health. But unchecked says of itself "never enumerated" and #830 deleted the recital, so the prose was the surviving half of a deleted rule. It now says: named where the answer's own wording would imply it was checked, or where it could change the recommendation.

AGENTS.md boundary 5 carries the same widening. Every pinned integrity substring — CLI-only, hand-assembly, third-party/cloud privacy, never-relabel — is untouched.

2. The increment gate reaches confidence and comparison, and a wait names its checkpoint

An owner amendment to docs/expression-contract.md §3 through §3.4's first lane, logged in §8. Not a new V/D/C ID — the freeze says a shape fix is either an amendment to the mother chapter or an exemplar, and this is the first.

The gate now asks: delete this block; can the reader still take the call, judge how far to trust it, and see why it beat the alternative they were weighing?

Paired with it: a wait stance owes what a directional call owes in its falsifier — the evidence that settles it and the next point it can be checked. A scheduled date is not itself that reason; an earnings date exists on every name every quarter.

This is not a volume licence, and the amendment says so in its own text. The four named bans are unchanged — restating, hedging, explaining a system default and inventing a scenario add nothing on any of the three. Confidence is a threshold, never an adverb; comparison is the one alternative in play, never a survey. No character-count ceiling: #543's stays deleted.

3. The record read reaches standing rules, and keeps what bounds a quote

SKILL.md "The user's own record comes first" and the mirrored agent-boundaries.md bullets now:

  • name the user's standing decision rules alongside the per-name record;
  • state where to start — the names in play and those rules, then what they name; no whole-folder sweep and no filename convention;
  • require the quote to carry what bounds it: a condition's threshold, a thesis's falsifier, whether a stance authorized acting yet. Quoting a line the user marked not-yet-actionable as an action basis is a misquote, not a summary;
  • prefer their newest statement, name a genuine conflict rather than silently picking a winner;
  • make their recorded preference the default and a departure arguable — from the evidence that changed, never by overwriting what they wrote and never by telling them they keep making the same mistake.

4. An unsettled consideration reaches the next one

review._unresolved_prior, _prior_decision's complement: at most one earlier consideration of the same ticker still at decision: "open", projected as the direction, the day it was asked, the user's own stored words when both were supplied, and the evaluation_id that closes it.

It carries no decision and no decided_on, because an open row is a question that was asked, never a decision and never proof of one — there is deliberately no key a reader could mistake for an answer the user never gave. A row with no stored context is still recalled, where the resolved projection drops it: the payload here is the open question, and the ticker, side, day and id are on any stored row.

Its one use is to ask what the user did — once, only when they have not already said, never a recap. Their answer goes back through --resolve, which is what makes it prior_decision next time.

Emitted beside the row, never stored on it, absent rather than null. It changes no number, no evaluation_id, no challenge; deleting the key returns the response byte-for-byte.

What was audited and left alone

Every remaining ceiling in the runtime tree, with why it stays:

Kept Where Why
One motive question per trade flows/light-capture.md The light tier's whole definition. Review lifecycle, not the decision route.
One exit-backlog sentence flows/weekly-review.md Stops the weekly brief becoming an interrogation list.
One condition per thesis references/condition-slots.md A storage rule for a stored slot, not a conversation rule.
One card invitation references/card-policy.md The checklist-shape defect #623 owns.
Three names in the weekly brief references/weekly-market-read.md A bounded prototype brief.
doctor on first run SKILL.md The engine fail-softs silently without its optional dependencies, dropping prices and market context. Verify once rather than mid-answer.
Research stop discipline references/market-lookup.md Already marginal-value against cost and latency, with no numeric ceiling. #827's requirement was already met here.
The four floors, and a one-sentence top docs/expression-contract.md §3.1 Reviewed and kept. The floors are the pyramid's delivery, not an extra template; a stance plus the reason that decides it does fit one sentence, and the corpus exemplars demonstrate it. No observed failure argues otherwise, and changing a rule with no named failure is what this pass is against.
The consider obligation floor (must_state / required_coverage) evaluation_challenge.build_challenge Untouched. It is computed per call, it is the one thing that stops a brief answer dropping a rule collision, and shrinking it is #27's separately approved next cut. A live call on the mock book emits nine owed facts, six available, one never rendered.
The receipt's resolution_presented event and its workflow_state vocabulary references/ux-receipt.md, tools/ux_receipt.py Untouched, and deliberately not extended. It already records that an invitation was shown and what it left the evaluation as, in the engine's own open/acted/declined/modified vocabulary. unresolved_prior feeds that same beat from the other direction; adding a field for it would be one written for a reader nobody built.

Two smaller corrections in the same pass: SKILL.md stopped telling a reader to "follow" each reference's exemplar and now says what to take from it (its shape, never its length or section pattern) — which is what the corpus itself asserts an exemplar is evidence of, and what freeform-answers.md already says about its own upper witness. And market-lookup.md's "Company research is in scope when…" paragraph was filed under the heading "When lookup does not happen"; it now has its own heading and decides scope by the same three-part test the amended gate uses.

Verification

  • Product group 49/49; whole registry 60/60, locally, on Python 3.14.

  • tests/test_consider.py — section R (R1–R8 plus the compatibility claims): 234/234 in that module.

  • Eighteen mutations run against _unresolved_prior and its call site, across two rounds. Fourteen were caught by the first section R; four survived and are what independent review found, each now with its own test that was re-run against the mutation to confirm it reddens. A fifth — moving the read below _append_evaluation_row — survives by design: the evaluation_id exclusion is what prevents self-recall, not the call's placement, so the code says that instead of borrowing a justification _prior_decision earns for a different reason.

  • Independent review of the diff was run and its seven findings are all addressed in the third commit. Six were mine and are fixed; the seventh is pre-existing and is filed as [bug] A corrupt evaluation_id of an unhashable type takes down the whole consider answer #846 rather than bundled.

  • The loaded instruction chain, measured per SHA (always-on = SKILL.md + agent-boundaries.md, the two files an installed host gets; routed = the five decision-route references):

    SHA Always-on Routed Sum
    5e0c7b6 14,113 119,337 133,450
    8c37b04 15,929 120,803 136,732
    this branch 18,735 123,022 141,757

    The always-on surface grew ~2.8 KB. That is a real cost and it is prose, not deletion; SKILL.md now sits at 11,866 of its 12,288-byte budget. The must_state reduction tracking: current context index — milestone, critical path, what an agent may pick up #27 has queued next is where the routed 123 KB is addressed, and it is deliberately not bundled here.

What independent review found, and what changed because of it

Seven findings, every one reproduced rather than reasoned. Six were defects in this branch; all six are fixed in the third commit.

  1. The amendment contradicted its own summary table. docs/expression-contract.md §3.1's Middle row still read "Only support that could change the decision" — the pre-amendment gate, fourteen lines above the amendment. A reader following the table deletes exactly the block the amendment exists to keep. The mother law was giving two different tests on one page. Fixed; the 2026-09-05 ruling row also moved to the end of §8, whose log is ascending.
  2. Four mutations survived section R. Dropping _fold_evaluations from the open reader let one consultation reach the agent as both "you reported acting on this" and "you never told me what you did" — the disjointness is the fold's, not the two decision tests', and nothing said so. Worse, the assertion that should have caught it was a coin flip: two ids minted the same day, ordered lexicographically, so it passed or failed with the calendar. The other three: a non-string evaluation_id crashed consider outright while its guard sat untested (max() on a one-row fixture never compares); a call site hard-coding "buy" reversed the same-side preference for every sell the user brings, with no sell-side current consultation anywhere in section R; and loosening decision != "open" projected a row carrying null as a question the user was asked and never answered. Each now has a test, verified against its own mutation.
  3. The recap gate became word-gameable. My narrowing — every line mentioning a recap must also say never — is defeated by this file's own house style: "Open every answer with a recap of each earlier consideration; never leave one out" passes a word-adjacency check while mandating the paragraph the gate forbids. The reviewer wrote that sentence and it passed. The whole-file ban is restored and the contract states the rule without the word.
  4. unresolved_prior had no relevance gate. Its only condition was they have not said yet, which contradicts AGENTS.md boundary 5 and SKILL.md's own "Ask only decision-changing questions, then stop" — and its twin prior_decision has carried that gate since [impl·M2] Decision Continuity v1 — reuse one resolved prior TradeEvaluation #609. Three surfaces, two rules. Added on both.
  5. Ranking a user's own holdings was reachable. "Which of these three should I trim?" is simultaneously a candidate comparison the agent may rank and an ordering of the user's own positions, which is the engine's — and nothing excluded the overlap. Split into its own bullet: the ordering is positions', the recommendation is the agent's, and a comparison does not become the agent's to compute because its candidates are names the user already holds.
  6. "portfolio-derived" distributed over cycle ID and ETF allocation exemption in the rewritten bullet, neither of which is a portfolio number — a reader could argue the rule no longer reached them. The list is restructured, and deriving a figure one arithmetic step from emitted figures is named as still the engine's.
  7. [bug] A corrupt evaluation_id of an unhashable type takes down the whole consider answer #846 — not this branch's. An evaluation_id of an unhashable type raises inside _fold_evaluations before any reader's checks run, taking down a live answer over a corrupt historical row. Reproduced on main@8c37b04 through prior_decision's identical fold, so it is filed rather than bundled: the fix belongs to the shared reader and its four callers.

One observation worth keeping even though it needs no change here: the product gate is blind to the mirrored-surface class. The second commit's five files were all still un-synced at 645dcb1 and --group product was 49/49 green.

What this does not prove

Privacy

Every added example is synthetic — the fictional-issuer convention, mock/sample_momentum.csv, invented note text. No holding, ticker, amount, date, note content or owner-identifying detail from the private folder is in any changed file, and the private harness was read but never copied.

🤖 Generated with Claude Code

test and others added 3 commits September 5, 2026 15:43
…omparison, and an unsettled consideration reaches the next one (refs #827, #844)

Four connected changes on one front. The always-on boundary file contradicted
the entry it loads beside; the shape law deleted the support a reader needs to
check a call; the record reader never named the user's standing rules; and the
continuity loop could not be closed on the route that opens it.

1. Scope each prohibition to the fact it protects. `agent-boundaries.md`
   forbade "rankings" and "derived analysis" unscoped and forbade arguing a
   considered trade "from anything but consider's output", while `SKILL.md`
   mandates candidate ranking, labelled judgment and a deciding reason from the
   user's own record. Now: a *portfolio-derived* number or ranking, a
   *portfolio* figure the engine did not emit, a considered trade's *portfolio
   consequence*. A risk the engine did not measure is named on materiality, not
   enumerated -- which is what `unchecked` already says of itself.

2. Amend the expression contract's increment gate (owner amendment through
   §3.4's first lane, logged in §8). "Does the decision change?" deleted every
   block that let a reader check the call or see why one candidate beat
   another. It now asks whether the reader can still take the call, judge how
   far to trust it, and see why it beat the alternative in play. Paired: a
   *wait* stance owes the evidence that settles it and the next checkable
   point, since a scheduled date is not a reason. The four bans, D1-D7, C1-C4,
   the registry freeze and the refusal of any character-count ceiling stand.

3. The record read reaches standing decision rules, states where to start
   (the names in play and those rules, then what they name -- no whole-folder
   sweep, no filename convention), and keeps what bounds a quote: a condition's
   threshold, a thesis's falsifier, whether a stance authorized acting. Prefer
   the newest statement, name a real conflict, argue a departure from their
   preference rather than overwriting it.

4. `review._unresolved_prior` projects at most one unsettled same-ticker
   consideration onto the `consider` response, carrying the `evaluation_id`
   that `--resolve` needs. That identifier was previously emitted only in the
   response that minted it, so a consideration left open on this route could
   never be closed on it -- and `_prior_decision` stayed permanently empty for
   that ticker. It carries no `decision` and no `decided_on`, because an open
   row is a question that was asked, never a decision.

Tests: `tests/test_consider.py` section R (R1-R8 plus compatibility), fourteen
mutations run against the new reader and its call site. Section Q's whole-file
recap ban is narrowed to require every mention be a prohibition, since the
contract now bans a recap outright on two fields. Product group 49/49 and the
whole registry 60/60 pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…827, #844)

Four mirrored surfaces the first commit left behind, and one misfiled
permission.

- The three README install paragraphs name the standing decision rules and the
  qualifier rule. That paragraph is what tells a user what the folder is read
  for, so it drifting is the same defect as the contract drifting.
- `skills/fomo-kernel/evals/evals.json` gains scene 5: the decision route,
  which the bank's four review-lifecycle scenes never covered. A returning user
  with an unsettled earlier consideration, a standing rule, and a note whose
  observation carries a not-yet-actionable qualifier. Its expectations name
  reading the standing rule and preserving the qualifier as what is under test,
  and leave the verdict to owner dogfood rather than the automated score.
- `references/market-lookup.md`'s "Company research is in scope when..."
  paragraph sat under the heading "When lookup does not happen" -- a permission
  filed under the heading that denies it, which is exactly the entry/reference
  ambiguity #827 exists to remove. It now has its own heading and decides scope
  by the same three-part test the amended increment gate uses, so what may be
  looked up and what reaches the answer are one rule rather than two.
- The maintainer guide's three 2026-09-05 rows name all of the above.

Product group 49/49.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…, four mutations survived, and a new field had no relevance gate (refs #827, #844, #846)

Seven findings, all confirmed by reproduction. Six are mine.

**The amendment contradicted its own summary table.** `docs/expression-contract.md`
§3.1's Middle row still read "Only support that could change the decision" --
the pre-amendment gate, fourteen lines above the amendment. A reader following
the table deletes exactly the block the amendment exists to keep. Fixed, and the
2026-09-05 ruling row moved to the end of §8, whose log is ascending.

**Four mutations survived section R.** Each now has a test, and each was
re-run against the new tests to confirm it reddens:

- Dropping `_fold_evaluations` from the open reader. A resolution appends a
  second row under the same `evaluation_id`, so one consultation reached the
  agent as both "you reported acting on this" and "you never told me what you
  did". The disjointness is the fold's, not the two `decision` tests'. The
  assertion that should have caught it was a coin flip: two ids minted the same
  day, ordered lexicographically, so it passed or failed with the calendar.
  Replaced with a seeded-date fixture that asserts the property directly.
- Dropping the `isinstance(evaluation_id, str)` guard. `max()` on a one-row
  fixture never compares, so the guard was present and untested while a
  non-string id crashed `consider` outright. R8 now pairs a broken row with a
  healthy one, both dated the same day, forcing the comparison onto identity.
- Hard-coding `"buy"` at the call site. No test ran a sell-side *current*
  consultation, so the same-side preference reversed for every sell the user
  brings, silently.
- Loosening `decision != "open"` to `not in PRIOR_DECISION_RESOLVED`, which
  projects a row carrying `null` or junk as a question the user was asked and
  never answered.

**The recap gate became word-gameable.** Narrowing it to "every line mentioning
a recap must also say never" is defeated by this file's own house style --
"Open every answer with a recap of each earlier consideration; never leave one
out" passes while mandating the paragraph the gate forbids. The whole-file ban
is restored and the contract states the rule without the word.

**`unresolved_prior` had no relevance gate.** Its only condition was "they have
not said yet", which contradicts `AGENTS.md` boundary 5 and `SKILL.md`'s own
"Ask only decision-changing questions, then stop" -- and its twin
`prior_decision` has had that gate since #609. Added on both surfaces.

**Ranking a user's own holdings was reachable.** "Which of these three should I
trim?" is both a candidate comparison the agent may rank and an ordering of the
user's own positions, which is the engine's. Split into its own bullet: the
ordering is `positions`', the recommendation is the agent's, and a comparison
does not become the agent's to compute because its candidates are names the
user already holds. Same bullet: "portfolio-derived" distributed over `cycle ID`
and `ETF allocation exemption`, neither of which is a portfolio number, so the
list is restructured; and deriving a figure one step from emitted figures is
named as still the engine's.

**#846 filed** for the one defect that is not this branch's: an `evaluation_id`
of an unhashable type raises inside `_fold_evaluations` before any reader's
checks run, taking down a live answer over a corrupt historical row. Reproduced
on `main@8c37b04` through `prior_decision`'s identical fold, so it is filed
rather than bundled -- the fix belongs to the shared reader and four callers.

Product group 49/49, whole registry 60/60, `test_consider.py` 234/234.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@atomchung

Copy link
Copy Markdown
Owner Author

CI on 91b937d: Actions run 33955286427 — product-contract and qa-eval-tooling pass on Python 3.11 and 3.12; network-smoke skipped. That is supported-version evidence for the code, not a live market or usefulness claim.

Not merging. The two things that would change the acceptance state are the owner's: a decision on this diff, and a named candidate SHA for #844 item 5 — for which FOMO_PIN has to move off b6bb069 first.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant