Repository navigation
feat(model-roles): add a per-role model map to CE config - #1821
kieranklaassen wants to merge 55 commits into
Conversation
Requirements, technical decisions, and implementation units for a per-role model map in CE config. Co-Authored-By: Claude Code <noreply@anthropic.com>
One stdlib-only script resolves the new model_roles config key into a per-role answer: grammar, role-by-role layering across the personal and team files, model family, and the review egress policy. It is byte-duplicated into each role skill, with a parity test, because skills cannot share files. Co-Authored-By: Claude Code <noreply@anthropic.com>
One byte-duplicated reference states how a role skill honors its model_roles entry: resolving the role, precedence, when a hand-off happens, the serving order, the single-model fallback ladder, and the Model role report line. Skill bodies stay untouched because most sit near the 8,000-byte prompt budget. Adds the Model role map and Review seat glossary entries. Co-Authored-By: Claude Code <noreply@anthropic.com>
Adds the Model roles section to the config template and its example copy, and to the configuration guide: entry grammar, role-by-role layering, the older keys an entry replaces, the two personal opt-outs, and the controls on review seats. A test resolves the template example through the resolver. Co-Authored-By: Claude Code <noreply@anthropic.com>
The elevation engine consults the role map entry before plan_model or brainstorm_model, and its Claude CLI worker takes an optional effort (CE_ELEVATION_EFFORT, default high). An entry with an effort always hands off, because no host exposes the session effort. The result envelope records requested_effort. Co-Authored-By: Claude Code <noreply@anthropic.com>
A debug entry moves the investigation to a read-only subagent and the test-first fix to a second one on the named model. The session keeps the causal-chain gate, the branch, the pre-fix scope record, verification, and the commit. The structured returns gain one optional model_role field, which the lfg gate accepts but does not require. Co-Authored-By: Claude Code <noreply@anthropic.com>
ce-simplify-code hands its apply-and-verify steps to one subagent on the named model; the three reviewers keep their tier. ce-compound has a subagent draft the body into scratch from a hand-off file, while the orchestrator still classifies and remains the only writer under the artifact root. Both fall back to the session and say so, and print the Model role line. Co-Authored-By: Claude Code <noreply@anthropic.com>
A work entry is resolved before the work_engine_* keys and becomes the only standing candidate: served by native subagents when the host can hand over the model and effort, otherwise through the engine route for its family at prefer. An entry with an effort is never collapsed to native, a personal work_engine_mode: off keeps a team entry off other models, and the fallback ladder steps effort down through the levels the adapter accepts. No return field is added. Co-Authored-By: Claude Code <noreply@anthropic.com>
A step whose entry was replaced by a live instruction or carrier prints no Model role line, and a route that started and then failed goes to the fallback ladder the same way an unservable entry does. Both came up as ambiguities while wiring the plan and work roles. Co-Authored-By: Claude Code <noreply@anthropic.com>
lfg resolves no role itself: each child reads the map. When a child prints a Model role line, lfg carries it unchanged into the step narration and the closing recap. Each role skill guide gains a short Model role section linking to the configuration guide. Co-Authored-By: Claude Code <noreply@anthropic.com>
/ce-setup models shows each role with its effective value and source, offers the model families the host and installed CLIs can serve, accepts a typed model marked unconfirmed, asks whether to write the team or personal file and states each reach, previews the block, and writes only on approval. The rule that setup never creates config.local.yaml is scoped to the template-copy step, since this flow may write the personal file. Co-Authored-By: Claude Code <noreply@anthropic.com>
A list on the doc-review or code-review role runs one reviewer per seat on every review, beside the persona review and in place of the single conditional cross-model pass. The trigger sits in the references every dispatching review reads, so seats do not depend on a judgment lens activating. Review roles fail closed: an unreadable map, a blocked seat, or a seat no route can serve runs nothing extra and is reported, never substituted. The review workers take CROSS_MODEL_SEAT: it names the artifact and reviewer per seat, and lets an explicit seat run in the host family or under an unknown host while independence_verified stays false. Co-Authored-By: Claude Code <noreply@anthropic.com>
With several review seats, two verified peers could promote a finding with no in-process reviewer agreeing. A peer corroborates an in-process reading, so promotion now needs at least one of each. The agent-mode return shape also names coverage.model_roles where the rest of that shape is defined. Co-Authored-By: Claude Code <noreply@anthropic.com>
Twelve decision cells cover resolution, layering, precedence, engine routing, the personal opt-out, review seats on a routine plan, and the review-mode policy. Eight restraint cells, one per role skill, fail when a checkout with no model_roles key runs the resolver or prints a Model role line. Tasks ask for each skill ordinary work and never name the map. A judged conversation scenario covers the setup flow. Co-Authored-By: Claude Code <noreply@anthropic.com>
--hosts cursor runs a cell through cursor-agent -p in a trusted, sandboxed workspace, using the flags the bundled workers already run. It is never a default peer, because it needs a signed-in cursor-agent and bills its own account. Adds the unattested-host row for the review-mode policy, which is the case Cursor exists to check. Co-Authored-By: Claude Code <noreply@anthropic.com>
The unattested-host row was listed out of order, so the sorted comparison failed. Co-Authored-By: Claude Code <noreply@anthropic.com>
When a personal work_engine_mode: off keeps a team work entry native, the resolver now says so in warnings. In an eval on Codex the host read only the boolean and attributed the native run to the role map itself; with the reason in the answer, no host has to re-read the personal file to report it. Co-Authored-By: Claude Code <noreply@anthropic.com>
Round one on Claude and Codex showed three cells graded something other than the behavior under test. The restraint rows failed Claude for narrating that it checked for a map and found none; they now look for the Model role report line itself. The plan and brainstorm rows pinned the Claude CLI as the first route, which a host whose subagent tool can set effort does not owe; they now require a hand-off. The debug task told a literal host to stop before reading the settings the skill names. Co-Authored-By: Claude Code <noreply@anthropic.com>
In one Codex eval run on a checkout with no model_roles key, the host loaded the shared reference and ran the resolver to check for the key. The result was the same, but a no-map checkout must not need Python. Every hook now says what skipping means: do not read the reference or run its resolver. The ce-debug hand-off also created its run directory before it knew a hand-off would happen; the directory and hand-off file now belong to the hand-off. Co-Authored-By: Claude Code <noreply@anthropic.com>
On Codex a decide-and-stop review cell reported TEAM: none and counted every seat as not run, because nothing had been dispatched yet. The fields now ask what the review would dispatch or leave out if it continued. Co-Authored-By: Claude Code <noreply@anthropic.com>
A live ce-doc-review run on the Cursor host served three seats as Cursor subagents on the ids its subagent tool lists. The case records how to repeat that check. It stays hand-run because the model list changes too often to pin in a graded fixture. Co-Authored-By: Claude Code <noreply@anthropic.com>
… path In eval round four Codex missed the role map twice. The hook named the files as `.compound-engineering/config.local.yaml` or `config.yaml`, so one run looked for `config.yaml` at the repo root. Both runs listed files with a search that skips hidden directories and concluded no config existed. The step then ran on the session model and printed no `Model role` line, which is the silent failure the map must not have. The hook now gives both full paths and says to open them by path. Co-Authored-By: Claude Code <noreply@anthropic.com>
After the hook told agents to open both config files, Codex read the entry itself in four of twenty entry cells and skipped the shared reference and the resolver. Its answers were right on simple fixtures, but hand-reading skips layering, validation, and the review policy. The hook now says to run the resolver and that its answer decides the role. Co-Authored-By: Claude Code <noreply@anthropic.com>
A misspelled `cross_model_review_mode` falls through to `auto`, which leaves other-provider review seats unblocked with nothing said. The resolver now warns on review roles, and the skills already print resolver warnings with the `Model role` line. Also adds tests for the fail-closed paths review roles rely on: structural errors in the map and a config file the resolver cannot read. Drops a plan id from a comment in the shipped script. Co-Authored-By: Claude Code <noreply@anthropic.com>
The `code-review` role is resolved where the reviewer team is dispatched, which only the full review path reaches. A lite or focused review skipped a configured seat list with nothing said, while two guides promised seats on every review. Those paths now state in Coverage that the role was not applied and that `depth:full` applies it, and the guides say the same. Adds an eval cell for that case and pins a review base in the code-review role cells, where Codex once stopped at scope on a repo with one commit. Co-Authored-By: Claude Code <noreply@anthropic.com>
The setup-flow test parsed frontmatter by hand, the native hand-off test carried a third copy of the prompt-size helper to repeat a bound the budget test owns, and the eval provenance check retyped the host list. Co-Authored-By: Claude Code <noreply@anthropic.com>
…oject Two Codex slips remained after the hook said the resolver decides. In 2 of 16 no-map cells Codex ran the resolver anyway, so the hook now states the skip before the action. In 2 of 40 entry cells Codex ran the resolver from the skill directory, where it finds no repository and answers `unset`; the shared reference now says to run it with the project as the working directory. The lite-review eval cell no longer pins the depth, which has its own rows. Co-Authored-By: Claude Code <noreply@anthropic.com>
…llow it Codex missed the model role map four ways across the hook's wording changes: a bare file name, a file search that skips the hidden directory, reading the entry instead of running the resolver, and running the resolver from the skill directory. The learning records the wording that passed and the eval counts for each version. Co-Authored-By: Claude Code <noreply@anthropic.com>
The simplify cell stops before the hand-off, so the answer is a prediction about the host's subagent tool. Codex said yes in three of six runs. The cell now grades the entry's model and effort and the shared reference being read. Co-Authored-By: Claude Code <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
An id such as `openai/<gpt id>` or `openai.<gpt id>` passed the entry grammar but landed in no family, so a review seat could not use the Codex worker and a work entry got no engine. Both review workers already accept those forms as a Codex model override. The resolver now classifies a provider-qualified GPT or o-series id as Codex. Other ids with a slash stay in no family. Co-Authored-By: Claude Code <noreply@anthropic.com>
…as the hook's trigger The shared model role reference loses filler, sentences repeated within the file, and three of five example lines. Every rule, the resolver command, the state table, and the working-directory sentence are unchanged. The resolver's docstring is trimmed with no behavior change. The hook now says the resolver runs whenever the `model_roles:` key exists, even with no entry for this role. With a map that held another role's entry only, Codex had skipped the resolver and once reported that other role's model. Over five rounds on the new wording Codex had no such miss. Content the feature adds drops from 80,071 to 78,793 bytes. Codex passed 89 of 95 cells over five rounds, Claude 21 of 21, and Cursor 20 of 20. Co-Authored-By: Claude Code <noreply@anthropic.com>
… and counts Adds the fifth fact the hook needs (the trigger is the key, not an entry for this role), replaces the example with the wording that shipped, and replaces the one-round row with eleven rounds on that wording and five on the last. Co-Authored-By: Claude Code <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 68db22d413
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
|
||
| ## Model Roles | ||
|
|
||
| When the invocation asks to set up model roles (the `models` argument, or the same request in words), read `references/model-roles-setup.md` from this skill's directory and follow it in place of Phases 1-2. It writes the `model_roles` map only after the user approves, and it reports into Phase 3 (Summary). |
There was a problem hiding this comment.
Validate model roles during the standard health check
When a developer edits model_roles manually and then runs ordinary ce-setup, only the models branch invokes the new resolver; scripts/check-health contains no model-role validation. For example, a checkout with model_roles: opus is reported as Project config healthy, even though every consuming skill resolves the block as invalid. Include resolver diagnostics in the standard health path so the setup health check can detect malformed instances of this new config surface.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 7c405c8b81
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
The resolver placed a provider-qualified GPT or o-series id in Codex only when the provider name was letters, so `azure2/<gpt id>` landed in no family while both review workers accept it. A name already in another family is now decided first, and any other id with `.` or `/` before a GPT or o-series name is Codex, the same form the workers accept. Co-Authored-By: Claude Code <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: bdfb70043a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| *) skip "host harness '$HOST_HARNESS' invalid (want codex|claude|grok|cursor|opencode|unknown); skipping cross-model pass" ;; | ||
| esac | ||
| [ "$HOST_PROVIDER" != "unknown" ] || skip "host serving family unattested; automatic cross-model review skipped" | ||
| [ -n "$SEAT" ] || [ "$HOST_PROVIDER" != "unknown" ] || skip "host serving family unattested; automatic cross-model review skipped" |
There was a problem hiding this comment.
Exempt configured seats from the automatic peer cap
When CROSS_MODEL_MAX_PEERS=0 is set, a configured non-native review seat reaches this seat-specific path but is still skipped by the later MAX_PEERS gate. The review-seat contract in references/cross-model-review.md explicitly excludes the single-peer pass's activation/run conditions and says every otherwise valid seat becomes a job, so an environment used to disable the automatic pass silently disables configured seats too. Make seat mode bypass the automatic peer-count gate; the doc-review worker has the same mismatch.
Useful? React with 👍 / 👎.
The resolver put a provider-qualified GPT or o-series id in Codex for every role. The review workers accept that form, but the work engine's Codex route and its controller accept an unqualified id only, so a `work` entry with one was sent to an engine whose preflight refused it. The resolver now reports no engine harness for such a `work` entry. The entry is then served by a native subagent or the session, and the Model role line says why. The work engine's own model validation is unchanged. Co-Authored-By: Claude Code <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5b875cf722
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
A `work` entry such as `opus[1m] high` resolved to the Claude engine, whose adapter and controller reject a bracketed qualifier, so preflight refused it. The resolver now reports no engine harness for any `work` entry the engine's routes do not accept: an id with a bracketed qualifier, or a provider- qualified Codex id. Such an entry runs on a native subagent or the session, and its Model role line says why. Co-Authored-By: Claude Code <noreply@anthropic.com>
The role hook sat at the start of Phase 1, and the trivial-bug fast path leaves Phase 0 for the fix-choice gate before reaching it. A diagnosis-only run on that path printed nothing about a configured `debug` role. The hook now sits before the fast path, and the hand-off paragraph says that the fast path has no investigation to hand off while the fix hand-off still applies. Codex resolved the role in two of three runs of the entry cell and ran no resolver in three of three no-map runs; Claude and Cursor passed both cells. Co-Authored-By: Claude Code <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e778fc30a9
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
A seat that requested one Claude model and was served by another still published its artifact, and synthesis read it as the configured reviewer. A review seat is never filled by another model, so in seat mode a receipt that names a different model now leaves the seat with no artifact, and the worker logs why. The seat is then reported as dropped. Outside a seat the single-peer pass still publishes and warns, as before. Co-Authored-By: Claude Code <noreply@anthropic.com>
… too The hand-off paragraph kept an exception: on the trivial-bug fast path the session stated the cause, although the `debug` role's deliverable is the diagnosis and the fix. The exception is gone. The entry's model produces the diagnosis on every path, and a trivial bug changes only how deep it goes: the subagent states the cause and the proposed fix and stops. The hand-off rules now sit before the fast path, where that path reads them. Claude and Cursor passed both debug cells. Codex passed the entry cell in two of three runs and the no-map cell in three of three. Co-Authored-By: Claude Code <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0aa1b7f63f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
… worker A requested full id was compared to the served id by bare prefix, so `claude-opus-5` accepted a receipt for `claude-opus-50-*` and the worker reported `receipt: matched` for another model. The worker now matches the requested stem itself or the stem followed by `-`, in both the receipt check and the choice of which served id to report. Alias matching is unchanged. Co-Authored-By: Claude Code <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ef7dba5b60
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…model The review workers accepted a Claude alias with a bracketed qualifier through a shell glob, which also let `opus[]`, `opus[a][b]`, and `opus[1m]junk]` through to a provider call that would fail. The workers now require exactly one trailing qualifier of letters and digits, the grammar the role resolver accepts, and check the id before it against the existing patterns. Co-Authored-By: Claude Code <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 64d1c3623b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
When the `debug` or `simplify` role is served by a native subagent, that subagent edits project files, and the steps it is handed tell it to follow the project's instructions. A subagent may start without the instructions the session has loaded, so both hand-off prompts now carry the active project instructions and subdirectory-scoped instructions that govern the files it may change, as text or as the paths of the files to read first. Co-Authored-By: Claude Code <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c2ecc942c9
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
A role hand-off to a native subagent sent a prompt that held neither the project's instructions nor a usable path to the skill files its steps name. A probe that printed the investigation hand-off prompt showed both missing on Codex and Cursor. The previous commit patched this for two editing hand-offs only. The shared model-roles reference now states the rule for every native subagent: its prompt gives the project instructions that govern the work, and, when its steps name a file of the skill, the skill directory's absolute path to resolve those names against. The debug investigation, the debug fix, the simplify apply, and the compound draft prompts each point at that rule where the prompt is composed, and the per-prompt wording from the previous commit is removed. Co-Authored-By: Claude Code <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a406bad411
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
The role resolver lowercased a model id before classifying it, so `Opus` or `GPT-6.1-Sol` came back as a valid Claude or Codex entry. Every route's own validator matches ids case-sensitively, so the route then refused the id: a review seat was dropped and a single-model role fell back, with nothing telling the user why. The resolver now classifies ids as written. An id that only a lowercase spelling would route is skipped with a warning that names that spelling. An id that no route claims in any spelling keeps its own. Co-Authored-By: Claude Code <noreply@anthropic.com>
…in-f97951 # Conflicts: # skills/ce-work/references/execution-engines.md
…in-f97951 # Conflicts: # skills/ce-code-review/scripts/cross-model-adversarial-review.sh # skills/ce-doc-review/scripts/cross-model-doc-review.sh # skills/ce-pov/scripts/cross-model-pov.sh # skills/ce-work/references/cross-model-execution.md # tests/skill-eval-cell/packs/ce-doc-review-cross-model.md # tests/skills/ce-code-review-cross-model-routes.test.ts # tests/skills/ce-doc-review-cross-model-routes.test.ts # tests/skills/ce-work-cross-model-routes.test.ts
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2311893876
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
A review seat is one model, so a seat whose receipt names another model is not served and writes no artifact. That check ran on the Claude route only, because Claude was the one route with a receipt. The workers now read a served-model receipt for Codex and native Grok as well, so a seat on either is held to it the same way. A Codex id may carry a provider namespace such as `openai/` that the served id does not, so the namespace is set aside before the two are compared. Co-Authored-By: Claude Code <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: eb5432ba99
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…ers see them The resolver's full summary, which `ce-setup` shows as the current routing, was wrong in two cases. An older key with a closed vocabulary was reported from the first file that set it, even when the value was unsupported. The review skills skip an unsupported `cross_model_peer` and read the next file, so the summary now does the same. A review role whose list held an invalid seat showed an empty or shortened value with no warning. The summary now keeps each invalid seat as written, with the reason it does not run. Co-Authored-By: Claude Code <noreply@anthropic.com>
Summary
A repo can now say which model, at which reasoning effort, produces each pipeline step's deliverable, in one map in CE config. Every skill honors its entry whether a developer invokes it or
lfgdoes, in Claude Code, Codex, and Cursor. Before this, three unrelated key families covered three of the eight steps, none took an effort for every step, and a review could add at most one outside model.A checkout with no
model_roles:key behaves as it did and runs nothing new.What a role entry does
brainstorm,planworkce-workdoc-review,code-reviewdebug,simplify,compound<model> [<effort>],inherit, or a list on the two review roles. It names a model, never a CLI flag or an app-specific id.lfg's hand-offs stay on the session model.plan_model,brainstorm_model,cross_model_peer,work_engine_*), which keep working when the role is unset.Model role plan: requested fable low; served fable (unverified) at low; route claude CLI.lfgrelays those lines and lists them in its recap./ce-setup modelsshows each role's effective value and where it comes from, asks which file to write, previews the change, and writes only after approval.Design decisions worth reviewing
findings-mechanics.pynow requires an in-process reading for a peer to corroborate.cross_model_review_mode: offandCROSS_MODEL_PEERSapply to every seat, however it is served. The resolver decides this once and reportsblocked_byper seat. A misspelled review mode now produces a warning instead of silently meaningauto.skills/ce-setup/scripts/model-role-resolve.py.**Model role.**paragraph, for example inskills/ce-plan/references/reasoning-elevation.md. Its history is under Validation.Session-settled decisions carried from planning: one shared map for
lfgand standalone skills (user-approved, over separate files); the map lives in the two existing repo config files (user-approved, over a user-wide file); a single role map with one grammar (user-approved, over extending three key families); the map works in every supported app and Cursor follows pstack's native-subagent mechanism (user-directed, over a Cursor-only map); an entry governs the step's deliverable (user-approved, over delegating wholelfgstages); fallback differs by role kind (user-approved); an entry with an effort always hands off (user-approved).Validation
Mechanical.
bun run release:validateandbun run plugin:validatepass.bun run teston the final tree: 4,423 pass, 15 fail. Fourteen failures are in the cross-model normalization tests and come from the authoring machine's jq 1.6; the same fourteen fail there on an untouched copy ofmain. The fifteenth isce-work workspace harness: makeRepo reseeds when the cached template directory has gone missing, which passes when its file runs alone. CI is the gate for both.Behavioral. 22 graded cells run a fresh agent with the on-disk skill injected (
bun run test:skill-eval-pack, post arm). Each entry cell has a matching no-map cell, so a cell fails both when the route is skipped with a map and when it is taken without one.cursor-agent)config.yaml, read the entry itself instead of running the resolver, ran the resolver from the skill directory, and ran it when no map existed. Each fix is its own commit, anddocs/solutions/skill-design/config-hook-states-path-decider-and-working-directory.mdrecords the counts per wording.ce-doc-reviewrun in Cursor served three seats as Cursor subagents on the ids its subagent tool lists. It is hand-run case 19 intests/skill-eval-cell/packs/ce-doc-review-cross-model.md.ce-setup/model-roles-flow): Claude passed every rubric item. Codex wrote the same correct file, but the harness records only Codex's last message per turn, so its role table and preview went ungraded.Unapplied review findings
cursoris listed, matching the resolver and the guide. Another provider's seat is still refused. Tests setCROSS_MODEL_SEATandCROSS_MODEL_PEERStogether in both worker suites.opus[1m]. Fixed after review: one trailing bracketed qualifier is accepted, the family comes from the part before it, and the review workers accept a bracketed Claude alias.model_roleswere reported only by--all. Every resolver answer now carries them, so a misspelled role is reported by the skill it was meant for.skills/ce-setup/scripts/check-health— the standard setup health check does not validatemodel_rolesA checkout with a malformed map, such as
model_roles: opus, still readsProject config healthy. The plan left this out on purpose: every governed step prints the resolver's errors for its role, each review names its recipients, and/ce-setup modelsshows effective values. A review bot asked for it. Building it means calling the bundled resolver with--allfrom the health script and reporting itserrorsas a failed check.CROSS_MODEL_MAX_PEERS=0also stops configured review seatsBoth review workers skip with
CROSS_MODEL_MAX_PEERS=0; cross-model pass disabled, in seat mode too, so the seat is dropped and reported. A review bot asked for seats to bypass that gate. It was left as is: the variable reads as a switch that stops review content leaving the machine, and a configured seat overriding it would send content the environment said not to send.A
workentry with a bracketed qualifier (opus[1m]) or a provider-qualified Codex id gets no engine and runs on a native subagent or the session. Supporting those on the engine means widening the model validation inskills/ce-work/scripts/cross-model-work.shand its controller, which this PR does not touch.skills/ce-code-review/references/depth-paths.md—code-reviewseats run only on the full review pathA small diff that takes the lite or focused path applies no
code-reviewentry. This follows the plan assumption "Seats run when a review dispatches its persona team", which the maintainer has not confirmed. This PR makes those paths say so in Coverage and corrects two guides that promised seats on every review. The alternative is to make a configured seat list lift every review to at least the focused path, which calls every seat on every small review.cross_model_review_mode: off, a seat in the attested host family is exempt whatever route serves it. On Cursor with an attested family, such a seat can go out through the standalone family CLI. This matches the by-provider rule in the plan; closing it means exempting only when the family's harness is the host harness.cross_model_review_mode: offreaches a worker-served seat only because the agent honors the resolver'sblocked_by. Neither worker reads the key.skills/ce-code-review/scripts/run-log.py:151(unchanged) records one seat's token usage when two Codex worker seats run.ce-babysit-prdoes not relay a child'sModel roleline. The older model keys (plan_model,brainstorm_model,cross_model_peer,cross_model_model,cross_model_effort,work_engine_*) stay, with no removal trigger. In non-interactive andmode:agentreviews the pre-send seat notice is not printed.ce-config-layersblock, in eleven files across ten skills, still names the second config file as bareconfig.yaml, the wording that made Codex miss the config here.Review run:
20261001-165637-ecfed90e, artifacts at/tmp/compound-engineering-501/ce-code-review/20261001-165637-ecfed90e(verdict: Ready with fixes; no P0 or P1). Reviewers: correctness, security, testing, maintainability, project standards, agent-native, learnings, and an independent adversarial pass requested from Codex (gpt-6-lunaatxhigh, served model unverified).Security Disclosure
scripts/model-role-resolve.pyin nine skills. It reads the two repo CE config files and theCROSS_MODEL_PEERSenvironment variable, runsgit rev-parse --show-toplevel, and prints JSON. It writes nothing and makes no network call.cross_model_review_mode: offandCROSS_MODEL_PEERSapply to every seat. WithCROSS_MODEL_SEATset, the two workers accept a target in the host's own family and an unknown host family; they still checkCROSS_MODEL_PEERS. In seat mode the workers clear a target in the attested host's own family, and Composer whencursoris listed, the same two cases the resolver clears.ce-setup modelscan write.compound-engineering/config.local.yamlorconfig.yaml, only after the user approves a preview.cursorhost runscursor-agentwith--trust --sandbox enabled --forceinside a throwaway workspace. It runs only when named with--hosts.Agent Disclosure
🤖 Generated with Claude Code