Skip to content

feat(model-roles): add a per-role model map to CE config - #1821

Open
kieranklaassen wants to merge 55 commits into
mainfrom
claude/pull-latest-main-f97951
Open

kieranklaassen wants to merge 55 commits into
mainfrom
claude/pull-latest-main-f97951

Conversation

@kieranklaassen

@kieranklaassen kieranklaassen commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

A repo can now say which model, at which reasoning effort, produces each pipeline step's deliverable, in one map in CE config. Every skill honors its entry whether a developer invokes it or lfg does, in Claude Code, Codex, and Cursor. Before this, three unrelated key families covered three of the eight steps, none took an effort for every step, and a review could add at most one outside model.

# .compound-engineering/config.yaml (team) or config.local.yaml (personal, wins per role)
model_roles:
  brainstorm: fable low
  plan: fable low
  doc-review: [grok-4.7 high, gpt-6.1-sol, opus medium]
  work: opus medium

A checkout with no model_roles: key behaves as it did and runs nothing new.

What a role entry does

Role Deliverable the entry governs How it is served
brainstorm, plan generated approaches, the authored plan native subagent, else the Claude CLI route
work the code one preferred engine candidate for ce-work
doc-review, code-review one independent review per list item (a seat) native subagent, else that skill's review worker
debug, simplify, compound diagnosis and fix, applied simplification, the learning's body native subagent only
  • An entry is <model> [<effort>], inherit, or a list on the two review roles. It names a model, never a CLI flag or an app-specific id.
  • Dialogue, orchestration, scouts, and lfg's hand-offs stay on the session model.
  • A live instruction or a caller's carrier outranks the entry. The entry outranks the older keys for that step (plan_model, brainstorm_model, cross_model_peer, work_engine_*), which keep working when the role is unset.
  • Every governed step prints one line: Model role plan: requested fable low; served fable (unverified) at low; route claude CLI. lfg relays those lines and lists them in its recap.
  • /ce-setup models shows each role's effective value and where it comes from, asks which file to write, previews the change, and writes only after approval.

Design decisions worth reviewing

  • Fallback differs by role kind. A single-model role that cannot be served falls back within the model's family, then to the session, and says so. A review seat that cannot be served is dropped and reported, never run on another model, because a substituted seat weakens the independence the list exists for.
  • Seats replace the single cross-model pass and run beside the persona reviewers. Agreement among seats alone never raises a finding's confidence. findings-mechanics.py now requires an in-process reading for a peer to corroborate.
  • cross_model_review_mode: off and CROSS_MODEL_PEERS apply to every seat, however it is served. The resolver decides this once and reports blocked_by per seat. A misspelled review mode now produces a warning instead of silently meaning auto.
  • An entry with an effort always hands off, because a subagent tool that cannot set effort would otherwise drop it.
  • One resolver script, copied byte-identical into nine skills, owns layering, validation, and the review policy. Skills cannot reference each other's files, so the copies are required and a parity test guards them. Review one copy: skills/ce-setup/scripts/model-role-resolve.py.
  • The hook wording was set by evals, not by reading. The reviewer would otherwise start in the wrong place: the sentence that matters is the **Model role.** paragraph, for example in skills/ce-plan/references/reasoning-elevation.md. Its history is under Validation.

Session-settled decisions carried from planning: one shared map for lfg and standalone skills (user-approved, over separate files); the map lives in the two existing repo config files (user-approved, over a user-wide file); a single role map with one grammar (user-approved, over extending three key families); the map works in every supported app and Cursor follows pstack's native-subagent mechanism (user-directed, over a Cursor-only map); an entry governs the step's deliverable (user-approved, over delegating whole lfg stages); fallback differs by role kind (user-approved); an entry with an effort always hands off (user-approved).

Validation

Mechanical. bun run release:validate and bun run plugin:validate pass. bun run test on the final tree: 4,423 pass, 15 fail. Fourteen failures are in the cross-model normalization tests and come from the authoring machine's jq 1.6; the same fourteen fail there on an untouched copy of main. The fifteenth is ce-work workspace harness: makeRepo reseeds when the cached template directory has gone missing, which passes when its file runs alone. CI is the gate for both.

Behavioral. 22 graded cells run a fresh agent with the on-disk skill injected (bun run test:skill-eval-pack, post arm). Each entry cell has a matching no-map cell, so a cell fails both when the route is skipped with a map and when it is taken without one.

Host Final content, cells passed Earlier rounds, all hook wordings
Claude Code 21 of 21 153 of 154
Codex 89 of 95 over five rounds 87 of 102 before the hook settled
Cursor (cursor-agent) 20 of 20 98 of 98
  • The evals changed the hook wording four times. Codex missed the config when the hook wrote a bare config.yaml, read the entry itself instead of running the resolver, ran the resolver from the skill directory, and ran it when no map existed. Each fix is its own commit, and docs/solutions/skill-design/config-hook-states-path-decider-and-working-directory.md records the counts per wording.
  • After the feature worked, its content was cut by 12% (89,754 to 78,793 bytes of skill prose and script) in four measured steps: text that restated the shared reference was removed from each skill, and the shared reference lost its filler. Each step was kept only after five Codex rounds, one Claude round, and one Cursor round. Codex passed 89 of 95 before the cuts and 89 of 95 after.
  • The cells stop at a decision. They do not run a fallback, dispatch a seat, or exercise a failed hand-off, so they show that resolution, routing, restraint, and reporting survived the cuts. Rules about the unexercised paths are now stated once, in the shared reference.
  • The six Codex misses in the last five rounds are one-off slips in different cells, plus the small-review Coverage note, which Codex states in about three runs of five.
  • A live ce-doc-review run in Cursor served three seats as Cursor subagents on the ids its subagent tool lists. It is hand-run case 19 in tests/skill-eval-cell/packs/ce-doc-review-cross-model.md.
  • Judged setup conversation (ce-setup/model-roles-flow): Claude passed every rubric item. Codex wrote the same correct file, but the harness records only Codex's last message per turn, so its role table and preview went ungraded.
  • Not run: Grok and OpenCode hosts, and the hand-run seat packs other than case 19.
  • Browser test: nothing to exercise. No web route changed.

Unapplied review findings

  • P2 — Worker allowlist dropped seats the resolver cleared. Fixed after review: in seat mode both workers now clear a seat in the attested host's own family and a Composer seat when cursor is listed, matching the resolver and the guide. Another provider's seat is still refused. Tests set CROSS_MODEL_SEAT and CROSS_MODEL_PEERS together in both worker suites.
  • P3 — Resolver rejected bracket-suffixed model ids such as opus[1m]. Fixed after review: one trailing bracketed qualifier is accepted, the family comes from the part before it, and the review workers accept a bracketed Claude alias.
  • Review comment — unknown keys under model_roles were reported only by --all. Every resolver answer now carries them, so a misspelled role is reported by the skill it was meant for.
  • Open decision — skills/ce-setup/scripts/check-health — the standard setup health check does not validate model_roles
    A checkout with a malformed map, such as model_roles: opus, still reads Project config healthy. The plan left this out on purpose: every governed step prints the resolver's errors for its role, each review names its recipients, and /ce-setup models shows effective values. A review bot asked for it. Building it means calling the bundled resolver with --all from the health script and reporting its errors as a failed check.
  • Open decision — CROSS_MODEL_MAX_PEERS=0 also stops configured review seats
    Both review workers skip with CROSS_MODEL_MAX_PEERS=0; cross-model pass disabled, in seat mode too, so the seat is dropped and reported. A review bot asked for seats to bypass that gate. It was left as is: the variable reads as a switch that stops review content leaving the machine, and a configured seat overriding it would send content the environment said not to send.
  • Open decision — qualified model ids on the work engine
    A work entry with a bracketed qualifier (opus[1m]) or a provider-qualified Codex id gets no engine and runs on a native subagent or the session. Supporting those on the engine means widening the model validation in skills/ce-work/scripts/cross-model-work.sh and its controller, which this PR does not touch.
  • Open decision — skills/ce-code-review/references/depth-paths.md — code-review seats run only on the full review path
    A small diff that takes the lite or focused path applies no code-review entry. This follows the plan assumption "Seats run when a review dispatches its persona team", which the maintainer has not confirmed. This PR makes those paths say so in Coverage and corrects two guides that promised seats on every review. The alternative is to make a configured seat list lift every review to at least the focused path, which calls every seat on every small review.
  • Review comment — a provider-qualified GPT or o-series id landed in no family. Fixed: the resolver now places it in Codex, matching what both review workers accept as a model override.
  • Residual — Under cross_model_review_mode: off, a seat in the attested host family is exempt whatever route serves it. On Cursor with an attested family, such a seat can go out through the standalone family CLI. This matches the by-provider rule in the plan; closing it means exempting only when the family's harness is the host harness.
  • Residual — cross_model_review_mode: off reaches a worker-served seat only because the agent honors the resolver's blocked_by. Neither worker reads the key.
  • Residual — skills/ce-code-review/scripts/run-log.py:151 (unchanged) records one seat's token usage when two Codex worker seats run.
  • Residual — On Codex, the resolver was run on a no-map checkout in 2 of 16 cells before the final wording, and 0 of 8 after it. One round is thin evidence.
  • Residual — ce-babysit-pr does not relay a child's Model role line. The older model keys (plan_model, brainstorm_model, cross_model_peer, cross_model_model, cross_model_effort, work_engine_*) stay, with no removal trigger. In non-interactive and mode:agent reviews the pre-send seat notice is not printed.
  • Follow-up, not in this PR — the shared ce-config-layers block, in eleven files across ten skills, still names the second config file as bare config.yaml, the wording that made Codex miss the config here.

Review run: 20261001-165637-ecfed90e, artifacts at /tmp/compound-engineering-501/ce-code-review/20261001-165637-ecfed90e (verdict: Ready with fixes; no P0 or P1). Reviewers: correctness, security, testing, maintainability, project standards, agent-native, learnings, and an independent adversarial pass requested from Codex (gpt-6-luna at xhigh, served model unverified).

Security Disclosure

  • New bundled script scripts/model-role-resolve.py in nine skills. It reads the two repo CE config files and the CROSS_MODEL_PEERS environment variable, runs git rev-parse --show-toplevel, and prints JSON. It writes nothing and makes no network call.
  • Review seats can send reviewed documents and diffs to additional providers, one per seat, through the existing review workers. cross_model_review_mode: off and CROSS_MODEL_PEERS apply to every seat. With CROSS_MODEL_SEAT set, the two workers accept a target in the host's own family and an unknown host family; they still check CROSS_MODEL_PEERS. In seat mode the workers clear a target in the attested host's own family, and Composer when cursor is listed, the same two cases the resolver clears.
  • ce-setup models can write .compound-engineering/config.local.yaml or config.yaml, only after the user approves a preview.
  • The eval harness's new cursor host runs cursor-agent with --trust --sandbox enabled --force inside a throwaway workspace. It runs only when named with --hosts.
  • No dependency changes.

Agent Disclosure


Compound Engineering

🤖 Generated with Claude Code

kieranklaassen and others added 29 commits October 1, 2026 14:07
Requirements, technical decisions, and implementation units for a per-role
model map in CE config.

Co-Authored-By: Claude Code <noreply@anthropic.com>
One stdlib-only script resolves the new model_roles config key into a
per-role answer: grammar, role-by-role layering across the personal and team
files, model family, and the review egress policy. It is byte-duplicated into
each role skill, with a parity test, because skills cannot share files.

Co-Authored-By: Claude Code <noreply@anthropic.com>
One byte-duplicated reference states how a role skill honors its model_roles
entry: resolving the role, precedence, when a hand-off happens, the serving
order, the single-model fallback ladder, and the Model role report line. Skill
bodies stay untouched because most sit near the 8,000-byte prompt budget.

Adds the Model role map and Review seat glossary entries.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Adds the Model roles section to the config template and its example copy, and
to the configuration guide: entry grammar, role-by-role layering, the older
keys an entry replaces, the two personal opt-outs, and the controls on review
seats. A test resolves the template example through the resolver.

Co-Authored-By: Claude Code <noreply@anthropic.com>
The elevation engine consults the role map entry before plan_model or
brainstorm_model, and its Claude CLI worker takes an optional effort
(CE_ELEVATION_EFFORT, default high). An entry with an effort always hands
off, because no host exposes the session effort. The result envelope records
requested_effort.

Co-Authored-By: Claude Code <noreply@anthropic.com>
A debug entry moves the investigation to a read-only subagent and the
test-first fix to a second one on the named model. The session keeps the
causal-chain gate, the branch, the pre-fix scope record, verification, and the
commit. The structured returns gain one optional model_role field, which the
lfg gate accepts but does not require.

Co-Authored-By: Claude Code <noreply@anthropic.com>
ce-simplify-code hands its apply-and-verify steps to one subagent on the named
model; the three reviewers keep their tier. ce-compound has a subagent draft
the body into scratch from a hand-off file, while the orchestrator still
classifies and remains the only writer under the artifact root. Both fall back
to the session and say so, and print the Model role line.

Co-Authored-By: Claude Code <noreply@anthropic.com>
A work entry is resolved before the work_engine_* keys and becomes the only
standing candidate: served by native subagents when the host can hand over the
model and effort, otherwise through the engine route for its family at prefer.
An entry with an effort is never collapsed to native, a personal
work_engine_mode: off keeps a team entry off other models, and the fallback
ladder steps effort down through the levels the adapter accepts. No return
field is added.

Co-Authored-By: Claude Code <noreply@anthropic.com>
A step whose entry was replaced by a live instruction or carrier prints no
Model role line, and a route that started and then failed goes to the fallback
ladder the same way an unservable entry does. Both came up as ambiguities
while wiring the plan and work roles.

Co-Authored-By: Claude Code <noreply@anthropic.com>
lfg resolves no role itself: each child reads the map. When a child prints a
Model role line, lfg carries it unchanged into the step narration and the
closing recap. Each role skill guide gains a short Model role section linking
to the configuration guide.

Co-Authored-By: Claude Code <noreply@anthropic.com>
/ce-setup models shows each role with its effective value and source, offers
the model families the host and installed CLIs can serve, accepts a typed
model marked unconfirmed, asks whether to write the team or personal file and
states each reach, previews the block, and writes only on approval. The rule
that setup never creates config.local.yaml is scoped to the template-copy
step, since this flow may write the personal file.

Co-Authored-By: Claude Code <noreply@anthropic.com>
A list on the doc-review or code-review role runs one reviewer per seat on
every review, beside the persona review and in place of the single conditional
cross-model pass. The trigger sits in the references every dispatching review
reads, so seats do not depend on a judgment lens activating. Review roles fail
closed: an unreadable map, a blocked seat, or a seat no route can serve runs
nothing extra and is reported, never substituted.

The review workers take CROSS_MODEL_SEAT: it names the artifact and reviewer
per seat, and lets an explicit seat run in the host family or under an
unknown host while independence_verified stays false.

Co-Authored-By: Claude Code <noreply@anthropic.com>
With several review seats, two verified peers could promote a finding with no
in-process reviewer agreeing. A peer corroborates an in-process reading, so
promotion now needs at least one of each. The agent-mode return shape also
names coverage.model_roles where the rest of that shape is defined.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Twelve decision cells cover resolution, layering, precedence, engine routing,
the personal opt-out, review seats on a routine plan, and the review-mode
policy. Eight restraint cells, one per role skill, fail when a checkout with
no model_roles key runs the resolver or prints a Model role line. Tasks ask
for each skill ordinary work and never name the map. A judged conversation
scenario covers the setup flow.

Co-Authored-By: Claude Code <noreply@anthropic.com>
--hosts cursor runs a cell through cursor-agent -p in a trusted, sandboxed
workspace, using the flags the bundled workers already run. It is never a
default peer, because it needs a signed-in cursor-agent and bills its own
account. Adds the unattested-host row for the review-mode policy, which is
the case Cursor exists to check.

Co-Authored-By: Claude Code <noreply@anthropic.com>
The unattested-host row was listed out of order, so the sorted comparison
failed.

Co-Authored-By: Claude Code <noreply@anthropic.com>
When a personal work_engine_mode: off keeps a team work entry native, the
resolver now says so in warnings. In an eval on Codex the host read only the
boolean and attributed the native run to the role map itself; with the reason
in the answer, no host has to re-read the personal file to report it.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Round one on Claude and Codex showed three cells graded something other than
the behavior under test. The restraint rows failed Claude for narrating that
it checked for a map and found none; they now look for the Model role report
line itself. The plan and brainstorm rows pinned the Claude CLI as the first
route, which a host whose subagent tool can set effort does not owe; they now
require a hand-off. The debug task told a literal host to stop before reading
the settings the skill names.

Co-Authored-By: Claude Code <noreply@anthropic.com>
In one Codex eval run on a checkout with no model_roles key, the host loaded
the shared reference and ran the resolver to check for the key. The result was
the same, but a no-map checkout must not need Python. Every hook now says what
skipping means: do not read the reference or run its resolver.

The ce-debug hand-off also created its run directory before it knew a hand-off
would happen; the directory and hand-off file now belong to the hand-off.

Co-Authored-By: Claude Code <noreply@anthropic.com>
On Codex a decide-and-stop review cell reported TEAM: none and counted every
seat as not run, because nothing had been dispatched yet. The fields now ask
what the review would dispatch or leave out if it continued.

Co-Authored-By: Claude Code <noreply@anthropic.com>
A live ce-doc-review run on the Cursor host served three seats as Cursor
subagents on the ids its subagent tool lists. The case records how to repeat
that check. It stays hand-run because the model list changes too often to pin
in a graded fixture.

Co-Authored-By: Claude Code <noreply@anthropic.com>
… path

In eval round four Codex missed the role map twice. The hook named the files
as `.compound-engineering/config.local.yaml` or `config.yaml`, so one run
looked for `config.yaml` at the repo root. Both runs listed files with a
search that skips hidden directories and concluded no config existed. The
step then ran on the session model and printed no `Model role` line, which is
the silent failure the map must not have.

The hook now gives both full paths and says to open them by path.

Co-Authored-By: Claude Code <noreply@anthropic.com>
After the hook told agents to open both config files, Codex read the entry
itself in four of twenty entry cells and skipped the shared reference and the
resolver. Its answers were right on simple fixtures, but hand-reading skips
layering, validation, and the review policy. The hook now says to run the
resolver and that its answer decides the role.

Co-Authored-By: Claude Code <noreply@anthropic.com>
A misspelled `cross_model_review_mode` falls through to `auto`, which leaves
other-provider review seats unblocked with nothing said. The resolver now
warns on review roles, and the skills already print resolver warnings with
the `Model role` line.

Also adds tests for the fail-closed paths review roles rely on: structural
errors in the map and a config file the resolver cannot read. Drops a plan id
from a comment in the shipped script.

Co-Authored-By: Claude Code <noreply@anthropic.com>
The `code-review` role is resolved where the reviewer team is dispatched,
which only the full review path reaches. A lite or focused review skipped a
configured seat list with nothing said, while two guides promised seats on
every review. Those paths now state in Coverage that the role was not applied
and that `depth:full` applies it, and the guides say the same.

Adds an eval cell for that case and pins a review base in the code-review
role cells, where Codex once stopped at scope on a repo with one commit.

Co-Authored-By: Claude Code <noreply@anthropic.com>
The setup-flow test parsed frontmatter by hand, the native hand-off test
carried a third copy of the prompt-size helper to repeat a bound the budget
test owns, and the eval provenance check retyped the host list.

Co-Authored-By: Claude Code <noreply@anthropic.com>
…oject

Two Codex slips remained after the hook said the resolver decides. In 2 of 16
no-map cells Codex ran the resolver anyway, so the hook now states the skip
before the action. In 2 of 40 entry cells Codex ran the resolver from the
skill directory, where it finds no repository and answers `unset`; the shared
reference now says to run it with the project as the working directory.

The lite-review eval cell no longer pins the depth, which has its own rows.

Co-Authored-By: Claude Code <noreply@anthropic.com>
…llow it

Codex missed the model role map four ways across the hook's wording changes:
a bare file name, a file search that skips the hidden directory, reading the
entry instead of running the resolver, and running the resolver from the
skill directory. The learning records the wording that passed and the eval
counts for each version.

Co-Authored-By: Claude Code <noreply@anthropic.com>
The simplify cell stops before the hand-off, so the answer is a prediction
about the host's subagent tool. Codex said yes in three of six runs. The cell
now grades the entry's model and effort and the shared reference being read.

Co-Authored-By: Claude Code <noreply@anthropic.com>
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-08T03:40:33.142457Z c3057e7 New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

kieranklaassen and others added 3 commits October 2, 2026 00:23
An id such as `openai/<gpt id>` or `openai.<gpt id>` passed the entry grammar
but landed in no family, so a review seat could not use the Codex worker and
a work entry got no engine. Both review workers already accept those forms
as a Codex model override. The resolver now classifies a provider-qualified
GPT or o-series id as Codex. Other ids with a slash stay in no family.

Co-Authored-By: Claude Code <noreply@anthropic.com>
…as the hook's trigger

The shared model role reference loses filler, sentences repeated within the
file, and three of five example lines. Every rule, the resolver command, the
state table, and the working-directory sentence are unchanged. The resolver's
docstring is trimmed with no behavior change.

The hook now says the resolver runs whenever the `model_roles:` key exists,
even with no entry for this role. With a map that held another role's entry
only, Codex had skipped the resolver and once reported that other role's
model. Over five rounds on the new wording Codex had no such miss.

Content the feature adds drops from 80,071 to 78,793 bytes. Codex passed 89
of 95 cells over five rounds, Claude 21 of 21, and Cursor 20 of 20.

Co-Authored-By: Claude Code <noreply@anthropic.com>
… and counts

Adds the fifth fact the hook needs (the trigger is the key, not an entry for
this role), replaces the example with the wording that shipped, and replaces
the one-round row with eleven rounds on that wording and five on the last.

Co-Authored-By: Claude Code <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 68db22d413

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread skills/ce-setup/scripts/model-role-resolve.py Outdated
Comment thread skills/ce-setup/SKILL.md

## Model Roles

When the invocation asks to set up model roles (the `models` argument, or the same request in words), read `references/model-roles-setup.md` from this skill's directory and follow it in place of Phases 1-2. It writes the `model_roles` map only after the user approves, and it reports into Phase 3 (Summary).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Validate model roles during the standard health check

When a developer edits model_roles manually and then runs ordinary ce-setup, only the models branch invokes the new resolver; scripts/check-health contains no model-role validation. For example, a checkout with model_roles: opus is reported as Project config healthy, even though every consuming skill resolves the block as invalid. Include resolver diagnostics in the standard health path so the setup health check can detect malformed instances of this new config surface.

Useful? React with 👍 / 👎.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 7c405c8b81

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread skills/ce-work/scripts/model-role-resolve.py Outdated
The resolver placed a provider-qualified GPT or o-series id in Codex only
when the provider name was letters, so `azure2/<gpt id>` landed in no family
while both review workers accept it. A name already in another family is
now decided first, and any other id with `.` or `/` before a GPT or o-series
name is Codex, the same form the workers accept.

Co-Authored-By: Claude Code <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: bdfb70043a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread skills/ce-work/scripts/model-role-resolve.py
*) skip "host harness '$HOST_HARNESS' invalid (want codex|claude|grok|cursor|opencode|unknown); skipping cross-model pass" ;;
esac
[ "$HOST_PROVIDER" != "unknown" ] || skip "host serving family unattested; automatic cross-model review skipped"
[ -n "$SEAT" ] || [ "$HOST_PROVIDER" != "unknown" ] || skip "host serving family unattested; automatic cross-model review skipped"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Exempt configured seats from the automatic peer cap

When CROSS_MODEL_MAX_PEERS=0 is set, a configured non-native review seat reaches this seat-specific path but is still skipped by the later MAX_PEERS gate. The review-seat contract in references/cross-model-review.md explicitly excludes the single-peer pass's activation/run conditions and says every otherwise valid seat becomes a job, so an environment used to disable the automatic pass silently disables configured seats too. Make seat mode bypass the automatic peer-count gate; the doc-review worker has the same mismatch.

Useful? React with 👍 / 👎.

Comment thread skills/ce-debug/references/investigate.md Outdated
The resolver put a provider-qualified GPT or o-series id in Codex for every
role. The review workers accept that form, but the work engine's Codex route
and its controller accept an unqualified id only, so a `work` entry with one
was sent to an engine whose preflight refused it. The resolver now reports
no engine harness for such a `work` entry. The entry is then served by a
native subagent or the session, and the Model role line says why.

The work engine's own model validation is unchanged.

Co-Authored-By: Claude Code <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 5b875cf722

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread skills/ce-code-review/scripts/cross-model-adversarial-review.sh
kieranklaassen and others added 2 commits October 2, 2026 00:55
A `work` entry such as `opus[1m] high` resolved to the Claude engine, whose
adapter and controller reject a bracketed qualifier, so preflight refused it.
The resolver now reports no engine harness for any `work` entry the engine's
routes do not accept: an id with a bracketed qualifier, or a provider-
qualified Codex id. Such an entry runs on a native subagent or the session,
and its Model role line says why.

Co-Authored-By: Claude Code <noreply@anthropic.com>
The role hook sat at the start of Phase 1, and the trivial-bug fast path
leaves Phase 0 for the fix-choice gate before reaching it. A diagnosis-only
run on that path printed nothing about a configured `debug` role. The hook
now sits before the fast path, and the hand-off paragraph says that the fast
path has no investigation to hand off while the fix hand-off still applies.

Codex resolved the role in two of three runs of the entry cell and ran no
resolver in three of three no-map runs; Claude and Cursor passed both cells.

Co-Authored-By: Claude Code <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e778fc30a9

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread skills/ce-debug/references/investigate.md Outdated
kieranklaassen and others added 2 commits October 2, 2026 01:03
A seat that requested one Claude model and was served by another still
published its artifact, and synthesis read it as the configured reviewer. A
review seat is never filled by another model, so in seat mode a receipt that
names a different model now leaves the seat with no artifact, and the worker
logs why. The seat is then reported as dropped. Outside a seat the
single-peer pass still publishes and warns, as before.

Co-Authored-By: Claude Code <noreply@anthropic.com>
… too

The hand-off paragraph kept an exception: on the trivial-bug fast path the
session stated the cause, although the `debug` role's deliverable is the
diagnosis and the fix. The exception is gone. The entry's model produces the
diagnosis on every path, and a trivial bug changes only how deep it goes: the
subagent states the cause and the proposed fix and stops. The hand-off rules
now sit before the fast path, where that path reads them.

Claude and Cursor passed both debug cells. Codex passed the entry cell in
two of three runs and the no-map cell in three of three.

Co-Authored-By: Claude Code <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0aa1b7f63f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread skills/ce-plan/scripts/elevation-dispatch.sh
… worker

A requested full id was compared to the served id by bare prefix, so
`claude-opus-5` accepted a receipt for `claude-opus-50-*` and the worker
reported `receipt: matched` for another model. The worker now matches the
requested stem itself or the stem followed by `-`, in both the receipt check
and the choice of which served id to report. Alias matching is unchanged.

Co-Authored-By: Claude Code <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ef7dba5b60

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread skills/ce-code-review/scripts/cross-model-adversarial-review.sh Outdated
…model

The review workers accepted a Claude alias with a bracketed qualifier through
a shell glob, which also let `opus[]`, `opus[a][b]`, and `opus[1m]junk]`
through to a provider call that would fail. The workers now require exactly
one trailing qualifier of letters and digits, the grammar the role resolver
accepts, and check the id before it against the existing patterns.

Co-Authored-By: Claude Code <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 64d1c3623b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread skills/ce-debug/references/fix.md Outdated
When the `debug` or `simplify` role is served by a native subagent, that
subagent edits project files, and the steps it is handed tell it to follow
the project's instructions. A subagent may start without the instructions
the session has loaded, so both hand-off prompts now carry the active
project instructions and subdirectory-scoped instructions that govern the
files it may change, as text or as the paths of the files to read first.

Co-Authored-By: Claude Code <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c2ecc942c9

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread skills/ce-debug/references/investigate.md Outdated
Comment thread skills/ce-debug/references/investigate.md Outdated
A role hand-off to a native subagent sent a prompt that held neither the
project's instructions nor a usable path to the skill files its steps name.
A probe that printed the investigation hand-off prompt showed both missing
on Codex and Cursor. The previous commit patched this for two editing
hand-offs only.

The shared model-roles reference now states the rule for every native
subagent: its prompt gives the project instructions that govern the work,
and, when its steps name a file of the skill, the skill directory's
absolute path to resolve those names against. The debug investigation, the
debug fix, the simplify apply, and the compound draft prompts each point at
that rule where the prompt is composed, and the per-prompt wording from the
previous commit is removed.

Co-Authored-By: Claude Code <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a406bad411

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread skills/ce-setup/scripts/model-role-resolve.py
kieranklaassen and others added 3 commits October 2, 2026 02:55
The role resolver lowercased a model id before classifying it, so `Opus` or
`GPT-6.1-Sol` came back as a valid Claude or Codex entry. Every route's own
validator matches ids case-sensitively, so the route then refused the id:
a review seat was dropped and a single-model role fell back, with nothing
telling the user why.

The resolver now classifies ids as written. An id that only a lowercase
spelling would route is skipped with a warning that names that spelling.
An id that no route claims in any spelling keeps its own.

Co-Authored-By: Claude Code <noreply@anthropic.com>
…in-f97951

# Conflicts:
#	skills/ce-work/references/execution-engines.md
…in-f97951

# Conflicts:
#	skills/ce-code-review/scripts/cross-model-adversarial-review.sh
#	skills/ce-doc-review/scripts/cross-model-doc-review.sh
#	skills/ce-pov/scripts/cross-model-pov.sh
#	skills/ce-work/references/cross-model-execution.md
#	tests/skill-eval-cell/packs/ce-doc-review-cross-model.md
#	tests/skills/ce-code-review-cross-model-routes.test.ts
#	tests/skills/ce-doc-review-cross-model-routes.test.ts
#	tests/skills/ce-work-cross-model-routes.test.ts

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 2311893876

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread skills/ce-code-review/scripts/cross-model-adversarial-review.sh
A review seat is one model, so a seat whose receipt names another model is
not served and writes no artifact. That check ran on the Claude route only,
because Claude was the one route with a receipt. The workers now read a
served-model receipt for Codex and native Grok as well, so a seat on either
is held to it the same way.

A Codex id may carry a provider namespace such as `openai/` that the served
id does not, so the namespace is set aside before the two are compared.

Co-Authored-By: Claude Code <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: eb5432ba99

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread skills/ce-setup/scripts/model-role-resolve.py Outdated
Comment thread skills/ce-setup/scripts/model-role-resolve.py Outdated
…ers see them

The resolver's full summary, which `ce-setup` shows as the current routing,
was wrong in two cases.

An older key with a closed vocabulary was reported from the first file that
set it, even when the value was unsupported. The review skills skip an
unsupported `cross_model_peer` and read the next file, so the summary now
does the same.

A review role whose list held an invalid seat showed an empty or shortened
value with no warning. The summary now keeps each invalid seat as written,
with the reason it does not run.

Co-Authored-By: Claude Code <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant