From 78407e14aacc58533a734ccede623b8dd16f52c4 Mon Sep 17 00:00:00 2001 From: tawanorg Date: Thu, 10 Sep 2026 16:07:56 +1000 Subject: [PATCH 1/4] feat: preserve software engineer agent and OCR delegation setup --- README.md | 15 +++++ bin/codex-sync | 3 +- codex/AGENTS.md | 4 ++ codex/agents/software-engineer-guide.md | 83 +++++++++++++++++++++++++ codex/agents/software-engineer.toml | 31 +++++++++ codex/config.toml | 7 +++ codex/external.json | 3 + tests/test_codex_sync.py | 10 ++- 8 files changed, 154 insertions(+), 2 deletions(-) create mode 100644 codex/agents/software-engineer-guide.md create mode 100644 codex/agents/software-engineer.toml diff --git a/README.md b/README.md index 85eab7a..5573c3c 100644 --- a/README.md +++ b/README.md @@ -452,6 +452,21 @@ configuration is tracked; plugin caches are not copied. Shared skills under Requires Python 3.11+ (`tomllib`). +**Implementation and review.** The personal `software_engineer` agent discovers +each project's conventions, implements scoped changes, and verifies behavior. +It inherits the parent session's model and available MCP connections, including +Serena, Context7 and Firecrawl. Its usage guide travels with it during export +and restore: [`codex/agents/software-engineer-guide.md`](codex/agents/software-engineer-guide.md). +In a new Codex session, ask `Use software_engineer to implement [task]`. + +`codex-sync install --external` also installs Open Code Review CLI **1.11.7** +and its native Codex plugin. Global instructions default OCR to **delegation +mode**, which uses the current Codex session for reasoning without a separate +OCR LLM endpoint. Ask `@Open Code Review review my current changes` in a new +session. Subscription-authenticated Codex uses its subscription allowance; +external services such as Firecrawl can consume separate credits. No OCR +credentials or plugin caches are tracked. + ## Not included, on purpose - **Docker Desktop** — its installer needs a sudo password, so this uses diff --git a/bin/codex-sync b/bin/codex-sync index bf0584a..d9f9aa0 100755 --- a/bin/codex-sync +++ b/bin/codex-sync @@ -136,7 +136,7 @@ def main(): config = tomllib.loads((live / "config.toml").read_text()) write(tracked / "config.toml", dumps(portable(capture(config), Path.home()))) shutil.copy2(live / "AGENTS.md", tracked / "AGENTS.md") - for agent in (live / "agents").glob("*.toml"): + for agent in [*(live / "agents").glob("*.toml"), *(live / "agents").glob("*.md")]: shutil.copy2(agent, tracked / "agents" / agent.name) for name in manifest["local_skills"]: shutil.copytree(live / "skills" / name, tracked / "skills" / name, dirs_exist_ok=True) @@ -163,6 +163,7 @@ def main(): backup_file(config_path) write(config_path, restored) sources = [tracked / "AGENTS.md", *(tracked / "agents").glob("*.toml"), + *(tracked / "agents").glob("*.md"), *(p for p in (tracked / "skills").rglob("*") if p.is_file())] for source in sources: dest = live / source.relative_to(tracked) diff --git a/codex/AGENTS.md b/codex/AGENTS.md index e4fefcb..752df89 100644 --- a/codex/AGENTS.md +++ b/codex/AGENTS.md @@ -81,6 +81,10 @@ different jobs. Use `/code-review` for correctness on a diff. Use the Matt Pocock one when the question is whether the change matches a written spec or the repo's documented standards. Say which you ran. +When Open Code Review is selected, use its `open-code-review-delegate` skill by +default. Do not run `ocr review` or configure an OCR LLM endpoint unless the +user explicitly requests OCR-managed review with their own API credentials. + ## Scope `critical-developer-mindset` declares itself always-on. Treat it as applying to diff --git a/codex/agents/software-engineer-guide.md b/codex/agents/software-engineer-guide.md new file mode 100644 index 0000000..a05329e --- /dev/null +++ b/codex/agents/software-engineer-guide.md @@ -0,0 +1,83 @@ +# Software engineer agent + +Start a new Codex session, then ask: + +> Use the software_engineer agent to implement [change]. Done means [observable behavior]. + +For a bug, include what happens, what should happen, and a reproduction if known. +Assign an existing worktree and file ownership when other sessions are editing. +For tiny edits, working directly in the parent avoids the extra agent context. + +## Installation and portability + +`software-engineer.toml` is a personal agent under the active Codex home's +`agents/` directory. To reuse it on another machine, copy that file to the +equivalent directory and start a new session. Project-scoped installation uses +`.codex/agents/software-engineer.toml`. Avoid defining the same role in both +places unless an intentional project override is desired. + +The agent inherits the parent model, reasoning effort, MCP connections, and +permission settings. It has no embedded provider credentials or project paths. +Its LLM work uses the parent session's authentication and billing arrangement; +with subscription-authenticated Codex it consumes that subscription's allowance. +Separate services can have their own billing. Inheritance does not make them free. + +## Tool choices + +| Need | Preferred tool | Fallback | +| --- | --- | --- | +| Locate code and callers | Serena symbol definitions/references | Scoped rg and file reads | +| Check a library API | Context7 or the project's dedicated docs MCP | Official version-matched docs | +| Public web research | Firecrawl MCP or CLI; narrow queries, cached results | Available approved web tool | +| Verify behavior | Project tests, lint, typecheck, browser/device tools | Report unverified behavior explicitly | +| Review substantial changes | Open Code Review delegation skill | Explicit self-review | + +Connections come from the host session. Copying the TOML file alone does not +install MCP servers or skills. Existing tools are reused, and missing optional +tools have fallbacks. Credentials stay in the host configuration/environment. +OCR-managed API calls require explicit authorization. Firecrawl's hosted service +can consume separate credits; local Serena navigation does not require an LLM API key. + +## Working style + +Understand the relevant path, implement a complete change, verify the observable +result. Small tasks proceed directly. Larger changes use a short working plan; +project-required financial, security, migration, and review gates still apply. +Routine test boundaries are inferred from approved behavior and existing patterns. +Specialist skills load only for the relevant task, rather than starting a fixed +chain of interviews and documents for every edit. + +The host can still bring a large tool/skill catalog into context. Selective tool +use limits retrieved output but does not remove that startup cost. Disable unused +plugins in the host deliberately if that becomes a measured problem. No percentage +token savings or enterprise reliability claim has been established for this agent. + +## Sources and design decisions + +Reviewed 2026-09-10. These are references, not installed nested coding agents. + +- [Matt Pocock skills](https://github.com/mattpocock/skills): small composable skills, public-interface tests, short feedback loops. Uses the already installed bundle; preserves the user's preference for minimal ceremony. +- [Serena](https://github.com/oraios/serena): semantic code navigation can retrieve definitions and references without loading entire files. Actual savings depend on task and language support. +- [Aider](https://github.com/Aider-AI/aider): immediate lint/test feedback is a useful implementation practice. Aider itself is a separate coding client and was not installed; its API setup is unnecessary for this Codex agent. +- [OpenHands](https://github.com/OpenHands/OpenHands): Agent Canvas supports local/remote agent backends and automations. Consider it for always-on infrastructure or team orchestration; those concerns are outside this personal agent's job. +- [Codex custom agents](https://learn.chatgpt.com/docs/agent-configuration/subagents): personal/project TOML agent definitions and inherited settings. + +## Evaluation + +A valid TOML file only establishes parseability. A meaningful trial must launch +the actual custom role, exercise an implementation with a failing regression test, +rerun it after the fix, and check unrelated work was preserved. A small fixture is +a smoke test, not a benchmark of production engineering quality. Repeat on actual +tasks before making claims about speed, cost, or reliability. + +2026-09-10 smoke result: a normal Codex CLI session successfully launched exactly +one `software_engineer` role on a disposable Node.js pagination fixture. The agent +reported 3 failing and 2 passing tests before the fix; independent reruns after +the fix passed all 5 tests. The unrelated notes file retained its SHA-256 checksum. +The agent configuration omitted model/reasoning/tool overrides. This trial did +not exercise Serena, Firecrawl, UI tools, or a complete OCR review. + +An earlier `codex exec --ephemeral` launch failed with `collab spawn failed: no +thread with id`. A regular session succeeded; use that path for this installation. +The host also warned that skill descriptions were shortened to fit its context +budget. These are observed limitations, not proven problems in other versions. diff --git a/codex/agents/software-engineer.toml b/codex/agents/software-engineer.toml new file mode 100644 index 0000000..3819274 --- /dev/null +++ b/codex/agents/software-engineer.toml @@ -0,0 +1,31 @@ +name = "software_engineer" +description = "Implement features, fix bugs, and refactor across projects: inspect the relevant code, deliver a complete change, and verify observable behavior. Uses available code-navigation and documentation tools selectively." +developer_instructions = """ +Own the assigned implementation through verification. Inherit the parent session's model, tools, permissions, and authorized scope. Repository instructions and domain constraints govern each project. + +Understand +- Read applicable repository instructions and relevant local architecture/decision records. Establish the working directory, branch/worktree, existing edits, stack, package manager, and actual validation commands from project files. Reuse fresh environment context from the parent; do not repeat completed discovery. +- Translate the request into observable acceptance criteria. For clear work, proceed immediately; for a change crossing several modules, keep a short working plan. Resolve routine choices from existing patterns. Ask only for a missing product decision, access, or authority that materially changes the result; continue independent work meanwhile. +- Trace the smallest relevant path from entry point through behavior and callers to tests. Finish discovery when the owning code, expected behavior, and a verification approach are known. Expand only to answer a concrete unresolved question. + +Implement +- Deliver the smallest complete change, including affected callers, error paths, and configuration/documentation where behavior requires it. Reuse existing abstractions; introduce new ones when they hide real complexity or variation. Preserve compatibility unless changing it is part of the request. +- For a bug, reproduce its specific symptom before fixing it where feasible. For behavior changes, use a short test/implementation feedback loop through a public interface with independently derived expected results. Prefer regression tests that fail on the original bug. Match verification effort to risk; mechanical copy/config edits need not acquire artificial unit tests. +- You are not alone in the codebase. Respect assigned file ownership, preserve others' edits, and follow project worktree rules. Escalate overlapping edits you cannot safely accommodate. Keep unrelated findings separate from the implementation. +- Use existing authorization for delivery steps. Commits, pushes, PRs, deployments, dependency/service additions, and external mutations must remain within the user's requested workflow. Stop for genuinely missing authority, not routine engineering decisions. + +Verify and hand off +- Run the relevant tests/checks discovered from the project, including required project gates. Exercise changed UI with the supported browser/device tools when available. For performance work, compare measurements against a baseline. +- Inspect the final diff for unintended changes and check the relevant failure modes: authorization, input validation, missing data, retries/concurrency, compatibility, and error reporting. These checks are internal work, not a mandatory approval ceremony. +- For substantial or high-risk changes, use Open Code Review's open-code-review-delegate skill when available, scoped to this change. OCR supplies file selection/rules; this Codex session supplies reasoning. Never fall back to OCR-managed LLM calls, ocr review/scan, or endpoint configuration without explicit user authorization. Account for excluded files in the applicable project review requirements. If OCR is unavailable, do an explicit self-review and state the limitation; do not imply independent review. +- Address in-scope findings and rerun affected checks. Completion means every acceptance criterion is verified or clearly marked unverified with a reason. Report the outcome, decisive validation evidence, and remaining blockers concisely. A passing version/config check is not proof of runtime behavior. + +Tools and context +- Discover available tools; inherit MCP connections rather than hardcoding servers, credentials, models, or project paths. If a tool is absent or fails, use an appropriate local fallback and name the limitation once. Missing optional tooling does not block implementation. +- Serena: use initial_instructions and confirm the active project before semantic navigation. Prefer symbol outlines, targeted definitions, and references over entire-file reads. Bound searches by path/symbol; use rg for literal text, config, or unsupported languages. Respect the host's editing-tool requirements. Do not change a shared server's active project underneath another worker. +- Context7 or a dedicated documentation MCP: verify unfamiliar or version-sensitive library APIs against the installed version. Prefer the project's dedicated framework tool when its instructions specify one. +- Firecrawl MCP or CLI: use for public web research and documentation gaps; retrieve the narrowest relevant page and reuse saved results. It may consume separate service credits. Keep private source, credentials, and client data out of external queries. Do not crawl entire sites for a single API question. +- Load applicable Matt Pocock skills by their current catalog names: tdd for test-first behavior changes, diagnosing-bugs for difficult regressions, codebase-design for consequential interface work. If unavailable, use the practices above. The user's preference is a lean workflow: infer routine test boundaries from the request and project conventions and proceed; reserve skill checkpoints for genuine product decisions or required project gates. +- Keep output bounded and preserve full logs as local artifacts when necessary. Read the relevant failures and summaries, expanding when needed; never silently discard findings. Reuse fresh evidence. Load only skill references needed for the current branch of work. +- Complete bounded assignments yourself. Return large-output specialist work to the parent for delegation when useful; avoid recursive agent fan-out. Delegate only under applicable user/project authorization, with explicit scope and file ownership. More tools or agents are not evidence of lower token use. +""" diff --git a/codex/config.toml b/codex/config.toml index 3538976..47b6d88 100644 --- a/codex/config.toml +++ b/codex/config.toml @@ -337,6 +337,10 @@ "source_type" = "git" "source" = "https://github.com/Egonex-AI/Understand-Anything.git" +["marketplaces"."open-code-review"] +"source_type" = "git" +"source" = "https://github.com/alibaba/open-code-review.git" + ["plugins"] ["plugins"."document-skills@anthropic-agent-skills"] @@ -350,3 +354,6 @@ ["plugins"."understand-anything@understand-anything"] "enabled" = true + +["plugins"."open-code-review-codex@open-code-review"] +"enabled" = true diff --git a/codex/external.json b/codex/external.json index 87f7c5d..7851fee 100644 --- a/codex/external.json +++ b/codex/external.json @@ -2,6 +2,9 @@ "local_skills": ["worktree", "worktree-cleanup", "ics-jira-dev-ready"], "removed_mcp_servers": ["graphify"], "install_commands": [ + ["npm", "install", "-g", "@alibaba-group/open-code-review@1.11.7"], + ["codex", "plugin", "marketplace", "add", "alibaba/open-code-review"], + ["codex", "plugin", "add", "open-code-review-codex@open-code-review"], ["npx", "-y", "skills", "add", "openai/skills", "--skill", "yeet", "gh-fix-ci", "--agent", "codex", "-y", "-g"], ["npx", "-y", "skills", "add", "mattpocock/skills", "--skill", "handoff", "--agent", "codex", "-y", "-g"], ["npx", "-y", "firecrawl-cli@1.23.3", "init", "--global", "--yes", "--agent", "codex", "--skip-auth"], diff --git a/tests/test_codex_sync.py b/tests/test_codex_sync.py index e3ef44d..7657aad 100644 --- a/tests/test_codex_sync.py +++ b/tests/test_codex_sync.py @@ -108,7 +108,15 @@ def test_fresh_restore_and_portability(self): self.assertTrue((live / "skills/worktree/SKILL.md").exists()) self.assertTrue((live / "skills/worktree-cleanup/references/docker.md").exists()) self.assertTrue((live / "skills/ics-jira-dev-ready/SKILL.md").exists()) - self.assertEqual(len(list((live / "agents").glob("*.toml"))), 3) + self.assertEqual(len(list((live / "agents").glob("*.toml"))), 4) + engineer = live / "agents/software-engineer.toml" + self.assertEqual(tomllib.loads(engineer.read_text())["name"], "software_engineer") + self.assertEqual((live / "agents/software-engineer-guide.md").read_bytes(), + (SCRIPT.parent.parent / "codex/agents/software-engineer-guide.md").read_bytes()) + self.assertTrue(restored["plugins"]["open-code-review-codex@open-code-review"]["enabled"]) + self.assertEqual(restored["marketplaces"]["open-code-review"]["source"], + "https://github.com/alibaba/open-code-review.git") + self.assertIn("open-code-review-delegate", (live / "AGENTS.md").read_text()) self.assertNotIn("{{HOME}}", (live / "config.toml").read_text()) portable = sync.portable(copy.deepcopy(restored), Path.home()) self.assertEqual(sync.portable(portable, Path.home(), restore=True), restored) From b5ef3f78c2863e82cd7cffe8d8807227b6224901 Mon Sep 17 00:00:00 2001 From: tawanorg Date: Thu, 10 Sep 2026 16:14:47 +1000 Subject: [PATCH 2/4] feat: sync selectable software engineer skill --- README.md | 6 ++++-- codex/external.json | 2 +- codex/skills/software-engineer/SKILL.md | 18 ++++++++++++++++++ .../software-engineer/agents/openai.yaml | 4 ++++ tests/test_codex_sync.py | 5 +++++ 5 files changed, 32 insertions(+), 3 deletions(-) create mode 100644 codex/skills/software-engineer/SKILL.md create mode 100644 codex/skills/software-engineer/agents/openai.yaml diff --git a/README.md b/README.md index 5573c3c..88e501d 100644 --- a/README.md +++ b/README.md @@ -423,7 +423,7 @@ The generated links are ignored by Git, so another machine gets its own paths. `codex-sync export` captures this user's model and web-search settings, status line, developer instructions, app preferences, portable notifications/hooks, MCP servers, Git marketplace/plugin preferences, global instructions and custom -agents. It also saves the local `worktree`, `worktree-cleanup` and +agents. It also saves the local `worktree`, `worktree-cleanup`, `software-engineer` and `ics-jira-dev-ready` skills. Run it after changing Codex settings, then review the diff. Home paths in configuration are stored as `{{HOME}}`; unrecognised top-level settings are reported for review instead of silently omitted. @@ -457,7 +457,9 @@ each project's conventions, implements scoped changes, and verifies behavior. It inherits the parent session's model and available MCP connections, including Serena, Context7 and Firecrawl. Its usage guide travels with it during export and restore: [`codex/agents/software-engineer-guide.md`](codex/agents/software-engineer-guide.md). -In a new Codex session, ask `Use software_engineer to implement [task]`. +In a new Codex session, select **Software Engineer** in the skills picker or +type `$software-engineer implement [task]`. This user-scope skill loads the +same agent instructions and can delegate to the custom role when useful. `codex-sync install --external` also installs Open Code Review CLI **1.11.7** and its native Codex plugin. Global instructions default OCR to **delegation diff --git a/codex/external.json b/codex/external.json index 7851fee..cb53b32 100644 --- a/codex/external.json +++ b/codex/external.json @@ -1,5 +1,5 @@ { - "local_skills": ["worktree", "worktree-cleanup", "ics-jira-dev-ready"], + "local_skills": ["worktree", "worktree-cleanup", "ics-jira-dev-ready", "software-engineer"], "removed_mcp_servers": ["graphify"], "install_commands": [ ["npm", "install", "-g", "@alibaba-group/open-code-review@1.11.7"], diff --git a/codex/skills/software-engineer/SKILL.md b/codex/skills/software-engineer/SKILL.md new file mode 100644 index 0000000..9175a5f --- /dev/null +++ b/codex/skills/software-engineer/SKILL.md @@ -0,0 +1,18 @@ +--- +name: software-engineer +description: Use the personal software_engineer implementation agent for a requested feature, bug fix, or refactor with behavior verification and minimal ceremony. +--- + +# Software Engineer + +This is the selectable user-scope entry point for the `software_engineer` agent. + +Read the `developer_instructions` from [the personal agent definition](../../agents/software-engineer.toml) before implementation. That file is the source of truth for the engineering workflow; this skill does not duplicate it. If it is missing, check the active Codex home's `agents/software-engineer.toml`; report a missing installation if neither resolves. + +Use the task and constraints supplied with this skill. If no task was supplied, ask what to implement rather than inventing work. + +When the custom `software_engineer` role is available and delegation is useful, dispatch a bounded implementation assignment with the project/worktree path, acceptance criteria, file ownership, and relevant existing evidence. Tell the worker that other sessions may have edits to preserve. The parent can check acceptance criteria and prepare independent verification while it runs; avoid editing the worker's files. Collect its result and verify completion. + +For a small task, or if the custom role is unavailable, perform the work in the current thread using the same agent instructions. State which execution mode is being used; do not claim a subagent launched if it did not. Inherit the current model, authentication, permissions, and available tools. Use Open Code Review in delegation mode as specified in the agent definition. + +Return the implemented outcome, actual verification results, and any remaining blocker. Follow the user's existing authorization for delivery steps. diff --git a/codex/skills/software-engineer/agents/openai.yaml b/codex/skills/software-engineer/agents/openai.yaml new file mode 100644 index 0000000..1413d82 --- /dev/null +++ b/codex/skills/software-engineer/agents/openai.yaml @@ -0,0 +1,4 @@ +interface: + display_name: "Software Engineer" + short_description: "Implement and verify features, fixes, and refactors" + default_prompt: "Use $software-engineer to implement my requested change and verify the result." diff --git a/tests/test_codex_sync.py b/tests/test_codex_sync.py index 7657aad..3c7202b 100644 --- a/tests/test_codex_sync.py +++ b/tests/test_codex_sync.py @@ -108,6 +108,11 @@ def test_fresh_restore_and_portability(self): self.assertTrue((live / "skills/worktree/SKILL.md").exists()) self.assertTrue((live / "skills/worktree-cleanup/references/docker.md").exists()) self.assertTrue((live / "skills/ics-jira-dev-ready/SKILL.md").exists()) + skill = live / "skills/software-engineer/SKILL.md" + self.assertTrue(skill.exists()) + self.assertTrue((skill.parent / "agents/openai.yaml").exists()) + self.assertEqual((skill.parent / "../../agents/software-engineer.toml").resolve(), + (live / "agents/software-engineer.toml").resolve()) self.assertEqual(len(list((live / "agents").glob("*.toml"))), 4) engineer = live / "agents/software-engineer.toml" self.assertEqual(tomllib.loads(engineer.read_text())["name"], "software_engineer") From 015cedcad3b15379b477141ee6d0741cc3a20145 Mon Sep 17 00:00:00 2001 From: tawanorg Date: Thu, 10 Sep 2026 16:15:38 +1000 Subject: [PATCH 3/4] chore: exclude project-specific Jira skill from shared dotfiles --- README.md | 4 +- codex/external.json | 2 +- codex/skills/ics-jira-dev-ready/SKILL.md | 66 ------------------- .../ics-jira-dev-ready/agents/openai.yaml | 4 -- tests/test_codex_sync.py | 2 +- 5 files changed, 4 insertions(+), 74 deletions(-) delete mode 100644 codex/skills/ics-jira-dev-ready/SKILL.md delete mode 100644 codex/skills/ics-jira-dev-ready/agents/openai.yaml diff --git a/README.md b/README.md index 88e501d..4c38427 100644 --- a/README.md +++ b/README.md @@ -423,8 +423,8 @@ The generated links are ignored by Git, so another machine gets its own paths. `codex-sync export` captures this user's model and web-search settings, status line, developer instructions, app preferences, portable notifications/hooks, MCP servers, Git marketplace/plugin preferences, global instructions and custom -agents. It also saves the local `worktree`, `worktree-cleanup`, `software-engineer` and -`ics-jira-dev-ready` skills. Run it after changing Codex settings, then review +agents. It also saves the local `worktree`, `worktree-cleanup` and +`software-engineer` skills. Run it after changing Codex settings, then review the diff. Home paths in configuration are stored as `{{HOME}}`; unrecognised top-level settings are reported for review instead of silently omitted. diff --git a/codex/external.json b/codex/external.json index cb53b32..21bdc5d 100644 --- a/codex/external.json +++ b/codex/external.json @@ -1,5 +1,5 @@ { - "local_skills": ["worktree", "worktree-cleanup", "ics-jira-dev-ready", "software-engineer"], + "local_skills": ["worktree", "worktree-cleanup", "software-engineer"], "removed_mcp_servers": ["graphify"], "install_commands": [ ["npm", "install", "-g", "@alibaba-group/open-code-review@1.11.7"], diff --git a/codex/skills/ics-jira-dev-ready/SKILL.md b/codex/skills/ics-jira-dev-ready/SKILL.md deleted file mode 100644 index ec6d747..0000000 --- a/codex/skills/ics-jira-dev-ready/SKILL.md +++ /dev/null @@ -1,66 +0,0 @@ ---- -name: ics-jira-dev-ready -description: Find my assigned Dev Ready Jira tickets in the active QB sprint using Atlassian MCP, with clickable links and ordering by product priority or complexity. Use when choosing what to work on next or asking for my sprint queue. ---- - -# ICS Jira Dev Ready - -Return a current, read-only work queue. Default to product priority; support easiest-first or hardest-first when requested. Listing or recommending work does not authorise ticket edits or starting implementation. - -## Personal defaults - -- Site: `https://coterieholdings.atlassian.net` -- Project: `QB`; board: `1`. -- Owner account from the user's board URL: `5aacdfe733719f2a5016620e`. -- Status: exactly `Dev Ready`. -- Board: `https://coterieholdings.atlassian.net/jira/software/projects/QB/boards/1`. - -Apply explicit user overrides. The supplied website filter also included PR Review, Dev In Progress and To Do; those are outside this queue unless requested. - -## Fetch the queue - -1. Use Atlassian MCP. Resolve the site's cloudId once with `getAccessibleAtlassianResources` when not already known. Pass cloudId on every site-scoped call. Fetch `atlassianUserInfo` once: use `currentUser()` when it matches the owner above; otherwise use the explicit owner account and disclose the mismatch. Resolve an explicitly requested different assignee instead of silently substituting the connected account. -2. Discover the read operations for board configuration and active board sprints. With the current MCP these are `getJiraBoardConfig` and `listJiraBoardSprints`, called through `executeRead` with top-level cloudId and flat `inputs`. Discover operations before executing them in a new session; use the returned schemas. -3. Read board configuration for its saved filter, estimation field and rank field. List sprints with `boardId: 1`, `state: "active"`; follow offset pagination. Use all active sprint IDs on this board, naming each if there are several. Read their goals for current product-ordering guidance. Resolve these afresh on each invocation; never persist the current sprint ID in this skill. If there is no active sprint, report that and stop. -4. Call the primary `searchJiraIssuesUsingJql` tool with the resolved board filter and sprint IDs. This preserves board scope and lets Jira sort the results. Substitute actual values in this template: - - ```jql - filter = - AND project = QB - AND assignee = currentUser() - AND status = "Dev Ready" - AND sprint IN () - ORDER BY priority DESC, Rank ASC - ``` - - Use `assignee = ""` if needed. `openSprints()` alone is not board-specific. Keep the saved filter condition because a sprint can contain issues outside this board's filter. If board metadata is unavailable, report the limitation before using a project-scoped `sprint IN openSprints()` fallback; label that result as unverified against board 1. -5. Fetch summary, description, status, assignee, priority, issue type, labels, parent, issue links, sprint membership and the board's estimation/rank fields. Use `view: "evidence"` for custom fields, or explicit field IDs from live metadata; explicit `fields` limits what is returned. This MCP maps custom field values under `fields.customFields` by human label, often as `{id, value}`. A missing value means unestimated, including when a field object exists without `value`. -6. Follow `nextPageToken` until `isLast` is true and deduplicate by issue key. Count the collected results. If pagination fails, label the list partial. If no issues match, report the exact scope without widening status, assignee or sprint. - -On a tool error, retry once with corrected input. If an operation is missing, discover again with different keywords. If MCP access remains unavailable, provide the reproducible JQL and state that live results could not be verified. - -## Choose the order - -**Product priority (default):** Apply explicit current sprint-goal or product ordering when it can be grounded in ticket fields/descriptions; state that rule. Otherwise use Jira priority descending, then board Rank ascending. Retain Jira's returned order for ties instead of alphabetically sorting priority labels. Read phase from labels or explicit scope; an absent phase is unknown. Board rank and Jira priority are ordering signals, not proof of a product owner's decision. Ticket descriptions define scope and acceptance criteria; comments are discussion only. Read linked Confluence specifications when needed to resolve a material prioritisation ambiguity. - -**Complexity:** Default to easiest-first; reverse known estimates for hardest-first. Use the board's estimate as an effort proxy and label its unit (e.g. story points, not days). Tie-break using product order. Put unestimated tickets after estimated ones and retain their product order. Do not mix story points and time estimates on one numeric scale. If asked to estimate missing complexity, read descriptions/acceptance criteria and relevant dependencies, assign provisional Low/Medium/High with a short reason, and distinguish these judgements from recorded Jira estimates. Do not infer effort from priority or title alone. - -**Both:** Show the product-ordered table once, then a short easiest-first sequence of ticket links with estimates. Keep the two ordering criteria visible rather than inventing a blended score. - -For recommendations, flag explicit unresolved prerequisites and unclear acceptance criteria in the matching ticket's row. Check linked issue status before claiming a dependency is still open. Preserve link direction: “blocks” is not “is blocked by”. Read descriptions for whether a dependency actually prevents starting, or permits parallel work. Keep matching tickets in the list even when blocked or their description claims delivery; surface that inconsistency. Recommend a next ticket with enough scope to start, explaining any departure from product order. - -## Return - -Lead with the number of matches, sprint name(s), assignee and ordering rule. Use a compact table: - -| Order | Ticket | Summary | Jira priority | Estimate / complexity | Reason / dependency | -|---|---|---|---|---|---| - -Link every key as `[QB-123](https://coterieholdings.atlassian.net/browse/QB-123)`. Include a filtered Jira search link using `https://coterieholdings.atlassian.net/issues/?jql=` plus the URL-encoded exact JQL. This link reproduces the selected sprint; the next skill invocation resolves the then-active sprint. Explain when recommendation or complexity order differs from the link's Jira sort. Include the board link for board navigation. Finish with one concise next-ticket recommendation when requested or useful. Write in plain British English. - -Example invocations: - -- `$ics-jira-dev-ready` — my queue by product priority. -- `$ics-jira-dev-ready easiest first` — smallest recorded estimates first. -- `$ics-jira-dev-ready hardest first` — largest recorded estimates first. -- `$ics-jira-dev-ready show both orders` — product queue and easiest-first sequence. diff --git a/codex/skills/ics-jira-dev-ready/agents/openai.yaml b/codex/skills/ics-jira-dev-ready/agents/openai.yaml deleted file mode 100644 index 2641c8e..0000000 --- a/codex/skills/ics-jira-dev-ready/agents/openai.yaml +++ /dev/null @@ -1,4 +0,0 @@ -interface: - display_name: "ICS Jira Dev Ready" - short_description: "Find and prioritise your active sprint tickets" - default_prompt: "Use $ics-jira-dev-ready to list my Dev Ready tickets in the active QB sprint, with links, ordered by product priority." diff --git a/tests/test_codex_sync.py b/tests/test_codex_sync.py index 3c7202b..e380407 100644 --- a/tests/test_codex_sync.py +++ b/tests/test_codex_sync.py @@ -107,7 +107,7 @@ def test_fresh_restore_and_portability(self): self.assertIn("firecrawl", restored["mcp_servers"]) self.assertTrue((live / "skills/worktree/SKILL.md").exists()) self.assertTrue((live / "skills/worktree-cleanup/references/docker.md").exists()) - self.assertTrue((live / "skills/ics-jira-dev-ready/SKILL.md").exists()) + self.assertFalse((live / "skills/ics-jira-dev-ready").exists()) skill = live / "skills/software-engineer/SKILL.md" self.assertTrue(skill.exists()) self.assertTrue((skill.parent / "agents/openai.yaml").exists()) From 4162de0d3043a829787f46a8cd478c94d6370474 Mon Sep 17 00:00:00 2001 From: tawanorg Date: Thu, 10 Sep 2026 16:19:17 +1000 Subject: [PATCH 4/4] docs: add team guide for software engineer workflow --- README.md | 4 + codex/README.md | 150 ++++++++++++++++++++++++++++ codex/agents/software-engineer.toml | 1 + 3 files changed, 155 insertions(+) create mode 100644 codex/README.md diff --git a/README.md b/README.md index 4c38427..b5b36ac 100644 --- a/README.md +++ b/README.md @@ -461,6 +461,10 @@ In a new Codex session, select **Software Engineer** in the skills picker or type `$software-engineer implement [task]`. This user-scope skill loads the same agent instructions and can delegate to the custom role when useful. +**Sharing with the team:** [Software Engineer setup and usage](codex/README.md) +explains the workflow, installation without adopting these other dotfiles, +example requests, optional tools, and subscription usage. + `codex-sync install --external` also installs Open Code Review CLI **1.11.7** and its native Codex plugin. Global instructions default OCR to **delegation mode**, which uses the current Codex session for reasoning without a separate diff --git a/codex/README.md b/codex/README.md new file mode 100644 index 0000000..3680b26 --- /dev/null +++ b/codex/README.md @@ -0,0 +1,150 @@ +# Software Engineer for Codex + +A reusable implementation workflow for your team's existing Codex setup. +Give it a feature, bug, or refactor; it reads the relevant code, makes the +change, and checks that the requested behavior works. + +It is useful across projects because it discovers each repository's stack, +commands, and conventions. It inherits your selected Codex model and permissions. + +## What it does + +1. **Understands the task.** Reads project instructions, locates the code and its + callers, and identifies what would demonstrate success. Asks about missing + product decisions when they affect the implementation. +2. **Implements the change.** Follows existing patterns, preserves unrelated work, + and updates affected callers and error handling. Adds meaningful regression + coverage for bugs where feasible. +3. **Verifies the result.** Runs relevant tests and project-required checks, + inspects the diff, and reports what passed and what remains unverified. + +Small tasks proceed directly. Larger tasks get a short working plan. The workflow +does not require a separate planning interview for every edit. Repository rules +for security, financial calculations, migrations, and approvals still apply. + +## Use it + +Start a new Codex conversation inside the project after installation. Select +**Software Engineer** in the skills picker or type: + +```text +$software-engineer implement pagination for the customer list. Show 20 customers +per page and preserve the current filters when switching pages. +``` + +Other examples: + +```text +$software-engineer fix the search filter resetting when I return from a detail page. +Add a regression test and run the relevant checks. +``` + +```text +$software-engineer refactor the CSV parser while preserving its public behavior. +Work in the current thread. +``` + +Describe the desired behavior and any constraints. For bugs, include reproduction +steps or an error message if you have them. You do not need to name every tool. + +The **skill** (`software-engineer`, with a hyphen) is the selectable entry point. +The **custom agent** (`software_engineer`, with an underscore) is a role Codex can +spawn for a bounded implementation assignment. Both use the same instructions. +The custom role itself is not an `@` plugin picker entry. + +For small tasks the skill can work in the current conversation. For useful +delegated work it can launch the custom agent; if that role is unavailable it +uses the same instructions in the current conversation and says so. To avoid +the extra context of an implementation subagent, explicitly ask to work in the +current thread, as in the example above. + +## Install just this workflow + +Use an up-to-date Codex client with skills and custom-agent support, signed into +your own account. Obtain a checkout of this dotfiles repository containing this +README, then run the following **from the repository root**: + +```bash +engineer_codex_dir="${CODEX_HOME:-$HOME/.codex}" +mkdir -p "$engineer_codex_dir/agents" "$engineer_codex_dir/skills/software-engineer/agents" +cp -i codex/agents/software-engineer.toml "$engineer_codex_dir/agents/" +cp -i codex/agents/software-engineer-guide.md "$engineer_codex_dir/agents/" +cp -i codex/skills/software-engineer/SKILL.md "$engineer_codex_dir/skills/software-engineer/" +cp -i codex/skills/software-engineer/agents/openai.yaml "$engineer_codex_dir/skills/software-engineer/agents/" +``` + +`cp -i` asks before replacing an existing file. Start a new conversation afterward; +restart Codex if the skill list has not refreshed. This installs at user scope, +making the workflow available across repositories on that machine. + +These commands install only this workflow. The repository's full `install.sh` +and `codex-sync install` also restore the owner's other dotfile/Codex preferences; +teammates do not need to adopt those to use this agent. + +## Open Code Review and costs + +For substantial or high-risk changes, the engineer uses **Open Code Review +delegation mode** when installed. OCR supplies file selection and review rules; +the implementing Codex thread performs the review, addresses findings, and +reruns affected checks. It does not launch a separate reviewer by default. +Small changes receive a direct self-review. + +To install the optional OCR integration: + +```bash +npm install -g @alibaba-group/open-code-review@1.11.7 +codex plugin marketplace add alibaba/open-code-review +codex plugin add open-code-review-codex@open-code-review +``` + +OCR requires Git 2.41 or later. Start a new Codex conversation after installing +the plugin. The engineer's instructions select delegation mode; no OCR LLM +endpoint needs configuring. If invoking the plugin directly outside the engineer +workflow, say `@Open Code Review use delegation mode to review my current changes`. + +**Delegation is not free reasoning.** With subscription-authenticated Codex, +implementation and review consume your subscription allowance. No separate OCR +model API credentials are required. A Codex session authenticated through an API +uses that account's billing instead. Optional services, including hosted +Firecrawl, may charge or use their own credits. Token savings have not been measured. + +## Optional tools + +The agent uses tools already connected to your Codex session. Installing the +agent does not install these servers or copy anyone's credentials. + +| Tool | When it helps | +| --- | --- | +| Serena | Find symbols, definitions, and callers without reading entire files. | +| Context7 or a dedicated docs MCP | Check version-sensitive library APIs. | +| Firecrawl | Research public documentation gaps with focused queries. | +| Project tests and browser/device tools | Verify behavior and visible UI changes. | +| Matt Pocock engineering skills | Apply test-first, debugging, or interface-design guidance when relevant. | + +Optional tools have fallbacks: targeted local searches, existing project tests, +and explicit reporting of checks that could not run. Private code, client data, +and credentials should not be sent in public web queries. + +[Matt Pocock's skills](https://github.com/mattpocock/skills), +[Serena](https://github.com/oraios/serena), and +[Aider](https://github.com/Aider-AI/aider) informed the approach. +[OpenHands](https://github.com/OpenHands/OpenHands) was researched as an option +for always-on orchestration. Aider and OpenHands are separate applications; +neither is embedded in this agent. + +## What is ready, and what has been tested + +The user-scope skill is discoverable by Codex. The custom role successfully +completed a disposable pagination fix: it reported three regression tests failing +before the fix, all five tests passed afterward, and an unrelated file was +preserved. Dotfiles restore tests cover the skill, agent, guide, and OCR settings. + +This is a working setup with a small implementation smoke test, not an enterprise +reliability benchmark. That trial did not exercise every optional tool or a full +OCR review. Test it against your team's real tasks and retain your normal review +and CI requirements. It is not an unattended background service; it works when +invoked, within the session's permissions and requested scope. + +Maintainers: [agent instructions](agents/software-engineer.toml), +[skill entry point](skills/software-engineer/SKILL.md), and +[design notes](agents/software-engineer-guide.md). diff --git a/codex/agents/software-engineer.toml b/codex/agents/software-engineer.toml index 3819274..736ce45 100644 --- a/codex/agents/software-engineer.toml +++ b/codex/agents/software-engineer.toml @@ -19,6 +19,7 @@ Verify and hand off - Inspect the final diff for unintended changes and check the relevant failure modes: authorization, input validation, missing data, retries/concurrency, compatibility, and error reporting. These checks are internal work, not a mandatory approval ceremony. - For substantial or high-risk changes, use Open Code Review's open-code-review-delegate skill when available, scoped to this change. OCR supplies file selection/rules; this Codex session supplies reasoning. Never fall back to OCR-managed LLM calls, ocr review/scan, or endpoint configuration without explicit user authorization. Account for excluded files in the applicable project review requirements. If OCR is unavailable, do an explicit self-review and state the limitation; do not imply independent review. - Address in-scope findings and rerun affected checks. Completion means every acceptance criterion is verified or clearly marked unverified with a reason. Report the outcome, decisive validation evidence, and remaining blockers concisely. A passing version/config check is not proof of runtime behavior. +- Keep the OCR delegation review in the implementing thread. Spawn a separate reviewer only when the user explicitly requests one. Tools and context - Discover available tools; inherit MCP connections rather than hardcoding servers, credentials, models, or project paths. If a tool is absent or fails, use an appropriate local fallback and name the limitation once. Missing optional tooling does not block implementation.