Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -265,6 +265,9 @@ EXECUTION_TIMEOUT_MS=
# "enforce" (default) fails spawn and build on oversized secret payloads;
# "warn" only logs. Cloudflare: not set by Terraform.
SECRETS_CAP_ENFORCEMENT=
# Session team enforcement: off, shadow, or on. Unset defaults to shadow.
# Private visibility applies in every mode. Cloudflare: var.teams_enforcement.
TEAMS_ENFORCEMENT=

# ---------------------------------------------------------------------------
# Logging
Expand Down
9 changes: 3 additions & 6 deletions .github/workflows/ci-python.yml
Original file line number Diff line number Diff line change
Expand Up @@ -220,16 +220,13 @@ jobs:
python-version: "3.12"
cache: "pip"

- name: Setup frozen image lock checker
- name: Setup uv
uses: astral-sh/setup-uv@v7
with:
version: "0.9.7"

- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install -e ../sandbox-runtime
pip install -e ".[dev]"
run: uv sync --frozen --extra dev

- name: Run tests
run: pytest tests/ -v
run: uv run --frozen --extra dev pytest tests/ -v
14 changes: 14 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -93,6 +93,8 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@v7
with:
persist-credentials: false

- name: Setup Node.js
uses: actions/setup-node@v7
Expand Down Expand Up @@ -151,6 +153,8 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@v7
with:
persist-credentials: false

- name: Setup Node.js
uses: actions/setup-node@v7
Expand All @@ -174,6 +178,8 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@v7
with:
persist-credentials: false

- name: Setup Node.js
uses: actions/setup-node@v7
Expand Down Expand Up @@ -221,6 +227,8 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@v7
with:
persist-credentials: false

- name: Setup Node.js
uses: actions/setup-node@v7
Expand Down Expand Up @@ -254,6 +262,8 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@v7
with:
persist-credentials: false

- name: Setup Node.js
uses: actions/setup-node@v7
Expand All @@ -279,6 +289,8 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@v7
with:
persist-credentials: false

- name: Setup Node.js
uses: actions/setup-node@v7
Expand Down Expand Up @@ -324,6 +336,8 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@v7
with:
persist-credentials: false

- name: Setup Node.js
uses: actions/setup-node@v7
Expand Down
14 changes: 14 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,10 +2,24 @@

New features, integrations, and notable improvements to Open-Inspect — newest first.

## September 29, 2026

### Added

`TEAMS_ENFORCEMENT` controls active-user session item routes (`/sessions/:id` and its subpaths)
using the persisted session row (`off`, `shadow` by default, or `on`). On those routes, private
visibility applies in every mode; team visibility and the delete ownership rule apply when `on`.
Workspace-wide session lists, bulk export, and WebSocket authorization follow in subsequent changes.
No route can make a session private or team-owned before those changes land.

## September 28, 2026

### Added

**Claude Sonnet 5.5.** Adds `anthropic/claude-sonnet-5-5` to the model picker and integrations, with
adaptive thinking controls from low through max. Claude Agent SDK 0.2.161 bundles Claude Code
2.1.284, which supports the new model.

OpenCode sessions using a connected ChatGPT subscription now report estimated model costs through
the existing session cost display and spending limit. These are API-price equivalents, not
additional subscription charges or an OpenAI invoice; estimates remain zero if catalog pricing is
Expand Down
11 changes: 3 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -209,19 +209,14 @@ await configureGitIdentity({

Choose the AI model that fits your task, with per-session reasoning effort controls:

| Provider | Models |
| ---------------- | ----------------------------------------------------------------------- |
| Anthropic | Claude Haiku 4.5, Sonnet 4.5/4.6/5, Opus 4.5/4.6/4.7/4.8/5, Fable 5/5.1 |
| OpenAI | GPT 5.4, GPT 5.5, 5.3 Codex, 5.3 Codex Spark |
| xAI / SuperGrok | Grok models (opt-in) |
| OpenCode Zen | Kimi K2.5/K2.6/K3, MiniMax M2.5, Qwen3.7 Max, GLM 5/5.1/5.2 (opt-in) |
| Z.AI Coding Plan | GLM 5.2/5.3 (opt-in) |
Anthropic and OpenAI models are enabled by default. xAI / SuperGrok, OpenCode Zen and Go, Z.AI
Coding Plan, and DeepSeek models are opt-in. See [Available Models](docs/AVAILABLE_MODELS.md) for
current model IDs, descriptions, and reasoning efforts.

OpenAI models work with your existing ChatGPT subscription via OAuth — no separate API key needed.
Anthropic models can run on the **Claude Agent** harness with a connected Claude subscription; see
[Using the Claude Agent Harness](docs/CLAUDE_AGENT.md). Grok models work with an eligible SuperGrok
subscription through control-plane-managed OAuth. See
**[docs/AVAILABLE_MODELS.md](docs/AVAILABLE_MODELS.md)** for the full model list and
**[docs/OPENAI_MODELS.md](docs/OPENAI_MODELS.md)** or **[docs/GROK_MODELS.md](docs/GROK_MODELS.md)**
for subscription setup instructions.

Expand Down
9 changes: 5 additions & 4 deletions docs/AVAILABLE_MODELS.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,7 +35,8 @@ Accounts) applies only on the Claude Agent harness; OpenCode sessions use `ANTHR
| `anthropic/claude-haiku-4-5` | Claude Haiku 4.5 | Fast and efficient | high, max | max |
| `anthropic/claude-sonnet-4-5` | Claude Sonnet 4.5 | Balanced performance | high, max | max |
| `anthropic/claude-sonnet-4-6` | Claude Sonnet 4.6 | Balanced, fast coding | low, medium, high, max | high |
| `anthropic/claude-sonnet-5` | Claude Sonnet 5 | Latest Sonnet, adaptive thinking | low, medium, high, xhigh, max | high |
| `anthropic/claude-sonnet-5` | Claude Sonnet 5 | Balanced performance, adaptive thinking | low, medium, high, xhigh, max | high |
| `anthropic/claude-sonnet-5-5` | Claude Sonnet 5.5 | Latest Sonnet, fast and intelligent | low, medium, high, xhigh, max | high |
| `anthropic/claude-opus-4-5` | Claude Opus 4.5 | Most capable | high, max | max |
| `anthropic/claude-opus-4-6` | Claude Opus 4.6 | Most capable, adaptive thinking | low, medium, high, max | high |
| `anthropic/claude-opus-4-7` | Claude Opus 4.7 | Most capable, adaptive thinking | low, medium, high, xhigh, max | high |
Expand All @@ -54,9 +55,9 @@ OpenAI models support connected ChatGPT provider accounts or `OPENAI_API_KEY` mo
| ---------------------- | ------------- | ---------------------------------------------- | ----------------------------------- | -------------- |
| `openai/gpt-5.4` | GPT 5.4 | Flagship model | none, low, medium, high, xhigh | Not set |
| `openai/gpt-5.5` | GPT 5.5 | Latest flagship model | none, low, medium, high, xhigh | Not set |
| `openai/gpt-5.6-sol` | GPT 5.6 Sol | Frontier model for complex professional work | none, low, medium, high, xhigh | Not set |
| `openai/gpt-5.6-terra` | GPT 5.6 Terra | Balanced, cost-efficient everyday work | none, low, medium, high, xhigh | Not set |
| `openai/gpt-5.6-luna` | GPT 5.6 Luna | Fast, cost-efficient high-volume workloads | none, low, medium, high, xhigh | Not set |
| `openai/gpt-5.6-sol` | GPT 5.6 Sol | Frontier model for complex professional work | none, low, medium, high, xhigh | medium |
| `openai/gpt-5.6-terra` | GPT 5.6 Terra | Balanced, cost-efficient everyday work | none, low, medium, high, xhigh | medium |
| `openai/gpt-5.6-luna` | GPT 5.6 Luna | Fast, cost-efficient high-volume workloads | none, low, medium, high, xhigh, max | medium |
| `openai/gpt-6-astra` | GPT-6 Astra | Most capable model for complex, demanding work | low, medium, high, xhigh, max | medium |
| `openai/gpt-6-sol` | GPT-6 Sol | Complex coding and agentic workflows | none, low, medium, high, xhigh, max | medium |
| `openai/gpt-6-luna` | GPT-6 Luna | Efficient model for focused, high-volume tasks | none, low, medium, high, xhigh, max | medium |
Expand Down
8 changes: 3 additions & 5 deletions docs/CLAUDE_AGENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -196,8 +196,6 @@ fix instead.
guard, the harness ignores the result of any turn it did not submit.
- **Follow-ups queue.** Both harnesses hold follow-up prompts until the running turn completes.
- **Image.** The sandbox image pins `claude-agent-sdk`, whose wheel bundles the `claude` binary. The
runtime manifest names the generation carrying the current pin under `harnessMinimumGeneration`,
so a Claude session never boots a prebuilt image from before that generation; this floor does not
touch OpenCode sessions' images or snapshots, since the global compatibility floor did not move.
Raise this floor whenever the SDK pin moves for a model the catalog advertises, otherwise a
session can be handed an older image whose bundled `claude` does not know that model.
runtime manifest's `harnessMinimumGeneration` controls which prepared images new Claude sessions
can use. Older images and resumed snapshots may lack newer models until rebuilt; a model request
can fail on a sandbox whose bundled CLI does not support it.
42 changes: 42 additions & 0 deletions docs/MODAL_DOCKER.md
Original file line number Diff line number Diff line change
Expand Up @@ -69,6 +69,48 @@ its generation is older than the materialization bound: the launch window plus t
older launch that materializes later may briefly block a replacement, but the single allocation name
and fenced credentials prevent overlapping work.

The authenticated `POST /api-resolve-vm-sandbox` endpoint is lookup-only. Its body contains exactly
`{"session_id":"...","sandbox_id":"..."}`; it accepts no launch settings or secrets. It finds the
running allocation by session name, checks the generation's ownership tags, and returns:

```json
{
"success": true,
"data": {
"sandbox_id": "generation-id",
"modal_object_id": "sb-real-id",
"code_server_url": null,
"code_server_password": null,
"vnc_url": null,
"vnc_password": null,
"ttyd_url": null,
"tunnel_urls": null,
"sandbox_backend": "modal-vm"
}
}
```

New VM allocations record versioned service flags and effective ports in provider-owned launch tags.
Only services enabled by these tags return URLs/passwords; extra tunnels use port-to-URL mappings.
Legacy allocations without these tags (or with unknown/incomplete metadata) resolve only the real
`modal_object_id`, not access credentials or tunnels. Resolve never infers enabled services from
environment variables, which may have contained user secrets on older allocations. Such VMs need a
new launch to recover interactive access. Resolve neither creates nor retires an allocation or
writes tunnel configuration. A stopped VM is not discoverable by name. Resolve does not return a
terminal access token; the control plane mints that token only when it still holds the generation's
sandbox auth token in memory.

Create, restore, and resolve report typed HTTP 409 error `detail` values:

- `not_visible`: resolve found no named allocation.
- `other_generation`: the ownership tags do not match.
- `window_closed`: create/restore missed the launch deadline with no owned allocation.
- `race_pending`: create/restore cannot yet see the winner after `AlreadyExistsError`, or resolve
found a VM whose enabled tunnel URLs are not all visible yet.

Unexpected provider errors remain 500. The pending-reference stop endpoint retains its separate
`pending_reference_not_visible` response.

## Switching backends

Changing `SANDBOX_PROVIDER` is an operator cutover, not session migration. Existing sessions and
Expand Down
64 changes: 58 additions & 6 deletions packages/control-plane/src/authorization/request-audit.ts
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,8 @@ export type RouteAuthorizationDecision =
kind: "allowed";
admission: "user" | "service" | "sandbox";
auditAllowed: boolean;
shadowReason?: string;
shadowDenials?: readonly { sessionId: string; reason: string }[];
})
| (AuthorizationDecisionEvidence & {
kind: "denied";
Expand All @@ -40,7 +42,7 @@ export type RouteAuthorizationDecision =
export function shouldAuditAllowedDecision(
decision: Extract<RouteAuthorizationDecision, { kind: "allowed" }>
): boolean {
return decision.auditAllowed;
return decision.auditAllowed || !!decision.shadowReason || !!decision.shadowDenials?.length;
}

/**
Expand All @@ -59,6 +61,7 @@ export async function auditRouteAuthorizationDecision(input: {
path: string;
response: Response;
decision: RouteAuthorizationDecision;
teamId?: string | null;
}): Promise<void> {
const principal = input.ctx.principal;
if (!principal) return;
Expand All @@ -77,6 +80,20 @@ export async function auditRouteAuthorizationDecision(input: {
const action = allowed
? AUTHORIZATION_DECISION_ACTIONS.allowed
: AUTHORIZATION_DECISION_ACTIONS.denied;
const shadowCode =
decision.kind === "allowed"
? decision.shadowDenials?.length
? "shadow_denied:batch"
: decision.shadowReason
? `shadow_denied:${decision.shadowReason}`
: null
: null;
const teamId =
input.teamId !== undefined
? input.teamId
: input.ctx.childSessionAdmission
? input.ctx.childSessionAdmission.row.ownerTeamId
: (input.ctx.sessionAdmission?.row.ownerTeamId ?? null);
const metadata = {
schema: AUTHORIZATION_DECISION_METADATA_SCHEMA,
httpMethod: input.method,
Expand All @@ -87,11 +104,14 @@ export async function auditRouteAuthorizationDecision(input: {
? { effectivePermissions: decision.effectivePermissions }
: {}),
...(requiredPermission ? { requiredPermission } : {}),
responseCode: decision.kind === "denied" ? decision.reasonCode : null,
responseCode: decision.kind === "denied" ? decision.reasonCode : shadowCode,
responseReason: decision.kind === "denied" ? decision.reason : null,
requestId: input.ctx.request_id,
traceId: input.ctx.trace_id,
...(decision.kind === "allowed" ? { admission: decision.admission } : {}),
...(decision.kind === "allowed" && decision.shadowDenials?.length
? { shadowDenials: decision.shadowDenials }
: {}),
...(principal.kind === "service" && principal.actor
? {
actor: {
Expand All @@ -110,8 +130,8 @@ export async function auditRouteAuthorizationDecision(input: {
`INSERT INTO authorization_audit_events
(id, occurred_at, request_id, principal_kind,
actor_user_id_snapshot, actor_service_snapshot, action, resource_type, resource_id,
reason_code, operation_result, metadata_json)
VALUES (?, ?, ?, ?, ?, ?, ?, 'http_route', ?, ?, ?, ?)`
reason_code, operation_result, metadata_json, team_id)
VALUES (?, ?, ?, ?, ?, ?, ?, 'http_route', ?, ?, ?, ?, ?)`
)
.bind(
crypto.randomUUID(),
Expand All @@ -122,9 +142,10 @@ export async function auditRouteAuthorizationDecision(input: {
principal.kind === "service" ? principal.service : null,
action,
input.path,
decision.kind === "allowed" ? "authorization_allowed" : decision.reasonCode,
decision.kind === "allowed" ? (shadowCode ?? "authorization_allowed") : decision.reasonCode,
allowed ? "applied" : "denied",
JSON.stringify(metadata)
JSON.stringify(metadata),
teamId
)
.run();
} catch (cause) {
Expand All @@ -137,3 +158,34 @@ export async function auditRouteAuthorizationDecision(input: {
});
}
}

export async function auditPrivateSessionBreakGlass(
ctx: RequestContext,
sessionId: string,
teamId: string | null
): Promise<void> {
const principal = ctx.principal;
const actorUserId = ctx.authorization?.userId;
if (!principal || !actorUserId) throw new Error("Missing private session break-glass actor");
await ctx.db
.prepare(
`INSERT INTO authorization_audit_events
(id, occurred_at, request_id, principal_kind, actor_user_id_snapshot,
actor_service_snapshot, action, resource_type, resource_id, team_id,
reason_code, operation_result, metadata_json)
VALUES (?, ?, ?, ?, ?, ?, 'session.private_break_glass', 'session', ?, ?, ?, 'applied', ?)`
)
.bind(
crypto.randomUUID(),
Date.now(),
ctx.request_id,
principal.kind,
actorUserId,
principal.kind === "service" ? principal.service : null,
sessionId,
teamId,
"session.private_break_glass",
JSON.stringify({ before: {}, requested: {}, after: {} })
)
.run();
}
Comment thread
coderabbitai[bot] marked this conversation as resolved.
Loading
Loading