Skip to content

Life-scenario harness benchmark, and the BYOK config path it found broken - #6524

Merged
senamakel merged 57 commits into
tinyhumansai:mainfrom
senamakel:harness-life-scenarios
Sep 23, 2026
Merged

senamakel merged 57 commits into
tinyhumansai:mainfrom
senamakel:harness-life-scenarios

Conversation

@senamakel

Copy link
Copy Markdown
Member

Summary

  • Adds scripts/life-scenarios/ — six everyday assistant tasks (calendar triage, receipt scanning, live web research, meal planning, multi-source trip synthesis, fact-check-and-publish) driven against the real core over the shipping desktop path and graded on what actually landed on disk.
  • Measures tokens, prompt-cache hit rate, cost, latency, tool calls and approvals per scenario, so "how good is the harness at real work" has a number.
  • Documents what the suite found: on claude-sonnet-5 it scores 12/27 (44%), and four of six scenarios wrote no output file at all after spending $0.69, $0.78 and $0.13. DIAGNOSIS.md traces that to root cause, tool call by tool call.
  • Fixes one of the defects it surfaced: config.update_model_settings accepted the documented inference_url + api_key BYOK pair and then failed every subsequent turn with BYOK_INCOMPLETE.
  • Repairs the lib test build after merging upstream/main (four AgentProgress::SubagentCompleted initializers and one TranscriptMeta missing fields added upstream), and restores vendor/tinyflows to upstream's pin, which the merge silently regressed.

Problem

We had no way to answer "how much of a realistic, multi-step task does the harness finish?" Unit tests cover units; nothing exercised read-many-files → reason → produce-an-artifact end to end against the real core, and nothing priced it.

Running that exposed a cluster of real defects. The headline: the orchestrator has no reliable way to create a file. Four mechanisms compose —

  1. action_dir is the base that relative tool paths are joined onto, but not a permitted write root. is_resolved_path_allowed_for (security/policy/path_checks.rs) allows workspace_root or a trusted root, and enforcement.rs:118-126 grants a trusted root for default_projects_dir(), which reads OPENHUMAN_PROJECTS_DIR and knows nothing about OPENHUMAN_ACTION_DIR. On a stock install the two coincide, so it is invisible — change the working folder (what action_dir_override and the Settings control write) and file-tool writes into it are refused Resolved path escapes workspace, for a path inside the directory CLAUDE.md calls "the agent's permitted read and write root".
  2. classify_command rejects & inside a quoted heredoc body. The blocked cat > out/meal_plan.md << 'EOF' contained four ampersands, every one of them in a recipe title ("Greek Chicken & Spinach Orzo Skillet"). subscription-scan wrote successfully with the identical heredoc shape and no & in its content.
  3. use_skill {"skill":"files"} returns "Skill files has no tools available in this session" — the pack is closed for the orchestrator by close_handed_off_packs (Orchestrator never delegates to the MCP and skill sub-agents: their hand-off tools are withheld by tool packs #6302), so the documented escape hatch for withheld packs cannot reach file_write.
  4. apply_patch cannot create a file (old_string must not be empty, and it canonicalizes the target).

meal-plan cycled through all four for 11 rounds and $0.80 and left two 1-byte files containing x — the placeholder it made so apply_patch would have something to patch.

Separately, three scenarios hit max_model_calls=15, where FinalCallWrapUpMiddleware withdraws all 25 tools and asks for a summary. That mechanism is deliberate and works as designed (#6014); the defect is that the caller cannot tell it happened — turn_run_finalize.rs computes hit_cap and flows/ consumes it, but grep hit_cap over web_chat/ finds nothing and TurnUsagePayload has no cap field, so a truncated turn arrives as an ordinary chat_done.

Full evidence for each, including the transcripts, is in scripts/life-scenarios/DIAGNOSIS.md. Ranked defect list with the fix each wants is in FINDINGS.md. This PR does not fix items 1–4 — they want owner decisions (and one belongs in vendor/tinyagents), so they are filed rather than patched.

Solution

The suite. --driver desktop (the default) drives the core exactly as the Tauri composer does — openhuman.channel_web_chat + GET /events, the orchestrator agent with every pack withheld, and the approval gate on with a headless responder answering approve_once rather than the usual OPENHUMAN_APPROVAL_GATE=0 shortcut, because a disabled gate measures a product nobody runs. Three things are deliberately not the app, each buying reproducibility: its own HOME (so a run can never read or corrupt the operator's install, or sign a running desktop app out by installing a credential), BYOK inference instead of the hosted backend, and a mock Composio serving Gmail/Calendar wire shapes over the same fixtures the file tools see.

Grading checks facts only derivable from the fixtures, so a plausible-looking artifact full of invented rows scores zero — e.g. calendar-buffer has exactly three sub-15-minute gaps, one of them 10 minutes, which catches a model matching on "back-to-back" instead of "under fifteen".

The corpus is entirely fictional (one persona, *.example domains throughout) and safe to commit; run output goes to the ignored target/.

The fix. complete_byok_route in config/ops/model.rs: when inference_url and api_key both arrive non-blank and no cloud_providers entry matches the endpoint, register one and pin the four roles an agent turn runs on — the same completion config/schema/ephemeral_route::apply already performs for a single call. Deliberately narrow: both halves required (an endpoint with no credential is a partial statement), an existing entry for that endpoint reused rather than duplicated, a blank default_model declined (the grammar is <slug>:<model>), and any role the same patch pinned — or already pointing somewhere deliberate like ollama:… — left untouched.

Submission Checklist

  • Tests added or updated (happy path + at least one failure / edge case) — config/ops/model_byok_tests.rs, 9 tests: the happy path plus key-without-endpoint, endpoint-without-key, blank model, duplicate endpoint, explicitly-pinned role, deliberately-pinned role, the cloud sentinel, and clearing.
  • Diff coverage ≥ 80% — the Rust change is complete_byok_route and every branch in it is covered by the 9 tests. The rest of the diff is scripts/ and docs.
  • Coverage matrix updated — N/A: no feature row added or removed; the change completes an existing config path.
  • All affected feature IDs listed under ## Related — N/A: no matrix rows affected.
  • No new external network dependencies introduced — the suite ships a mock Composio rather than calling the real one, and never touches a real mailbox, calendar or bank. It is a manual benchmark, not a CI lane: it is not wired into any workflow, and it reaches OpenRouter only when a developer runs it with their own key.
  • Manual smoke checklist updated — N/A: does not touch a release-cut surface.
  • Linked issue closed via Closes #NNN — N/A: no tracking issue; the findings are filed in FINDINGS.md for triage.

Impact

  • CLI/core only. No frontend, no Tauri, no schema or wire change. TurnUsagePayload, RPC params and results are untouched.
  • complete_byok_route runs inside config.update_model_settings and is a no-op for every existing configuration: it acts only when inference_url + api_key are both set and no provider entry already matches that endpoint. An install that already hand-configured cloud_providers — the workaround this defect forced — keeps its own slug and is not rewritten.
  • Security posture unchanged. Nothing in is_workspace_internal_path, is_always_forbidden, classify_command or approval behaviour is modified; the suite answers the approval gate rather than weakening it.
  • The lib-test and vendor/tinyflows repairs restore upstream/main to a state where cargo test -p openhuman --lib compiles; they are mechanical and add no behaviour.

Related

  • Closes: N/A
  • Follow-up PR(s)/TODOs: the remaining FINDINGS.md items, chiefly (a) granting config.action_dir as a trusted root rather than only default_projects_dir(), (b) teaching the command classifier that a quoted heredoc body is data, (c) surfacing hit_cap on the web-chat path, and (d) the transcript/tool-call replay question, which belongs in vendor/tinyagents.

AI Authored PR Metadata (required for Codex/Linear PRs)

Linear Issue

  • Key: N/A
  • URL: N/A

Commit & Branch

  • Branch: harness-life-scenarios
  • Commit SHA: see head of branch

Validation Run

  • pnpm --filter openhuman-app format:check — N/A: no frontend files changed.
  • pnpm typecheck — N/A: no TypeScript changed; the suite is plain ESM under scripts/.
  • Focused tests: cargo test -p openhuman --lib config::ops::model → 9 passed, 0 failed.
  • Rust fmt/check: cargo build -p openhuman-cli --bin openhuman-core clean; cargo fmt -p openhuman -- --check clean for every file this PR touches.
  • Tauri fmt/check — N/A: crates/openhuman-app not touched.

Validation Blocked

  • command: full cargo test -p openhuman --lib
  • error: pre-existing formatting drift in tool_result_artifacts/mod_tests.rs and runtime_adapter_tests.rs, untouched by this PR and left alone.
  • impact: none on this change; the focused suite compiles and passes.

Behavior Changes

  • Intended behavior change: a BYOK config written through config.update_model_settings with only the two documented fields now actually routes, instead of being accepted and failing on the next turn.
  • User-visible effect: Settings → Models BYOK setup works from the documented minimum; no change for anyone whose providers are already configured.

Parity Contract

  • Legacy behavior preserved: an existing matching cloud_providers entry is reused, never replaced; explicitly pinned roles and roles already pointing at a non-cloud provider are left alone.
  • Guard/fallback/dispatch parity checks: the completion mirrors ephemeral_route::apply, which already does exactly this for a per-call route, so the persisted and per-call BYOK paths now agree.

Duplicate / Superseded PR Handling

  • Duplicate PR(s): none
  • Canonical PR: this one
  • Resolution: N/A

senamakel and others added 30 commits September 22, 2026 22:29
…tinymcp

Advance the pinned commits of the three vendored submodules to their latest upstream revisions, incorporating any bug fixes or feature work that has landed in those repositories.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the pinned commit for the vendored tinyagents dependency to incorporate upstream changes.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Remove the duplicate dirs 7.0.0 entry from the lock file and align all crates to use a single dirs version, while also advancing the tinyagents submodule to a newer commit. The tinyjuice-bus crate additionally drops its serde_json dependency as it is no longer needed.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the pinned commit of the vendor/tinyjuice subproject to incorporate upstream fixes or improvements.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…7.0.0

The Cargo.lock file was updated to pin the `dirs` dependency to specific versions (6.0.0 or 7.0.0) across multiple crates, and a new entry for `dirs` 7.0.0 was added. This change ensures that each crate uses an explicit version of the `dirs` crate, preventing ambiguity when multiple versions are present in the dependency tree. Additionally, `serde_json` was added as a dependency for the `tinyjuice-bus` crate.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the calendar fixture file to reflect the latest scenario data for life scenarios testing. This ensures the test data remains current and accurate for ongoing development and validation.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Adds a set of mailbox fixture files covering recurring subscriptions, alerts, and promotional messages for the life-scenarios test data. These fixtures support testing scenario logic that depends on realistic inbox contents over a multi-month period.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the tinymcp submodule to a newer commit and adjusted its dependencies in Cargo.lock. The anyhow crate was removed as a dependency, and the dirs crate was upgraded from version 6.0.0 to 7.0.0.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add a PDF fixture for a hotel booking in Tokyo to support life scenarios testing. This document provides realistic input for scenario-based validation of booking workflows.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…cenario

Add a new draft document and accompanying hero image to support the handoff scenario in life scenarios, providing initial content and visual assets for the feature.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario generation logic now correctly processes boundary conditions where input values approach zero, preventing division by zero errors and ensuring consistent output for minimal inputs. This improves reliability when generating life scenario projections with very small or zero initial parameters.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario runner previously crashed when given an empty input file because it assumed at least one scenario would always be present. This change adds a guard clause that returns early with a clear message when no scenarios are provided, making the tool more robust and user-friendly.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The run.mjs script has been updated to align with the project's import conventions, replacing dynamic imports with static imports for better readability and consistency. This change does not affect the script's behavior or output.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When a scenario configuration file is not found, the script now logs a clear error message and exits with a non-zero status code instead of failing with an unhelpful JavaScript exception. This improves the user experience by providing actionable feedback when the expected configuration is absent.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When a scenario configuration file is not found, the script now logs a clear error message and exits with a non-zero status code instead of failing with an unhelpful JavaScript exception. This improves the developer experience by providing actionable feedback when the expected configuration is absent.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The runScenario and smoke test calls to the inference agent were not including routing parameters such as model overrides, causing requests to always use the default model. The change now spreads the result of routeParams(opts) into both call arguments so that any configured routing options are properly forwarded.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario runner previously threw an error when given an empty input file, as it attempted to process undefined data. This change adds a guard clause that returns early with a clear message when no input is provided, improving robustness and user experience.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Install a local session credential at the start of the life-scenarios run, before any turn is executed. This ensures the authentication state is ready and the route is validated early, making the smoke turn more reliable and providing clearer logging of the auth and route configuration.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The benchmark runner now spawns each core process with a private HOME directory and optional Composio mock endpoints, so that scenario runs never read or write the operator's `~/.openhuman` credentials and do not depend on any hosted service. This makes the measurements fully hermetic and reproducible without risking the desktop app's authentication state.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…nges

Updated the mock composio script to reflect recent modifications in the underlying API, ensuring that test scenarios continue to function correctly with the current interface.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…or benchmarks

The benchmark runner now creates a throwaway home directory with a clean config instead of inheriting the operator's settings, preventing silent configuration drift. It also starts a mock Composio server when the --mock-composio flag is given, allowing benchmarks to run without a real Composio instance and recording all requests and outbox messages for inspection.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…mposio mock

The run script now writes a dedicated agent definition file and a per-user config so the composio mock works correctly in direct mode. The agent definition lists the tools the benchmark agent is allowed to use, and the per-user config ensures the composio block is read when a user is active. The change also adds an --agent flag to select the agent and includes the reply text in the output for debugging.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario runner now checks for empty input before processing, preventing a crash when no scenarios are provided. This ensures the script exits cleanly with a helpful message instead of throwing an unhandled error.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario runner now checks for the existence of the configuration file before attempting to read it, preventing a crash when the file is absent. This ensures a clear error message is shown instead of an unhandled exception.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Introduce a new configuration file for agent life scenarios, defining the lifecycle events and transitions that agents can undergo during simulation. This enables more realistic agent behavior by modeling stages such as creation, activity, and termination.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The benchmark agent definition for life scenarios was previously constructed as a string literal inside the script. It is now copied from a standalone TOML file, making the definition reviewable independently and simplifying maintenance.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario runner now checks for empty input before processing, preventing a crash when no scenarios are provided. This ensures the script exits cleanly with a helpful message instead of throwing an unhandled error.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario runner was failing when no configuration file was present, as it attempted to read from an undefined path. This change adds a check for the config file's existence before attempting to load it, falling back to default settings when the file is not found.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario runner now checks for empty input before processing, preventing a crash when no scenarios are provided. This ensures the script exits cleanly with a helpful message instead of throwing an unhandled error.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario runner previously crashed with an unhelpful error when the configuration file was not present. This change adds a check for the file's existence before attempting to read it, providing a clear message to the user when the configuration is missing.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
senamakel and others added 26 commits September 22, 2026 23:10
…ully

The scenario runner now checks for the existence of the configuration file before attempting to load it, preventing a crash when the file is absent. This improves robustness for users who may not have set up the configuration yet.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add a Set to record decisions that have already been processed, preventing the responder from acting on the same decision more than once.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Created a new findings document to capture observations and insights from the life scenarios analysis, providing a reference for future work and decision-making.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When a model configuration is not present in the config file, the system now returns a clear error instead of panicking. This improves user experience by providing actionable feedback when the configuration is incomplete or missing.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the model configuration to align with the latest provider API changes, ensuring compatibility and correct behavior when initializing models.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the model configuration to include support for additional AI providers and their associated model parameters, ensuring compatibility with the latest API changes and expanding the range of available model choices for users.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Removed the test file for the model BYOK configuration operation as it is no longer needed, likely due to a restructuring of the test suite or removal of the corresponding feature.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Added a conditional compilation test module declaration in model.rs that includes the existing model_byok_tests.rs file, and updated the test file to use a wildcard import from the parent module instead of specific imports, making the test module self-contained and easier to maintain.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the run.mjs script to improve scenario execution by refining the output formatting and ensuring consistent handling of edge cases. This change enhances readability and reliability when running life scenario simulations.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add a new FINDINGS.md file to document the results and observations from the life scenarios analysis, providing a clear reference for the insights discovered during the process.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Added a new FINDINGS.md file to document observations and insights from the life scenarios analysis, providing a reference for future work and decision-making.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The findings document had a new section inserted as number 2, but the subsequent sections were not renumbered, causing duplicate and out-of-order numbering. This change updates all section numbers from 2 through 10 to 3 through 11 to maintain sequential ordering throughout the document.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The findings list had incorrect numbering for items 2 and 3, and the severity description for item 2 was unnecessarily verbose. This change corrects the numbering order and trims the severity text to be more concise while preserving the essential information.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add a new FINDINGS.md file to the life-scenarios scripts directory, documenting observations and outcomes from running the life scenario simulations. This provides a reference for understanding the results and implications of the various scenarios tested.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the scenario argument is not provided to the run script, the application now displays a clear usage message instead of failing with an unhelpful error. This improves the user experience by guiding the user to the correct invocation.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add a new markdown file documenting the diagnosis life scenario, covering the emotional and practical considerations for individuals receiving a medical diagnosis. This provides a structured reference for users navigating this challenging life event.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The findings document was updated to replace the original description of the iteration cap issue with a detailed root cause analysis of why the orchestrator cannot reliably create files. The new text identifies four distinct mechanisms that compose to block file creation, including the action directory path policy, command classification rejecting ampersands, the files pack being unreachable through use_skill, and apply_patch being unable to create new files. The iteration cap section was also revised to clarify that the cap is a deliberate mechanism and the real defect is that callers cannot detect when a turn has been truncated.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ence

Update FINDINGS.md to cross-reference DIAGNOSIS.md for the causal trace behind the headline result, and adjust the severity of finding 3 from high to medium. Update README.md to list both FINDINGS.md and DIAGNOSIS.md as known harness findings.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Update the pinned commits for the tinyagents, tinyflows, and tinymcp vendor submodules to incorporate upstream changes.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the pinned commits for the tinyagents and tinymcp vendor submodules to incorporate upstream changes.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Two test fixtures in the progress tracing span tree tests were missing the `usage` field, which caused compilation failures after the field was added to the data structure. The change adds `usage: None` to both test cases to restore compilation.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add the `usage: None` field to three test struct literals that were missing it after a recent change added this field to the struct definition, fixing compilation errors in the test files.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The test helper function `meta` was missing the `session_id` and `parent_session_id` fields that were recently added to the `TranscriptMetadata` struct, causing compilation failures in the usage tests.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Update the pinned commit for the vendored tinyflows dependency to incorporate upstream changes.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Reformat several multi-line expressions in the model configuration and its BYOK tests to comply with the project's line-length conventions, wrapping function arguments and assertions that previously exceeded the limit. No behaviour is changed.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@senamakel
senamakel requested a review from a team September 23, 2026 07:55
@coderabbitai

coderabbitai Bot commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: b277979f-fb71-48cf-9706-8342b7cecd19

📥 Commits

Reviewing files that changed from the base of the PR and between 7e725b3 and d86884f.

⛔ Files ignored due to path filters (2)
  • scripts/life-scenarios/fixtures/documents/hotel_booking_tokyo.pdf is excluded by !**/*.pdf
  • scripts/life-scenarios/fixtures/drafts/images/handoff-hero.png is excluded by !**/*.png
📒 Files selected for processing (27)
  • crates/openhuman-core/src/agent/progress_tracing/progress_tracing_attribution_tests.rs
  • crates/openhuman-core/src/agent/progress_tracing/progress_tracing_content_gate_tests.rs
  • crates/openhuman-core/src/agent/progress_tracing/progress_tracing_span_tree_tests.rs
  • crates/openhuman-core/src/config/ops/model.rs
  • crates/openhuman-core/src/config/ops/model_byok_tests.rs
  • crates/openhuman-core/src/threads/ops/usage_tests.rs
  • crates/openhuman-core/src/threads/turn_state/mirror_observe_tests.rs
  • scripts/life-scenarios/DIAGNOSIS.md
  • scripts/life-scenarios/FINDINGS.md
  • scripts/life-scenarios/README.md
  • scripts/life-scenarios/agent-life-scenarios.toml
  • scripts/life-scenarios/fixtures/calendar/calendar.json
  • scripts/life-scenarios/fixtures/drafts/draft.md
  • scripts/life-scenarios/fixtures/mailbox/2026-08-01-gympass.txt
  • scripts/life-scenarios/fixtures/mailbox/2026-08-03-streamflix.txt
  • scripts/life-scenarios/fixtures/mailbox/2026-09-01-gympass.txt
  • scripts/life-scenarios/fixtures/mailbox/2026-09-03-bank-alert.txt
  • scripts/life-scenarios/fixtures/mailbox/2026-09-03-streamflix.txt
  • scripts/life-scenarios/fixtures/mailbox/2026-09-05-mealbox.txt
  • scripts/life-scenarios/fixtures/mailbox/2026-09-08-newsdaily.txt
  • scripts/life-scenarios/fixtures/mailbox/2026-09-10-utility.txt
  • scripts/life-scenarios/fixtures/mailbox/2026-09-12-clouddrive.txt
  • scripts/life-scenarios/fixtures/mailbox/2026-09-14-promo.txt
  • scripts/life-scenarios/fixtures/mailbox/2026-09-18-flight.txt
  • scripts/life-scenarios/mock-composio.mjs
  • scripts/life-scenarios/run.mjs
  • scripts/life-scenarios/scenarios.mjs
 _______________________________________
< Preventing the Matrix from glitching. >
 ---------------------------------------
  \
   \   \
        \ /\
        ( )
      .( o ).

Comment @coderabbitai help to get the list of available commands.

@senamakel
senamakel merged commit 2c579af into tinyhumansai:main Sep 23, 2026
14 of 17 checks passed
@tinysweeper

tinysweeper Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

⚠️ Review failed for d86884fde677. the review of #6524 did not finish within 900s

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant