Life-scenario harness benchmark, and the BYOK config path it found broken - #6524
Merged
Merged
Conversation
…tinymcp Advance the pinned commits of the three vendored submodules to their latest upstream revisions, incorporating any bug fixes or feature work that has landed in those repositories. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the pinned commit for the vendored tinyagents dependency to incorporate upstream changes. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Remove the duplicate dirs 7.0.0 entry from the lock file and align all crates to use a single dirs version, while also advancing the tinyagents submodule to a newer commit. The tinyjuice-bus crate additionally drops its serde_json dependency as it is no longer needed. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the pinned commit of the vendor/tinyjuice subproject to incorporate upstream fixes or improvements. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…7.0.0 The Cargo.lock file was updated to pin the `dirs` dependency to specific versions (6.0.0 or 7.0.0) across multiple crates, and a new entry for `dirs` 7.0.0 was added. This change ensures that each crate uses an explicit version of the `dirs` crate, preventing ambiguity when multiple versions are present in the dependency tree. Additionally, `serde_json` was added as a dependency for the `tinyjuice-bus` crate. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the calendar fixture file to reflect the latest scenario data for life scenarios testing. This ensures the test data remains current and accurate for ongoing development and validation. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Adds a set of mailbox fixture files covering recurring subscriptions, alerts, and promotional messages for the life-scenarios test data. These fixtures support testing scenario logic that depends on realistic inbox contents over a multi-month period. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the tinymcp submodule to a newer commit and adjusted its dependencies in Cargo.lock. The anyhow crate was removed as a dependency, and the dirs crate was upgraded from version 6.0.0 to 7.0.0. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add a PDF fixture for a hotel booking in Tokyo to support life scenarios testing. This document provides realistic input for scenario-based validation of booking workflows. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…cenario Add a new draft document and accompanying hero image to support the handoff scenario in life scenarios, providing initial content and visual assets for the feature. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario generation logic now correctly processes boundary conditions where input values approach zero, preventing division by zero errors and ensuring consistent output for minimal inputs. This improves reliability when generating life scenario projections with very small or zero initial parameters. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario runner previously crashed when given an empty input file because it assumed at least one scenario would always be present. This change adds a guard clause that returns early with a clear message when no scenarios are provided, making the tool more robust and user-friendly. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The run.mjs script has been updated to align with the project's import conventions, replacing dynamic imports with static imports for better readability and consistency. This change does not affect the script's behavior or output. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When a scenario configuration file is not found, the script now logs a clear error message and exits with a non-zero status code instead of failing with an unhelpful JavaScript exception. This improves the user experience by providing actionable feedback when the expected configuration is absent. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When a scenario configuration file is not found, the script now logs a clear error message and exits with a non-zero status code instead of failing with an unhelpful JavaScript exception. This improves the developer experience by providing actionable feedback when the expected configuration is absent. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The runScenario and smoke test calls to the inference agent were not including routing parameters such as model overrides, causing requests to always use the default model. The change now spreads the result of routeParams(opts) into both call arguments so that any configured routing options are properly forwarded. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario runner previously threw an error when given an empty input file, as it attempted to process undefined data. This change adds a guard clause that returns early with a clear message when no input is provided, improving robustness and user experience. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Install a local session credential at the start of the life-scenarios run, before any turn is executed. This ensures the authentication state is ready and the route is validated early, making the smoke turn more reliable and providing clearer logging of the auth and route configuration. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The benchmark runner now spawns each core process with a private HOME directory and optional Composio mock endpoints, so that scenario runs never read or write the operator's `~/.openhuman` credentials and do not depend on any hosted service. This makes the measurements fully hermetic and reproducible without risking the desktop app's authentication state. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…nges Updated the mock composio script to reflect recent modifications in the underlying API, ensuring that test scenarios continue to function correctly with the current interface. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…or benchmarks The benchmark runner now creates a throwaway home directory with a clean config instead of inheriting the operator's settings, preventing silent configuration drift. It also starts a mock Composio server when the --mock-composio flag is given, allowing benchmarks to run without a real Composio instance and recording all requests and outbox messages for inspection. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…mposio mock The run script now writes a dedicated agent definition file and a per-user config so the composio mock works correctly in direct mode. The agent definition lists the tools the benchmark agent is allowed to use, and the per-user config ensures the composio block is read when a user is active. The change also adds an --agent flag to select the agent and includes the reply text in the output for debugging. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario runner now checks for empty input before processing, preventing a crash when no scenarios are provided. This ensures the script exits cleanly with a helpful message instead of throwing an unhandled error. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario runner now checks for the existence of the configuration file before attempting to read it, preventing a crash when the file is absent. This ensures a clear error message is shown instead of an unhandled exception. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Introduce a new configuration file for agent life scenarios, defining the lifecycle events and transitions that agents can undergo during simulation. This enables more realistic agent behavior by modeling stages such as creation, activity, and termination. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The benchmark agent definition for life scenarios was previously constructed as a string literal inside the script. It is now copied from a standalone TOML file, making the definition reviewable independently and simplifying maintenance. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario runner now checks for empty input before processing, preventing a crash when no scenarios are provided. This ensures the script exits cleanly with a helpful message instead of throwing an unhandled error. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario runner was failing when no configuration file was present, as it attempted to read from an undefined path. This change adds a check for the config file's existence before attempting to load it, falling back to default settings when the file is not found. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario runner now checks for empty input before processing, preventing a crash when no scenarios are provided. This ensures the script exits cleanly with a helpful message instead of throwing an unhandled error. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The scenario runner previously crashed with an unhelpful error when the configuration file was not present. This change adds a check for the file's existence before attempting to read it, providing a clear message to the user when the configuration is missing. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ully The scenario runner now checks for the existence of the configuration file before attempting to load it, preventing a crash when the file is absent. This improves robustness for users who may not have set up the configuration yet. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add a Set to record decisions that have already been processed, preventing the responder from acting on the same decision more than once. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Created a new findings document to capture observations and insights from the life scenarios analysis, providing a reference for future work and decision-making. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When a model configuration is not present in the config file, the system now returns a clear error instead of panicking. This improves user experience by providing actionable feedback when the configuration is incomplete or missing. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the model configuration to align with the latest provider API changes, ensuring compatibility and correct behavior when initializing models. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the model configuration to include support for additional AI providers and their associated model parameters, ensuring compatibility with the latest API changes and expanding the range of available model choices for users. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Removed the test file for the model BYOK configuration operation as it is no longer needed, likely due to a restructuring of the test suite or removal of the corresponding feature. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Added a conditional compilation test module declaration in model.rs that includes the existing model_byok_tests.rs file, and updated the test file to use a wildcard import from the parent module instead of specific imports, making the test module self-contained and easier to maintain. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the run.mjs script to improve scenario execution by refining the output formatting and ensuring consistent handling of edge cases. This change enhances readability and reliability when running life scenario simulations. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add a new FINDINGS.md file to document the results and observations from the life scenarios analysis, providing a clear reference for the insights discovered during the process. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Added a new FINDINGS.md file to document observations and insights from the life scenarios analysis, providing a reference for future work and decision-making. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The findings document had a new section inserted as number 2, but the subsequent sections were not renumbered, causing duplicate and out-of-order numbering. This change updates all section numbers from 2 through 10 to 3 through 11 to maintain sequential ordering throughout the document. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The findings list had incorrect numbering for items 2 and 3, and the severity description for item 2 was unnecessarily verbose. This change corrects the numbering order and trims the severity text to be more concise while preserving the essential information. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add a new FINDINGS.md file to the life-scenarios scripts directory, documenting observations and outcomes from running the life scenario simulations. This provides a reference for understanding the results and implications of the various scenarios tested. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the scenario argument is not provided to the run script, the application now displays a clear usage message instead of failing with an unhelpful error. This improves the user experience by guiding the user to the correct invocation. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add a new markdown file documenting the diagnosis life scenario, covering the emotional and practical considerations for individuals receiving a medical diagnosis. This provides a structured reference for users navigating this challenging life event. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The findings document was updated to replace the original description of the iteration cap issue with a detailed root cause analysis of why the orchestrator cannot reliably create files. The new text identifies four distinct mechanisms that compose to block file creation, including the action directory path policy, command classification rejecting ampersands, the files pack being unreachable through use_skill, and apply_patch being unable to create new files. The iteration cap section was also revised to clarify that the cap is a deliberate mechanism and the real defect is that callers cannot detect when a turn has been truncated. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ence Update FINDINGS.md to cross-reference DIAGNOSIS.md for the causal trace behind the headline result, and adjust the severity of finding 3 from high to medium. Update README.md to list both FINDINGS.md and DIAGNOSIS.md as known harness findings. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Update the pinned commits for the tinyagents, tinyflows, and tinymcp vendor submodules to incorporate upstream changes. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the pinned commits for the tinyagents and tinymcp vendor submodules to incorporate upstream changes. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Two test fixtures in the progress tracing span tree tests were missing the `usage` field, which caused compilation failures after the field was added to the data structure. The change adds `usage: None` to both test cases to restore compilation. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add the `usage: None` field to three test struct literals that were missing it after a recent change added this field to the struct definition, fixing compilation errors in the test files. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The test helper function `meta` was missing the `session_id` and `parent_session_id` fields that were recently added to the `TranscriptMetadata` struct, causing compilation failures in the usage tests. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Update the pinned commit for the vendored tinyflows dependency to incorporate upstream changes. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Reformat several multi-line expressions in the model configuration and its BYOK tests to comply with the project's line-length conventions, wrapping function arguments and assertions that previously exceeded the limit. No behaviour is changed. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Contributor
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: ⛔ Files ignored due to path filters (2)
📒 Files selected for processing (27)
Comment |
Tiny Sweeper review
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
scripts/life-scenarios/— six everyday assistant tasks (calendar triage, receipt scanning, live web research, meal planning, multi-source trip synthesis, fact-check-and-publish) driven against the real core over the shipping desktop path and graded on what actually landed on disk.claude-sonnet-5it scores 12/27 (44%), and four of six scenarios wrote no output file at all after spending $0.69, $0.78 and $0.13.DIAGNOSIS.mdtraces that to root cause, tool call by tool call.config.update_model_settingsaccepted the documentedinference_url+api_keyBYOK pair and then failed every subsequent turn withBYOK_INCOMPLETE.upstream/main(fourAgentProgress::SubagentCompletedinitializers and oneTranscriptMetamissing fields added upstream), and restoresvendor/tinyflowsto upstream's pin, which the merge silently regressed.Problem
We had no way to answer "how much of a realistic, multi-step task does the harness finish?" Unit tests cover units; nothing exercised read-many-files → reason → produce-an-artifact end to end against the real core, and nothing priced it.
Running that exposed a cluster of real defects. The headline: the orchestrator has no reliable way to create a file. Four mechanisms compose —
action_diris the base that relative tool paths are joined onto, but not a permitted write root.is_resolved_path_allowed_for(security/policy/path_checks.rs) allowsworkspace_rootor a trusted root, andenforcement.rs:118-126grants a trusted root fordefault_projects_dir(), which readsOPENHUMAN_PROJECTS_DIRand knows nothing aboutOPENHUMAN_ACTION_DIR. On a stock install the two coincide, so it is invisible — change the working folder (whataction_dir_overrideand the Settings control write) and file-tool writes into it are refusedResolved path escapes workspace, for a path inside the directoryCLAUDE.mdcalls "the agent's permitted read and write root".classify_commandrejects&inside a quoted heredoc body. The blockedcat > out/meal_plan.md << 'EOF'contained four ampersands, every one of them in a recipe title ("Greek Chicken & Spinach Orzo Skillet").subscription-scanwrote successfully with the identical heredoc shape and no&in its content.use_skill {"skill":"files"}returns "Skillfileshas no tools available in this session" — the pack is closed for the orchestrator byclose_handed_off_packs(Orchestrator never delegates to the MCP and skill sub-agents: their hand-off tools are withheld by tool packs #6302), so the documented escape hatch for withheld packs cannot reachfile_write.apply_patchcannot create a file (old_stringmust not be empty, and it canonicalizes the target).meal-plancycled through all four for 11 rounds and $0.80 and left two 1-byte files containingx— the placeholder it made soapply_patchwould have something to patch.Separately, three scenarios hit
max_model_calls=15, whereFinalCallWrapUpMiddlewarewithdraws all 25 tools and asks for a summary. That mechanism is deliberate and works as designed (#6014); the defect is that the caller cannot tell it happened —turn_run_finalize.rscomputeshit_capandflows/consumes it, butgrep hit_capoverweb_chat/finds nothing andTurnUsagePayloadhas no cap field, so a truncated turn arrives as an ordinarychat_done.Full evidence for each, including the transcripts, is in
scripts/life-scenarios/DIAGNOSIS.md. Ranked defect list with the fix each wants is inFINDINGS.md. This PR does not fix items 1–4 — they want owner decisions (and one belongs invendor/tinyagents), so they are filed rather than patched.Solution
The suite.
--driver desktop(the default) drives the core exactly as the Tauri composer does —openhuman.channel_web_chat+GET /events, the orchestrator agent with every pack withheld, and the approval gate on with a headless responder answeringapprove_oncerather than the usualOPENHUMAN_APPROVAL_GATE=0shortcut, because a disabled gate measures a product nobody runs. Three things are deliberately not the app, each buying reproducibility: its ownHOME(so a run can never read or corrupt the operator's install, or sign a running desktop app out by installing a credential), BYOK inference instead of the hosted backend, and a mock Composio serving Gmail/Calendar wire shapes over the same fixtures the file tools see.Grading checks facts only derivable from the fixtures, so a plausible-looking artifact full of invented rows scores zero — e.g.
calendar-bufferhas exactly three sub-15-minute gaps, one of them 10 minutes, which catches a model matching on "back-to-back" instead of "under fifteen".The corpus is entirely fictional (one persona,
*.exampledomains throughout) and safe to commit; run output goes to the ignoredtarget/.The fix.
complete_byok_routeinconfig/ops/model.rs: wheninference_urlandapi_keyboth arrive non-blank and nocloud_providersentry matches the endpoint, register one and pin the four roles an agent turn runs on — the same completionconfig/schema/ephemeral_route::applyalready performs for a single call. Deliberately narrow: both halves required (an endpoint with no credential is a partial statement), an existing entry for that endpoint reused rather than duplicated, a blankdefault_modeldeclined (the grammar is<slug>:<model>), and any role the same patch pinned — or already pointing somewhere deliberate likeollama:…— left untouched.Submission Checklist
config/ops/model_byok_tests.rs, 9 tests: the happy path plus key-without-endpoint, endpoint-without-key, blank model, duplicate endpoint, explicitly-pinned role, deliberately-pinned role, thecloudsentinel, and clearing.complete_byok_routeand every branch in it is covered by the 9 tests. The rest of the diff isscripts/and docs.N/A: no feature row added or removed; the change completes an existing config path.## Related—N/A: no matrix rows affected.N/A: does not touch a release-cut surface.Closes #NNN—N/A: no tracking issue; the findings are filed in FINDINGS.md for triage.Impact
TurnUsagePayload, RPC params and results are untouched.complete_byok_routeruns insideconfig.update_model_settingsand is a no-op for every existing configuration: it acts only wheninference_url+api_keyare both set and no provider entry already matches that endpoint. An install that already hand-configuredcloud_providers— the workaround this defect forced — keeps its own slug and is not rewritten.is_workspace_internal_path,is_always_forbidden,classify_commandor approval behaviour is modified; the suite answers the approval gate rather than weakening it.vendor/tinyflowsrepairs restoreupstream/mainto a state wherecargo test -p openhuman --libcompiles; they are mechanical and add no behaviour.Related
FINDINGS.mditems, chiefly (a) grantingconfig.action_diras a trusted root rather than onlydefault_projects_dir(), (b) teaching the command classifier that a quoted heredoc body is data, (c) surfacinghit_capon the web-chat path, and (d) the transcript/tool-call replay question, which belongs invendor/tinyagents.AI Authored PR Metadata (required for Codex/Linear PRs)
Linear Issue
Commit & Branch
harness-life-scenariosValidation Run
pnpm --filter openhuman-app format:check—N/A: no frontend files changed.pnpm typecheck—N/A: no TypeScript changed; the suite is plain ESM under scripts/.cargo test -p openhuman --lib config::ops::model→ 9 passed, 0 failed.cargo build -p openhuman-cli --bin openhuman-coreclean;cargo fmt -p openhuman -- --checkclean for every file this PR touches.N/A: crates/openhuman-app not touched.Validation Blocked
command:fullcargo test -p openhuman --liberror:pre-existing formatting drift intool_result_artifacts/mod_tests.rsandruntime_adapter_tests.rs, untouched by this PR and left alone.impact:none on this change; the focused suite compiles and passes.Behavior Changes
config.update_model_settingswith only the two documented fields now actually routes, instead of being accepted and failing on the next turn.Parity Contract
cloud_providersentry is reused, never replaced; explicitly pinned roles and roles already pointing at a non-cloudprovider are left alone.ephemeral_route::apply, which already does exactly this for a per-call route, so the persisted and per-call BYOK paths now agree.Duplicate / Superseded PR Handling