Skip to content

docs(performance): refresh whole-app audit and track optimization batch - #281

Closed
sambitcreate wants to merge 6 commits into
mainfrom
feature/performance-audit
Closed

sambitcreate wants to merge 6 commits into
mainfrom
feature/performance-audit

Conversation

@sambitcreate

Copy link
Copy Markdown
Owner

The July performance plan contains findings that newer work has already fixed. This refresh records 40 current optimization opportunities across desktop rendering/processes, iOS, Android and shared runtime, with source evidence, existing mitigations, risks and validation plans from five specialist audits.

Adds mobile optimism matrices, a before/after ledger and the first four independent implementation lanes. Includes a reproducible synthetic Remote-stream baseline: 64 appends trigger 64 full snapshot calls; runtime timings are exploratory and no GPU/battery savings are claimed. No application behavior changes.

Validation: Markdown source/file anchors and relative links checked; inventory JSON parsed; git diff --check; unchanged baseline Remote stream suite 59/59 passed. The implementation lanes and subsequent PR/CI review status are tracked in docs/audits/performance-2026-09-27/batch-1.md.

@very-hermes-bot

very-hermes-bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Hermes Review Bot

Confidence: 5

Engine: agy/gemini-3.8-flash-high
Review mode: full
Head: 0a3fc7eca020154b7b411614b8b4c9ca3e54c2f8
Generated: 2026-09-28T03:27:17+00:00
Reviews: 1

Summary

Refreshes Aiden's whole-app performance audit against baseline commit a9baa4aa3027893e5455043083465c34b4c8b4ac (v0.50.0), replacing historical and now-mitigated July 2026 planning assumptions with 40 source-confirmed optimization opportunities across five distinct domains. It provides comprehensive specialist reports covering desktop rendering/GPU, background helper processes, iOS, Android, and shared runtime/transport, accompanied by mobile optimism matrices, a before/after ledger, and an environment inventory. A synthetic microbenchmark fixture and deterministic counters for Remote-stream snapshot checking (docs/audits/performance-2026-09-27/evidence/stream-baseline.mts) demonstrate the 64-to-0 snapshot reduction achieved by sibling implementation PR #285 without altering production application code. Maintainers should double-check that subsequent execution lanes continue to track their branch heads and CI verification status in docs/audits/performance-2026-09-27/batch-1.md as individual optimization PRs land.

Confidence Score: 5/5

5/5 — All 16 changed documentation, memory, and evidence files were fully inspected; all 167 source line anchors across 8 audit documents were programmatically verified against the codebase, JSON manifests parse cleanly, the benchmark script passes TypeScript/ESM validation, and no production code or contracts are modified.

📁 Important Files Changed
  • docs/audits/performance-2026-09-27/README.md: Central performance audit ledger, recommended optimization sequence, mobile optimism guidelines, and living before/after tracking table.
  • docs/audits/performance-2026-09-27/desktop-gpu.md: 10 desktop rendering and GPU findings (DG-01–DG-10) addressing transcript mounting, offscreen HTML frames, scroll-follow coalescing, and code syntax highlighting.
  • docs/audits/performance-2026-09-27/desktop-processes.md: 8 Electron main/helper findings (DP-01–DP-08) covering background browser tab throttling, voice engine timeout handling, MCP quit barrier membership, and schedule catch-up concurrency.
  • docs/audits/performance-2026-09-27/ios.md: 7 iOS client findings (IOS-01–IOS-07) and an optimistic interaction matrix covering SSE queue backpressure, cache eviction bounds, and decoupling fresh transcript publication from catalog fetches.
  • docs/audits/performance-2026-09-27/android.md: 8 Android native findings (AND-01–AND-08) and an optimistic interaction matrix addressing UI-thread disk persistence, selected-image preparation off Main, composed invisible tab gating, and lazy row virtualization.
  • docs/audits/performance-2026-09-27/shared-data-runtime.md: 7 storage and runtime findings (SDR-1–SDR-7) covering Remote stream aggregate accounting, shared chat index scans, and Pi compaction/session memory.
  • docs/audits/performance-2026-09-27/batch-1.md: Tracking and orchestration document for Batch 1 follow-up PRs (docs(performance): refresh whole-app audit and track optimization batch #281, perf(browser): throttle idle guests with owned activity exceptions #282, perf(ios): publish refreshed chats independently of model catalogs #283, perf(android): prepare selected images off Main with bounded ownership #284, perf(remote): avoid full journal snapshots during event accounting #285) with remediation logs and gate criteria.
  • docs/audits/performance-2026-09-27/evidence/stream-baseline.mts: Reproducible synthetic microbenchmark script measuring Remote stream append snapshot calls using AidenRemoteStreamService.
  • docs/audits/performance-2026-09-27/evidence/stream-before.json & stream-after.json: Captured benchmark outputs comparing 64 snapshot calls in the baseline versus 0 after accounting optimizations.
  • docs/audits/performance-2026-09-27/baseline-inventory.json: Source code file inventory and environment baseline metadata at commit a9baa4aa3027893e5455043083465c34b4c8b4ac.
  • docs/plans/performance-stability-efficiency-plan.md & docs/plans/README.md: Updates linking the historical performance master plan to the new September 2026 whole-app audit.
  • .memory/performance-audit-20260927.md & .papercuts/troubleshooting.md: Project memory and developer troubleshooting logs detailing multi-specialist worktree constraints, Android test harness lifetime edge cases, and GitHub CLI workflow log limitations.

Findings

No findings.

Sequence Diagram

Unchanged or not applicable for this change.

[]

Last reviewed commit: 0a3fc7eca020
Reviews (1) · Comment /hermes review to trigger a new review · /hermes review full for full re-review

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ℹ️ One minor reproducibility suggestion inline.

Reviewed changes I reviewed the full five-area performance audit, its evidence and measurement protocol, the optimization ledger and lanes, and the accompanying plan and memory updates. No application behavior changes are included.

  • Source audit — adds 40 ranked desktop, iOS, Android, and shared-runtime findings with current mitigations, risks, and validation plans.
  • Evidence and execution tracking — adds an inventory, mobile optimism matrices, a measurement protocol, Remote-stream baseline helper/results, a before/after ledger, and four implementation lanes.
  • Plan context — points the existing performance plan and index at the September audit and records the audit status.

Pullfrog  | Fix all ➔ | Fix 👍s ➔ | View workflow run | Using GPT Luna | 𝕏

Comment thread docs/audits/performance-2026-09-27/evidence/stream-baseline.mts

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ℹ️ The incremental documentation and evidence updates introduce no new review findings.

Reviewed changes I re-reviewed the new implementation follow-through and measurement updates since the prior Pullfrog review.

  • Recorded implementation follow-through — updated publication, independent-review, and validation status for the browser, iOS, Android, and Remote lanes.
  • Preserved Android validation limits — documented the test-harness cleanup while keeping the initial-load timeout unresolved and distinguishing diagnostic results from a clean-source full-suite pass.
  • Added Remote after-run evidence — recorded deterministic snapshot-call and byte-size results while keeping timing exploratory and GPU/energy savings unclaimed.

Pullfrog  | Fix it ➔ | View workflow run | Using GPT Luna | 𝕏

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes I re-reviewed the documentation and evidence updates since the prior Pullfrog review.

  • Recorded hosted follow-up — updated the audit memory and troubleshooting notes with first-pass review findings, CI status, and Android validation limitations.
  • Made the stream fixture reproducible — documented the installed tsx command, target revisions, output handling, and rerun results while preserving the original exploratory timing evidence.

Pullfrog  | View workflow run | Using GPT Luna | 𝕏

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes I reviewed the documentation and validation updates added since the prior Pullfrog review.

  • Recorded reviewed follow-ups — updated the audit ledger, project memory, troubleshooting notes, and batch tracker with independently reviewed browser and iOS remediation plus corresponding validation evidence.
  • Clarified remaining validation limits — distinguished Android teardown cleanup from its unresolved initial-load timeout and captured the 03:07 UTC exact-head CI checkpoint.

Pullfrog  | View workflow run | Using GPT Luna | 𝕏

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes I reviewed the cancellation-safe Android teardown follow-up recorded since the prior Pullfrog review.

  • Recorded the Android cleanup follow-up — updated project memory, troubleshooting notes, and the batch tracker with publication 632f11a57, the non-cancellable teardown barrier, regression and independent-review evidence, and the still-unexplained startup timeout and superseded packaging failure.

Pullfrog  | View workflow run | Using GPT Luna | 𝕏

@sambitcreate

Copy link
Copy Markdown
Owner Author

Performance batch follow-up complete — 2026-09-28 04:57 UTC

All five current PR heads have successful applicable checks, with no unresolved review threads or pending requested reviews. Each implementation and material follow-up received fresh-context GPT-6 Astra medium review. Nothing merged or released.

PR Verified head Checks
#281 0a3fc7eca020154b7b411614b8b4c9ca3e54c2f8 23/23 successful
#282 deec0c010740e8bf81d41fb93f5c8550f03208c7 23/23 successful
#283 cb53bbe6a8bed4f31299a0be9363671b92ec5008 23/23 successful
#284 632f11a57b593d8bf1f77344403de088968ccfb9 23/23 successful
#285 c4f195f4a5fd98dfd0a54b141e2a503e3c6b4797 22/22 successful

The audit iOS simulator job passed on its single permitted retry. The original testProgressObservationReleasesCompletedHandleAndCanRestart failure remains an unresolved flaky-test observation; the unchanged local probe and hosted retry do not establish its cause or a code fix. The Android local initial-HTTP setup timeout also remains unexplained despite green hosted CI. Proven teardown fixes are recorded separately. Physical-device battery/GPU/thermal/latency improvements remain unmeasured.

The audit and lane documents contain before/after mechanism evidence, validation and limits. The final polling ledger and papercut additions are retained in the local audit worktree; this comment preserves final hosted status without creating a new commit that would invalidate the checked heads. The 15-minute follow-up is being paused at this successful checkpoint.

@sambitcreate

Copy link
Copy Markdown
Owner Author

Closing per maintainer decision; not landing the performance audit doc in 0.51.0.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants