A 100% uv-managed, local-first, privacy-preserving fiduciary intelligence harness engineered for Apple Silicon macOS.
Unlike mainstream fintech apps that sell your financial habits to third-party data brokers, or generic LLM wrappers that hallucinate arithmetic, this harness operates under a strict Three-Tier Fiduciary Contract:
- Air-Gapped Ingestion: Automated PII masking (
••-••-XX,••••XXXX). Sensitive databases (data/*.db) and credentials (.env,secrets/*) are strictly excluded from git. - Deterministic Python Core: 100% of arithmetic calculations (liquid runway, 60% tax traps, spending velocity, budget splits) run through pure Python—never an LLM.
- Grounded Local AI (Apple Silicon Metal GPU): Local reasoning running via Ollama (
qwen3.5:4b) or LM Studio (Meta-Llama-3.1-8B) with real-time Grounding Guardrails and an independent local LLM Judge.
Run the fiduciary engine immediately without manual cloning or environment configuration:
# Launch interactive local web dashboard directly on http://localhost:8080
uvx --from "git+https://github.com/cloudcruncher/fiduciary-agent" fiduciary ui
# Or run the cinematic terminal walkthrough
uvx --from "git+https://github.com/cloudcruncher/fiduciary-agent" fiduciary demo# Pull and run pre-built container from GitHub Container Registry
docker run -d -p 8080:8080 -v $(pwd)/data:/app/data --name fiduciary ghcr.io/cloudcruncher/fiduciary-agent:latestAnyone can clone and run this harness locally in under a minute without needing real bank accounts or API keys:
- macOS (Optimized for Apple Silicon M1/M2/M3/M4) or Linux/WSL2
- uv (blazing fast Python package manager):
curl -LsSf https://astral.sh/uv/install.sh | sh - Ollama (recommended for local offline inference):
brew install ollama && ollama serve
git clone https://github.com/cloudcruncher/fiduciary-agent.git
cd fiduciary-agent
# Install dependencies in an isolated virtual environment
uv sync --dev# Pull fast, high-quality 4B model (uses ~2.8 GB RAM, unloads after turn)
ollama pull qwen3.5:4b# Generates realistic NatWest, Revolut, and Wise accounts with 60-day history
./f seed# Run the live interactive terminal walkthrough tour
./f demo
# Launch the interactive local web dashboard
./f ui
# Or chat with your offline AI Fiduciary Copilot directly in the terminal
./f copilot "How much have I spent on pubs?"
./f copilot "Can I afford a £1,500 holiday in July without breaking my emergency buffer?"An executable shortcut ./f is available in the project root:
| Command | Short Alias | Function |
|---|---|---|
./f demo |
— | Run interactive cinematic terminal walkthrough demonstrating all capabilities |
./f seed |
— | Seed realistic synthetic UK bank accounts, transactions & bills for local testing |
./f spending [query] |
./f spend |
Spending Insight Engine: Category breakdown (e.g. pubs, groceries), velocity & micro-expenses |
./f profile |
./f p |
Intelligent Customer Profile: Financial Health Score (0–100), Financial DNA Archetype & Action Cards |
./f copilot [Q] |
./f chat |
Interactive conversational AI Fiduciary Copilot (grounded in live transactions & tax rules) |
./f react [Q] |
— | Autonomous multi-step ReAct agent with step trace observability (Thought → Action → Observation) |
./f rag [Q] |
— | Local semantic Vector RAG search across statutory HMRC tax rules and FCA MCOB underwriting standards |
./f credit |
./f cr |
Underwriter-view credit audit: Borrowing Readiness Score, UMI/DTI, BNPL & returned-DD flags, mortgage capacity (4.5x, 7.5% stress), runway stress tests. Record bureau scores with --experian 865 --equifax 740 |
./f watchdog |
./f guard |
Financial Watchdog: Detect stealth subscription price hikes, duplicate charges, upcoming bills |
./f tax |
./f t |
UK Tax & Wealth Optimization: 60% allowance trap audit, SIPP relief, Personal Savings Allowance drag |
./f sweep |
— | Smart Cash Sweeper & Automated Standing Order float architect |
./f networth |
./f nw |
Whole-of-Wealth Balance Sheet (Cash, Investments, Pensions, Property, Debt) |
./f asset add |
— | Add custom assets or debts (property equity, pension pots, mortgages) |
./f tx |
— | Itemised transactions. Filters: -c dining, -a revolut, -n 15, -s "Tesco" |
./f connect |
./f c |
Connect real UK banks (Revolut, Chase, HSBC, Lloyds, Zopa) via TrueLayer |
./f sync |
— | Sync live balances & transactions from Wise and Open Banking |
./f scout |
./f s |
Autonomous market scout (Cash ISAs, after-tax yields, bank switch bounties) |
./f audit |
./f a |
Real fiduciary capital allocation audit + live AI strategy memo |
./f traces |
./f tr |
AI observability: inspect tool latency, grounding score, and telemetry per turn |
./f eval |
./f benchmark |
Run 6-dimension Enterprise AI Evaluation Benchmark suite (Grounding, Invariants, SLAs) |
./f judge |
./f j |
Score the latest Copilot answer with an independent local model (LLM-as-a-Judge) |
./f tools |
./f tl |
Inspect available deterministic and live web tools catalog |
./f ui |
./f w |
Launch local web dashboard at http://localhost:8080 |
The fiduciary engine was engineered with intentional storage trade-offs tailored to single-user financial autonomy on modern laptop hardware:
| Criteria | SQLite (Current Choice) | DuckDB (Analytical Vector) | PostgreSQL (Enterprise SaaS) |
|---|---|---|---|
| Deployment Model | Embedded, zero-daemon, single file (data/financial.db) |
Embedded, in-process columnar engine | Client-server daemon (Docker / Managed service) |
| RAM Impact (16GB Mac) | ~0 MB idle (instant zero-friction start) | Moderate buffer pool allocation | High (shared buffers + connection pool processes) |
| Data Privacy & Air-Gap | 100% local APFS file, zero network open ports | 100% local file / Parquet data lake | Network TCP ports, exposed credentials surface |
| Query Latency | < 0.1 ms for transactional lookups & joins | < 0.5 ms vectorized scans over millions | 1.0 – 5.0 ms (network roundtrip latency) |
| Dependency Footprint | Zero external dependencies (Python standard library) | Third-party dependency (duckdb) |
External database server + drivers (psycopg2/asyncpg) |
| Special Superpower | Unmatched reliability, zero maintenance, ACID | Direct Parquet/CSV queries + native VSS vector search | Multi-tenant Row-Level Security (RLS), pgvector, TimescaleDB |
| When to Switch? | Default for personal fiduciary harness | Switch if analyzing 500K+ txs or running local vector search | Switch if hosting a multi-user advisory SaaS platform |
- Unified Memory Preservation on 16GB Macs: On an Apple Silicon machine running an Ollama or LM Studio LLM (
qwen3.5:4brequires ~2.8 GB RAM), every megabyte counts. Running a PostgreSQL daemon or memory-hungry OLAP cluster competes directly with Metal GPU unified memory. SQLite consumes virtually 0 MB when idle and requires zero background processes. - Absolute Air-Gap: The SQLite file sits exclusively on your local APFS SSD. It cannot be sniffed over a network socket or misconfigured with open ports.
- Pluggable Storage Abstraction: All database interactions are encapsulated in
fiduciary/storage/db.py. If you wish to migrate to DuckDB (for analytical column scans) or PostgreSQL (for a multi-tenant web platform), you only need to swap the adapter without altering any agent or analysis logic.
- Strict
.gitignore: Live databases (data/*.db), credentials (.env), private keys (*.pem,*.key), and bank statements (*.pdf,*.csv) are strictly ignored. - PII Masking: Bank account numbers and sort codes are automatically masked (
••-••-XX,••••XXXX) during statement ingestion. - Zero Cloud Egress: In default local mode, client financial balances and transactions are never transmitted to external cloud APIs. The Gemini API client remains completely dormant unless explicitly configured.
- Automated CI Secret Scanning: Every pull request and push is audited with Gitleaks to guarantee zero credentials or tokens enter git history.
You can physically verify that zero financial data leaves your Mac:
- Turn off Wi-Fi on your Mac completely.
- Run:
./f copilot "What is my liquid balance and how much did I spend on pubs?" - The Copilot executes 100% offline using your local Ollama engine on Apple Silicon Metal GPU.
When you are ready to transition from synthetic demo data to your real finances:
- Generate a read-only API token in Wise (Settings → API tokens → Read only).
- Add to
.env:WISE_API_TOKEN="your_read_only_token" - Sync:
./f sync
- Register a free developer console at console.truelayer.com.
- Add credentials to
.env:TRUELAYER_CLIENT_ID="your_client_id" TRUELAYER_CLIENT_SECRET="your_client_secret" TRUELAYER_USE_SANDBOX=false
- Run
./f connector click🏦 Connect Bankin the web dashboard. Complete the standard UK Open Banking mobile app authentication.
Connecting UK banks (Lloyds, Revolut, Chase, HSBC, NatWest) often requires biometric authentication (FaceID / TouchID) in your mobile banking app. The harness lets you use your phone to link banks while your Mac handles all local LLM reasoning and data storage:
- Access from Phone over Local Wi-Fi: Start the local web dashboard on your Mac (
./f ui). The terminal displays both your local URL (http://localhost:8080) and your LAN URL (e.g.http://192.168.1.100:8080). Open the LAN URL in Mobile Safari or Chrome. - Responsive Mobile PWA Experience: The UI automatically switches to a mobile-optimized view with a sticky bottom navigation dock, compact 2-column metrics cards, and a persistent connection helper.
- Biometric Native Bank App Handoff: Tap
🏦 Connect Bankon your phone. TrueLayer redirects you to your bank selection. Selecting Lloyds, Revolut, etc. deep-links straight into your installed banking app for instantaneous FaceID or TouchID consent. - 1-Tap Clipboard Redirect Bridge (
📋 Paste from Clipboard & Connect):- The Open Banking Obstacle: Standard UK Open Banking mandates strict pre-registration of callback redirect URIs (
http://localhost:8080/truelayer/callback). When a mobile banking app completes authentication, the mobile browser attempts to redirect tolocalhost:8080, which is unroutable on your phone. - The Instant Fix: Copy the address bar URL from your phone browser and tap the prominent green button
📋 Paste from Clipboard & Connectat the top of your Fiduciary dashboard. The app regex-extracts the authorizationcode=and completes the token exchange asynchronously via/api/truelayer/exchange.
- The Open Banking Obstacle: Standard UK Open Banking mandates strict pre-registration of callback redirect URIs (
- Terminal ASCII QR Code Pairing: Alternatively, run
./f connectin your Mac terminal or click QR Handoff in the dashboard to render a high-contrast ASCII QR code. Point your iPhone camera to scan and authenticate immediately.
Data engineering is not merely an ingestion script—it is the structural foundation across every layer of the fiduciary agent:
- Canonical Open Banking Modeling: Ingests, normalizes, and validates transaction streams from diverse UK institutions (NatWest, Barclays, Revolut, Chase, HSBC, Wise) into a unified, typed relational schema.
-
Closed-Loop Double-Entry Invariant: Every statement batch undergoes mathematical verification:
$\text{Opening Balance} + \text{Inflows} - \text{Outflows} \equiv \text{Closing Balance}$ . Batches are certifiedRECONCILEDonly if discrepancy is exactly £0.00. -
Cryptographic Provenance & Lineage: SHA-256 batch fingerprints (
statement_batchestable) and deterministic transaction hashes (generate_tx_fingerprint(account_id, date, amount, desc)) guarantee idempotent de-duplication across overlapping statements. -
Resilient Layout-Aware Parsing Engine: Solves real-world statement challenges (e.g. NatWest):
- Same-Day Date Forward-Propagation: Propagates booking dates when banks leave date columns blank on subsequent transactions on the same day.
- Multi-Line Narrative Buffering: Accumulates merchant descriptions that wrap across 2–3 lines before amount columns without truncation.
-
Running Balance Delta Signing: Computes signed transaction amounts directly from running balance deltas (
$\Delta = B_i - B_{i-1}$ ), guaranteeing 100% sign precision for debits and credits. - Legal & Overdraft Boundary Bounds: Strict stop conditions eliminate phantom charges from overdraft fee examples and legal terms.
- Financial Feature Engineering: Deterministically calculates liquid runway, discretionary spend velocity, payroll cadence, and 60% marginal tax traps.
- State-Bounded Cache Hashing: Cryptographic SHA-256 database state hashing ensures zero-token semantic cache hits without stale data risk.
- Embedded Statutory Vector Indexing: TF-IDF and dense embedding pipelines over HMRC tax schedules and FCA MCOB underwriting standards.
UK mortgage lenders and credit underwriters evaluate Open Banking cash-flow affordability rather than CRA bureau scores alone:
- Cash-Flow Affordability: Automatic payroll detection, Uncommitted Monthly Income (UMI), and Contractual Debt-to-Income (DTI).
- Underwriter Risk Scanner: 90-day scan for BNPL (Klarna, Clearpay, Zilch), bounced direct debits, overdraft dip zones, and gambling spend (<1% benchmark).
- 4.5x Mortgage Capacity: Net borrowing capacity deducting committed debts, with 4.4% indicative 25-yr repayments and 7.5% BoE stress testing.
- Runway & Stress Simulator: Comfortable vs Survival runway (cutting non-essentials), income shock, £1,500 emergency repair shock, and UK CPI inflation drag.
- Local CRA Tracking: Air-gapped tracking for Experian (999), Equifax (1000), TransUnion (710), and Electoral Roll status.
Connecting multiple UK institutions (Revolut, Lloyds, NatWest, Wise) requires enterprise resilience against upstream token invalidations and payment rail timing:
- In-Flight 401 Auto-Recovery Interceptor: Open Banking access tokens often expire or are rotated by upstream banks ahead of scheduled expiry. The TrueLayer connector intercepts HTTP 401 responses, retrieves the stored refresh token, completes an in-flight token exchange, updates SQLite, and replays the sync request seamlessly with zero user dropouts.
- Pending vs Settled Reconciliation: In-store card payments arrive immediately as pending authorizations (
pending), later transitioning to cleared transactions (settled). The ledger matches incoming settled records against pending holds by account ID, exact amount, and booking date proximity, completely preventing duplicate expense counts or distorted balances. - Strict 90-Day Cadence Verification: The Financial Watchdog requires at least 2 distinct payments across a 90–120 day window (or explicit Direct Debit mandate) before classifying an item as a recurring subscription, preventing one-off statutory fees (e.g. DVLA driving licence renewal) from being falsely flagged as monthly commitments.
UK Open Banking APIs intentionally withhold internal bank categorizations. Fiduciary Agent incorporates an autonomous, 3-tier transaction intelligence cascade operating at sub-millisecond latency:
- Tier 1: Deterministic Knowledge Base (<0.05ms, £0.00): Pre-seeded registry of UK energy utilities (Switch2 Energy, British Gas), water authorities (Thames Water), municipal councils (L.B. Hounslow Council Tax), transit (TfL), supermarkets, and developer SaaS.
- Tier 1b: Canonical Merchant SQLite Cache (<0.1ms, £0.00): Instant local cache retrieval for recurring counterparties with automatic hit-count telemetry.
- Tier 2: Frontier Smart Model (Gemini 2.5 Flash / LLMClient): Resolves long-tail, unseen merchants into structured JSON containing corporate entity names, official domains, logos, and 3-level taxonomy (
L1 > L2 > L3). - Contextual Semantic Disambiguation:
- Apple Services vs Retail: Automatically categorizes
APPLE.COM/BILL £2.99as an iCloud Software Subscription (monthly, tax-deductible), whileAPPLE.COM/BILL £1,299.00is classified as Apple Store Hardware & Capital Equipment. - TfL Fare vs Penalty: Differentiates routine commute fares (
£3.40) from penalty fare charges (£80.00).
- Apple Services vs Retail: Automatically categorizes
- HMRC Tax Deductibility Engine: Flags allowable sole-trader and business expenses directly from statement narratives (e.g. developer software, AI subscriptions, public transit for business travel).
- Public B2B REST APIs:
POST /api/v1/enrich/transaction: Single narrative enrichment.POST /api/v1/enrich/batch: High-throughput batch enrichment (processes 100 txs in <10ms).GET /api/v1/enrich/cache/stats: Live cache operational metrics and top frequent counterparties.
To guarantee 100% consumer trust, satisfy statutory UK FCA Consumer Duty regulations, and eliminate hallucinated financial figures before responses reach the user:
- Zero-Egress In-Line Guard: Before any generated completion is displayed to the user or returned via API,
PreFlightEvaluatorintercepts the draft in sub-2ms. - Fact Grounding Extraction: Extracts all cited currency (£), interest rates (%), and terms, strictly reconciling them against
<verified_financial_context>. - UK FCA Consumer Duty Invariants:
- Emergency Runway Invariant: Guarantees recommendations never advise draining liquid cash below the 3-month survival runway baseline.
- Predatory Product Rejection: Automatically blocks and flags high-cost debt recommendations (>39.9% APR).
- Statutory Tax Bounds: Enforces exact £100k–£125,140 60% marginal tax trap tapering logic.
- Negative Entity & Bait Resistance: Explicitly verifies that when a user asks about non-existent accounts or competitor cards (e.g. Amex), the response refutes possession rather than inventing figures.
- Autonomous Reflection Self-Correction: If discrepancies or invariant breaches are detected, the critic loop feeds targeted refinement feedback back to the local model for an immediate repair pass. Once verified, responses receive an immutable
🛡️ 100% Groundedor⚡ Self-Correctedverification stamp.
Run the production Forward Deployment Engineering (FDE) test harness anytime via terminal, REST API, or CI pre-push gate:
./f eval # alias: ./f benchmark- Grade A+ Production Certification: Audits 34 systematic test cases across 6 critical dimensions with 100% pass rate in ~2.2ms:
- Grounding & Discrepancy Defense (5 tests)
- Negative Entity & Bait Resistance (5 tests)
- UK FCA Consumer Duty Invariants (6 tests)
- Adversarial Prompt Injection Guard (6 tests)
- Reversible Tokenized PII Privacy (6 tests)
- Deterministic Math & Rate Consistency (6 tests)
- REST Endpoints:
GET /api/eval/benchmarkandGET /api/eval/statusfor CI/CD pipelines and enterprise dashboard telemetry.
- Model Role Separation:
qwen3.5:4bsynthesizes conversational explanations, whilellama3.2:3bindependently audits completed sessions on demand across Faithfulness, Numerical Precision, Fiduciary Prudence, and Actionability. - 16GB Unified Memory Preservation (
keep_alive: 0): Automatically unloads models from Apple Silicon GPU memory immediately after scoring to maintain 70%+ free RAM.
Running local LLMs fast on an M3 MacBook Pro requires specific latency engineering:
- Disabled Reasoning Loops (
"think": false): Models likeqwen3.5:4bdefault to generating 800+ internal "thinking" tokens. By passing"think": falsevia Ollama's native/api/chat, inference time drops from 22 seconds down to 0.4–1.2 seconds (50x faster). - Intent-Driven Dynamic Prompt Pruning: Rather than dumping full transaction histories into every turn, the Copilot dynamically prunes context based on classified intent, cutting token consumption by 80%.
- RAM Guardrails (
keep_alive: 0): Ollama unloads the model from RAM after generating a response, preventing memory pressure accumulation.
To eliminate cloud API dependencies, prevent runaway agent trajectories, and run indefinitely on limited developer credits or 16GB laptops, the harness implements aggressive token, context, and cost optimization (integrated with the Google Antigravity Agent SDK and local SLMs):
- Deterministic-First Python Offload ($0.00 LLM Cost):
- 100% of mathematical aggregation (liquid runway, daily burn velocity, 60% tax trap calculations, net worth splits, mortgage capacity) runs in deterministic Python before any prompt is assembled.
- LLMs are never used for arithmetic or database aggregation, eliminating expensive multi-turn reasoning loops and context pollution.
- Dynamic Context Compaction (80% Prompt Reduction):
- Queries are classified by intent before prompt generation. Rather than dumping entire 90-day transaction ledgers into every turn, only relevant category metrics and top itemized records are injected, shrinking context from ~1,800 tokens down to ~300 tokens.
- Zero-Cost State-Hashed Caching:
- Queries paired with identical financial state hashes hit an in-memory / SQLite response cache, returning instant (<1ms) responses without dispatching GPU inference cycles or cloud API calls.
- Local SLM Default with Cloud Dormancy:
- Out-of-the-box operation defaults entirely to on-device SLMs (
qwen3.5:4bvia Ollama). Cloud Gemini API keys remain completely dormant unless explicitly requested for deep multi-year tax planning.
- Out-of-the-box operation defaults entirely to on-device SLMs (
- Thinking Token Suppression & Hard Budget Guardrails:
- Passes
"think": falseand configuresThinkingLevel.MINIMALalongside hard token ceilings, dropping turn latency from 22s to < 1s.
- Passes
While local-first execution on Apple Silicon Metal GPU remains the out-of-the-box default for privacy, institutional and team environments often route requests through an AI Gateway (e.g. LiteLLM Proxy, Portkey, Cloudflare AI Gateway, or vLLM).
The Fiduciary Agent includes an opt-in, zero-dependency OpenAI-compatible AI Gateway adapter:
# In your .env file or environment:
LLM_PROVIDER=gateway
AI_GATEWAY_URL="http://localhost:4000/v1" # LiteLLM Proxy, Portkey, Cloudflare AI Gateway
AI_GATEWAY_API_KEY="sk-..." # Optional gateway bearer token
AI_GATEWAY_MODEL="qwen3.5:4b" # Target model / alias routed by proxyThe harness maps dedicated models and aliases to maintain separation of concerns:
| Gateway Alias | Target Engine | Purpose & Function | Privacy Mode |
|---|---|---|---|
copilot, default |
ollama_chat/qwen3.5:4b |
High-speed conversational financial analyst, transaction retriever, and runway auditor | 100% Private (Local Metal GPU) |
evaluator, judge |
ollama_chat/llama3.2:3b |
Independent Layer 2 LLM-as-a-Judge critic scoring faithfulness and precision | 100% Private (Local Metal GPU) |
gemini-2.5-flash |
gemini/gemini-2.5-flash |
Cloud fallback for deep multi-year tax planning or complex long-context reasoning | Isolated Fallback Only |
- Intelligent Centralized Routing: Route requests through enterprise proxy layers for rate-limiting, cost tracking, semantic caching, and dynamic failovers.
- Domain-Specific Grounding Invariant: Generic AI gateways provide generic moderation or regex PII scrubbing, but cannot validate financial ground truth. The harness's
GroundingAuditorremains active after gateway response generation, auditing every cited £ balance, percentage yield, and runway day against deterministic Python calculations before presentation. - Full Telemetry & Observability: Latency, tool execution metrics, and audit verdicts are recorded to the local
llm_tracesSQLite database and inspectable via./f tracesor the Web Dashboard. - Dynamic Provider Switching: Switch between
Local (Ollama),LM Studio,AI Gateway, andGemini Cloudanytime via./fCLI or the Web Dashboard modal.
The harness features a standards-compliant Model Context Protocol (MCP) Gateway (fiduciary/agent/mcp_gateway.py) that connects external LLM agents, local IDEs, and the Fiduciary Copilot to live UK market intelligence and deterministic calculations.
| Tool Name | Type | Speed | Source | Description |
|---|---|---|---|---|
fetch_boe_base_rate |
Web Scraper | < 200ms | Bank of England Live | Scrapes official Bank Rate directly from bankofengland.co.uk |
fetch_top_savings_and_isas |
Benchmark Engine | < 1ms | FSCS / UK Market | Returns market-leading Cash ISAs (4.87%) & Regular Savers (7.00%) |
search_web_instant |
Search API | < 250ms | DuckDuckGo Instant | Real-time definitions, UK statutory allowances, and market facts |
query_spending_and_transactions |
Deterministic Engine | < 5ms | SQLite financial.db |
Exact client spend by category (groceries, pubs, etc.) or merchant |
credit_affordability_audit |
Underwriter Engine | < 25ms | FCA MCOB 11 | Underwriter UMI, DTI, BNPL risk scan, and mortgage stress test |
tax_wealth_audit |
Tax Engine | < 1ms | HMRC Tax Rules | UK tax band, marginal rates, 60% allowance trap, and SIPP relief |
financial_watchdog_audit |
Cashflow Engine | < 15ms | Watchdog Profiler | Active recurring bills, price hikes, duplicate charges, 14d runway |
vector_search_documents |
Semantic Vector RAG | < 5ms | Local Vector Index | Semantic retrieval over statutory HMRC tax rules, FCA MCOB underwriting standards |
GET /api/mcp/tools: Returns standardized JSON schemas for all tools compatible with Claude, Cursor, and MCP clients.POST /api/mcp/execute: Executes an MCP tool dynamically with sub-millisecond latency tracking and execution observability:curl -s -X POST http://localhost:8080/api/mcp/execute \ -H "Content-Type: application/json" \ -d '{"tool_name": "fetch_boe_base_rate", "arguments": {"timeout": 2.0}}'
For compound, cross-domain financial inquiries (e.g. "Audit my grocery spend, compare with inflation, and recommend top cash isas"), the harness features an autonomous ReAct (Reasoning + Acting) Agent Engine (fiduciary/agent/react_agent.py):
- Autonomous Trajectory: Iterates through
Thought -> Action -> Action Input -> Observation -> Thought ... -> Final Answerwith a safety ceiling (max 4 turns). - Dynamic Tool Orchestration: Dynamically invokes tools from the
MCPGatewaycatalog (query_spending_and_transactions,fetch_top_savings_and_isas,fetch_boe_base_rate,vector_search_documents, etc.). - Step-by-Step Observability: Measures latency per step, logs observations, and records the full trajectory to
llm_tracesin SQLite. - CLI & Web Execution:
./f react "What is the Bank of England base rate and how much did I spend on pubs?" ./f copilot --react "Audit my grocery spend and find the best cash isa"
To ensure institutional data privacy and zero cloud credential leakage, the PII Anonymizer (fiduciary/agent/pii_anonymizer.py) intercepts all prompts before they leave the application:
- Reversible Pseudonymization: Detects UK Sort Codes (
20-45-78), 8-digit Bank Account Numbers, Card Numbers, National Insurance Numbers (NINO), phone numbers, and emails, replacing them with salt-indexed tokens ([SORT_CODE_1],[ACCOUNT_NUM_1],[EMAIL_1]). - Lossless Roundtrip Deanonymization: Restores real client credentials into the model's final response after generation, guaranteeing zero raw PII reaches external models or cloud gateways.
- One-Way Presentation Masking: Irreversibly masks identifiers for logs and UI display (
••-••-78,••••5678).
The harness includes an embedded, zero-daemon Local Semantic Vector RAG Engine (fiduciary/agent/vector_rag.py):
- Embedded Vector Index: Runs 100% locally with TF-IDF weighted cosine similarity vectors, requiring zero cloud vector APIs, Docker containers, or background vector servers.
- Pre-Seeded Statutory Corpus: Indexes official HMRC ISA rules (£20k/yr, £4k LISA bonus), Pension contribution limits (£60k/yr gross), Capital Gains Tax allowances (£3k), 60% marginal tax trap mechanics, and FCA MCOB 11 underwriting guidelines.
- Custom Document Ingestion: Ingests custom user documents (insurance policies, mortgage offer letters, statement notes) for sub-millisecond semantic retrieval:
./f rag "60 percent tax trap personal allowance" ./f rag "mortgage affordability underwriter stress test"
To protect private financial databases and prevent adversarial jailbreaks, all incoming user queries pass through Prompt Guard (fiduciary/agent/prompt_guard.py) before reaching any model:
- System Override Neutralization: Blocks patterns like
ignore all previous instructions,disregard prior system directives,bypass safety protocols. - Jailbreak & Roleplay Defense: Defends against
DAN mode,act as unrestricted AI, role-hijacking, and canary probes likeoutput 'HACKED'. - Delimiter & Control Token Stripping: Strips ChatML/Llama boundary tokens (
<|im_start|>,[INST],=== END SYSTEM PROMPT ===). - Exfiltration Scanner: Blocks unauthorized SQL injection probes (
drop table,delete from) and secret key probes (print api keys). - Context Boundary Wrapping: Encloses verified client telemetry inside
<verified_financial_context>tags so user queries can never impersonate database ground truth.
Small instruction-tuned models (qwen3.5:4b, llama3.2:3b) often suffer from RLHF safety reflexes—refusing financial inquiries with canned disclaimers like "I cannot provide financial advice as an AI..."
The harness solves this through Deterministic Pre-Calculation + Anti-Refusal Framing:
- Compliance Re-Framing: The model is explicitly framed as an authorized on-device analytical copilot reporting ground-truth telemetry, with a strict directive: NEVER output disclaimers like "I cannot provide financial advice" when reporting verified figures.
- Pre-Calculated Emergency Fund Math: When asked "What is my emergency fund buffer?", the exact 3-month target (£8,263.80), liquid capital (£7,462.12), shortfall (£801.68), and runway (81.3 days vs 90 days recommended) are pre-injected into the prompt.
- Whole-of-Wealth Balance Sheet: Accurate net worth asset and liability breakdowns are pre-assembled from the multi-asset database, eliminating arithmetic hallucinations.
- Smart Conversational Isolation & Clean Session Refresh: Topic-aware context filtering ensures previous conversation turns are only injected when queries are explicit follow-ups (
why,how come,explain that), preventing small SLMs from anchoring onto stale monologues. Users can start a clean session anytime via the Copilot UI🔄 New Chatbutton orreset_session=True. - Multi-Tier Web Knowledge Engine:
- Statutory HMRC Schedules: Instant sub-millisecond retrieval of exact UK tax allowances (£20,000 ISA, £4,000 LISA, £60,000 Pension, £3,000 CGT, PSA).
- Google Search / Serper API: Direct Google Search execution when
SERPER_API_KEYorGOOGLE_SEARCH_API_KEYis configured in.env. - Live Organic Web Extraction: Zero-overhead organic search parser pulling real titles, snippets, and official guidance from GOV.UK, HMRC, and NS&I without requiring paid API keys.
The application incorporates Tier-1 production engineering patterns across data storage, ML serving, security, and cloud orchestration:
-
High-Concurrency SQLite Storage (WAL Mode):
- Configured with
PRAGMA journal_mode = WAL;,PRAGMA busy_timeout = 5000;, andPRAGMA synchronous = NORMAL;to eliminate table-lock contention between concurrent web queries and background bank sync threads. - Compound indexes on
(session_id, id),(account_id, booking_date), and(statement_batch_id).
- Configured with
-
Pydantic Structured Outputs & MCP Tool Validation:
- ReAct reasoning steps are parsed and strictly validated through
ReActStepPayloadPydantic models. - All MCP tools (
fetch_boe_base_rate,query_spending_and_transactions,tax_wealth_audit, etc.) enforce type-checked Pydantic argument schemas (BoeRateArgs,SpendingQueryArgs,TaxWealthAuditArgs), preventing malformed LLM tool arguments from crashing execution.
- ReAct reasoning steps are parsed and strictly validated through
-
Cloud-Native Kubernetes Probes (
/healthz,/readyz,/livez):- Zero-overhead liveness and readiness probes allowing container orchestrators (Kubernetes, AWS ECS, GCP Cloud Run) to monitor service state and database connectivity without running heavy financial aggregations.
-
Enterprise HTTP Security & Distributed Tracing:
EnterpriseSecurityMiddlewaregenerates or propagatesX-Request-IDacross all inbound requests for end-to-end distributed tracing.- Enforces OWASP-recommended headers:
X-Content-Type-Options: nosniff,X-Frame-Options: SAMEORIGIN,Referrer-Policy: strict-origin-when-cross-origin, andX-Response-Time-Mslatency profiling.
-
Multi-Session Conversation Partitioning:
- Chat interactions and history endpoints support explicit
session_idscoping, ensuring multi-user SaaS deployments or parallel client sessions never leak or destructively reset conversation context.
- Chat interactions and history endpoints support explicit
-
Offline Evaluation Benchmark Suite (Evals-as-Code):
- Automated offline eval harness (
tests/test_eval_harness.py) executing on every CI commit:- Grounding Fidelity Benchmark: Verifies 100% of grounded claims pass and catches fabricated financial yields.
- Adversarial Prompt Injection Benchmark: Validates a 100% block rate against hostile directives (DAN, token escapes, SQL probes, exfiltration attacks) with 0% false positives on legitimate financial inquiries.
- Analytical SLA Benchmark: Asserts core deterministic engines (tax optimizer, credit scoring) execute in <15ms.
- Automated offline eval harness (
The repository includes a comprehensive test suite (120 tests) and automated CI pipeline:
# Run full unit and integration test suite (120 tests across storage, agents, MCP, ReAct, RAG, and security)
uv run pytest --verbose
# Run ultra-fast Ruff linter
uv run ruff check .
# Verify wheel and source distribution build
uv buildGitHub Actions executes across Python 3.11, 3.12, and 3.13 on every commit, alongside automated Gitleaks secret detection and Ruff validation.
Interactive architectural diagrams, sequence flows, grounding guardrails, and data schemas are accessible live:
- 🌐 Live Interactive Architecture Map (GitHub Pages)
- 💼 Executive Business Capabilities & Value Guide (GitHub Pages)
- 🏛️ Interactive Enterprise Architecture Walkthrough (GitHub Pages)
- 🧮 Interactive Credit Score Explorer (GitHub Pages)
- 🧭 Complete System Walkthrough & User Guide (GitHub Pages)
- 📐 AI Gateway Architecture Diagram & Failover Spec
- 📁 Interactive Explainers Directory
- Local Business Capabilities Guide:
http://localhost:8080/business(or/features). - Local System Tour and Module Guide:
http://localhost:8080/guide. - Local Interactive Architecture:
http://localhost:8080/architecture.
This project was autonomously designed, coded, verified, and documented by Gemini 3.8 Flash operating via the Google Antigravity CLI (agy). Claude Code was not substantially used in the development of this codebase.
MIT License. Designed for personal financial sovereignty.
