Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
37 commits
Select commit Hold shift + click to select a range
3915351
Rearchitect for doc/code parity; evidence-gated defaults; E2E fixes
PromtEngineer Aug 9, 2026
c62014f
Add evaluation harness: gold set, retrieval metrics, judge, smoke, de…
PromtEngineer Aug 9, 2026
4ed293d
Add research evidence, evidence-based roadmap, and design rationale
PromtEngineer Aug 9, 2026
c373825
Roadmap Phase 4: evidence-filtered adoption plan for agentic-file-sea…
PromtEngineer Aug 9, 2026
f54992e
Phase 4: flag-gated adoptions from agentic-file-search, decided by me…
PromtEngineer Aug 10, 2026
e3c3d60
eval: acquisition corpus, final-list metrics, Phase-4 benchmark record
PromtEngineer Aug 10, 2026
c2d9f48
docs: Phase-4 verdicts, token accounting, filters, and parity fixes
PromtEngineer Aug 10, 2026
007d0b6
fix: size Ollama num_ctx per request — ends silent 8k front-truncation
PromtEngineer Aug 12, 2026
2a0774f
eval: 4.1 escalation re-run post-truncation-fix — reject as default
PromtEngineer Aug 12, 2026
184dcc7
eval: Anthropic-backend judge option + verifier-suffix stripping + ha…
PromtEngineer Aug 13, 2026
33adb7c
eval: validate Sonnet-class judge via subagent voters — 38/38, zero s…
PromtEngineer Aug 13, 2026
513344f
fix: think:false on synthesis streams; quote-safe FTS with dense fall…
PromtEngineer Aug 13, 2026
7d71051
fix: chunker dropped half of every long document
PromtEngineer Aug 13, 2026
dc539d7
eval: unseen-corpus RFC dataset — 23 real cross-referencing documents…
PromtEngineer Aug 13, 2026
80d5215
eval: RFC shakedown results — retrieval passes on unseen docs, synthe…
PromtEngineer Aug 13, 2026
f5c587e
synthesis: strict grounded prompt + temperature-0 decode (A/B arm C);…
PromtEngineer Aug 13, 2026
d6c4bc3
eval: strict compose prompt tested and rejected (arm E) — revert, kee…
PromtEngineer Aug 13, 2026
aafc193
Dedupe retrieval legs + budget synthesis context (7/24 -> 16/24)
PromtEngineer Aug 15, 2026
e70611b
Reranker as final-stage selection: Qwen3-Reranker-4B on by default (a…
PromtEngineer Aug 15, 2026
bd8ad26
Pooled decomposition + deterministic decompose (ties composer, wins o…
PromtEngineer Aug 15, 2026
a3f999a
Label synthesis snippets with their source document (17/24 -> 21/24)
PromtEngineer Aug 15, 2026
36c4b6e
eval: cross-bench validation — RFC-tuned changes generalize; hr -3 di…
PromtEngineer Aug 15, 2026
243f7be
eval: completeness-clause A/B (arm J) — reverted, does not fix hr tar…
PromtEngineer Aug 16, 2026
ebcc88b
fix: create FTS index on late-chunk tables + pin 4b judge to temp 0
PromtEngineer Aug 16, 2026
c1400ca
feat: decomposer sees last assistant answer on multi-turn; single-tur…
PromtEngineer Aug 16, 2026
3a8ebd9
fix: apply code-review fix-set — correctness, security, determinism, CI
PromtEngineer Aug 17, 2026
3f78be4
docs: sync all docs with shipped behavior (arm G/H defaults, status c…
PromtEngineer Aug 17, 2026
e0f2399
eval: fix-set impact measured — authored +4, rfc 21->19 trade, hr_h05…
PromtEngineer Aug 17, 2026
f697536
fix: dev-watcher stream aborts, invisible avatars; feat: Quick Chat s…
PromtEngineer Aug 18, 2026
ac38f51
fix: markdown answers rendered with doubled whitespace
PromtEngineer Aug 18, 2026
75b5348
feat: component ablation study — latechunk + verification off by default
PromtEngineer Aug 19, 2026
52553f9
eval: resolve-only + MaxSim experiments — neither adopted; multi-turn…
PromtEngineer Aug 19, 2026
ceefed0
eval: multi-vector first-stage retrieval experiment — not adopted
PromtEngineer Aug 20, 2026
e363345
eval: third-RRF-leg multi-vector experiment — not adopted (Sonnet-jud…
PromtEngineer Aug 20, 2026
d7d1205
eval: paraphrase-robustness study — MV third leg wins +2 on reworded …
PromtEngineer Aug 20, 2026
8ba3888
eval: candidate-pool union fusion — +4 real on paraphrased queries, -…
PromtEngineer Aug 20, 2026
fc994c7
docs: sync README with current defaults and eval state
PromtEngineer Aug 21, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
101 changes: 101 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
@@ -0,0 +1,101 @@
# localGPT environment configuration
#
# Copy to .env and edit. Every variable below is read by code; the value shown
# after "default:" is what the code uses when the variable is unset.
#
# cp .env.example .env
#
# For Docker use docker.env instead (it is passed with --env-file and also
# supplies build-time values for the frontend).

# ---------------------------------------------------------------------------
# Services
# ---------------------------------------------------------------------------

# Ollama server. Read by backend/ollama_client.py and rag_system/main.py.
# In Docker this becomes http://host.docker.internal:11434.
# default: http://localhost:11434
OLLAMA_HOST=http://localhost:11434

# Base URL of the RAG API. backend/server.py builds /chat and /index from it.
# In Docker compose this is http://rag-api:8001.
# default: http://localhost:8001
RAG_API_URL=http://localhost:8001

# Browser-facing URLs, read by the Next.js frontend (src/lib/api.ts).
# NEXT_PUBLIC_* values are inlined at build time, so change them before `npm run build`.
# default: http://localhost:8000
NEXT_PUBLIC_API_URL=http://localhost:8000
# default: http://localhost:8001
NEXT_PUBLIC_RAG_API_URL=http://localhost:8001

# ---------------------------------------------------------------------------
# Storage
# ---------------------------------------------------------------------------

# SQLite database holding sessions, messages and index metadata.
# Set this to a shared path so the backend and the RAG API use one file.
# default: backend/chat_data.db (local) or /app/backend/chat_data.db (Docker)
# DB_PATH=backend/chat_data.db

# LanceDB vector store. Defaults to the `storage.lancedb_uri` of the active
# pipeline config in rag_system/main.py.
# default: ./lancedb
# LANCEDB_PATH=./lancedb

# ---------------------------------------------------------------------------
# Models
# ---------------------------------------------------------------------------

# Answer generation (Ollama). Options: qwen3.6:27b (high-end, ~17GB), qwen3.5:4b (light).
# default: qwen3.5:9b
GENERATION_MODEL=qwen3.5:9b

# Routing, triage, query decomposition, contextual enrichment and verification
# (Ollama). Light option: qwen3.5:2b.
# default: qwen3.5:4b
ENRICHMENT_MODEL=qwen3.5:4b

# Embeddings (HuggingFace). The default is MIT-licensed, 1.2GB, 1024-dim, and
# measured best on our gold set (eval/DECISIONS.md).
# Option: Qwen/Qwen3-Embedding-4B for multilingual / long-context corpora.
# Changing this requires re-indexing every existing index - the stored vectors
# belong to the old model's vector space. localGPT records the embedding model
# on each table and refuses to query it with a different one.
# default: microsoft/harrier-oss-v1-0.6b
EMBEDDING_MODEL=microsoft/harrier-oss-v1-0.6b

# Reranker (HuggingFace). Only loaded when reranking is switched on - the
# default profile ships with it OFF, because the first stage above already
# outranks the cheap cross-encoder (eval/DECISIONS.md). When you do switch it
# on (UI "AI reranker" toggle, or reranker.enabled in the profile) this is the
# model that gets loaded, lazily.
# Options: BAAI/bge-reranker-v2-m3 (low latency, only pays off with a weaker
# embedder), answerdotai/answerai-colbert-small-v1, Qwen/Qwen3-Reranker-0.6B.
# default: Qwen/Qwen3-Reranker-4B
RERANKER_MODEL=Qwen/Qwen3-Reranker-4B

# ---------------------------------------------------------------------------
# Optional tuning
# ---------------------------------------------------------------------------

# Seconds the backend waits for a RAG API chat response.
# default: 600
# RAG_API_TIMEOUT=600

# Seconds the backend waits for a RAG API indexing run.
# default: 3600
# RAG_API_INDEX_TIMEOUT=3600

# LLM backend selector for rag_system (`ollama` or `watsonx`).
# WatsonX additionally needs the WATSONX_* variables - see env.example.watsonx.
# default: ollama
# LLM_BACKEND=ollama

# HuggingFace token, only needed for gated model downloads.
# HF_TOKEN=

# Ollama divides its context window (and any per-request num_ctx) across its
# parallel slots — with 2 slots a 32k request is served as a 16k window and
# oversized prompts are silently front-truncated. For single-user localGPT:
# OLLAMA_NUM_PARALLEL=1
2 changes: 1 addition & 1 deletion .github/ISSUE_TEMPLATE/bug_report.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,7 +53,7 @@ Please include relevant error messages or logs:

## 🔧 Configuration
- Deployment method: [Docker / Direct Python]
- Models used: [e.g. qwen3:0.6b, qwen3:8b]
- Models used: [e.g. qwen3.5:9b, qwen3.5:4b]
- Document types: [e.g. PDF, DOCX, TXT]

## 📎 Additional Context
Expand Down
25 changes: 25 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
name: CI

on:
push:
branches: [main]
pull_request:

jobs:
gateway-routing:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4

- uses: actions/setup-python@v5
with:
python-version: '3.12'

- name: Install test dependencies
run: pip install requests python-dotenv

# Plain-script test (no pytest/unittest cases); exits non-zero on failure.
# backend/server.py guards its rag_system import, so only the two light
# dependencies above are needed.
- name: Run gateway routing test
run: python backend/test_gateway_routing.py
20 changes: 16 additions & 4 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -63,16 +63,28 @@ logs/

# SQLite or other database files
*.db
#backend/*.db
# backend/chat_history.db
backend/chroma_db/
backend/chroma_db/**

# Document and user-uploaded files (PDFs, images, etc.)
rag_system/documents/
*.pdf

# Ensure docker.env remains tracked
# Ensure docker.env and .env.example remain tracked
!docker.env
!backend/chat_data.db
!.env.example

# Phase 0 evaluation harness (eval/) — rebuildable artefacts only.
# The corpora, gold set, scripts and BASELINE.md are tracked.
eval/.eval_indexes/
eval/results/
!eval/corpora/*.pdf
# The Phase 4 cross-reference corpus lives one directory deeper.
!eval/corpora/acquisition/*.pdf


# local virtualenv
.venv/

# agent worktree / session state
.claude/
Loading
Loading