Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
57 changes: 56 additions & 1 deletion .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -28,10 +28,20 @@ API_KEY=
# MODEL_BASE_URL=http://172.17.0.1:8000/v1 # serving model only (agent, A/B, rollouts)
# BASE_URL=http://172.17.0.1:11434/v1 # or everything: teacher + judge too

# Routing model + thresholds. The defaults are calibrated TOGETHER on a drafted routing eval;
# Routing model + thresholds. The defaults are calibrated TOGETHER on a drafted routing eval.
# CPU stays the safe default. NVIDIA hosts can opt into the larger Apache-2.0 Qwen3-Embedding-8B
# GPU server with:
# docker compose -f docker-compose.yml -f compose.gpu-embeddings.yml \
# --profile gpu-embeddings up -d embed-gpu mcp
# The overlay sets these four values; set them directly only for an external compatible server:
# EMBED_BACKEND=remote
# EMBED_BASE_URL=http://embed-gpu:8080/v1
# EMBED_REMOTE_MODEL=Qwen/Qwen3-Embedding-8B-GGUF
# EMBED_TIMEOUT_SECONDS=120
# if you override EMBED_MODEL, recalibrate the three scores with it (cosine distributions differ
# per model. For the previous default BAAI/bge-small-en-v1.5 use 0.65 / 0.45 / 0.93):
# EMBED_MODEL=onnx-community/Qwen3-Embedding-0.6B-ONNX # or any fastembed model name
# ROUTER_BODY_CHARS=1000 # approved body prefix embedded beside name + description; max 4000
# MIN_SCORE=0.53 # at/above -> routable match; below -> related band or novel
# RELATED_SCORE=0.37 # floor of the "related" (compose/extend) band; below it a task
# # is novel (weak/strong escalation)
Expand All @@ -57,6 +67,51 @@ API_KEY=
# MINE_MAX_JUDGE_CALLS=24 # new representative judge calls per run; <=0 removes the cap
# MINE_CLUSTER_THRESHOLD=0.90 # task cosine at/above this shares a representative verdict

# Vault publisher handoff. Approval writes runs/publications at mode 0700, and the publisher
# (ops/systemd/ingot-publisher.service) reads those receipts from the host as an ordinary user.
# When both run on one machine, set these to that user so the two agree on who owns the receipts.
# Leaving them unset keeps the UI container as root, which is right for every deployment with no
# host publisher. Get them wrong and the failure is silent in the console: approvals queue, the
# publisher lists an empty directory, and nothing publishes. The publisher logs it at each poll.
# INGOT_UID=1000
# INGOT_GID=1000

# Where mutable state lives: the served library, the review queue, publication receipts, evidence,
# snapshots, and eval task sets. Defaults to $XDG_STATE_HOME/ingot (else ~/.local/state/ingot), so
# a pip-installed Ingot keeps nothing inside site-packages where an upgrade would discard it.
# Compose sets each path explicitly to what it mounted. `ingot status` prints every resolved path,
# where it came from, and whether it is writable.
# INGOT_HOME moves all of them at once; the specific settings override it one at a time.
# INGOT_HOME=~/.local/state/ingot
# INGOT_LIBRARY=/srv/ingot/library # SKILLS_DIR is the deprecated name for this
# INGOT_RUNS=/srv/ingot/runs
# INGOT_TASKS=/srv/ingot/tasks

# Published host ports. Both stay on loopback; only the host side moves. Set them when this box
# already runs something on 8000 or 8080, including a second Ingot stack. `0` asks the kernel for
# a free port, which is what scripts/managed_smoke.sh does: it reaches every container through
# `docker compose exec` and needs no host port at all.
# INGOT_MCP_PORT=8000
# INGOT_UI_PORT=8080

# Publication backend. `local` (the default) publishes into ./vault, a Git repository on this
# machine: no network, no GitHub account, no `gh`. `forge` makes a merged pull request the
# publication authority instead, which anchors activation off-box at the cost of the air gap and
# requires compose.forge.yaml plus a ./vault that is a clone of the repository below. Setting the
# forge variables without selecting the backend is inert, and the publisher says so at startup.
# INGOT_PUBLISH_BACKEND=local
# INGOT_FORGE_REPOSITORY=owner/repo
# INGOT_FORGE_REMOTE=origin
# INGOT_FORGE_BRANCH=main

# Delivery targets: where an approved revision is installed once the vault carries it. The vault is
# the managed-MCP library, so agents using Ingot's MCP server are already served; this is for an
# agent that reads a native skill directory on disk instead. Comma-separated `name=kind:path`;
# `filesystem` is the only kind you configure (the vault target is always present). The publisher
# creates each root at startup and refuses to start if it cannot. Names are yours -- Ingot knows
# nothing about what reads the directory.
# INGOT_DELIVERY_TARGETS=claude=filesystem:~/.claude/skills,codex=filesystem:~/.codex/skills

# Change-control UI login. Three modes; AUTH_MODE picks one (compose default: password).
# See docs/sso.md. To share the UI beyond this machine, use the TLS front door:
# `docker compose --profile lan up -d proxy` (docs/security.md "Network exposure").
Expand Down
35 changes: 34 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,39 @@ jobs:
- name: Build image
run: docker build -t ingot-mcp .
- name: Run tests
run: docker run --rm -v "$PWD:/app" -w /app ingot-mcp python -m pytest tests -q
# git comes from the image now: the publisher container drives real repositories, so it is
# a runtime dependency rather than something only the tests need.
run: >
docker run --rm -v "$PWD:/app" -w /app ingot-mcp
python -m pytest tests -q
- name: Smoke test Compose and Langfuse TLS
run: ./scripts/compose_smoke.sh

# The control-plane claim is that no non-publisher service can change what is served. Everything
# else that checks it -- tests/test_compose_managed.py, `docker compose config` -- reads the
# configuration. This job is the only one that watches the kernel refuse the write, so it is the
# one that has to pass before that claim is repeated anywhere. Make it a required check.
managed:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: One writer, and it is the publisher
run: ./scripts/managed_smoke.sh
- name: A removed read-only mount must fail the smoke test
# A check that cannot fail proves nothing. This deliberately breaks the invariant and
# requires the script to notice, so a later edit that quietly drops a `:ro` cannot leave a
# green job behind it.
run: |
python3 - <<'PY'
import pathlib
path = pathlib.Path("docker-compose.yml")
text = path.read_text()
broken = text.replace("./vault:/app/skills:ro", "./vault:/app/skills", 1)
assert broken != text, "no read-only served mount left to break"
path.write_text(broken)
PY
if MANAGED_SMOKE_PROJECT=ingot-managed-negative ./scripts/managed_smoke.sh; then
echo "the smoke test passed with a writable served mount; it is checking nothing"
exit 1
fi
git checkout -- docker-compose.yml
41 changes: 41 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -18,10 +18,51 @@ __pycache__/
.pytest_cache/
.hf_cache/
.venv/
/.venv-harbor-gateway/
.worktrees/

# packaging build artifacts (`pip install -e .`)
*.egg-info/
build/
dist/

# fetched skills are optional & not redistributed here (scripts/fetch_skills.sh)
skills/*
!skills/.gitkeep

# the managed stack's skill vault: its own Git repository, created by `ingot vault init`
vault/

# eval task sets are runtime artifacts (auto-drafted or user-authored), not shipped opinions
optimize/tasks/*.yaml
# ...except a hand-authored one, which is the measuring instrument rather than its output: the
# seeded working trees and weighted checklists are the experiment's design, and a matrix produced
# by a task set that no longer exists in the tree cannot be reproduced or argued with.
!optimize/tasks/build-loop.yaml
!optimize/tasks/adversarial-council-review.yaml
!optimize/tasks/assumption-audit.yaml
!optimize/tasks/auditing-economic-claims.yaml
!optimize/tasks/auditing-system-claims.yaml
!optimize/tasks/decomposing-skill-libraries.yaml
!optimize/tasks/forward-intro.yaml
!optimize/tasks/isolated-integration-fixtures.yaml
!optimize/tasks/live-caller-gate.yaml
!optimize/tasks/measurement-integrity.yaml
!optimize/tasks/memory-defrag.yaml
!optimize/tasks/memory-notes.yaml
!optimize/tasks/memory-reflect.yaml
!optimize/tasks/operating-accountability-loop.yaml
!optimize/tasks/oss-ready.yaml
!optimize/tasks/product-marketing.yaml
!optimize/tasks/prose-style-hemingway.yaml
!optimize/tasks/copywriting.yaml
!optimize/tasks/routing-economic-evidence.yaml
!optimize/tasks/skill-security.yaml
!optimize/tasks/skill-retrospective.yaml
!optimize/tasks/agentic-action-safety.yaml
!optimize/tasks/turning-buyer-notes-into-decisions.yaml
!optimize/tasks/unattended-overnight-ops.yaml
!optimize/tasks/writing-clearly-and-concisely.yaml
!optimize/tasks/linting-implementation-plans.yaml
!optimize/tasks/op-credentials.yaml
!optimize/tasks/slancha-cred.yaml
33 changes: 18 additions & 15 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
@@ -1,10 +1,10 @@
# Architecture

Ingot is a local-first change-control system for agent instructions. A skill folder is the unit of
change; every version of it is content-addressed, every proposed change is quarantined until a human
approves it, and every promotion is atomic and reversible. Routing exists to serve the approved
revision to an agent. Ingot is not a multi-tenant service, and the default Compose deployment
exposes only loopback ports.
Ingot is a local-first, air-gappable control plane for a team's skill library. A skill folder is the
unit of change; every version of it is content-addressed, every proposed change is quarantined until
a human approves it, and every promotion is atomic and reversible. Routing exists to serve the
approved revision to an agent. Ingot is not a multi-tenant service, and the default Compose
deployment exposes only loopback ports.

## The change-control pipeline

Expand Down Expand Up @@ -36,19 +36,22 @@ exposes only loopback ports.

1. The bundled agent sends task and execution context to `route_and_load` over MCP.
2. The router refreshes the skill registry when files change, filters by harness, platform, scope,
tools, MCPs, activation, and trust, then ranks compatible descriptions. Description embeddings
are cached across refreshes.
tools, MCPs, activation, and trust, then ranks compatible skills by the stronger cosine score
from the description or a bounded document containing name, description, and approved
harness-specific body. Both representations share a 4,096-vector least-recently-used cache
across refreshes; description-only vectors remain authoritative for collision detection.
3. One response is authoritative for the direct `match` or explicit `related_match`, loaded body,
revision, root, body-free alternatives, and `novel` escalation signal. A related match is loaded
for compose-or-extend use. The agent uses the weak model unless `novel` is true.
revision, root, body-free alternatives, component scores, and `novel` escalation signal. A
related match is loaded for compose-or-extend use. The agent uses the weak model unless `novel`
is true.
4. The run is recorded to Langfuse (the default evals backend, or a Langfuse-compatible endpoint
`LANGFUSE_*` points at); mining reads it back and has no local fallback. Hosted model calls use
the configured OpenAI-compatible endpoint. OpenRouter calls always request ZDR providers.

## SkillOpt optimization

SkillOpt optimization is a core product capability. It proposes changes but never activates them.
Runs start in the background (`optimize.loop`) or on demand from the UI, and every result enters the
Runs start in the background (`ingot.optimize.loop`) or on demand from the UI, and every result enters the
same human review path.

Mining reads every usable Langfuse trace by default, with `--limit N` available only as an explicit
Expand Down Expand Up @@ -102,7 +105,7 @@ text components (`OPTIMIZE_COMPONENTS=body,file:<path>`), diffed for review and

## Stores and ownership

`skills/` contains active skills. `optimize/tasks/` contains eval sets. `runs/pending/` contains one
`skills/` contains active skills. `ingot/optimize/tasks/` contains eval sets. `runs/pending/` contains one
active review slot per skill, with displaced candidates archived. `runs/revisions/` contains
rollback snapshots, plus a `.snapshots.json` index per skill that records when each revision was
last snapshotted; it sits beside the snapshot directories, never inside one, so a rollback restores
Expand Down Expand Up @@ -163,13 +166,13 @@ Promotion stages changes and restores the prior directory if the swap fails. Eve
the displaced revision. Restore it from the UI's History section, or with:

```bash
docker compose run --rm --entrypoint python optimize -m optimize.promote rollback SKILL REVISION
docker compose run --rm --entrypoint python optimize -m ingot.optimize.promote rollback SKILL REVISION
```

The `optimize` service's entrypoint is `python -m optimize.ab`, so the entrypoint override is what
makes the arguments reach `optimize.promote`.
The `optimize` service's entrypoint is `python -m ingot.optimize.ab`, so the entrypoint override is what
makes the arguments reach `ingot.optimize.promote`.

Operators should back up `skills/`, `runs/`, and `optimize/tasks/`. Container databases require
Operators should back up `skills/`, `runs/`, and `ingot/optimize/tasks/`. Container databases require
normal volume backup procedures.

## Trust boundaries
Expand Down
20 changes: 17 additions & 3 deletions Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,13 @@ FROM python:3.12-slim-bookworm

WORKDIR /app

# The publisher's vault is a Git repository and every publication is a worktree, a commit and a
# fast-forward. slim does not ship git, so without this the one service that owns the served
# library cannot start.
RUN apt-get update \
&& apt-get install -y --no-install-recommends git \
&& rm -rf /var/lib/apt/lists/*

COPY requirements.txt .
RUN pip install --no-cache-dir pip==26.1.2 \
&& pip install --no-cache-dir -r requirements.txt
Expand Down Expand Up @@ -32,12 +39,19 @@ ARG FALLBACK_EMBED_MODEL=BAAI/bge-small-en-v1.5
RUN python -c "from fastembed import TextEmbedding; TextEmbedding('${FALLBACK_EMBED_MODEL}')"
ENV BAKED_FALLBACK_EMBED_MODEL=${FALLBACK_EMBED_MODEL}

COPY mcp_server ./mcp_server
# `mcp_server` and `optimize` live under `ingot/` and arrive with the COPY below.
COPY agent ./agent
COPY optimize ./optimize
COPY ui ./ui
COPY skills ./skills

# The `ingot` console script. `--no-deps` keeps the installed set exactly the pinned
# requirements.txt above instead of re-resolving it from pyproject.toml, and `-e` points the script
# at the /app copies the services already run with `python -m`, so there is only ever one copy of
# the code in the image.
COPY pyproject.toml README.md ./
COPY ingot ./ingot
RUN pip install --no-cache-dir --no-deps -e .

ENV PYTHONUNBUFFERED=1

CMD ["python", "-m", "mcp_server.server"]
CMD ["python", "-m", "ingot.mcp_server.server"]
2 changes: 1 addition & 1 deletion PRODUCTION_SETUP.md
Original file line number Diff line number Diff line change
Expand Up @@ -209,7 +209,7 @@ Codex, with the exact rule in [Make skill loading part of the agent instructions
## Operations

Back up all named Langfuse datastore volumes and the repository's `skills/`, `runs/`, and
`optimize/tasks/` directories. Pin image versions, review upgrades before applying them, and test
`ingot/optimize/tasks/` directories. Pin image versions, review upgrades before applying them, and test
restore procedures. Monitor container health and disk usage:

```bash
Expand Down
Loading
Loading