Skip to content

AMUL LLM: llm-gateway backend, API keys, and HTTP routes - #7

Merged
warheart1984-ctrl merged 2 commits into
mainfrom
feat/llm-gateway-adapter
Oct 2, 2026
Merged

warheart1984-ctrl merged 2 commits into
mainfrom
feat/llm-gateway-adapter

Conversation

@warheart1984-ctrl

@warheart1984-ctrl warheart1984-ctrl commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

What

  • llm-gateway backend. With JARVIS_LLM_API=llm-gateway, core_model_generate calls POST {JARVIS_LLM_URL}/chat/complete (warheart1984-ctrl/llm-gateway) with temperature / max_tokens nested under params. It reads the answer from content, falling back to reasoning, and reports backend="llm-gateway". The default stays openai (unchanged behaviour).
  • Backend credential. JARVIS_LLM_API_KEY or JARVIS_LLM_API_KEY_FILE is sent as Authorization: Bearer … on generation and on /models discovery, in both modes. With neither set, no header is sent, as before.
  • httpx is now a runtime dependency. amul_llm imports it to reach the backend, but it was only in [dev].
  • HTTP routes. GET /api/jarvis/llm/status and POST /api/jarvis/llm/generate (body: PromptContract, response: the replay record). Both are behind JARVIS_API_KEY through ApiKeyMiddleware. An unknown mode override returns 400.
  • Memory recall in generate. recall defaults to true. The route recalls through emr_recall, so the evidence floor and the conflict membrane both apply. Recalled memories go ahead of the caller's context, each with its id, type, status and confidence, and the model is told they're recorded claims, not verified truth. Unresolved conflicts are stated even when the membrane holds every claim back, so the model says they conflict instead of guessing. The replay record gains recall (query, memory ids, abstention, conflict subjects), so R-B still holds. Options: recall_query, recall_intent, subjects, max_memories, truth_scope.
  • Confidence fix. generate() only scored backend == "openai-compat" as a real model, so an llm-gateway answer got a lower confidence than the same answer from an OpenAI-compatible server. Both count now (MODEL_BACKENDS).

Why

  • The OpenAI-only adapter couldn't use llm-gateway. It spoke only the OpenAI dialect and never sent a key, so against llm-gateway every call failed. The except Exception fallback then answered from the echo stub, with no error visible to the caller.
  • The Docker image could never reach a backend. It runs pip install ., which doesn't install httpx, so every generation hit ImportError and fell back to the stub. That's fixed by the dependency change, independently of the gateway work.
  • Nothing in the service called amul_llm. Only the tests did. The routes make the governed loop reachable.

Verified

  • pytest -q: 230 passed. That includes 5 new adapter tests in tests/test_amul_llm.py, with httpx.post patched so there's no network, covering request shape, headers, the reasoning fallback, a 402 degrading to the stub, and key-file precedence. It also includes 5 new route tests in tests/test_llm_api.py, covering status, the replay record and its log line, the 400 on an unknown mode, 422 validation, and the 401 without the key. Five more recall tests cover: a matching memory goes into context first and into the replay log, an unrelated query abstains, recall: false, a conflict is stated without picking a side, and grounded confidence for the gateway backend.
  • Live, against a running llm-gateway (tenant memory, Groq gpt-oss-20b): /api/jarvis/llm/generate returned a real model answer with model_version=groq/gpt-oss-20b. The gateway's metrics counted the requests under the tenant, and the routes returned 401 without the key.

For the reviewer

  • Recall is lexical, and full questions can miss. Live, against the real ledger, "Where does llm-gateway run, and on which port?" abstained (top-score-below-floor), while "llm-gateway port" and "where does the llm-gateway run" recalled the memory. With nothing recalled, the model invented a port. Callers that know their keywords should pass recall_query. Better retrieval (neural embeddings are declared in the RAG maturity table) is the real fix, and it's out of scope here.

  • Silent fallback unchanged. On any backend failure the adapter still falls back to the echo stub. The replay record shows it (metadata.model_version == "echo-stub-v0"), but nothing is logged. Worth a follow-up if you want failures to be loud.

  • 33 pre-existing failures in agent-hooks/ tests. agent-hooks/ contains older copies of app/ modules and tests. Run directly, those tests fail 33 of 145 on main too. CI runs only tests/. Not touched here.

  • RAG adapter still OpenAI-only. amul_rag.llm_generate (JARVIS_RAG_LLM_URL) is a separate path. It doesn't get the gateway mode or the key in this PR.

🤖 Generated with Claude Code

The core-model adapter only spoke the OpenAI /chat/completions dialect and
sent no credential, so it could not use a keyed backend; against
llm-gateway every call failed and degraded silently to the echo stub.

- JARVIS_LLM_API=llm-gateway: POST {url}/chat/complete with params nested
  under `params` (max_tokens is required there to reserve cost), answer
  from `content`, falling back to `reasoning`. backend="llm-gateway".
- JARVIS_LLM_API_KEY / JARVIS_LLM_API_KEY_FILE: sent as a Bearer token on
  generation and model discovery, in both modes. Unset: no header, as before.
- httpx becomes a runtime dependency: amul_llm imports it to call the
  backend, so a plain `pip install .` (the Docker image) always fell back
  to the stub.
- GET /api/jarvis/llm/status and POST /api/jarvis/llm/generate expose the
  governed loop over HTTP. Nothing in the service called amul_llm before.
  Both sit behind JARVIS_API_KEY via ApiKeyMiddleware; an unknown mode
  override is a 400.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

POST /api/jarvis/llm/generate now recalls from the ledger before
generating (`recall`, default true), through emr_recall so the evidence
floor and the conflict membrane apply. Recalled memories go ahead of the
caller's context with their id, type, status and confidence, and the model
is told they are recorded claims, not verified truth. Unresolved conflicts
are stated even when the membrane holds every claim back, so the model
hears that they conflict instead of guessing.

The replay record gains `recall` (query, memory ids, abstention, conflict
subjects), keeping a recalled generation replayable (R-B). Options:
recall_query, recall_intent, subjects, max_memories, truth_scope.

Also: generate() scored only backend="openai-compat" as a real model, so
an llm-gateway answer got a lower confidence than the same answer from an
OpenAI-compatible server. Both are MODEL_BACKENDS now.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

@warheart1984-ctrl
warheart1984-ctrl merged commit c1893fd into main Oct 2, 2026
2 checks passed
@warheart1984-ctrl
warheart1984-ctrl deleted the feat/llm-gateway-adapter branch October 2, 2026 02:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant