AMUL LLM: llm-gateway backend, API keys, and HTTP routes - #7
Merged
Merged
Conversation
The core-model adapter only spoke the OpenAI /chat/completions dialect and
sent no credential, so it could not use a keyed backend; against
llm-gateway every call failed and degraded silently to the echo stub.
- JARVIS_LLM_API=llm-gateway: POST {url}/chat/complete with params nested
under `params` (max_tokens is required there to reserve cost), answer
from `content`, falling back to `reasoning`. backend="llm-gateway".
- JARVIS_LLM_API_KEY / JARVIS_LLM_API_KEY_FILE: sent as a Bearer token on
generation and model discovery, in both modes. Unset: no header, as before.
- httpx becomes a runtime dependency: amul_llm imports it to call the
backend, so a plain `pip install .` (the Docker image) always fell back
to the stub.
- GET /api/jarvis/llm/status and POST /api/jarvis/llm/generate expose the
governed loop over HTTP. Nothing in the service called amul_llm before.
Both sit behind JARVIS_API_KEY via ApiKeyMiddleware; an unknown mode
override is a 400.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
Your trial has ended. Reactivate Greptile to resume code reviews.
POST /api/jarvis/llm/generate now recalls from the ledger before generating (`recall`, default true), through emr_recall so the evidence floor and the conflict membrane apply. Recalled memories go ahead of the caller's context with their id, type, status and confidence, and the model is told they are recorded claims, not verified truth. Unresolved conflicts are stated even when the membrane holds every claim back, so the model hears that they conflict instead of guessing. The replay record gains `recall` (query, memory ids, abstention, conflict subjects), keeping a recalled generation replayable (R-B). Options: recall_query, recall_intent, subjects, max_memories, truth_scope. Also: generate() scored only backend="openai-compat" as a real model, so an llm-gateway answer got a lower confidence than the same answer from an OpenAI-compatible server. Both are MODEL_BACKENDS now. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
Your trial has ended. Reactivate Greptile to resume code reviews.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
JARVIS_LLM_API=llm-gateway,core_model_generatecallsPOST {JARVIS_LLM_URL}/chat/complete(warheart1984-ctrl/llm-gateway) withtemperature/max_tokensnested underparams. It reads the answer fromcontent, falling back toreasoning, and reportsbackend="llm-gateway". The default staysopenai(unchanged behaviour).JARVIS_LLM_API_KEYorJARVIS_LLM_API_KEY_FILEis sent asAuthorization: Bearer …on generation and on/modelsdiscovery, in both modes. With neither set, no header is sent, as before.httpxis now a runtime dependency.amul_llmimports it to reach the backend, but it was only in[dev].GET /api/jarvis/llm/statusandPOST /api/jarvis/llm/generate(body:PromptContract, response: the replay record). Both are behindJARVIS_API_KEYthroughApiKeyMiddleware. An unknownmodeoverride returns 400.generate.recalldefaults to true. The route recalls throughemr_recall, so the evidence floor and the conflict membrane both apply. Recalled memories go ahead of the caller'scontext, each with its id, type, status and confidence, and the model is told they're recorded claims, not verified truth. Unresolved conflicts are stated even when the membrane holds every claim back, so the model says they conflict instead of guessing. The replay record gainsrecall(query, memory ids, abstention, conflict subjects), so R-B still holds. Options:recall_query,recall_intent,subjects,max_memories,truth_scope.generate()only scoredbackend == "openai-compat"as a real model, so an llm-gateway answer got a lower confidence than the same answer from an OpenAI-compatible server. Both count now (MODEL_BACKENDS).Why
except Exceptionfallback then answered from the echo stub, with no error visible to the caller.pip install ., which doesn't installhttpx, so every generation hitImportErrorand fell back to the stub. That's fixed by the dependency change, independently of the gateway work.amul_llm. Only the tests did. The routes make the governed loop reachable.Verified
pytest -q: 230 passed. That includes 5 new adapter tests intests/test_amul_llm.py, withhttpx.postpatched so there's no network, covering request shape, headers, the reasoning fallback, a 402 degrading to the stub, and key-file precedence. It also includes 5 new route tests intests/test_llm_api.py, covering status, the replay record and its log line, the 400 on an unknown mode, 422 validation, and the 401 without the key. Five more recall tests cover: a matching memory goes into context first and into the replay log, an unrelated query abstains,recall: false, a conflict is stated without picking a side, and grounded confidence for the gateway backend.memory, Groqgpt-oss-20b):/api/jarvis/llm/generatereturned a real model answer withmodel_version=groq/gpt-oss-20b. The gateway's metrics counted the requests under the tenant, and the routes returned 401 without the key.For the reviewer
Recall is lexical, and full questions can miss. Live, against the real ledger,
"Where does llm-gateway run, and on which port?"abstained (top-score-below-floor), while"llm-gateway port"and"where does the llm-gateway run"recalled the memory. With nothing recalled, the model invented a port. Callers that know their keywords should passrecall_query. Better retrieval (neural embeddings are declared in the RAG maturity table) is the real fix, and it's out of scope here.Silent fallback unchanged. On any backend failure the adapter still falls back to the echo stub. The replay record shows it (
metadata.model_version == "echo-stub-v0"), but nothing is logged. Worth a follow-up if you want failures to be loud.33 pre-existing failures in
agent-hooks/tests.agent-hooks/contains older copies ofapp/modules and tests. Run directly, those tests fail 33 of 145 onmaintoo. CI runs onlytests/. Not touched here.RAG adapter still OpenAI-only.
amul_rag.llm_generate(JARVIS_RAG_LLM_URL) is a separate path. It doesn't get the gateway mode or the key in this PR.🤖 Generated with Claude Code