Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 20 additions & 0 deletions docs/examples/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,26 @@ the problem you're trying to solve.
Survive a simulated mid-pipeline crash and resume from the saved
checkpoint, then re-resume under an upgraded state schema with
a v1->v2 migration backfilling new fields.
- [**Structured-output reask**](structured-output-reask.md). Extract
a mission record under an output-token ceiling too tight for a
complete answer, then recover by correcting the model and raising
the ceiling on the retry. Three modes show why one without the
other does not recover.

### Retrieval

- [**Retrieval-augmented answering**](retrieval-rag.md). Answer a
question from a small corpus using embed-then-rerank before
generation: cosine similarity for recall over the whole corpus,
a cross-encoder for precision over the shortlist.

### Providers

- [**Provider extras**](provider-extras.md). Reach two OpenAI
request fields openarmature does not model, and see the three
things the `extras` container refuses. Runs with no credentials
against a stub transport, because the outbound request body is
the subject.

### Observability

Expand Down
132 changes: 132 additions & 0 deletions docs/examples/provider-extras.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,132 @@
# Provider extras

!!! info "Source"
[https://github.com/LunarCommand/openarmature-python/blob/main/examples/provider-extras/main.py](https://github.com/LunarCommand/openarmature-python/blob/main/examples/provider-extras/main.py){target="_blank" rel="noopener"}

Reach a vendor knob openarmature does not model, and watch the
guardrails that stop you reaching the wrong one.

## Overview

You classify lunar telemetry alerts in bulk. Two of the things you
want are provider-specific rather than portable, so there is no
first-class field for either: `service_tier` to take the cheaper,
slower lane for a batch nobody is waiting on, and `logit_bias` to
stop the model emitting a severity label your team retired last
quarter. Both are real OpenAI request fields. Neither means
anything on another provider.

`extras` is where those go. It is a named container on the runtime
config, and whatever you put in it rides to the wire untouched.
That is the whole feature, and it exists so a provider-specific
knob does not require either a fork or a framework release.

The interesting half is what it refuses. A field openarmature
already models is managed, and putting it in `extras` as well is an
error rather than an override, because two sources of truth for one
wire field is a bug you want at the call site and not in a trace
three days later.

## What it teaches

- `RuntimeConfig(extras={...})` carrying anything the framework does
not model. `service_tier` and `logit_bias` arrive on the request
body verbatim.
- **A key naming a field this call produced is rejected.** Setting
`temperature=0.2` and also `extras={"temperature": 0.9}` raises
`ProviderInvalidRequest` naming the key and both values.
- **Managed means "produced by this call", not "nameable".** A
sampling field you leave unset is not managed on that call, so
`extras={"temperature": 0.9}` with no `temperature=` rides
through. This is the one that surprises people, and it is what
makes `extras` usable as an escape hatch for a field the framework
models but you did not set.
- **A matching value is a no-op, not an error.** Sending `0.2` in
both places is not ambiguous, so it is allowed.
- **Structural keys are managed unconditionally.** `model`,
`messages`, `tools` and `tool_choice` are managed whether or not
the mapping produced the field, so a *conflicting* value rejects
even on a call whose body carries nothing of that name. A matching
value is still a no-op, by the same rule as above. The asymmetry is
deliberate: it is what stops an `extras` tool array from reaching
the wire on a no-tools call without passing tool validation.
- **`stop` merges instead of colliding**, because it realizes the
same wire field as `stop_sequences`. Both lists arrive,
concatenated and de-duplicated.

## How to run

```bash
uv run python examples/provider-extras/main.py
```

**No credentials, no endpoint.** The demo installs a stub transport
and prints the outbound request body, because the shape of that body
*is* the subject: whether a knob reached the wire, and what happened
when it collided with one the framework manages. A real endpoint
answers neither question any better, and every refusal happens
before a request is sent.

To watch a real provider accept the knobs, drop the `transport=`
argument from `_provider()` and supply a real `base_url` and
`api_key`.

## The graph

```mermaid
flowchart TD
start([start])
classify[classify]
show_guardrails[show_guardrails]
stop([end])

start --> classify --> show_guardrails --> stop
```

`classify` runs the alerts with the two passthrough knobs set.
`show_guardrails` then attempts the collisions and records what
came back, so the accepted and refused cases print side by side from
one run.

## Reading the output

```
=== openarmature provider-extras demo ===
alerts: 3

classified:
[watch] Regolith intake auger current 18% above nominal for 40 seconds, ...
[watch] South-pole relay lost carrier for 3 frames during Earth occultat...
[watch] Battery bus B cell 4 reading 0.2V under its siblings at end of c...

the knobs that reached the wire:
service_tier = 'flex'
logit_bias = {'24886': -100}
temperature = 0.0

what extras refuses:
temperature in both: refused, extras key 'temperature' conflicts with the
mapping-managed wire field 'temperature' (managed value 0.2, extras value
0.9); a managed field cannot be overridden via extras
tools via extras: refused, extras key 'tools' conflicts with the
mapping-managed wire field 'tools' (managed value None, extras value
<list of 1>); a managed field cannot be overridden via extras
stop merges rather than collides: accepted
```

Three things in that output are the point.

- **The wire block** is read off the body the stub captured, not
from the config. `service_tier` and `logit_bias` are there because
nothing manages them; `temperature` is there because the mapping
produced it.
- **`managed value None`** on the `tools` refusal is the structural
rule showing its work. The call declared no tools, so the managed
value is `None`, and a list conflicts with that. Every other entry
in the table is unmanaged on a call that did not produce it, so the
same extras key would have ridden through.
- **`stop` accepted** is the merge arm. It is the only managed field
in the OpenAI mapping that combines rather than collides.

The error messages name the key and both values, so the fix is
readable from the message without reaching for the mapping table.
143 changes: 143 additions & 0 deletions docs/examples/retrieval-rag.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,143 @@
# Retrieval-augmented answering

!!! info "Source"
[https://github.com/LunarCommand/openarmature-python/blob/main/examples/retrieval-rag/main.py](https://github.com/LunarCommand/openarmature-python/blob/main/examples/retrieval-rag/main.py){target="_blank" rel="noopener"}

Answer a question about the Moon from a small corpus of passages,
using the two-stage retrieval pattern before generation.

## Overview

You have eight passages of lunar reference material and a question.
Stuffing all eight into the prompt would work at this size and stop
working at a thousand, so the pipeline narrows instead: find the
plausibly-relevant passages cheaply, reorder them accurately, and
ground the answer in the few that survive.

Four steps, and the middle two are the interesting ones.

1. **Index**, once and offline. Batch-embed every passage. `embed`
over a list returns one vector per input in input order, so the
index lines up positionally with the corpus and no separate id
map is needed.
2. **Retrieve**, per query. Embed the question, rank the corpus by
cosine similarity, keep the top four. Cheap and broad: it buys
recall, not precision.
3. **Rerank** those four with a cross-encoder, which scores each
candidate against the query directly rather than comparing two
independently-computed vectors. More accurate and more
expensive, which is why it runs over four candidates instead of
the whole corpus. Two survive.
4. **Generate** from the two reranked passages.

Retrieval gives recall, reranking gives precision, and the split is
what makes the cost work: the expensive comparison runs over a
shortlist the cheap one produced.

## What it teaches

- `OpenAIEmbeddingProvider` from `openarmature.retrieval`. Batch
`embed` for the index, single `embed` for the query. One vector
per input, in input order, both times.
- `EmbeddingRuntimeConfig(input_type=...)`, the query-versus-document
knob. On OpenAI it is a wire no-op because the model is symmetric,
and it is set anyway: the same call selects the correct
representation on an asymmetric provider, so the pipeline moves to
TEI, Cohere or Jina without a code change. Setting a field that
does nothing today is what keeps it portable.
- `CohereRerankProvider.rerank`, returning `ScoredDocument` results
sorted by relevance.
- **Mapping results back by index, not by text.** Cohere does not
echo the document body, so `ScoredDocument.document` is `None` and
the only way home is `ScoredDocument.index`, which points into the
candidate list you passed. The example translates that back to a
corpus index. A pipeline that matched on returned text would work
against a provider that echoes and break against one that does not.
- `RerankRuntimeConfig(return_documents=True)` asking for the echo
where a provider supports it.
- **Retrieval providers driven inside node bodies**, so their
`EmbeddingEvent` and `RerankEvent` reach an attached observer the
same way an LLM completion does. The offline index build runs
outside the graph deliberately, and emits nothing: there is no
invocation to attribute it to.
- An `OpenAIProvider` answer node grounded in the reranked passages.

## How to run

```bash
uv sync --group examples
OPENAI_API_KEY=sk-... COHERE_API_KEY=... \
uv run python examples/retrieval-rag/main.py
```

Both keys are required: OpenAI serves the embeddings and the answer,
Cohere serves the rerank.

| Variable | Default | Notes |
| --- | --- | --- |
| `OPENAI_API_KEY` | required | embeddings and the answer |
| `OPENAI_BASE_URL` | `https://api.openai.com` | host root, no `/v1` |
| `OPENAI_EMBED_MODEL` | `text-embedding-3-small` | |
| `OPENAI_CHAT_MODEL` | `gpt-4o-mini` | |
| `COHERE_API_KEY` | required | rerank |
| `COHERE_RERANK_MODEL` | `rerank-v3.5` | |

`OPENAI_BASE_URL` takes the host root. The provider appends the
`/v1` routes itself, so a URL that already ends in `/v1` is
rejected rather than producing a doubled path.

## The graph

```mermaid
flowchart TD
start([start])
retrieve[retrieve]
rerank[rerank]
answer[answer]
stop([end])

start --> retrieve --> rerank --> answer --> stop
```

Linear, because each stage narrows the input to the next. The index
build is not a node: it runs once before the first `invoke`, outside
any graph.

State carries corpus indices rather than passage text between
stages. `candidate_indices` after retrieval, `ranked_indices` after
reranking, both best-first, the second a reordered and trimmed
subset of the first.

## Reading the output

```
indexed 8 passages

[obs] embed retrieve: 1 in / 1536d
[obs] rerank rerank: 4 in / top 2
Q: Why did Apollo 13 not land on the Moon?
retrieved: [3, 0, 6, 1]
reranked: [3, 6]
A: <a grounded one-paragraph answer>
```

The indices are what to watch, and the two lines together show the
rerank doing its job.

- **`indexed 8 passages`** is the offline build. It prints no
observer line, because it ran outside the graph.
- **`[obs] embed retrieve: 1 in / 1536d`** is the per-query embed,
inside the node, so it reaches the observer. One input, and the
dimensionality the model returned.
- **`[obs] rerank rerank: 4 in / top 2`** is `_RETRIEVE_K` in and
`_RERANK_K` out.
- **`retrieved`** is cosine order, best-first. **`reranked`** is a
subset in a different order. If the reranked list is the first
two of the retrieved list unchanged, the cross-encoder agreed with
cosine on that query, which happens and is not a failure. The
interesting case is the one above, where a passage cosine ranked
third is promoted over the two above it.

The specific indices and the answer text depend on the models you
point at, so treat the numbers as shape rather than as expected
output.
Loading
Loading