Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 17 additions & 0 deletions .changeset/example-gateway.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
---
"resilix": patch
---

Adds a runnable example: `pnpm example:gateway`.

A simulated LLM provider degrades from ~140ms to seconds **at a flat error rate**, starts
returning `429`s, then recovers, while ~25 requests/second flow through a full pipeline across
three tenants. It prints a per-second table so the adaptation is visible rather than described:
between 7s and 19s latency rises **23× while failures stay at 2**, and the limiter walks
concurrency from 25 down to 5 and sheds the excess before the timeouts start.

Covers policy ordering, the verdict model, per-model isolation keys, tenant fairness, criticality
shedding, a shared retry budget, `ctx.mark()` for time-to-first-token, and `RejectedError.reason`.

It is type-checked and run in CI (nine seconds at 4× speed) — an example that does not compile is
worse than no example, and it is the first code anyone reads.
5 changes: 5 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -27,5 +27,10 @@ jobs:
- run: pnpm test:paths
- run: pnpm test:perf
- run: pnpm test:compat

# The example is the first code most people read, so a broken one is worse
# than none. Nine seconds at 4x speed traverses all four provider phases —
# enough to prove it runs, not enough to slow the build.
- run: RESILIX_EXAMPLE_MS=9000 RESILIX_EXAMPLE_SPEED=4 pnpm example:gateway
- run: pnpm build
- run: pnpm check:package
11 changes: 11 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,17 @@ because customers submitted bad input. resilix classifies outcomes into verdicts

<!-- #endregion why -->

## See it work

```bash
pnpm example:gateway
```

A simulated provider degrades from ~140ms to seconds **at a flat error rate**, then recovers.
Watch the limiter walk concurrency down before any failures appear —
[examples/llm-gateway](examples/llm-gateway), written up at
[resilix.js.org/guide/example](https://resilix.js.org/guide/example).

<!-- #region quickstart -->
## Quick start

Expand Down
2 changes: 1 addition & 1 deletion biome.json
Original file line number Diff line number Diff line change
Expand Up @@ -52,7 +52,7 @@
},
"overrides": [
{
"include": ["src/**/*.test.ts", "scripts/**"],
"include": ["src/**/*.test.ts", "scripts/**", "examples/**"],
"linter": {
"rules": {
"style": {
Expand Down
1 change: 1 addition & 0 deletions docs/.vitepress/config.ts
Original file line number Diff line number Diff line change
Expand Up @@ -134,6 +134,7 @@ export default defineConfig({
items: [
{ text: "What resilix is", link: "/guide/" },
{ text: "Getting started", link: "/guide/getting-started" },
{ text: "A worked example", link: "/guide/example" },
{ text: "The verdict model", link: "/guide/verdicts" },
],
},
Expand Down
47 changes: 47 additions & 0 deletions docs/guide/example.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,47 @@
---
description: "A runnable LLM gateway example: watch the adaptive limiter walk concurrency down as an upstream degrades at a flat error rate, then recover."
---

# A worked example

Prose about adaptive concurrency limiting is hard to believe. This is the same thing as a program
you can run:

```bash
git clone https://github.com/lintdeveloper/resilix
cd resilix && pnpm install
pnpm example:gateway
```

A simulated provider degrades from ~140ms to seconds **at a flat error rate**, starts returning
`429`s, then recovers. Around 25 requests/second flow through a full pipeline across three tenants.

## The output

```
t phase limit inflight p90ms | ok 4xx shed fail
7s healthy 25 0 143 | 148 27 0 0
14s degrading 23 56 1186 | 190 34 68 2
19s degrading 21 42 3358 | 203 38 190 2
23s overloaded 5 36 7880 | 215 40 272 12
```

Between 7s and 19s, **latency rises 23× while failures stay at 2**. A failure-rate circuit breaker
sees nothing in that window — it is watching errors, and there are none. The limiter is watching
latency, so it walks concurrency from 25 down to 5 and sheds the excess before the timeouts start.

The `4xx` column is the [verdict model](./verdicts) doing its job: fifty healthy rejections from a
validating upstream, none of which counted against the breaker.

## What each file shows

| File | Use case |
|---|---|
| `gateway.ts` | policy ordering, verdicts, isolation key, tenant fairness, criticality, a shared retry budget |
| `run.ts` | `ctx.mark()` for time-to-first-token, and `RejectedError.reason` |
| `upstream.ts` | why concurrency is the right lever — latency there rises *with* concurrency |

Source: [examples/llm-gateway](https://github.com/lintdeveloper/resilix/tree/main/examples/llm-gateway)

It is a simulation with a seeded PRNG, tuned so the transitions are visible in under a minute —
the *shape* of the behaviour, not numbers to quote.
51 changes: 51 additions & 0 deletions examples/llm-gateway/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
# Example — an LLM gateway

```bash
pnpm example:gateway
```

A simulated provider that **degrades from ~140ms to seconds at a flat error rate**, starts pushing
back with `429`s, then recovers. Roughly 25 requests/second are offered through a resilix pipeline
across three tenants, one in five of them background work.

It runs for 44 seconds. `RESILIX_EXAMPLE_MS` and `RESILIX_EXAMPLE_SPEED` compress it — CI runs it
for nine seconds at 4× purely to prove it still works.

## What to watch

The **`limit`** column against the **`p90ms`** and **`fail`** columns:

```
t phase limit inflight p90ms | ok 4xx shed fail
7s healthy 25 0 143 | 148 27 0 0
14s degrading 23 56 1186 | 190 34 68 2 ← 8x slower, still 2 failures
19s degrading 21 42 3358 | 203 38 190 2 ← 23x slower, still 2 failures
23s overloaded 5 36 7880 | 215 40 272 12
```

Between 7s and 19s latency rises **23×** while failures stay at **2**. That is the incident this
library was built for, and a failure-rate circuit breaker sees nothing in that window — it is
watching errors, and there are none. The limiter is watching latency, so it walks concurrency down
from 25 to 21 to 5 and sheds the excess *before* the timeouts start.

The **`4xx` column** is the other half. Fifty of those arrive across the run — a validating
upstream rejecting bad prompts. Every one is `answered`: healthy, never counted against the
breaker. A library whose failure predicate is *"did the promise reject?"* opens the circuit on
them, which is a self-inflicted outage on a working provider.

## Which use case is where

| File | What it demonstrates |
|---|---|
| `gateway.ts` | policy **ordering**, and why cheapest-refusal-first is not arbitrary |
| `gateway.ts` | **verdicts** — `classifyHttp`, so a `422` is `answered` and a `429` is `overload` |
| `gateway.ts` | **isolation key** per model, **tenant** fairness, **priority** so background sheds first |
| `gateway.ts` | a **shared retry budget** — one instance for the process, not one per pipeline |
| `run.ts` | **`ctx.mark()`** — latency is time to first token, not time to drain |
| `run.ts` | **`RejectedError.reason`**, so "why was I refused?" always has an answer |
| `upstream.ts` | why **concurrency** is the right lever: latency here rises with concurrency |

## What it is not

Not a benchmark. The provider is a simulation with a seeded PRNG, tuned so the transitions are
visible in under a minute. It shows the *shape* of the behaviour, not numbers you should quote.
74 changes: 74 additions & 0 deletions examples/llm-gateway/gateway.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
/**
* The gateway: one pipeline, every policy resilix ships, wired the way you would
* actually wire them.
*
* Order matters and is not arbitrary. Cheapest and most decisive refusals go
* first, so an obviously-doomed call never occupies a slot it would only have to
* release:
*
* throttler → the upstream is refusing most of what we send; shed a fraction
* breaker → the upstream looks wholly down; fail fast
* limiter → it is up, but this is more concurrency than it can absorb
*/
import {
type Priority,
breaker,
budget,
classifyHttp,
limiter,
pipeline,
throttler,
} from "../../src/index.ts";
import type { Reply } from "./upstream.ts";

export interface Job {
/** Which model — the isolation key. One bad model must not shed the others. */
model: string;
/** Which customer, for fairness under pressure. */
tenant: string;
/** Background work is shed before anything a user is waiting on. */
background: boolean;
}

/**
* ONE budget for the whole process. A per-pipeline cap cannot bound system-wide
* retry amplification, which is the entire point of having one.
*/
export const retryBudget = budget({ ratio: 0.1 });

export const gateway = pipeline<Job>({
key: (job) => job.model,
tenant: (job) => job.tenant,
priority: (job): Priority => (job.background ? "bulk" : "critical"),

// A Response is a value, not a throw, so the classifier must see it either
// way. This is what makes a 422 `answered` rather than a failure.
classify: classifyHttp,

policies: [
throttler(),
breaker({
// ~3x the healthy p95. No default exists for this on purpose: "slow" is
// meaningless without your own baseline, and a wrong guess is worse than
// a required argument.
slowCallMs: 600,
slowCallRate: 0.5,
window: { calls: 60, minCalls: 8, maxAgeMs: 20_000 },
openForMs: 4_000,
consecutiveBackstop: 12,
}),
limiter({ initialLimit: 16, minLimit: 4 }),
],

// Bounds the WHOLE sequence, not each attempt. Most libraries bound each
// attempt, so a caller asking for 12s can wait maxAttempts x (12s + backoff).
timeoutMs: 12_000,
retry: { maxAttempts: 3, jitter: "full", budget: retryBudget },
});

/** Shape a provider reply the way the classifier expects to see it. */
export const toResponse = (reply: Reply): Response =>
new Response(null, {
status: reply.status,
headers: reply.retryAfterS ? { "retry-after": String(reply.retryAfterS) } : undefined,
});
125 changes: 125 additions & 0 deletions examples/llm-gateway/run.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,125 @@
/**
* Drive traffic through the gateway and print what the policies are doing.
*
* pnpm example:gateway
*
* Watch the `limit` column. The provider degrades from ~120ms to seconds at a
* flat error rate; the limiter sees latency rise and walks the concurrency down
* before the errors start, and walks it back up on recovery. That adaptation is
* the thing that has no equivalent in npm, and it is hard to believe from prose.
*/
import { RejectedError, type RejectionReason } from "../../src/index.ts";
import { type Job, gateway, toResponse } from "./gateway.ts";
import { FakeProvider, sleep } from "./upstream.ts";

const MODEL = "gpt-oss-120b";
const TENANTS = ["acme", "globex", "initech"] as const;
/** Default long enough to show degradation AND recovery; CI overrides it. */
const RUN_MS = Number(process.env.RESILIX_EXAMPLE_MS ?? "44000");

const provider = new FakeProvider();
const tally = { ok: 0, answered: 0, shed: 0, failed: 0 };
const shedBy = new Map<RejectionReason, number>();

async function once(job: Job): Promise<void> {
try {
const reply = await gateway.execute(job, async (ctx) => {
const r = await provider.call();
// Time to first token, not time to drain. Judge a stream end-to-end and a
// healthy 45-second completion looks like saturation.
ctx.mark();
return toResponse(r);
});
if (reply.status === 200) tally.ok++;
else if (reply.status < 500) tally.answered++;
else tally.failed++;
} catch (error) {
if (error instanceof RejectedError) {
tally.shed++;
shedBy.set(error.reason, (shedBy.get(error.reason) ?? 0) + 1);
} else {
tally.failed++;
}
}
}

function header(): void {
console.log(
"\n t phase limit inflight p90ms | ok 4xx shed fail | shed reason",
);
console.log(` ${"─".repeat(88)}`);
}

function row(second: number): void {
const m = gateway.metrics().find((x) => x.policy === "limiter")?.values ?? {};
const top = [...shedBy.entries()].sort((a, b) => b[1] - a[1])[0];
const cell = (v: number, w: number) => String(Math.round(v)).padStart(w);
const cols = [
` ${String(second).padStart(2)}s ${provider.phase().padEnd(11)}`,
cell(m.limit ?? 0, 6),
cell(m.inFlight ?? 0, 10),
cell(m.recentMs ?? 0, 8),
" |",
cell(tally.ok, 6),
cell(tally.answered, 5),
cell(tally.shed, 6),
cell(tally.failed, 6),
` | ${top ? `${top[0]} x${top[1]}` : "—"}`,
];
console.log(cols.join(""));
}

async function main(): Promise<void> {
console.log(" resilix — LLM gateway example");
console.log(" A provider that degrades at a FLAT error rate, then recovers.");
console.log(" Watch `limit` fall before the failures start, and rise again after.");
header();

const started = Date.now();
const inFlight = new Set<Promise<void>>();
let tick = 0;

const printer = setInterval(() => row(++tick), 1_000);

while (Date.now() - started < RUN_MS) {
// ~25 requests/sec offered load, mixed tenants, one in five is background.
for (let i = 0; i < 5; i++) {
const job: Job = {
model: MODEL,
tenant: TENANTS[Math.floor(Math.random() * TENANTS.length)] ?? "acme",
background: Math.random() < 0.2,
};
const p = once(job).finally(() => inFlight.delete(p));
inFlight.add(p);
}
await sleep(200);
}

clearInterval(printer);
await Promise.allSettled([...inFlight]);

const total = tally.ok + tally.answered + tally.shed + tally.failed;
console.log(` ${"─".repeat(88)}`);
console.log(`\n ${total} requests offered\n`);
console.log(` ${String(tally.ok).padStart(5)} succeeded`);
console.log(
` ${String(tally.answered).padStart(5)} answered 4xx — healthy, never opened the circuit`,
);
console.log(
` ${String(tally.shed).padStart(5)} shed by resilix, fast, without touching the provider`,
);
console.log(` ${String(tally.failed).padStart(5)} failed`);
console.log("\n shed by reason:");
for (const [reason, n] of [...shedBy.entries()].sort((a, b) => b[1] - a[1])) {
console.log(` ${String(n).padStart(5)} ${reason}`);
}
console.log(`
The 4xx column is the point of the verdict model: a boolean
'did it reject?' breaker would have opened the circuit on those.
`);
}

main().catch((error: unknown) => {
console.error(error);
process.exitCode = 1;
});
Loading
Loading