An API gateway you can run on a laptop with one command. It sits in front of a small catalogue API and does what a real gateway does: it spreads traffic over three backends, rate limits each client, caches reads, keeps answering when its database is down, records what happened, and hands slow work to a background queue.
It is a system-design project, so the reasoning matters as much as the code. This repository has three documents:
| Read | For |
|---|---|
| README.md (this file) | What it is, what to know first, and how to run it |
| SUMMARY.md | How it works, explained from scratch in plain language |
| DESIGN.md | Why it is built this way: every decision with its pros, cons and alternatives, plus capacity math, failure modes and measurements |
What it does
| Capability | How |
|---|---|
| Load balancing | Nginx round robin over three identical FastAPI instances, with failover and connection caps |
| Rate limiting | Token bucket in Redis, shared by every instance. Three tiers, per-route overrides, Retry-After on every 429 |
| Caching | Read-through Redis cache with tag invalidation on writes and stale copies for outages |
| Resilience | Circuit breaker and a hard timeout on the database. Cached answers survive a Postgres outage |
| Background jobs | Redis Streams queue, two workers, retries, dead-letter stream, backpressure |
| Analytics | Every request becomes a Kafka event, totalled per client per minute in Postgres |
| Auth | Signed JWTs with tiers, plus a client registry |
| Observability | Prometheus metrics and a provisioned Grafana dashboard |
| Demo console | A web page that runs each behaviour on purpose and shows the real responses |
Stack: Python 3.13, FastAPI, Nginx, Redis, PostgreSQL, Kafka, Prometheus, Grafana, k6, pytest, Docker Compose, GitHub Actions.
Measured results (one 16-core laptop with everything sharing it, so compare rows to each other)
| Question | Answer |
|---|---|
| What does one backend carry? | About 400 requests per second cleanly from the database, about 800 with the cache |
| What do three backends carry? | 1,600 uncached with no errors. 3,193 of 3,200 offered with the cache |
| Kill a backend under load | 0 errors in 24,001 requests. The cost was a latency spike of about 1.1 seconds |
| Does the limiter match its policy? | An anonymous client offered 2,000 requests was allowed 118. The policy predicts about 110 |
| Pause Postgres | Cached items are served stale, everything else gets a 503 within a second, and it recovers by itself |
Things worth knowing before you rely on it
- The 100 ms p99 latency target is met at 800 requests per second and missed at 1,600 and above.
DESIGN.mdsays so and lists what is unexplained. - API keys identify a caller but prove nothing, and the demo credentials are public. This is a local demo, not a hardened service.
- Queued jobs do not survive a Redis restart, because Redis runs without persistence here.
- The full list of known weaknesses, each with its fix, is in
DESIGN.md.
| Requirement | Notes |
|---|---|
| Docker Desktop with Compose v2 | Windows, macOS or Linux. Give Docker at least 4 GB of memory (developed with 8 GB) |
| Free ports | 8080 (gateway and console), 3000 (Grafana), 9090 (Prometheus) |
| About 2 GB of disk | For the images. The first start downloads them, which takes a few minutes |
| Python 3.12 or newer | Only for the tests and helper scripts. Not needed to run the gateway |
You do not need to install Redis, Postgres, Kafka or k6. Everything runs in containers, including the load generator.
From the project folder:
cp .env.example .env
docker compose up -d --build --waitOn Windows PowerShell the first line is Copy-Item .env.example .env.
--wait returns only when every service reports healthy. The first run pulls images and builds, so allow a few minutes. Later starts take under a minute.
docker compose ps
curl -i http://127.0.0.1:8080/items/1You should see about 13 containers that are Up, and a 200 OK with these headers:
| Header | Meaning |
|---|---|
X-Instance: api-2 |
Which backend answered. It changes on each request |
X-Cache: MISS |
Not cached yet. The second identical request says HIT |
X-RateLimit-Limit: 10, X-RateLimit-Remaining: 9 |
Your token bucket |
Use 127.0.0.1, not localhost. On some Windows machines Docker Desktop's IPv6 forwarding hangs, and localhost tries IPv6 first.
| Page | URL | What it is for |
|---|---|---|
| Console | http://127.0.0.1:8080/console/ | Five experiments and a live request log. The quickest way to see everything work |
| Grafana dashboard | http://127.0.0.1:3000 | Live throughput, latency, cache hit rate, throttling, backend spread and breaker state. Login admin / admin, or view anonymously |
| API docs | http://127.0.0.1:8080/docs | Interactive Swagger page for the public routes |
Open Dashboards, FlowGate, FlowGate gateway in Grafana. It is empty until traffic arrives. To fill it, set up Python once and run the demo traffic generator:
python -m venv .venv
. .venv/Scripts/activate # macOS or Linux: source .venv/bin/activate
pip install -r requirements-dev.txt
python loadtest/traffic.py 120 # two minutes of mixed trafficOn PowerShell, activate with .venv\Scripts\Activate.ps1.
In the console. Choose who to send as, then pick an experiment. Every request lands in the log at the bottom.
- Rate limiter. Send a burst and watch the bucket drain and refill against the dashed line the policy predicts.
- Cache. Fetch an item twice, update it, fetch again: MISS, HIT, then MISS.
- Backends. Start polling, then run
docker compose stop api-2in a terminal. Its lane goes quiet while the other two carry on. - Job queue. Submit reports and follow them from queued to done.
- Database outage. Run
docker compose pause postgresin a terminal and watch cached answers turn stale.
The console never controls Docker. It shows the commands and you run them, so the page watches a real failure.
From a terminal. Reads are public and writes need a token:
TOKEN=$(curl -s -X POST http://127.0.0.1:8080/auth/token \
-H 'content-type: application/json' \
-d '{"client_id":"demo-pro","client_secret":"pro-secret"}' \
| python -c "import sys,json;print(json.load(sys.stdin)['access_token'])")
curl -i -X PUT http://127.0.0.1:8080/items/1 \
-H "Authorization: Bearer $TOKEN" -H 'content-type: application/json' \
-d '{"name":"renamed","price_cents":100,"stock":5}'PowerShell:
$t = (Invoke-RestMethod -Method Post http://127.0.0.1:8080/auth/token -ContentType application/json -Body '{"client_id":"demo-pro","client_secret":"pro-secret"}').access_token
Invoke-RestMethod -Method Put http://127.0.0.1:8080/items/1 -Headers @{Authorization="Bearer $t"} -ContentType application/json -Body '{"name":"renamed","price_cents":100,"stock":5}'The demo logins are demo-pro / pro-secret and demo-free / free-secret, kept in config/clients.yaml. They are for local use only.
Hit the rate limit by sending 30 anonymous requests at once. Expect a mix of 200 and 429, each 429 carrying Retry-After: 1:
seq 30 | xargs -P 30 -I{} curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:8080/items/1 | sort | uniq -cSubmit a slow report and poll it. The status moves from queued to done in about two seconds:
curl -s -X POST http://127.0.0.1:8080/reports \
-H "Authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"min_price_cents":5000}'
curl -s http://127.0.0.1:8080/jobs/<job_id> -H "Authorization: Bearer $TOKEN"Kill a backend, keep calling, then bring it back:
docker compose stop api-2
for i in 1 2 3 4 5 6; do curl -s http://127.0.0.1:8080/health; echo; done
docker compose start api-2To watch the database failure sequence with timings, run python loadtest/chaos_db.py. It takes a minute or two.
| Goal | Command |
|---|---|
| Stop everything, keep the data | docker compose stop |
| Start it again | docker compose up -d --wait |
| Remove containers, keep the data | docker compose down |
| Remove containers and all data | docker compose down -v |
Change a setting in .env |
Edit it, then docker compose up -d --force-recreate |
| Rebuild after changing code | docker compose up -d --build --wait |
| Follow the logs | docker compose logs -f api-1 |
If you have make, make up, make down, make reset, make test and make lint do the same things. make help lists them.
Set in .env, then recreate the containers.
| Setting | Default | Effect |
|---|---|---|
RATE_LIMIT_ENABLED |
true |
Turn the limiter off, to measure raw capacity |
CACHE_ENABLED |
true |
Turn the cache off |
KAFKA_ENABLED |
true |
Stop sending access events |
BREAKER_ENABLED |
true |
Turn the database circuit breaker off |
JWT_SECRET |
a development value | Signs the tokens. Change it for anything beyond local use |
POSTGRES_USER, POSTGRES_PASSWORD, POSTGRES_DB |
flowgate |
Database credentials |
Per-route rules (who needs a token, which tier, each route's own limit and cache lifetime) are in config/routes.yaml.
| Symptom | Cause and fix |
|---|---|
A request to localhost hangs, or a page will not load |
Use 127.0.0.1. See step 2 |
port is already allocated on start |
Something else uses 8080, 3000 or 9090. Stop it, or change the left side of that ports: entry in docker-compose.yml |
Everything returns 429 immediately |
You are anonymous and the burst is 10. That is the limiter working. Send X-API-Key: anything for a bigger bucket, or use a token |
| A backend keeps restarting right after Docker restarted | It is waiting for Postgres and retries for up to a minute. Check docker compose ps again shortly |
docker compose up --wait times out on Kafka |
Kafka is the slowest to start, especially the first time. Run docker compose ps and retry |
| Docker says the engine is unable to start | Restart Docker Desktop. Volumes survive and services restart by themselves |
| The dashboard shows no data | No traffic yet. Run python loadtest/traffic.py 120 |
Browser console errors mentioning grafana-lokiexplore-app |
A bundled Grafana plugin failing to load. Unrelated and harmless |
Running k6 by hand from Git Bash fails with a C:/Program Files/Git/... path |
Git Bash rewrites the path. Run export MSYS_NO_PATHCONV=1 first, or use the driver below |
python -m pytest -q # 89 unit tests, no services needed
python -m pytest -q -m integration # 8 tests against the running stackThe unit tests use fakeredis (with Lua support) and in-memory stores from tests/fakes.py, so they run anywhere. The integration tests need the stack from step 1. GitHub Actions runs lint, formatting, both suites and a full stack start.
The driver reconfigures the stack between runs, runs k6 inside the compose network, and puts the stack back to its normal settings when it finishes. It needs the Python setup from step 3.
python loadtest/run.py matrix # 1 vs 3 backends, cache off vs on, Kafka overhead (about 15 minutes)
python loadtest/run.py failure # kill a backend mid-test and restart it (about 3 minutes)
python loadtest/run.py ratelimit # check the limiter against its policy (about 3 minutes)
python loadtest/run.py report # print the saved results as tablesClose other heavy programs first, since the numbers depend on how much CPU Docker gets. The results already in loadtest/results/ came from a 16-core Windows machine, and DESIGN.md explains what they show and how noisy they are.
app/ the Python service
api/ the routes
gateway/ the request pipeline: identity, limiter, cache, metrics
data/ Postgres stores, the circuit breaker
background/ Kafka events, analytics, the job queue and worker
console/ the demo page (HTML, CSS, JavaScript modules)
config/ route rules and the demo client list
nginx/ load balancer config
observability/ Prometheus config and the Grafana dashboard
loadtest/ k6 scripts, the test driver, chaos and traffic helpers, saved results
tests/ unit/ and integration/
assets/ screenshots used in this file
SUMMARY.md has a file-by-file map. Continue there for how it works, or in DESIGN.md for why.

