Summary
claude-code-proxy currently exposes the Anthropic-compatible data-plane endpoints without validating any inbound client/proxy token. The same handler then selects the operator's server-side OpenAI/Gemini/Anthropic API key from environment variables and sends the request upstream.
That is fine when the service is strictly loopback-only, but the current server/Docker/README examples bind or publish 0.0.0.0:8082. If a user follows those examples on a VM, NAS, Docker host, or cloud box, anyone who can reach the port can spend the operator's upstream LLM keys through /v1/messages.
Relevant code
The only global middleware logs the request and then calls the next handler; it does not authenticate the caller:
# server.py
401 @app.middleware("http")
402 async def log_requests(request: Request, call_next):
403 # Get request details
404 method = request.method
405 path = request.url.path
407 # Log only basic request details at debug level
408 logger.debug(f"Request: {method} {path}")
411 response = await call_next(request)
413 return response
The main message endpoint is public from the proxy's point of view:
# server.py
1214 @app.post("/v1/messages")
1215 async def create_message(request: MessagesRequest, raw_request: Request):
Inside that handler, the proxy chooses server-side provider keys:
# server.py
1243 # Determine which API key to use based on the model
1244 if request.model.startswith("openai/"):
1245 litellm_request["api_key"] = OPENAI_API_KEY
...
1263 litellm_request["api_key"] = GEMINI_API_KEY
...
1266 litellm_request["api_key"] = ANTHROPIC_API_KEY
The token counting endpoint is also exposed without an inbound proxy token:
# server.py
1567 @app.post("/v1/messages/count_tokens")
1568 async def count_tokens(request: TokenCountRequest, raw_request: Request):
The default direct run path binds all interfaces:
# server.py
1707 print("Run with: uvicorn server:app --reload --host 0.0.0.0 --port 8082")
1711 uvicorn.run(app, host="0.0.0.0", port=8082, log_level="error")
The Docker image and README publish the same port on all host interfaces:
# Dockerfile
14 # Start the proxy
15 EXPOSE 8082
16 CMD uv run uvicorn server:app --host 0.0.0.0 --port 8082 --reload
# README.md
72 services:
74 proxy:
75 image: ghcr.io/1rgs/claude-code-proxy:latest
77 env_file: .env
78 ports:
79 - 8082:8082
# README.md
84 docker run -d --env-file .env -p 8082:8082 ghcr.io/1rgs/claude-code-proxy:latest
Why this matters
For an LLM proxy, there are two different credentials:
- upstream provider credentials, owned by the proxy operator;
- inbound proxy credentials, used to decide which clients may call the proxy.
Right now the first exists, but the second appears to be absent. When the listener is exposed beyond localhost, the proxy can become an unauthenticated relay backed by the operator's provider accounts. That can cause unexpected LLM spend, quota exhaustion, account abuse, and confusing upstream rate-limit/auth failures for the legitimate user.
Suggested fix
Consider adding fail-closed inbound authentication for data-plane endpoints:
- add a
PROXY_API_KEY or ANTHROPIC_AUTH_TOKEN setting and require Authorization: Bearer <token> or x-api-key: <token> on /v1/messages and /v1/messages/count_tokens;
- reject startup when binding a non-loopback host without an inbound proxy token, or print a very prominent warning and require an explicit opt-in;
- change Docker examples to bind loopback by default, for example
127.0.0.1:8082:8082, unless a proxy token is configured;
- add regression tests for missing, wrong, and correct inbound proxy tokens.
Happy to clarify or test a patch if useful.
Summary
claude-code-proxycurrently exposes the Anthropic-compatible data-plane endpoints without validating any inbound client/proxy token. The same handler then selects the operator's server-side OpenAI/Gemini/Anthropic API key from environment variables and sends the request upstream.That is fine when the service is strictly loopback-only, but the current server/Docker/README examples bind or publish
0.0.0.0:8082. If a user follows those examples on a VM, NAS, Docker host, or cloud box, anyone who can reach the port can spend the operator's upstream LLM keys through/v1/messages.Relevant code
The only global middleware logs the request and then calls the next handler; it does not authenticate the caller:
The main message endpoint is public from the proxy's point of view:
Inside that handler, the proxy chooses server-side provider keys:
The token counting endpoint is also exposed without an inbound proxy token:
The default direct run path binds all interfaces:
The Docker image and README publish the same port on all host interfaces:
# Dockerfile 14 # Start the proxy 15 EXPOSE 8082 16 CMD uv run uvicorn server:app --host 0.0.0.0 --port 8082 --reload# README.md 84 docker run -d --env-file .env -p 8082:8082 ghcr.io/1rgs/claude-code-proxy:latestWhy this matters
For an LLM proxy, there are two different credentials:
Right now the first exists, but the second appears to be absent. When the listener is exposed beyond localhost, the proxy can become an unauthenticated relay backed by the operator's provider accounts. That can cause unexpected LLM spend, quota exhaustion, account abuse, and confusing upstream rate-limit/auth failures for the legitimate user.
Suggested fix
Consider adding fail-closed inbound authentication for data-plane endpoints:
PROXY_API_KEYorANTHROPIC_AUTH_TOKENsetting and requireAuthorization: Bearer <token>orx-api-key: <token>on/v1/messagesand/v1/messages/count_tokens;127.0.0.1:8082:8082, unless a proxy token is configured;Happy to clarify or test a patch if useful.