An online judge built from scratch in Kotlin.
Import a Codeforces problem, write C++, Java or Python, and get a verdict from a sandboxed judge.
The source code is private. This repository holds the documentation, architecture and media. Recruiters and reviewers can email me for access. The demo backend runs on a free Render instance: the first request after a quiet period takes about a minute.
- Import any Codeforces problem by its code (
4A,255A,1352C). The backend scrapes the statement, the real time and memory limits and the sample tests. - Write a solution in the Monaco editor (C++17, Java 21 or Python 3).
- Execute. The API stores the submission and answers at once; a judge worker compiles it once, runs every test in a sandbox and records the verdict.
- Read the result: one node per test case (passed, failed, not run), CPU time and peak memory against the problem's limits, and expected vs received output for the failing test.
| Accepted | Wrong answer |
|---|---|
![]() |
![]() |
More in the screenshot gallery.
The judge harness compiles each submission once with a timeout, then runs every test case:
- Limits per run: CPU time, memory (
RLIMIT_AS; Java gets-Xmxbecause the JVM reserves large address ranges), a 16 MB output cap and no core dumps. A wall-clock cap catches programs that sleep or block. - Real metrics: peak memory is the program's own resident set size from
wait4(), not the judge's heap. CPU time decides TLE, so a slow host does not fail a correct program. - Verdicts from the exit status: a warning on stderr is not a runtime error; a non-zero exit or a signal is.
- No pipe deadlocks: stdin, stdout and stderr go through files, so a program printing megabytes never stalls.
Eight verdicts: ACCEPTED, WRONG_ANSWER, TLE, MEMORY_LIMIT_EXCEEDED, OUTPUT_LIMIT_EXCEEDED, RUNTIME_ERROR,
COMPILE_ERROR, INTERNAL_ERROR.
| Mode | Where | Isolation |
|---|---|---|
| Docker | self-hosted (docker compose) |
one container per submission: --network none, read-only filesystem, memory / CPU / PID limits, every capability dropped except the three the harness needs, no-new-privileges; the program runs as an unprivileged user that cannot touch the compiled binary |
| Process | the free live demo on Render (no Docker available) | the same per-run limits as a child process with an empty environment; database and Redis secrets are moved out of the environment before any program runs |
PostgreSQL is the source of truth and Redis only wakes the workers up:
- Workers claim the oldest queued submission with
FOR UPDATE SKIP LOCKED, so two workers never take the same one. - While judging, a worker refreshes a lease (
claimed_at) every 10 seconds. If it or the whole service dies, a reaper puts the submission back in the queue after 60 seconds, and fails it after two attempts so a submission that crashes the judge cannot loop forever. - A lost Redis message, or Redis being down, costs at most one polling interval.
- Every error ends in a verdict, never in a row stuck in
RUNNING.
Per-client rate limits on submissions and imports, a cap on queue depth, a 64 KB source limit, a language check and
404 for unknown submissions. APP_ROLE=api|worker|all and a configurable worker count let the API and the workers
run on separate machines.
- 42 sandbox tests across both modes: infinite loops, sleeps, C++/Java/Python memory hogs, 10 MB of output, endless output, segfaults, compile-time bombs, a fork bomb, network access and attempts to overwrite the compiled program or read secrets.
- 13 end-to-end API checks per mode, plus crash recovery (killing the service mid-judge) and running with Redis down.
flowchart LR
Client["Web (Wasm) · Android · Desktop"] -->|REST| API[Ktor API]
API -->|INSERT QUEUED| DB[(PostgreSQL)]
API -.->|LPUSH wake-up| Redis[(Redis)]
Redis -.->|BRPOP| Worker[Judge workers]
Worker -->|claim: FOR UPDATE SKIP LOCKED| DB
Worker -->|job| Sandbox["Sandbox harness<br/>Docker or process mode"]
Sandbox -->|per-test results| Worker
Worker -->|verdict + metrics| DB
Client -->|poll GET /submission/id| API
| Document | What it covers |
|---|---|
| Architecture | components, the life of a submission, failure handling |
| Engineering decisions | why each design choice, and what it costs |
| API | the 8 endpoints with examples |
| Database | schema, the lease columns, migrations |
| Project structure | Gradle modules and packages |
| Setup and deployment | local run, Docker mode, the free Render + Neon + Upstash + Netlify setup |
| Layer | Technologies |
|---|---|
| Client | Kotlin Multiplatform, Compose Multiplatform (Web via WebAssembly, Android, Desktop), Monaco editor, Ktor client |
| API | Kotlin, Ktor (Netty), kotlinx.serialization, rate limiting, Jsoup (Codeforces import) |
| Judge | Python harness with POSIX resource limits; g++ (C++17), OpenJDK 21, Python 3.12 |
| Data | PostgreSQL (Neon), Exposed, HikariCP, Redis (Upstash) via Lettuce |
| Deployment | Docker, Render (backend, free), Netlify (web), docker compose for the full Docker sandbox |
Harshvardhan Singh, B.Tech CSE at IIIT Bhopal · GitHub · LinkedIn · hvsr29march2004@gmail.com



