Skip to content

Repository files navigation

University of Alberta Logo

Aleph — The Science Inference Cluster

License: MIT Kubernetes GPU Scheduling Serving Models Docker Hub

Over 100 model deployments — one endpoint, one key, from protein folds to LLMs.

One point. Every model. Infinite unity.

Deployed on the University of Alberta / AMII Vulcan environment for multi-model GPU inference

Maintained by: Rahim Khoja (khoja1@ualberta.ca) and Karim Ali (kali2@ualberta.ca)

Get started · Models · Architecture · Documentation

Live catalog · Browser chat · API access guide


Overview

Aleph is a local inference server. It serves AI models much like a web server serves content: send a request over HTTP and receive text, an image, a prediction, or another model-specific result. The hosted models run on hardware we operate on Vulcan.

Compatible tools use the same API formats they already support: change the server address, supply an Aleph API key, and name the model. A notebook, research pipeline, browser application, or agent can use the service without managing its own model weights, software environment, or GPU allocation.

Researchers share running model servers, avoiding repeated setup and loading for each workflow. The model stays loaded while in use; idle models can release their GPU resources while retaining their weights on persistent storage. The researcher chooses the right model and evaluates its results; Aleph handles serving it.

Capabilities

  • Science and language models — protein structure, genomics, materials, weather, medical imaging, and other scientific tasks alongside chat, vision, and audio.
  • Compatible APIs — OpenAI-style chat and embeddings, Anthropic-style messages, and model-specific science endpoints. Check each card for supported inputs and features.
  • Flexible runtimes — KServe manages deployment; vLLM, Text Embeddings Inference, NVIDIA NIM, or custom servers execute the models. A new API may need a gateway adapter; adding a supported model needs its deployment and card.
  • GPU sharing and scaling — small models can share a GPU through HAMi, while larger models can request multiple devices. Adding workers expands the pool; starting more model copies shares demand within that pool. Both need capacity.
  • Always-on or on-demand — selected models keep a running copy; others scale to zero when idle. A wake-up can return 503 with retry guidance, or a capacity refusal when the gateway cannot find a suitable placement.
  • Authentication and accounting — Tyk handles API keys and rate limits; the gateway records caller identity, token counts, latency, status, and allocated resources, and exposes aggregate metrics. See Logging and metrics for examples, content exclusions, and retention.
  • Persistent weights — NFS-backed storage allows models to reuse downloaded weights across pod and node replacement.

Get started

I want to… Start here
Find a model and see examples Live model catalog
Chat in a browser Open WebUI
Get API access Alliance Aleph guide
Connect a notebook, SDK, or agent Endpoints and client configuration
Deploy Aleph Deployment quickstart
Add a model Model deployment and validation

Choose a model from the catalog, use its supported API, and supply your Aleph key. Routing follows the model name you request. Each model card describes its inputs, examples, and limitations.

The hosted service is a proof of concept with shared capacity and no SLA.

Model catalog

Aleph hosts over 100 model deployments across scientific and language domains:

Domain Examples
Protein / Structural biology AlphaFold2, Boltz-2, ESMFold, ESM2, ESM-C 300M, ProstT5, LigandMPNN, DiffDock, SaProt
Genomics / DNA / RNA Nucleotide Transformer, DNABERT-2, GENA-LM, Borzoi, Enformer, Caduceus, RNAbert
Materials / Chemistry MACE-MH-1, MACE-MP, CHGNet, ChemBERTa, MatterSim, CrystalLLM, ChemGPT
Weather / Climate Aurora, GraphCast, FourCastNet3, Pangu-Weather, NeuralGCM, ClimaX, FengWu
Astronomy AstroCLIP, AstroPT, AstroSage, Zoobot
Medical / Imaging MedGemma, BiomedCLIP, TotalSegmentator, MedSAM, ClinicalBERT
Vision / 3D FLUX.1, Kandinsky 3, DUSt3R, MASt3R, YOLOv8, Mask R-CNN, Depth Anything
Time-series / Audio Chronos-Bolt, TimesFM, TTM, XTTS-v2, BirdNET, CLAP
Language models Gemma 3/4, Qwen 3/3.5/3.6, GLM-4/Z1, GPT OSS 20B/120B, DeepSeek R1, Command-R
Science NLP SciBERT, BioGPT, SciNCL, SpecTer2, OceanGPT, GeoGalactica, OpenBioLLM

A repository entry does not guarantee that the model is currently served. Use the live catalog for availability and each model's README for its test status and limitations. Authenticated GET /v1/models lists chat models; add ?all=true for the full catalog. To contribute a deployment, follow Add a model.

Architecture

Aleph has two main layers: a gateway that accepts and routes requests, and a Kubernetes serving platform that runs the selected models. Warewulf provisions the cluster; shared NFS storage keeps persistent data outside the pods.

Request flow

Solid arrows show request traffic. Dotted connections show supporting state or configuration.

flowchart TD
    client["Applications, research jobs and agents"]
    ingress["Traefik · HTTPS ingress"]
    auth["Tyk · API keys and rate limits"]
    gateway["Aleph gateway · Model routing"]
    mesh["Istio · Internal model routes"]
    activator["Knative activator"]
    predictor["Model predictor pod"]
    redis[("Redis · Key and rate-limit state")]
    discovery["Kubernetes API · Cards and deployment state"]

    client --> ingress --> auth --> gateway --> mesh
    auth -.-> redis
    discovery -.-> gateway
    mesh -->|Ready revision| predictor
    mesh -->|When activation is required| activator
    activator --> predictor
Loading

Traefik terminates TLS and routes requests to Tyk. Tyk authenticates API calls, applies limits, and supplies caller identity to the FastAPI gateway. The gateway translates supported APIs and routes by model name using cards and deployment state discovered through Kubernetes watches.

The gateway checks cold-start capacity before forwarding. A model waking up may return 503 with retry guidance; a request may also be refused when suitable capacity is unavailable. Knative can include its activator for activation and buffering. Ready revisions can receive traffic directly through the model route.

Provisioning and placement

flowchart TD
    source["Node image, Aleph overlays and site configuration"]
    warewulf["Warewulf · Provisioning"]
    control["Control-plane VMs · RKE2 servers"]
    workers["GPU workers · RKE2 agents"]
    serving["KServe and Knative · Model lifecycle"]
    placement["Kubernetes and HAMi · GPU placement"]
    pods["Model predictor pods"]
    nfs[("Shared NFS server")]
    state["Redis data and gateway usage records"]

    source --> warewulf
    warewulf --> control
    warewulf --> workers
    control --> serving
    control --> placement
    serving -->|Services and replicas| pods
    placement -->|Scheduling and GPU allocation| pods
    workers -->|Host| pods
    pods -.->|Model PVCs · weights and caches| nfs
    state -.->|Separate PVCs| nfs
Loading

Ingress and the gateway run on control-plane nodes; GPU predictors run on workers. KServe and Knative manage services, revisions, and replica counts. Kubernetes and HAMi place the pods and allocate shared GPU capacity or whole devices. These controllers manage inference workloads; they are not extra request hops. The dotted connections show persistent storage mounts.

Supporting components

Component Role
MetalLB Advertises the public service IP over L2 for Traefik's LoadBalancer Service.
cert-manager + Let's Encrypt Issue and renew the TLS certificate using ACME HTTP-01. Traefik handles HTTPS and HTTP redirection.
Serving runtimes A predictor's queue-proxy forwards requests to vLLM, TEI, NVIDIA NIM, or a custom serving container.
GPU worker stack RKE2, containerd, NVIDIA drivers and container runtime, HAMi device plugin and monitoring. Network and RDMA support are covered in the system guide.
Shared NFS storage Separate PVCs and directories hold Redis data, gateway usage records, and model weights, caches, and environments.
Usage and metrics The gateway writes per-replica JSONL usage files and exposes aggregate counters at /metrics. Tyk's administrative audit and physical GPU telemetry are separate.

Persistent files survive model scale-to-zero and node replacement; they still need backups. Selected models can remain running, while others release GPU resources when idle. More replicas or workers require available capacity.

Deployment details: provisioning and storage, serving and scaling, worker system, and logging and metrics.

Gateway releases are built by CI. Production deployments pin an image version or digest; restarting a Deployment alone does not change that pin. Keep deployed configuration and boot sources aligned.

Documentation

Guide Purpose
Quickstart Deploy an instance
Add a model Our deployment and validation workflow
Endpoints API paths and client configuration
API keys Authentication and identity
Logging and metrics Recorded data, examples, retention, and usage reports
Warewulf Overlays, site settings, and storage
Kubernetes Serving components, placement, and model lifecycle
System Boot integration, node services, and GPU/RDMA support
Gateway reference Routing and model-card behavior
Model definitions Deployment manifests, model cards, and tests
Gateway source Implementation and tests

🔗 References


🤝 Support

Many Bothans died to bring us this information. This project is provided as-is, but reasonable questions may be answered based on my coffee intake or mood. ;)

Feel free to open an issue or email khoja1@ualberta.ca or kali2@ualberta.ca for U of A related deployments.

📜 License

This project is released under the MIT License — use it, modify it, distribute it, include it in proprietary software. Keep the copyright notice. That's it.

Full license text: MIT License

🧠 About University of Alberta Research Computing

The Research Computing Group supports high-performance computing, data-intensive research, and advanced infrastructure for researchers at the University of Alberta and across Canada.

We help design and operate compute environments that power innovation — from AI training clusters to national research infrastructure.

About

A science inference cluster — every model, from protein folding to LLMs, behind one OpenAI/Anthropic-compatible endpoint and a single key. Built to run beside HPC so agentic AI researchers can drive simulations and model inference together. Self-deploying on RKE2 with KServe/Knative + HAMi GPU.

Topics

Resources

Stars

10 stars

Watchers

0 watching

Forks

Contributors

Languages