Skip to content
View YanissAmz's full-sized avatar

Block or report YanissAmz

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
YanissAmz/README.md

Yaniss Amazouz

AI Engineer // LLM Inference & Systems — Télécom SudParis (IP Paris)

I bring big models to small machines. Upstream llama.cpp contributor, GGUF publisher (64k+ downloads), running 300B+ MoE models 24/7 on a two-node home fleet (Strix Halo 128 GB + RTX 3090).

Upstream work

Home inference lab

Repo What
strix-halo-llm-serving Speculative decoding + kernel work on AMD Strix Halo: DeepSeek-V4-Flash, GLM-5.3-Flash, Qwen3.8 — paired duels, both legs reported
escha-port Porting the ESCHAM 2-bit format (QTIP trellis) into llama.cpp — CUDA/HIP kernels, ×33 prefill, honest head-to-head vs Q4_K_XL
qwdense-turbo 2.28× faster Claude Code on a local Qwen3.6-27B int4 — MTP speculative decoding, tool calling, corruption guards
qwdense-llamacpp llama.cpp native Qwen3.6 MTP lab for long-context agentic coding

Production

  • GETA Solutions — B2C/B2B e-commerce platform (truck tires, parts, transport), designed, built and operated solo: Django 5.2, Stripe (cards, SEPA, Alma), PostgreSQL, TecDoc — live at getasolutions.fr, €18k revenue in August 2026.

Most loved

  • hyte-y70-touch-dashboard ⭐ most-starred — the only Linux dashboard for the HYTE Y70 Touch touchscreen: 10 swipe pages, reverse-engineered LED control.
  • federated-learning-privacy — FedAvg + gradient-inversion attacks + DP defenses, with the honest finding that naive Central DP collapses utility on small federations.
  • efficient-llm-pipeline — TurboQuant KV-cache compression + LoRA on Phi-4-mini, documenting three architecture-dependent findings the original paper doesn't mention.

Stack

Inference — C++, llama.cpp, GGUF, quantization, speculative decoding, vLLM, SGLang, CUDA, ROCm, Vulkan · ML — Python, PyTorch, RAG, Hugging Face · Software — Django, FastAPI, PostgreSQL, Docker, GitHub Actions, Azure


LinkedIn · Hugging Face · getasolutions.fr

Pinned Loading

  1. YanissAmz YanissAmz Public

    GitHub profile README

  2. getasolutions-portfolio getasolutions-portfolio Public

    Public showcase of GETA Solutions — production B2C/B2B e-commerce platform for truck tires, parts, transport (live at getasolutions.fr). Source code is private; this repo holds the README with the …