Pinned Loading
-
Tinydeepseek
Tinydeepseek Public26M parameter language model with Mixture of Experts architecture, Multi-Latent Attention and efficient expert routing. Only activates ~13M parameters per token.
Jupyter Notebook 1
-
qwen-depth-distill
qwen-depth-distill PublicPrune Qwen2.5-0.5B from 24 to 7 layers, distill it back on 800M tokens, and test what the loss curve can and can't tell you.
Python
-
-
agentic_screening
agentic_screening PublicHybrid resume screening (vector + cross-encoder + LLM, RRF-fused) with live voice interviews
Python
-
-
simple_GRPO
simple_GRPO PublicForked from lsdefine/simple_GRPO
A very simple GRPO implement for reproducing r1-like LLM thinking.
Python
If the problem persists, check the GitHub status page or contact support.


