Skip to content

inference-learnUnderstand LLM Inference Engines from Scratch

Load Qwen2.5-0.5B with ~2700 lines of pure C. No PyTorch, just libc. Every component written from scratch.

MIT Licensed