Open Documentation

Model Architecture & Internals

How transformers actually work, built from scratch rather than taken on faith. Tokenizers, the KV cache, speculative decoding, and what a small model has genuinely learned once training finishes.