Inside the mind of LLMs
This is an upcoming series of articles about the workings of large language models (LLMs). Expect it to be fairly technical, but I will try to keep it accessible to people with a basic understanding of computer science, technology and mathematics.
Planned topics:
- Tokenizers, Tokens, Vocabulary
- Building Blocks: Query, Key, Value, Attention, Non-linearities, Layer Norm, Batch Norm
- The Evolution of Attention: From the "Attention is all you need" paper to the latest Qwen 4 architecture
- Logits and Next-Token Prediction
- Memory: KV Cache, MMap, FreeTokens
- Constrained Decoding
- Inference: Prefill, Decode, Batching, MTP, and hardware mechanics
- Quantization formats: FP16, BF16, FP8, INT8, INT4, NVFP4, and more
- Benchmarks: what they measure
- Teaching a model: SFT, instruction-tuning, domain-adaptation
- Fine tuning efficiently: full-finetuning, PEFT, LoRA, QLoRA
- Reinforcement Learning
- Distillation
- How reasoning works
- Vision transformers
- Diffusion models
- Emergent characteristics
- Key papers to read