Deep Learning Optimizers: Deriving Momentum, RMSProp, and AdamW mathematically
Derive and implement modern deep learning optimizers (SGD, Momentum, RMSProp, and AdamW) from scratch in Python.
Search for a command to run...

Series
This series breaks down complex mathematical theories and algorithms into simple, intuitive explanations with practical examples, making AI accessible to everyone from beginners to aspiring data scientists.
Derive and implement modern deep learning optimizers (SGD, Momentum, RMSProp, and AdamW) from scratch in Python.
Derive and implement the backpropagation algorithm from scratch in Python using raw matrix calculus.
Why every neural network classifier converts its raw output through exponential normalization before calling it a prediction

From cosine similarity to scaled dot-product attention — how one operation powers modern AI
How modern AI skips millions of training hours by reusing pretrained knowledge — and when to freeze, fine-tune, or train from scratch

From the query-key-value formulation to multi-head attention — the mechanism that made transformers dominate AI
