Blog

writing on AI, systems, and things i find interesting.

Technical

Flash Attention: The Hidden Gem A one-liner in PyTorch. Thousands of lines of math underneath. Here is what is actually going on.
The Stride Model: How Tensors Represent Reshape and Transpose Once you can see the flat array underneath, reshape and transpose stop being two separate things to memorise.

Experience

Training GPT-2: My Experience From applying AI to understanding it. Covers kernels, flash attention, and a lot of PyTorch relearning.