Archives
- 06 Aug Tight, Reproducible, and Wrong: Performance Analysis on an Apple M4
- 30 Jul How JAX Shards a Computation Across a Mesh
- 29 Jul When XLA Isn't Enough: Pallas, Mosaic, and Triton
- 28 Jul What JAX Traces, and What It Refuses
- 27 Jul Programming the TPU: What Its Open-Source Compiler Already Tells You
- 25 Jul XLA Up Close: The Performance Bargain and Its Rigid Price
- 23 Jul A Tour of XLA: The Compiler Beneath JAX, TensorFlow, and PyTorch
- 21 Jul What torch.compile Sees, and What It's Blind To
- 19 Jul Triton: The Compiler That Pretends to Be a Library
- 14 Jul Anatomy of a CUDA Binary
- 13 Jul Where Should Your Code Live?
- 13 Jul Algorithms + Data Structures = Programs
- 12 Jul The Anatomy of an Instruction Pipeline Hazard
- 24 Jun The Mathematics of Dataflow Analysis: Lattices, Fixpoints, and Semirings
- 24 Jun Three Ways to Take a Gradient: Tape, Trace, and Source Transform
- 23 Jun A tour of MLIR: The Dialect Stack Everyone Depends On
- 22 Jun Crossing the Boundary: Custom Kernels and the C++/Python ABI in vLLM
- 21 Jun Why VLIW Architecture is Popular Again
- 19 Jun Building vLLM from Source: A Field Guide (with all the pitfalls)
- 17 Jun vLLM's op IR, or: where the inference engine meets the compiler
- 16 Jun Loop Unrolling in the ML Era
- 13 Jun "Hello, World!" in a Heterogeneous System
- 12 Jun Hardening the ELF: Understanding RELRO and GOT Overwrites
- 10 Jun The Hidden Complexity of the Simplest C Program