There Is No Such Thing as a Passive Read
Software engineers operate under a comforting fiction: reading a value is free. In C++, we decorate pointers with const to assure the compiler that inspection produces no side effects. In function...
Software engineers operate under a comforting fiction: reading a value is free. In C++, we decorate pointers with const to assure the compiler that inspection produces no side effects. In function...
Pull up the die photo of any high-performance CPU from the last fifteen years and point at the parts that do arithmetic. They are small. The integer ALUs and the vector units are a modest slice of ...
I keep noticing the same pattern working with models on real code. The Rust I get back is better than the C++ I get back, and the Lean is better than either. The Rust compiles and does what I asked...
At Hot Chips this week, OpenAI presented Jalapeño, an inference ASIC co-designed with Broadcom.1 The architecture slide circulating since the presentation frames a request pipeline split three ways...
Prefix caching can cause the same prompt to produce different logits on a cache hit versus a cache miss. In many operational contexts, this is quickly categorized as a defect. However, a bug strict...
Each of the previous three posts assumed the next layer of the stack would save it. Part two showed that KV caches lack an interchange format, assuming the bytes would flow if two vendors merely ag...
Compilers are commonly described as embarrassingly parallel because translation units are independent of one another, and machines have many cores; yet the frontend that turns source into IR (the A...
The previous post ended on a pattern: a memory hierarchy works because one component can see the whole path, and disaggregation splits that visibility across two vendors. Scheduling has the same sh...
The last post argued that a KV cache has no interchange format. Grant one anyway. Suppose the two vendors agree on layout, dtype, block size, scale placement and every other axis in that list. The ...
The last post argued that prefill and decode want different computers, and that the industry has started buying them separately. This one is about the handoff, which sounds like the easy part but i...