anandps

Anand Pratap Singh

I make inference fast in Mojo at Modular. Attention, matmul, quantization and collectives, across accelerators, and whatever else stands between a kernel win and an end-to-end one. Before that, turbulence models learned from data, and a PhD in aerospace at Michigan.

Every kernel I touch lives somewhere under this. The job is working out which line is holding it down, and what it costs to move.

matrix peak vector peak memory bound fuse · stage in LDS arithmetic intensity · FLOP / byte attainable FLOP / s
Fig. 1 One bandwidth slope, two peaks. Raising arithmetic intensity walks a kernel right, past the knee, onto the upper roof.