Blog

Between the model and the machine

Sparse acceleration, GPU kernels, and hardware-software co-design. Mostly pieces I went looking for and could not find written down anywhere.

Two lotteries, and the space between them

You can prune more than half of the parameters inside a trained neural network and it still works. So why does it not get proportionally faster? A look at why sparsity kept losing on real hardware, what finally changed, and who gets to decide what comes next.

Sparsity Accelerators Co-design Sparse attention
Read the post

More coming. If you want to be told when, the easiest way is to follow me on GitHub.