Two lotteries, and the space between them
You can prune more than half of the parameters inside a trained neural network and it still works. So why does it not get proportionally faster? A look at why sparsity kept losing on real hardware, what finally changed, and who gets to decide what comes next.
Sparsity
Accelerators
Co-design
Sparse attention
Read the post