model-compression
an archive of posts with this tag
| Jun 11, 2026 | Inference speed isn't just a model-size problem |
|---|---|
| Apr 30, 2026 | Why joint compression beats doing it in sequence |
| Mar 21, 2026 | Quantization: fewer bits per weight, without losing the model |
| Feb 04, 2026 | Pruning: removing what a network doesn't need |