AutoAWQ
Pythonpushed almost 2 years ago
AutoAWQ implements the AWQ algorithm for 4-bit quantization with a 2x speedup during inference. Documentation:
AI summary
Model optimizer
An optimization package for 4-bit quantized models
- stars
- 1.8K
- forks
- 220
- watching
- 15
- awesome list
- 1
Featured in 1 awesome list
Each link jumps to the spot where the list mentions AutoAWQ.