neural-compressor
by intel
Pythonpushed almost 2 years ago
SOTA low-bit LLM quantization (INT8/FP8/INT4/FP4/NF4) & sparsity; leading model compression techniques on TensorFlow, PyTorch, and ONNX Runtime
AI summary
Model optimizer
Tools and techniques for optimizing large language models on various frameworks and hardware platforms.
- stars
- 2.3K
- forks
- 257
- watching
- 33
- awesome lists
- 2
Featured in 2 awesome lists
Each link jumps to the spot where the list mentions neural-compressor.