exllama
by turboderp
Pythonpushed about 3 years ago
A more memory-efficient rewrite of the HF transformers implementation of Llama for use with quantized weights.
AI summary
GPU-based chat model
A re-implementation of Llama for efficient use with quantized weights on modern GPUs.
- stars
- 2.8K
- forks
- 220
- watching
- 37
- awesome list
- 1
Featured in 1 awesome list
Each link jumps to the spot where the list mentions exllama.