airllm

Inference optimizer

Optimizes large language model inference on limited GPU resources

AirLLM 70B inference with single 4GB GPU

GitHub

5k stars
128 watching
437 forks
Language: Jupyter Notebook
last commit: almost 2 years ago
Linked from 1 awesome list

chinese-llmchinese-nlpfinetunegenerative-aiinstruct-gptinstruction-setllamallmloraopen-modelsopen-sourceopen-source-modelsqlora

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
modeltc/lightllmA Python-based framework for serving large language models with low latency and high scalability.2,691
ggerganov/llama.cppEnables LLM inference with minimal setup and high performance on various hardware platforms69,185
mit-han-lab/llm-awqAn open-source software project that enables efficient and accurate low-bit weight quantization for large language models.2,593
internlm/lmdeployA toolkit for optimizing and serving large language models4,854
opengvlab/llama-adapterAn implementation of a method for fine-tuning language models to follow instructions with high efficiency and accuracy5,775
vllm-project/vllmAn inference and serving engine for large language models31,982
optimalscale/lmflowA toolkit for fine-tuning and inferring large machine learning models8,312
fminference/flexllmgenGenerates large language model outputs in high-throughput mode on single GPUs9,236
alpha-vllm/llama2-accessoryAn open-source toolkit for pretraining and fine-tuning large language models2,732
young-geng/easylmA framework for training and serving large language models using JAX/Flax2,428
nomic-ai/gpt4allAn open-source Python client for running Large Language Models (LLMs) locally on any device.71,176
haotian-liu/llavaA system that uses large language and vision models to generate and process visual instructions20,683
sjtu-ipads/powerinferAn efficient Large Language Model inference engine leveraging consumer-grade GPUs on PCs8,011
thudm/glm-130bAn open-source implementation of a large bilingual language model pre-trained on vast amounts of text data.7,672
tloen/alpaca-loraTuning a large language model on consumer hardware using low-rank adaptation18,710