FlexLLMGen

Batch processor

Generates large language model outputs in high-throughput mode on single GPUs

Running large language models on a single GPU for throughput-oriented scenarios.

Archived

GitHub

9k stars
112 watching
553 forks
Language: Python
last commit: almost 2 years ago
deep-learninggpt-3high-throughputlarge-language-modelsmachine-learningoffloadingopt

Related projects:

RepositoryDescriptionStars
sjtu-ipads/powerinferAn efficient Large Language Model inference engine leveraging consumer-grade GPUs on PCs8,011
sgl-project/sglangA fast serving framework for large language models and vision language models.6,551
huggingface/text-generation-inferenceA toolkit for deploying and serving Large Language Models (LLMs) for high-performance text generation9,456
brexhq/prompt-engineeringGuides software developers on how to effectively use and build systems around Large Language Models like GPT-4.8,487
lyogavin/airllmOptimizes large language model inference on limited GPU resources5,446
modeltc/lightllmA Python-based framework for serving large language models with low latency and high scalability.2,691
google/big-benchA benchmark designed to probe large language models and extrapolate their future capabilities through a diverse set of tasks.2,899
optimalscale/lmflowA toolkit for fine-tuning and inferring large machine learning models8,312
thudm/glm-130bAn open-source implementation of a large bilingual language model pre-trained on vast amounts of text data.7,672
microsoft/flamlAutomates machine learning workflows and optimizes model performance using large language models and efficient algorithms3,968
aksnzhy/xlearnA high-performance machine learning package with linear models and factorization machines.3,087
qwenlm/qwen2.5A large language model series with various sizes and variants for text generation and understanding.10,959
dair-ai/ml-papers-explainedAn explanation of key concepts and advancements in the field of Machine Learning7,352
mlabonne/llm-courseA comprehensive course and resource package on building and deploying Large Language Models (LLMs)40,053
young-geng/easylmA framework for training and serving large language models using JAX/Flax2,428