Megatron-LM

LLM trainer

A framework for training large language models using scalable and optimized GPU techniques

Ongoing research training transformer models at scale

GitHub

11k stars
165 watching
2k forks
Language: Python
last commit: almost 2 years ago
Linked from 4 awesome lists

large-language-modelsmodel-paratransformers

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
eleutherai/gpt-neoxProvides a framework for training large-scale language models on GPUs with advanced features and optimizations.6,997
microsoft/megatron-deepspeedResearch tool for training large transformer language models at scale1,926
pytorch/torchtitanA native PyTorch library for training large language models using distributed parallelism and optimization techniques.2,765
google-research/vision_transformerProvides pre-trained models and code for training vision transformers and mixers using JAX/Flax10,620
bigscience-workshop/megatron-deepspeedA collection of tools and scripts for training large transformer language models at scale1,342
facebookresearch/metaseqA codebase for working with Open Pre-trained Transformers, enabling deployment and fine-tuning of transformer models on various platforms.6,519
microsoft/lmopsA research initiative focused on developing fundamental technology to improve the performance and efficiency of large language models.3,747
haotian-liu/llavaA system that uses large language and vision models to generate and process visual instructions20,683
nvidia/fastertransformerA high-performance transformer-based NLP component optimized for GPU acceleration and integration into various frameworks.5,937
kimiyoung/transformer-xlImplementations of a neural network architecture for language modeling3,619
opennmt/ctranslate2A high-performance inference engine for transformer models3,467
llava-vl/llava-nextDevelops large multimodal models for various computer vision tasks including image and video analysis3,099
huggingface/trlA library designed to train transformer language models with reinforcement learning using various optimization techniques and fine-tuning methods.10,308
huggingface/transformersA collection of pre-trained machine learning models for various natural language and computer vision tasks, enabling developers to fine-tune and deploy these models on their own projects.136,357
alpha-vllm/llama2-accessoryAn open-source toolkit for pretraining and fine-tuning large language models2,732