LLaVA

Visual Instruction System

A system that uses large language and vision models to generate and process visual instructions

[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.

GitHub

21k stars
159 watching
2k forks
Language: Python
last commit: about 2 years ago
chatbotchatgptfoundation-modelsgpt-4instruction-tuningllamallama-2llama2llavamulti-modalitymultimodalvision-language-modelvisual-language-learning

Related projects:

RepositoryDescriptionStars
llava-vl/llava-nextDevelops large multimodal models for various computer vision tasks including image and video analysis3,099
instruction-tuning-with-gpt-4/gpt-4-llmThis project generates instruction-following data using GPT-4 to fine-tune large language models for real-world tasks.4,244
opengvlab/llama-adapterAn implementation of a method for fine-tuning language models to follow instructions with high efficiency and accuracy5,775
pku-yuangroup/video-llavaA deep learning framework for generating videos from text inputs and visual features.3,071
hiyouga/llama-factoryA tool for efficiently fine-tuning large language models across multiple architectures and methods.36,219
dvlab-research/mgmAn open-source framework for training large language models with vision capabilities.3,229
alpha-vllm/llama2-accessoryAn open-source toolkit for pretraining and fine-tuning large language models2,732
salt-nlp/llavarAn open-source project that enhances visual instruction tuning for text-rich image understanding by integrating GPT-4 models with multimodal datasets.259
facico/chinese-vicunaAn instruction-following Chinese LLaMA-based model project aimed at training and fine-tuning models on specific hardware configurations for efficient deployment.4,152
damo-nlp-sg/video-llamaAn audio-visual language model designed to understand and respond to video content with improved instruction-following capabilities2,842
luodian/otterA multi-modal AI model developed for improved instruction-following and in-context learning, utilizing large-scale architectures and various training datasets.3,570
qwenlm/qwen-vlA large vision language model with improved image reasoning and text recognition capabilities, suitable for various multimodal tasks5,179
sgl-project/sglangA fast serving framework for large language models and vision language models.6,551
eleutherai/lm-evaluation-harnessProvides a unified framework to test generative language models on various evaluation tasks.7,200
tloen/alpaca-loraTuning a large language model on consumer hardware using low-rank adaptation18,710