LLaVA-NeXT

Multimodal model developer

Develops large multimodal models for various computer vision tasks including image and video analysis

GitHub

3k stars
37 watching
266 forks
Language: Python
last commit: almost 2 years ago

Related projects:

RepositoryDescriptionStars
haotian-liu/llavaA system that uses large language and vision models to generate and process visual instructions20,683
pku-yuangroup/video-llavaA deep learning framework for generating videos from text inputs and visual features.3,071
opengvlab/llama-adapterAn implementation of a method for fine-tuning language models to follow instructions with high efficiency and accuracy5,775
damo-nlp-sg/video-llamaAn audio-visual language model designed to understand and respond to video content with improved instruction-following capabilities2,842
dvlab-research/mgmAn open-source framework for training large language models with vision capabilities.3,229
alpha-vllm/llama2-accessoryAn open-source toolkit for pretraining and fine-tuning large language models2,732
llava-vl/llava-interactive-demoAn all-in-one demo for interactive image processing and generation353
llava-vl/llava-plus-codebaseA platform for training and deploying large language and vision models that can use tools to perform tasks717
scisharp/llamasharpAn efficient C#/.NET library for running Large Language Models (LLMs) on local devices2,750
wisconsinaivision/vip-llavaA system designed to enable large multimodal models to understand arbitrary visual prompts302
eleutherai/lm-evaluation-harnessProvides a unified framework to test generative language models on various evaluation tasks.7,200
hiyouga/llama-factoryA tool for efficiently fine-tuning large language models across multiple architectures and methods.36,219
optimalscale/lmflowA toolkit for fine-tuning and inferring large machine learning models8,312
qwenlm/qwen2-vlA multimodal large language model series developed by the Qwen team to understand and process images, videos, and text.3,613
nvidia/megatron-lmA framework for training large language models using scalable and optimized GPU techniques10,804