LAVIS

vision-language toolkit

A library that provides pre-trained models and frameworks for multimodal vision-language intelligence tasks such as image captioning and visual question answering.

LAVIS - A One-stop Library for Language-Vision Intelligence

GitHub

10k stars
97 watching
978 forks
Language: Jupyter Notebook
last commit: almost 2 years ago
Linked from 3 awesome lists

deep-learningdeep-learning-libraryimage-captioningmultimodal-datasetsmultimodal-deep-learningsalesforcevision-and-languagevision-frameworkvision-language-pretrainingvision-language-transformervisual-question-anwsering

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
haotian-liu/llavaA system that uses large language and vision models to generate and process visual instructions20,683
nvidia/nemoA scalable generative AI framework for creating and deploying large language models and multimodal models12,438
luogen1996/lavinAn open-source implementation of a vision-language instructed large language model513
freedomintelligence/allavaA collection of datasets and models designed to support the training of lite vision-language models.249
dvlab-research/mgmAn open-source framework for training large language models with vision capabilities.3,229
llava-vl/llava-nextDevelops large multimodal models for various computer vision tasks including image and video analysis3,099
vision-cair/minigpt-4Enabling vision-language understanding by fine-tuning large language models on visual data.25,490
eleutherai/lm-evaluation-harnessProvides a unified framework to test generative language models on various evaluation tasks.7,200
qwenlm/qwen2-vlA multimodal large language model series developed by the Qwen team to understand and process images, videos, and text.3,613
qwenlm/qwen-vlA large vision language model with improved image reasoning and text recognition capabilities, suitable for various multimodal tasks5,179
jy0205/lavitA unified framework for training large language models to understand and generate visual content544
google-research/big_visionSupports large-scale vision model training on GPU machines or Google Cloud TPUs using scalable input pipelines.2,439
optimalscale/lmflowA toolkit for fine-tuning and inferring large machine learning models8,312
nvlabs/prismerA deep learning framework for training multi-modal models with vision and language capabilities.1,299
sgl-project/sglangA fast serving framework for large language models and vision language models.6,551