PVIT

Visual Instruction Model

A project that extends large language models by integrating an additional region-level vision encoder to improve visual instruction tuning.

Repository of paper: Position-Enhanced Visual Instruction Tuning for Multimodal Large Language Models

GitHub

37 stars
2 watching
2 forks
Language: Python
last commit: almost 3 years ago

Related projects:

RepositoryDescriptionStars
baai-dcai/visual-instruction-tuningA dataset and model designed to scale visual instruction tuning using language-only GPT-4 models.164
aidc-ai/parrotA method and toolkit for fine-tuning large language models to perform visual instruction tasks in multiple languages.34
vt-nlp/multiinstructA multimodal benchmark dataset designed to evaluate the performance of vision-language foundation models through instruction tuning.134
vpgtrans/vpgtransTransfers visual prompt generators across large language models to reduce training costs and enable customization of multimodal LLMs270
wisconsinaivision/vip-llavaA system designed to enable large multimodal models to understand arbitrary visual prompts302
salt-nlp/llavarAn open-source project that enhances visual instruction tuning for text-rich image understanding by integrating GPT-4 models with multimodal datasets.259
whai362/pvtAn implementation of Pyramid Vision Transformers for image classification, object detection, and semantic segmentation tasks1,745
x2fd/lvis-instruct4vA dataset of fine-grained visual instructions generated by prompting a large language model with images from another dataset131
rucaibox/comvintCreating synthetic visual reasoning instructions to improve the performance of large language models on image-related tasks18
vlf-silkie/vlfeedbackAn annotated preference dataset and training framework for improving large vision language models.88
jy0205/lavitA unified framework for training large language models to understand and generate visual content544
jshilong/gpt4roiTraining and deploying large language models on computer vision tasks using region-of-interest inputs517
opendatalab/vigcAutonomously generates high-quality image-text instruction fine-tuning datasets91
vchitect/vbenchA benchmark suite for evaluating the performance of video generative models643
pvlib/pvlib-pythonA Python library for simulating photovoltaic energy system performance and modeling solar energy systems.1,228