ViP-LLaVA

Visual Prompt Model

A system designed to enable large multimodal models to understand arbitrary visual prompts

[CVPR2024] ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts

GitHub

302 stars
5 watching
21 forks
Language: Python
last commit: about 2 years ago
chatbotclipcvpr2024foundation-modelsgpt-4gpt-4-visionllamallama2llavamulti-modalvision-languagevisual-prompting

Related projects:

RepositoryDescriptionStars
llava-vl/llava-interactive-demoAn all-in-one demo for interactive image processing and generation353
llava-vl/llava-plus-codebaseA platform for training and deploying large language and vision models that can use tools to perform tasks717
mlpc-ucsd/blivaA multimodal LLM designed to handle text-rich visual questions270
dvlab-research/llama-vidAn image-based language model that uses large language models to generate visual and text features from videos748
vpgtrans/vpgtransTransfers visual prompt generators across large language models to reduce training costs and enable customization of multimodal LLMs270
airaria/visual-chinese-llama-alpacaDevelops a multimodal Chinese language model with visual capabilities429
gordonhu608/mqt-llavaA vision-language model that uses a query transformer to encode images as visual tokens and allows flexible choice of the number of visual tokens.101
yfzhang114/llava-alignDebiasing techniques to minimize hallucinations in large visual language models75
360cvgroup/360vlA large multi-modal model developed using the Llama3 language model, designed to improve image understanding capabilities.32
ailab-cvc/seedAn implementation of a multimodal language model with capabilities for comprehension and generation585
nvlabs/prismerA deep learning framework for training multi-modal models with vision and language capabilities.1,299
milvlg/prophetAn implementation of a two-stage framework designed to prompt large language models with answer heuristics for knowledge-based visual question answering tasks.270
baaivision/eveA PyTorch implementation of an encoder-free vision-language model that can be fine-tuned for various tasks and modalities246
deepseek-ai/deepseek-vlA multimodal AI model that enables real-world vision-language understanding applications2,145
lxtgh/omg-segDevelops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.1,336