MQT-LLaVA

Visual encoder

A vision-language model that uses a query transformer to encode images as visual tokens and allows flexible choice of the number of visual tokens.

[NeurIPS 2024] Matryoshka Query Transformer for Large Vision-Language Models

GitHub

101 stars
13 watching
11 forks
Language: Python
last commit: about 2 years ago

Related projects:

RepositoryDescriptionStars
wisconsinaivision/vip-llavaA system designed to enable large multimodal models to understand arbitrary visual prompts302
microsoft/vision-longformerAn implementation of a vision transformer architecture designed for high-resolution image encoding with multiple efficient attention mechanisms243
nvlabs/prismerA deep learning framework for training multi-modal models with vision and language capabilities.1,299
lxtgh/omg-segDevelops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.1,336
opengvlab/visionllmA large language model designed to process and generate visual information956
dvlab-research/llama-vidAn image-based language model that uses large language models to generate visual and text features from videos748
llava-vl/llava-plus-codebaseA platform for training and deploying large language and vision models that can use tools to perform tasks717
pku-yuangroup/moe-llavaA large vision-language model using a mixture-of-experts architecture to improve performance on multi-modal learning tasks2,023
salt-nlp/llavarAn open-source project that enhances visual instruction tuning for text-rich image understanding by integrating GPT-4 models with multimodal datasets.259
alibaba/conv-llavaThis project presents an optimization technique for large-scale image models to reduce computational requirements while maintaining performance.106
yfzhang114/llava-alignDebiasing techniques to minimize hallucinations in large visual language models75
whai362/pvtAn implementation of Pyramid Vision Transformers for image classification, object detection, and semantic segmentation tasks1,745
llava-vl/llava-interactive-demoAn all-in-one demo for interactive image processing and generation353
lavi-lab/visual-tableA project that generates visual representations tailored for general visual reasoning, leveraging hierarchical scene descriptions and instance-level world knowledge.14
byungkwanlee/moaiImproves performance of vision language tasks by integrating computer vision capabilities into large language models314