MQT-LLaVA
by gordonhu608
[NeurIPS 2024] Matryoshka Query Transformer for Large Vision-Language Models
AI summary
Visual encoder
A vision-language model that uses a query transformer to encode images as visual tokens and allows flexible choice of the number of visual tokens.
- stars
- 101
- forks
- 11
- watching
- 13
Similar projects
Found by comparing what the projects do, not just their names.
Visual Prompt Model
A system designed to enable large multimodal models to understand arbitrary visual prompts
Image encoder
An implementation of a vision transformer architecture designed for high-resolution image encoding with multiple efficient attention mechanisms
nvlabs/prismer1.3K
Vision-Language Model
A deep learning framework for training multi-modal models with vision and language capabilities.
lxtgh/omg-seg1.3K
Visual Model
Develops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.
Visual decoder
A large language model designed to process and generate visual information
Video image processor
An image-based language model that uses large language models to generate visual and text features from videos
Model trainer
A platform for training and deploying large language and vision models that can use tools to perform tasks
Mixture of Experts Model
A large vision-language model using a mixture-of-experts architecture to improve performance on multi-modal learning tasks
Visual Instruction Tuning
An open-source project that enhances visual instruction tuning for text-rich image understanding by integrating GPT-4 models with multimodal datasets.
Model optimizer
This project presents an optimization technique for large-scale image models to reduce computational requirements while maintaining performance.
Model debiasing
Debiasing techniques to minimize hallucinations in large visual language models
whai362/pvt1.7K
PVT
An implementation of Pyramid Vision Transformers for image classification, object detection, and semantic segmentation tasks
Image processor
An all-in-one demo for interactive image processing and generation
Visual representation generator
A project that generates visual representations tailored for general visual reasoning, leveraging hierarchical scene descriptions and instance-level world knowledge.
Vision Language Integrator
Improves performance of vision language tasks by integrating computer vision capabilities into large language models