VCoder

Perception adapter

An adapter for improving large language models at object-level perception tasks with auxiliary perception modalities

VCoder: Versatile Vision Encoders for Multimodal Large Language Models, arXiv 2023 / CVPR 2024

GitHub

266 stars
9 watching
15 forks
Language: Python
last commit: over 2 years ago

Related projects:

RepositoryDescriptionStars
lhoyer/micAn unsupervised domain adaptation method that uses contextual information to improve performance on visual recognition tasks271
shi-labs/gfr-dsodImproving Object Detection from Scratch via Gated Feature Reuse65
vchitect/vbenchA benchmark suite for evaluating the performance of video generative models643
roboflow/maestroA tool to streamline fine-tuning of multimodal models for vision-language tasks1,415
wasidennis/adaptsegnetThis project implements a deep learning-based approach to adapt semantic segmentation models from one domain to another.851
vision-cair/longvuAn artificial intelligence system designed to understand and describe long-form video content329
yiyangzhou/lureAnalyzing and mitigating object hallucination in large vision-language models to improve their accuracy and reliability.136
yunxinli/lingcloudEnhances language models by incorporating human-like eyes to improve visual comprehension and interaction with external world48
gordonhu608/mqt-llavaA vision-language model that uses a query transformer to encode images as visual tokens and allows flexible choice of the number of visual tokens.101
byungkwanlee/moaiImproves performance of vision language tasks by integrating computer vision capabilities into large language models314
vlf-silkie/vlfeedbackAn annotated preference dataset and training framework for improving large vision language models.88
tianyi-lab/hallusionbenchAn image-context reasoning benchmark designed to challenge large vision-language models and help improve their accuracy259
thecodrr/vspeechProvides an interface to Mozilla's DeepSpeech TensorFlow-based Speech-to-Text library using V bindings.49
cvondrick/vaticTools for efficiently scaling up video annotation using crowdsourced marketplaces.609
sergioburdisso/pyss3A Python package implementing an interpretable machine learning model for text classification with visualization tools336