VCoder
by SHI-Labs
VCoder: Versatile Vision Encoders for Multimodal Large Language Models, arXiv 2023 / CVPR 2024
AI summary
Perception adapter
An adapter for improving large language models at object-level perception tasks with auxiliary perception modalities
- stars
- 266
- forks
- 15
- watching
- 9
Similar projects
Found by comparing what the projects do, not just their names.
lhoyer/mic271
Context-aware adaptor
An unsupervised domain adaptation method that uses contextual information to improve performance on visual recognition tasks
Object Detector
Improving Object Detection from Scratch via Gated Feature Reuse
Video model evaluator
A benchmark suite for evaluating the performance of video generative models
roboflow/maestro1.4K
fine-tuner
A tool to streamline fine-tuning of multimodal models for vision-language tasks
Domain adaptation model
This project implements a deep learning-based approach to adapt semantic segmentation models from one domain to another.
Video describer
An artificial intelligence system designed to understand and describe long-form video content
Model validator
Analyzing and mitigating object hallucination in large vision-language models to improve their accuracy and reliability.
Visual enhancer for LLMs
Enhances language models by incorporating human-like eyes to improve visual comprehension and interaction with external world
Visual encoder
A vision-language model that uses a query transformer to encode images as visual tokens and allows flexible choice of the number of visual tokens.
Vision Language Integrator
Improves performance of vision language tasks by integrating computer vision capabilities into large language models
Vision language model trainer
An annotated preference dataset and training framework for improving large vision language models.
Benchmark
An image-context reasoning benchmark designed to challenge large vision-language models and help improve their accuracy
Speech decoder
Provides an interface to Mozilla's DeepSpeech TensorFlow-based Speech-to-Text library using V bindings.
Video annotator
Tools for efficiently scaling up video annotation using crowdsourced marketplaces.
Text classifier
A Python package implementing an interpretable machine learning model for text classification with visualization tools