VEGA
by zhourax
AI summary
Multimodal evaluation framework
Develops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.
- stars
- 33
- forks
- 2
- watching
- 1
Similar projects
Found by comparing what the projects do, not just their names.
tsb0601/mmvp296
Visual model evaluation
An evaluation framework for multimodal language models' visual capabilities using image and question benchmarks.
Model evaluator
Evaluates the capabilities of large multimodal models using a set of diverse tasks and metrics
Image captioner
An end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.
Multimodal benchmarking
Evaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.
yuxie11/r2d2157
Vision-Language Framework
A framework for large-scale cross-modal benchmarks and vision-language tasks in Chinese
Multimodal model
A large multimodal model for visual question answering, trained on a dataset of 2.1B image-text pairs and 8.2M instruction sequences.
Vision-Language Model Framework
Implementing a unified modal learning framework for generative vision-language models
nvlabs/prismer1.3K
Vision-Language Model
A deep learning framework for training multi-modal models with vision and language capabilities.
Multimodal model developer
Develops large multimodal models for high-resolution understanding and analysis of text, images, and other data types.
Multimodal model framework
A framework for grounding language models to images and handling multimodal inputs and outputs
Vision Language Model
An implementation of a vision language model designed for mobile devices, utilizing a lightweight downsample projector and pre-trained language models.
Multimodal alignment model
Extending pretraining models to handle multiple modalities by aligning language and video representations
Multimodal learner
Develops a multimodal vision-language model to enable machines to understand complex relationships between instructions and images in various tasks.
Data processing framework
This project provides tools and frameworks to mitigate hallucinatory toxicity in visual instruction data, allowing researchers to fine-tune MLLM models on specific datasets.
Visual search framework
PyTorch implementation of guided visual search mechanism for multimodal LLMs