VEGA

Multimodal evaluation framework

Develops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.

GitHub

33 stars
1 watching
2 forks
Language: Python
last commit: about 2 years ago

Related projects:

RepositoryDescriptionStars
tsb0601/mmvpAn evaluation framework for multimodal language models' visual capabilities using image and question benchmarks.296
yuweihao/mm-vetEvaluates the capabilities of large multimodal models using a set of diverse tasks and metrics274
yuliang-liu/monkeyAn end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.1,849
multimodal-art-projection/omnibenchEvaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.15
yuxie11/r2d2A framework for large-scale cross-modal benchmarks and vision-language tasks in Chinese157
xverse-ai/xverse-v-13bA large multimodal model for visual question answering, trained on a dataset of 2.1B image-text pairs and 8.2M instruction sequences.78
shizhediao/davinciImplementing a unified modal learning framework for generative vision-language models43
nvlabs/prismerA deep learning framework for training multi-modal models with vision and language capabilities.1,299
yfzhang114/slimeDevelops large multimodal models for high-resolution understanding and analysis of text, images, and other data types.143
kohjingyu/fromageA framework for grounding language models to images and handling multimodal inputs and outputs478
meituan-automl/mobilevlmAn implementation of a vision language model designed for mobile devices, utilizing a lightweight downsample projector and pre-trained language models.1,076
pku-yuangroup/languagebindExtending pretraining models to handle multiple modalities by aligning language and video representations751
haozhezhao/micDevelops a multimodal vision-language model to enable machines to understand complex relationships between instructions and images in various tasks.337
yuqifan1117/hallucidoctorThis project provides tools and frameworks to mitigate hallucinatory toxicity in visual instruction data, allowing researchers to fine-tune MLLM models on specific datasets.41
penghao-wu/vstarPyTorch implementation of guided visual search mechanism for multimodal LLMs541