Multi-Modality-Arena

Model arena

An evaluation platform for comparing multi-modality models on visual question-answering tasks

Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4, LLaMA-Adapter V2, LLaVA, BLIP-2, and many more!

GitHub

478 stars
6 watching
36 forks
Language: Python
last commit: over 2 years ago
chatchatbotchatgptgradiolarge-language-modelsllmsmulti-modalityvision-language-modelvqa

Related projects:

RepositoryDescriptionStars
open-mmlab/multimodal-gptTrains a multimodal chatbot that combines visual and language instructions to generate responses1,478
opengvlab/lammA framework and benchmark for training and evaluating multi-modal large language models, enabling the development of AI agents capable of seamless interaction between humans and machines.305
mbzuai-oryx/groundinglmmAn end-to-end trained model capable of generating natural language responses integrated with object segmentation masks for interactive visual conversations797
openbmb/viscpmA family of large multimodal models supporting multimodal conversational capabilities and text-to-image generation in multiple languages1,098
mbzuai-oryx/video-chatgptA video conversation model that generates meaningful conversations about videos using large vision and language models1,246
nvlabs/eagleDevelops high-resolution multimodal LLMs by combining vision encoders and various input resolutions549
xverse-ai/xverse-v-13bA large multimodal model for visual question answering, trained on a dataset of 2.1B image-text pairs and 8.2M instruction sequences.78
multimodal-art-projection/omnibenchEvaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.15
open-compass/mmbenchA collection of benchmarks to evaluate the multi-modal understanding capability of large vision language models.168
tsb0601/mmvpAn evaluation framework for multimodal language models' visual capabilities using image and question benchmarks.296
mlpc-ucsd/blivaA multimodal LLM designed to handle text-rich visual questions270
ailab-cvc/seedAn implementation of a multimodal language model with capabilities for comprehension and generation585
ucsc-vlaa/sight-beyond-textAn implementation of a multimodal LLM training paradigm to enhance truthfulness and ethics in language models19
wisconsinaivision/vip-llavaA system designed to enable large multimodal models to understand arbitrary visual prompts302
mlo-lab/muviA software framework for multi-view latent variable modeling with domain-informed structured sparsity27