MMBench

Multi-modal model evaluation suite

A collection of benchmarks to evaluate the multi-modal understanding capability of large vision language models.

Official Repo of "MMBench: Is Your Multi-modal Model an All-around Player?"

GitHub

168 stars
3 watching
10 forks
last commit: about 2 years ago

Related projects:

RepositoryDescriptionStars
open-compass/vlmevalkitAn evaluation toolkit for large vision-language models1,514
multimodal-art-projection/omnibenchEvaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.15
openbmb/viscpmA family of large multimodal models supporting multimodal conversational capabilities and text-to-image generation in multiple languages1,098
open-compass/lawbenchEvaluates the legal knowledge of large language models using a custom benchmarking framework.273
opengvlab/multi-modality-arenaAn evaluation platform for comparing multi-modality models on visual question-answering tasks478
will-singularity/skywork-mmAn empirical study aiming to develop a large language model capable of effectively integrating multiple input modalities23
openm3d/m3dbenchAn open-source software project providing a comprehensive 3D instruction-following dataset with multi-modal prompts for training large language models.58
fuxiaoliu/mmcDevelops a large-scale dataset and benchmark for training multimodal chart understanding models using large language models.87
tsb0601/mmvpAn evaluation framework for multimodal language models' visual capabilities using image and question benchmarks.296
sail-sg/mmcbenchA benchmarking framework designed to evaluate the robustness of large multimodal models against common corruption scenarios27
mickcrosse/mtrf-toolboxA MATLAB package for modeling and analyzing multivariate neural responses to dynamic stimuli.85
open-mmlab/mmhuman3dProvides a modular framework and tools for working with 3D human parametric models in computer vision and graphics1,253
open-mmlab/mmactionAn open-source toolbox for action understanding from video data using PyTorch.1,863
ailab-cvc/seed-benchA benchmark for evaluating large language models' ability to process multimodal input322
pleisto/yuren-baichuan-7bA multi-modal large language model that integrates natural language and visual capabilities with fine-tuning for various tasks73