MM-Vet

Model evaluator

Evaluates the capabilities of large multimodal models using a set of diverse tasks and metrics

MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities (ICML 2024)

GitHub

274 stars
2 watching
11 forks
Language: Python
last commit: almost 2 years ago

Related projects:

RepositoryDescriptionStars
zhourax/vegaDevelops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.33
tsb0601/mmvpAn evaluation framework for multimodal language models' visual capabilities using image and question benchmarks.296
yuliang-liu/monkeyAn end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.1,849
chenllliang/mmevalproA benchmarking framework for evaluating Large Multimodal Models by providing rigorous metrics and an efficient evaluation pipeline.22
mshukor/evalign-iclEvaluating and improving large multimodal models through in-context learning21
freedomintelligence/mllm-benchEvaluates and compares the performance of multimodal large language models on various tasks56
allenai/olmo-evalA framework for evaluating language models on NLP tasks326
haozhezhao/micDevelops a multimodal vision-language model to enable machines to understand complex relationships between instructions and images in various tasks.337
evolvinglmms-lab/lmms-evalTools and evaluation framework for accelerating the development of large multimodal models by providing an efficient way to assess their performance2,164
yfzhang114/slimeDevelops large multimodal models for high-resolution understanding and analysis of text, images, and other data types.143
fuxiaoliu/mmcDevelops a large-scale dataset and benchmark for training multimodal chart understanding models using large language models.87
mikegu721/xiezhibenchmarkAn evaluation suite to assess language models' performance in multi-choice questions93
yuliang-liu/multimodalocrAn evaluation benchmark for OCR capabilities in large multmodal models.484
tiger-ai-lab/uniirTrains and evaluates a universal multimodal retrieval model to perform various information retrieval tasks.114
felixgithub2017/mmcuMeasures the understanding of massive multitask Chinese datasets using large language models87