CMMMU

Multimodal QA Benchmark

A benchmark for evaluating the performance of multimodal question answering models on diverse domains and data types

GitHub

46 stars
2 watching
1 forks
Language: Python
last commit: about 2 years ago

Related projects:

RepositoryDescriptionStars
felixgithub2017/mmcuMeasures the understanding of massive multitask Chinese datasets using large language models87
damo-nlp-sg/m3examA benchmark for evaluating large language models in multiple languages and formats93
multimodal-art-projection/omnibenchEvaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.15
aifeg/benchlmmAn open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models84
ailab-cvc/seed-benchA benchmark for evaluating large language models' ability to process multimodal input322
cmawer/reproducible-modelA project demonstrating how to create a reproducible machine learning model using Python and version control86
qcri/llmebenchA benchmarking framework for large language models81
mna/gocostmodelA benchmarking package for the Go language.61
junyangwang0410/amberAn LLM-free benchmark suite for evaluating MLLMs' hallucination capabilities in various tasks and dimensions98
mlcommons/inferenceMeasures the performance of deep learning models in various deployment scenarios.1,256
mikegu721/xiezhibenchmarkAn evaluation suite to assess language models' performance in multi-choice questions93
mariomka/regex-benchmarkA benchmarking project comparing the performance of different programming languages' regex engines315
cmu-safari/prim-benchmarksA benchmarking suite for evaluating the performance of memory-centric computing architectures142
freedomintelligence/mllm-benchEvaluates and compares the performance of multimodal large language models on various tasks56
bradyfu/video-mmeComprehensive benchmark for evaluating multi-modal large language models on video analysis tasks422