CMMMU
AI summary
Multimodal QA Benchmark
A benchmark for evaluating the performance of multimodal question answering models on diverse domains and data types
- stars
- 46
- forks
- 1
- watching
- 2
Similar projects
Found by comparing what the projects do, not just their names.
Chinese understanding benchmark
Measures the understanding of massive multitask Chinese datasets using large language models
LM Benchmark
A benchmark for evaluating large language models in multiple languages and formats
Multimodal benchmarking
Evaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.
Visual Model Benchmark
An open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models
Multimodal LLM test suite
A benchmark for evaluating large language models' ability to process multimodal input
ML reproducibility
A project demonstrating how to create a reproducible machine learning model using Python and version control
LLM benchmarker
A benchmarking framework for large language models
Go benchmarks
A benchmarking package for the Go language.
MLLM benchmark
An LLM-free benchmark suite for evaluating MLLMs' hallucination capabilities in various tasks and dimensions
Model benchmarking suite
Measures the performance of deep learning models in various deployment scenarios.
Questionnaire
An evaluation suite to assess language models' performance in multi-choice questions
Regex benchmark
A benchmarking project comparing the performance of different programming languages' regex engines
Memory-centric computing benchmarks
A benchmarking suite for evaluating the performance of memory-centric computing architectures
Model evaluator
Evaluates and compares the performance of multimodal large language models on various tasks
Video analysis benchmark
Comprehensive benchmark for evaluating multi-modal large language models on video analysis tasks