MLLM-Bench
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
AI summary
Model evaluator
Evaluates and compares the performance of multimodal large language models on various tasks
- stars
- 56
- forks
- 3
- watching
- 10
Similar projects
Found by comparing what the projects do, not just their names.
Model evaluation toolkit
Tools and evaluation framework for accelerating the development of large multimodal models by providing an efficient way to assess their performance
Model Evaluator
A benchmarking framework for evaluating Large Multimodal Models by providing rigorous metrics and an efficient evaluation pipeline.
Multimodal model evaluator
Evaluating and improving large multimodal models through in-context learning
LLM evaluation resource
A repository of papers and resources for evaluating large language models.
Multimodal LLM test suite
A benchmark for evaluating large language models' ability to process multimodal input
Model evaluator
A tool to automate the evaluation of large language models in Google Colab using various benchmarks and custom parameters.
Evaluation framework
An evaluation toolkit for large vision-language models
MLLM benchmark
An LLM-free benchmark suite for evaluating MLLMs' hallucination capabilities in various tasks and dimensions
Chinese understanding benchmark
Measures the understanding of massive multitask Chinese datasets using large language models
Model evaluator
Evaluates the capabilities of large multimodal models using a set of diverse tasks and metrics
Visual Model Benchmark
An open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models
maluuba/nlg-eval1.4K
Model evaluator
A toolset for evaluating and comparing natural language generation models
Evaluation metrics library
Provides implementations of various supervised machine learning evaluation metrics in multiple programming languages.
Model evaluator
An evaluation framework for large language models trained with instruction tuning methods
Bias datasets
Compiles bias evaluation datasets and provides access to original data sources for large language models