MMEvalPro
by chenllliang
Source code for MMEvalPro, a more trustworthy and efficient benchmark for evaluating LMMs
AI summary
Model Evaluator
A benchmarking framework for evaluating Large Multimodal Models by providing rigorous metrics and an efficient evaluation pipeline.
- stars
- 22
- forks
- 2
- watching
- 1
Similar projects
Found by comparing what the projects do, not just their names.
Model evaluation toolkit
Tools and evaluation framework for accelerating the development of large multimodal models by providing an efficient way to assess their performance
Multimodal model evaluator
Evaluating and improving large multimodal models through in-context learning
Evaluation pipeline
A framework for evaluating language models on NLP tasks
Model evaluator
Evaluates and compares the performance of multimodal large language models on various tasks
Model evaluator
A tool to automate the evaluation of large language models in Google Colab using various benchmarks and custom parameters.
Model Evaluator
An evaluation framework for machine learning models and datasets, providing standardized metrics and tools for comparing model performance.
Multimodal model evaluator
An open-source benchmark and evaluation tool for assessing multimodal large language models' performance in embodied decision-making tasks
Model evaluator
A tool for evaluating and visualizing machine learning model performance
maluuba/nlg-eval1.4K
Model evaluator
A toolset for evaluating and comparing natural language generation models
tsb0601/mmvp296
Visual model evaluation
An evaluation framework for multimodal language models' visual capabilities using image and question benchmarks.
Model evaluator
Evaluates the capabilities of large multimodal models using a set of diverse tasks and metrics
Model evaluator
A community-developed tool for evaluating climate models and providing diagnostic metrics.
Evaluation framework
An evaluation toolkit for large vision-language models
Model Evaluator
Evaluates language models using standardized benchmarks and prompting techniques.
LLM evaluator
An all-in-one toolkit for evaluating Large Language Models (LLMs) across multiple backends.