lmms-eval

Model evaluation toolkit

Tools and evaluation framework for accelerating the development of large multimodal models by providing an efficient way to assess their performance

Accelerating the development of large multimodal models (LMMs) with one-click evaluation module - lmms-eval.

GitHub

2k stars
3 watching
168 forks
Language: Python
last commit: almost 2 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
freedomintelligence/mllm-benchEvaluates and compares the performance of multimodal large language models on various tasks56
chenllliang/mmevalproA benchmarking framework for evaluating Large Multimodal Models by providing rigorous metrics and an efficient evaluation pipeline.22
mlgroupjlu/llm-eval-surveyA repository of papers and resources for evaluating large language models.1,450
allenai/olmo-evalA framework for evaluating language models on NLP tasks326
mlabonne/llm-autoevalA tool to automate the evaluation of large language models in Google Colab using various benchmarks and custom parameters.566
mshukor/evalign-iclEvaluating and improving large multimodal models through in-context learning21
open-compass/vlmevalkitAn evaluation toolkit for large vision-language models1,514
declare-lab/instruct-evalAn evaluation framework for large language models trained with instruction tuning methods535
prometheus-eval/prometheus-evalAn open-source framework that enables language model evaluation using Prometheus and GPT4820
esmvalgroup/esmvaltoolA community-developed tool for evaluating climate models and providing diagnostic metrics.230
h2oai/h2o-llm-evalAn evaluation framework for large language models with Elo rating system and A/B testing capabilities50
maluuba/nlg-evalA toolset for evaluating and comparing natural language generation models1,350
huggingface/lightevalAn all-in-one toolkit for evaluating Large Language Models (LLMs) across multiple backends.879
evolvinglmms-lab/longvaAn open-source project that enables the transfer of language understanding to vision capabilities through long context processing.347
modelscope/evalscopeA framework for efficiently evaluating and benchmarking large models308