EvALign-ICL

Multimodal model evaluator

Evaluating and improving large multimodal models through in-context learning

[ICLR2024] (EvALign-ICL Benchmark) Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning

GitHub

21 stars
2 watching
0 forks
Language: Python
last commit: over 2 years ago

Related projects:

RepositoryDescriptionStars
pkunlp-icler/pca-evalAn open-source benchmark and evaluation tool for assessing multimodal large language models' performance in embodied decision-making tasks99
ys-zong/vl-iclA benchmarking suite for multimodal in-context learning models31
chenllliang/mmevalproA benchmarking framework for evaluating Large Multimodal Models by providing rigorous metrics and an efficient evaluation pipeline.22
freedomintelligence/mllm-benchEvaluates and compares the performance of multimodal large language models on various tasks56
x-plug/mplug-halowlEvaluates and mitigates hallucinations in multimodal large language models82
ailab-cvc/seed-benchA benchmark for evaluating large language models' ability to process multimodal input322
evolvinglmms-lab/lmms-evalTools and evaluation framework for accelerating the development of large multimodal models by providing an efficient way to assess their performance2,164
yuweihao/mm-vetEvaluates the capabilities of large multimodal models using a set of diverse tasks and metrics274
uw-madison-lee-lab/cobsatProvides a benchmarking framework and dataset for evaluating the performance of large language models in text-to-image tasks30
yuliang-liu/multimodalocrAn evaluation benchmark for OCR capabilities in large multmodal models.484
lancopku/iaisThis project proposes a novel method for calibrating attention distributions in multimodal models to improve contextualized representations of image-text pairs.30
multimodal-art-projection/omnibenchEvaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.15
open-compass/vlmevalkitAn evaluation toolkit for large vision-language models1,514
declare-lab/instruct-evalAn evaluation framework for large language models trained with instruction tuning methods535
esmvalgroup/esmvaltoolA community-developed tool for evaluating climate models and providing diagnostic metrics.230