EvALign-ICL
by mshukor
[ICLR2024] (EvALign-ICL Benchmark) Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning
AI summary
Multimodal model evaluator
Evaluating and improving large multimodal models through in-context learning
- stars
- 21
- forks
- 0
- watching
- 2
Similar projects
Found by comparing what the projects do, not just their names.
Multimodal model evaluator
An open-source benchmark and evaluation tool for assessing multimodal large language models' performance in embodied decision-making tasks
Learning benchmark
A benchmarking suite for multimodal in-context learning models
Model Evaluator
A benchmarking framework for evaluating Large Multimodal Models by providing rigorous metrics and an efficient evaluation pipeline.
Model evaluator
Evaluates and compares the performance of multimodal large language models on various tasks
Hallucination tester
Evaluates and mitigates hallucinations in multimodal large language models
Multimodal LLM test suite
A benchmark for evaluating large language models' ability to process multimodal input
Model evaluation toolkit
Tools and evaluation framework for accelerating the development of large multimodal models by providing an efficient way to assess their performance
Model evaluator
Evaluates the capabilities of large multimodal models using a set of diverse tasks and metrics
ML Model Benchmarker
Provides a benchmarking framework and dataset for evaluating the performance of large language models in text-to-image tasks
OCR Benchmark
An evaluation benchmark for OCR capabilities in large multmodal models.
Attention calibrator
This project proposes a novel method for calibrating attention distributions in multimodal models to improve contextualized representations of image-text pairs.
Multimodal benchmarking
Evaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.
Evaluation framework
An evaluation toolkit for large vision-language models
Model evaluator
An evaluation framework for large language models trained with instruction tuning methods
Model evaluator
A community-developed tool for evaluating climate models and providing diagnostic metrics.