MM-Vet
by yuweihao
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities (ICML 2024)
AI summary
Model evaluator
Evaluates the capabilities of large multimodal models using a set of diverse tasks and metrics
- stars
- 274
- forks
- 11
- watching
- 2
Similar projects
Found by comparing what the projects do, not just their names.
Multimodal evaluation framework
Develops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.
tsb0601/mmvp296
Visual model evaluation
An evaluation framework for multimodal language models' visual capabilities using image and question benchmarks.
Image captioner
An end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.
Model Evaluator
A benchmarking framework for evaluating Large Multimodal Models by providing rigorous metrics and an efficient evaluation pipeline.
Multimodal model evaluator
Evaluating and improving large multimodal models through in-context learning
Model evaluator
Evaluates and compares the performance of multimodal large language models on various tasks
Evaluation pipeline
A framework for evaluating language models on NLP tasks
Multimodal learner
Develops a multimodal vision-language model to enable machines to understand complex relationships between instructions and images in various tasks.
Model evaluation toolkit
Tools and evaluation framework for accelerating the development of large multimodal models by providing an efficient way to assess their performance
Multimodal model developer
Develops large multimodal models for high-resolution understanding and analysis of text, images, and other data types.
Chart model trainer
Develops a large-scale dataset and benchmark for training multimodal chart understanding models using large language models.
Questionnaire
An evaluation suite to assess language models' performance in multi-choice questions
OCR Benchmark
An evaluation benchmark for OCR capabilities in large multmodal models.
Retrieval model trainer
Trains and evaluates a universal multimodal retrieval model to perform various information retrieval tasks.
Chinese understanding benchmark
Measures the understanding of massive multitask Chinese datasets using large language models