BenchLMM

Visual Model Benchmark

An open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models

[ECCV 2024] BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models

GitHub

84 stars
0 watching
6 forks
Language: Python
last commit: about 2 years ago
benchmarkcvdatasetlarge-language-modelslarge-multimodal-models

Related projects:

RepositoryDescriptionStars
ailab-cvc/seed-benchA benchmark for evaluating large language models' ability to process multimodal input322
damo-nlp-sg/m3examA benchmark for evaluating large language models in multiple languages and formats93
qcri/llmebenchA benchmarking framework for large language models81
ucsc-vlaa/vllm-safety-benchmarkA benchmark for evaluating the safety and robustness of vision language models against adversarial attacks.72
junyangwang0410/amberAn LLM-free benchmark suite for evaluating MLLMs' hallucination capabilities in various tasks and dimensions98
freedomintelligence/mllm-benchEvaluates and compares the performance of multimodal large language models on various tasks56
szilard/benchm-mlA benchmark for evaluating machine learning algorithms' performance on large datasets1,874
multimodal-art-projection/omnibenchEvaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.15
ailab-cvc/seedAn implementation of a multimodal language model with capabilities for comprehension and generation585
bradyfu/video-mmeComprehensive benchmark for evaluating multi-modal large language models on video analysis tasks422
felixgithub2017/mmcuMeasures the understanding of massive multitask Chinese datasets using large language models87
i-gallegos/fair-llm-benchmarkCompiles bias evaluation datasets and provides access to original data sources for large language models115
mlcommons/inferenceMeasures the performance of deep learning models in various deployment scenarios.1,256
tsb0601/mmvpAn evaluation framework for multimodal language models' visual capabilities using image and question benchmarks.296
lxtgh/omg-segDevelops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.1,336