AMBER

MLLM benchmark

An LLM-free benchmark suite for evaluating MLLMs' hallucination capabilities in various tasks and dimensions

An LLM-free Multi-dimensional Benchmark for Multi-modal Hallucination Evaluation

GitHub

98 stars
1 watching
2 forks
Language: Python
last commit: over 2 years ago

Related projects:

RepositoryDescriptionStars
x-plug/mplug-halowlEvaluates and mitigates hallucinations in multimodal large language models82
junyangwang0410/haelmA framework for detecting hallucinations in large language models17
damo-nlp-sg/m3examA benchmark for evaluating large language models in multiple languages and formats93
vectara/hallucination-leaderboardCompares performance of large language models on generating coherent summaries from short documents1,281
freedomintelligence/mllm-benchEvaluates and compares the performance of multimodal large language models on various tasks56
aifeg/benchlmmAn open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models84
multimodal-art-projection/omnibenchEvaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.15
tianyi-lab/hallusionbenchAn image-context reasoning benchmark designed to challenge large vision-language models and help improve their accuracy259
ailab-cvc/seed-benchA benchmark for evaluating large language models' ability to process multimodal input322
uw-madison-lee-lab/cobsatProvides a benchmarking framework and dataset for evaluating the performance of large language models in text-to-image tasks30
km1994/llmsninestorydemontowerExploring various LLMs and their applications in natural language processing and related areas1,854
bradyfu/woodpeckerA method to correct hallucinations in multimodal large language models without requiring retraining617
oval-group/mloggerA lightweight logger for machine learning experiments127
pleisto/yuren-baichuan-7bA multi-modal large language model that integrates natural language and visual capabilities with fine-tuning for various tasks73
szilard/benchm-mlA benchmark for evaluating machine learning algorithms' performance on large datasets1,874