HallusionBench

Benchmark

An image-context reasoning benchmark designed to challenge large vision-language models and help improve their accuracy

[CVPR'24] HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning Benchmark Challenging for GPT-4V(ision), LLaVA-1.5, and Other Multi-modality Models

GitHub

259 stars
4 watching
7 forks
Language: Python
last commit: almost 2 years ago
benchmarkbenchmarksgpt-4gpt-4vhallucinationlarge-language-modelslarge-vision-language-modelsllavallmlmmvlms

Related projects:

RepositoryDescriptionStars
bradyfu/woodpeckerA method to correct hallucinations in multimodal large language models without requiring retraining617
yfzhang114/llava-alignDebiasing techniques to minimize hallucinations in large visual language models75
amazon-science/refcheckerAutomates fine-grained hallucination detection in large language model outputs325
x-plug/mplug-halowlEvaluates and mitigates hallucinations in multimodal large language models82
nvlabs/bongard-hoiA benchmarking tool and software framework for evaluating few-shot visual reasoning capabilities in computer vision models.64
yiyangzhou/lureAnalyzing and mitigating object hallucination in large vision-language models to improve their accuracy and reliability.136
1zhou-wang/memvrAn implementation of a method to mitigate hallucinations in large language models using visual re-tracing28
fuxiaoliu/lrv-instructionA research project focused on mitigating hallucinations in large multi-modal models by improving instruction tuning through robust training methods.262
junyangwang0410/amberAn LLM-free benchmark suite for evaluating MLLMs' hallucination capabilities in various tasks and dimensions98
lalbj/paiImproves the performance of large language models by intervening in their internal workings to reduce hallucinations83
bcdnlp/faithscoreEvaluates answers generated by large vision-language models to assess hallucinations27
vectara/hallucination-leaderboardCompares performance of large language models on generating coherent summaries from short documents1,281
qcri/llmebenchA benchmarking framework for large language models81
junyangwang0410/haelmA framework for detecting hallucinations in large language models17
yuqifan1117/hallucidoctorThis project provides tools and frameworks to mitigate hallucinatory toxicity in visual instruction data, allowing researchers to fine-tune MLLM models on specific datasets.41