hallucination-leaderboard

Model Comparison

Compares performance of large language models on generating coherent summaries from short documents

Leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents

GitHub

1k stars
37 watching
50 forks
Language: Python
last commit: almost 2 years ago
Linked from 1 awesome list

generative-aihallucinationsllm

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
junyangwang0410/amberAn LLM-free benchmark suite for evaluating MLLMs' hallucination capabilities in various tasks and dimensions98
tianyi-lab/hallusionbenchAn image-context reasoning benchmark designed to challenge large vision-language models and help improve their accuracy259
bradyfu/woodpeckerA method to correct hallucinations in multimodal large language models without requiring retraining617
junyangwang0410/haelmA framework for detecting hallucinations in large language models17
x-plug/mplug-halowlEvaluates and mitigates hallucinations in multimodal large language models82
chiragbadhe/lensanalytics-v1A leaderboard application using public data from the Lens Protocol API to rank notable profiles5
fuxiaoliu/lrv-instructionA research project focused on mitigating hallucinations in large multi-modal models by improving instruction tuning through robust training methods.262
bcdnlp/faithscoreEvaluates answers generated by large vision-language models to assess hallucinations27
amazon-science/refcheckerAutomates fine-grained hallucination detection in large language model outputs325
yfzhang114/llava-alignDebiasing techniques to minimize hallucinations in large visual language models75
lalbj/paiImproves the performance of large language models by intervening in their internal workings to reduce hallucinations83
m1guelpf/lens-leaderboardA leaderboard app using public data from the Lens Protocol API to rank notable profiles31
bronyayang/halle_controlControlling object hallucination in large multimodal models28
victordibia/llmxAn API that provides a unified interface to multiple large language models for chat fine-tuning79
damo-nlp-sg/m3examA benchmark for evaluating large language models in multiple languages and formats93