Bingo

Model evaluation tool

An analysis project investigating limitations of visual language models in understanding and processing images with potential biases and interference challenges.

GitHub

53 stars
3 watching
1 forks
last commit: over 2 years ago

Related projects:

RepositoryDescriptionStars
mikegu721/xiezhibenchmarkAn evaluation suite to assess language models' performance in multi-choice questions93
zzhanghub/eval-co-sodAn evaluation tool for co-saliency detection tasks97
mbzuai-oryx/groundinglmmAn end-to-end trained model capable of generating natural language responses integrated with object segmentation masks for interactive visual conversations797
open-compass/mmbenchA collection of benchmarks to evaluate the multi-modal understanding capability of large vision language models.168
mlgroupjlu/llm-eval-surveyA repository of papers and resources for evaluating large language models.1,450
cluebenchmark/supercluelybA benchmarking platform for evaluating Chinese general-purpose models through anonymous, random battles143
open-compass/vlmevalkitAn evaluation toolkit for large vision-language models1,514
felixgithub2017/mmcuMeasures the understanding of massive multitask Chinese datasets using large language models87
agrigpts/agrigptsDeveloping large language models for agricultural applications to improve crop yields and support rural development.22
yuweihao/mm-vetEvaluates the capabilities of large multimodal models using a set of diverse tasks and metrics274
cgnorthcutt/cleanlabA tool for evaluating and improving the fairness of machine learning models57
freedomintelligence/mllm-benchEvaluates and compares the performance of multimodal large language models on various tasks56
masaiahhan/correlationqaAn investigation into the relationship between misleading images and hallucinations in large language models8
applieddatasciencepartners/xgboostexplainerProvides tools to understand and interpret the decisions made by XGBoost models in machine learning253
tsb0601/mmvpAn evaluation framework for multimodal language models' visual capabilities using image and question benchmarks.296