Bingo
by gzcch
AI summary
Model evaluation tool
An analysis project investigating limitations of visual language models in understanding and processing images with potential biases and interference challenges.
- stars
- 53
- forks
- 1
- watching
- 3
Similar projects
Found by comparing what the projects do, not just their names.
Questionnaire
An evaluation suite to assess language models' performance in multi-choice questions
Evaluation toolkit
An evaluation tool for co-saliency detection tasks
Multimodal conversational model
An end-to-end trained model capable of generating natural language responses integrated with object segmentation masks for interactive visual conversations
Multi-modal model evaluation suite
A collection of benchmarks to evaluate the multi-modal understanding capability of large vision language models.
LLM evaluation resource
A repository of papers and resources for evaluating large language models.
Model benchmark
A benchmarking platform for evaluating Chinese general-purpose models through anonymous, random battles
Evaluation framework
An evaluation toolkit for large vision-language models
Chinese understanding benchmark
Measures the understanding of massive multitask Chinese datasets using large language models
Crop models
Developing large language models for agricultural applications to improve crop yields and support rural development.
Model evaluator
Evaluates the capabilities of large multimodal models using a set of diverse tasks and metrics
Model fairness tool
A tool for evaluating and improving the fairness of machine learning models
Model evaluator
Evaluates and compares the performance of multimodal large language models on various tasks
Image model flaw detection
An investigation into the relationship between misleading images and hallucinations in large language models
Model interpreter
Provides tools to understand and interpret the decisions made by XGBoost models in machine learning
tsb0601/mmvp296
Visual model evaluation
An evaluation framework for multimodal language models' visual capabilities using image and question benchmarks.