TouchStone

Vision Model Evaluator

A tool to evaluate vision-language models by comparing their performance on various tasks such as image recognition and text generation.

Touchstone: Evaluating Vision-Language Models by Language Models

GitHub

79 stars
3 watching
0 forks
Language: Python
last commit: over 2 years ago

Related projects:

RepositoryDescriptionStars
openai/simple-evalsEvaluates language models using standardized benchmarks and prompting techniques.2,059
huggingface/evaluateAn evaluation framework for machine learning models and datasets, providing standardized metrics and tools for comparing model performance.2,063
allenai/olmo-evalA framework for evaluating language models on NLP tasks326
pkunlp-icler/pca-evalAn open-source benchmark and evaluation tool for assessing multimodal large language models' performance in embodied decision-making tasks99
edublancas/sklearn-evaluationA tool for evaluating and visualizing machine learning model performance3
open-compass/vlmevalkitAn evaluation toolkit for large vision-language models1,514
modelscope/evalscopeA framework for efficiently evaluating and benchmarking large models308
vchitect/vbenchA benchmark suite for evaluating the performance of video generative models643
truskovskiyk/nima.pytorchAssesses and evaluates images using deep learning models335
tsb0601/mmvpAn evaluation framework for multimodal language models' visual capabilities using image and question benchmarks.296
zhourax/vegaDevelops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.33
ucsc-vlaa/vllm-safety-benchmarkA benchmark for evaluating the safety and robustness of vision language models against adversarial attacks.72
chenllliang/mmevalproA benchmarking framework for evaluating Large Multimodal Models by providing rigorous metrics and an efficient evaluation pipeline.22
huggingface/lightevalAn all-in-one toolkit for evaluating Large Language Models (LLMs) across multiple backends.879
vishaal27/sus-xThis is an open-source project that proposes a novel method to train large-scale vision-language models with minimal resources and no fine-tuning required.94