MMCBench

Model robustness tester

A benchmarking framework designed to evaluate the robustness of large multimodal models against common corruption scenarios

GitHub

27 stars
5 watching
0 forks
Language: Python
last commit: over 2 years ago

Related projects:

RepositoryDescriptionStars
freedomintelligence/mllm-benchEvaluates and compares the performance of multimodal large language models on various tasks56
felixgithub2017/mmcuMeasures the understanding of massive multitask Chinese datasets using large language models87
hendrycks/robustnessEvaluates and benchmarks the robustness of deep learning models to various corruptions and perturbations in computer vision tasks.1,030
borealisai/advertorchA toolbox for researching and evaluating robustness against attacks on machine learning models1,311
0x0mar/smodA modular framework for testing and exploiting Modbus protocol vulnerabilities in industrial control systems74
ailab-cvc/seed-benchA benchmark for evaluating large language models' ability to process multimodal input322
open-compass/mmbenchA collection of benchmarks to evaluate the multi-modal understanding capability of large vision language models.168
robustbench/robustbenchA standardized benchmark for measuring the robustness of machine learning models against adversarial attacks682
sww9370/rocbertA pre-trained Chinese language model designed to be robust against maliciously crafted texts15
chenllliang/mmevalproA benchmarking framework for evaluating Large Multimodal Models by providing rigorous metrics and an efficient evaluation pipeline.22
guanghelee/neurips19-certificates-of-robustnessProvides a framework for computing tight certificates of adversarial robustness for randomly smoothed classifiers.17
tsb0601/mmvpAn evaluation framework for multimodal language models' visual capabilities using image and question benchmarks.296
fuxiaoliu/mmcDevelops a large-scale dataset and benchmark for training multimodal chart understanding models using large language models.87
google-research/robustness_metricsA toolset to evaluate the robustness of machine learning models466
vernamlab/medusaAutomated attack synthesis tool for discovering vulnerabilities in CPU architecture and cryptographic protocols18