vllm-safety-benchmark

Vision model safety test

A benchmark for evaluating the safety and robustness of vision language models against adversarial attacks.

[ECCV 2024] Official PyTorch Implementation of "How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs"

GitHub

72 stars
4 watching
3 forks
Language: Python
last commit: almost 3 years ago
adversarial-attacksbenchmarkdatasetsllmmultimodal-llmrobustnesssafetyvision-language-model

Related projects:

RepositoryDescriptionStars
aifeg/benchlmmAn open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models84
ucsc-vlaa/sight-beyond-textAn implementation of a multimodal LLM training paradigm to enhance truthfulness and ethics in language models19
ys-zong/vlguardImproves safety and helpfulness of large language models by fine-tuning them using safety-critical tasks47
safellama/plexiglassA toolkit to detect and protect against vulnerabilities in Large Language Models.122
leondz/lm_risk_cardsA set of tools and guidelines for assessing the security vulnerabilities of language models in AI applications28
pku-alignment/safety-gymnasiumA unified benchmark for safe reinforcement learning algorithms and environments.410
howiehwong/trustllmA toolkit for assessing trustworthiness in large language models491
hendrycks/robustnessEvaluates and benchmarks the robustness of deep learning models to various corruptions and perturbations in computer vision tasks.1,030
opengvlab/visionllmA large language model designed to process and generate visual information956
byungkwanlee/collavoDevelops a PyTorch implementation of an enhanced vision language model93
mlpc-ucsd/blivaA multimodal LLM designed to handle text-rich visual questions270
baaivision/eveA PyTorch implementation of an encoder-free vision-language model that can be fine-tuned for various tasks and modalities246
dvlab-research/lisaA system that uses large language models to generate segmentation masks for images based on complex queries and world knowledge.1,923
kaiyangzhou/dassl.pytorchA PyTorch toolbox for supporting research and development of domain adaptation, generalization, and semi-supervised learning methods in computer vision.1,236
ailab-cvc/seed-benchA benchmark for evaluating large language models' ability to process multimodal input322