FoolyourVLLMs

Attack framework

An attack framework to manipulate the output of large language models and vision-language models

[ICML 2024] Fool Your (Vision and) Language Model With Embarrassingly Simple Permutations

GitHub

14 stars
1 watching
2 forks
Language: Python
last commit: almost 3 years ago
adversarial-attacksllmsmcqvision-and-language

Related projects:

RepositoryDescriptionStars
yunqing-me/attackvlmAn adversarial attack framework on large vision-language models165
ys-zong/vlguardImproves safety and helpfulness of large language models by fine-tuning them using safety-critical tasks47
ys-zong/vl-iclA benchmarking suite for multimodal in-context learning models31
hfzhang31/a3flA framework for attacking federated learning systems with adaptive backdoor attacks23
ethz-spylab/rlhf_trojan_competitionDetecting backdoors in language models to prevent malicious AI usage109
yuxie11/r2d2A framework for large-scale cross-modal benchmarks and vision-language tasks in Chinese157
jeremy313/fl-wbcA defense mechanism against model poisoning attacks in federated learning37
junyizhu-ai/r-gapA tool to demonstrate and analyze attacks on private data in machine learning models using gradients34
zjunlp/knowlmA framework for training and utilizing large language models with knowledge augmentation capabilities1,251
weisong-ucr/mab-malwareAn open-source reinforcement learning framework to generate adversarial examples for malware classification models.41
lhfowl/robbing_the_fedThis implementation allows an attacker to directly obtain user data from federated learning gradient updates by modifying the shared model architecture.23
yiyangzhou/lureAnalyzing and mitigating object hallucination in large vision-language models to improve their accuracy and reliability.136
kaiyuanzh/flipA framework for defending against backdoor attacks in federated learning systems48
yuliang-liu/monkeyAn end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.1,849
jind11/textfoolerA tool for generating adversarial examples to attack text classification and inference models496