Awesome Lists

awesome-llm-security

by corca-ai

awesome listpushed almost 2 years ago

A curation of awesome tools, documents and projects about LLM Security.

AI summary

LLM Security Resources

Curation of resources and tools related to security vulnerabilities in large language models.

stars
985
forks
99
watching
33
awesome lists
2
entries
96
View on GitHub

Embed the badge

Show how many awesome lists link to your project. The count updates automatically.

Awesome Lists badge
Markdown
[![Awesome Lists Badge](https://awesome.facts.dev/shield/corca-ai/awesome-llm-security/links.svg)](https://awesome.facts.dev/awesome/corca-ai/awesome-llm-security)
HTML
<a href="https://awesome.facts.dev/awesome/corca-ai/awesome-llm-security"><img src="https://awesome.facts.dev/shield/corca-ai/awesome-llm-security/links.svg" alt="Awesome Lists Badge" /></a>
Image URL
https://awesome.facts.dev/shield/corca-ai/awesome-llm-security/links.svg

What's in the list

96 links in 12 sections, with live GitHub stats.activeno commit in 2y

Papers / White-box attack

  • [paper]

    "Visual Adversarial Examples Jailbreak Large Language Models", 2023-06, AAAI(Oral) 24, ,

  • [paper]

    "Are aligned neural networks adversarially aligned?", 2023-06, NeurIPS(Poster) 23, ,

  • [paper]

    "(Ab)using Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs", 2023-07,

  • [paper]

    "Universal and Transferable Adversarial Attacks on Aligned Language Models", 2023-07, ,

  • [paper]

    "Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models", 2023-07, ,

  • [paper]

    "Image Hijacking: Adversarial Images can Control Generative Models at Runtime", 2023-09, ,

  • [paper]

    "Weak-to-Strong Jailbreaking on Large Language Models", 2024-04, ,

Papers / Black-box attack

  • [paper]

    "Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection", 2023-02, AISec@CCS 23

  • [paper]

    "Jailbroken: How Does LLM Safety Training Fail?", 2023-07, NeurIPS(Oral) 23,

  • [paper]

    "Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models", 2023-07,

  • [paper]

    "Effective Prompt Extraction from Language Models", 2023-07, ,

  • [paper]

    "Multi-step Jailbreaking Privacy Attacks on ChatGPT", 2023-04, EMNLP 23, ,

  • [paper]

    "LLM Censorship: A Machine Learning Challenge or a Computer Security Problem?", 2023-07,

  • [paper]

    "Jailbreaking chatgpt via prompt engineering: An empirical study", 2023-05,

  • [paper]

    "Prompt Injection attack against LLM-integrated Applications", 2023-06,

  • [paper]

    "MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots", 2023-07, ,

  • [paper]

    "GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher", 2023-08, ICLR 24, ,

  • [paper]

    "Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities", 2023-08,

  • [paper]

    "Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs", 2023-08,

  • [paper]

    "Detecting Language Model Attacks with Perplexity", 2023-08,

  • [paper]

    "Open Sesame! Universal Black Box Jailbreaking of Large Language Models", 2023-09, ,

  • [paper]

    "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!", 2023-10, ICLR(oral) 24,

  • [paper]

    "AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models", 2023-10, ICLR(poster) 24, , ,

  • [paper]

    "Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations", 2023-10, CoRR 23, ,

  • [paper]

    "Multilingual Jailbreak Challenges in Large Language Models", 2023-10, ICLR(poster) 24,

  • [paper]

    "Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation", 2023-11, SoLaR(poster) 24,

  • [paper]

    "DeepInception: Hypnotize Large Language Model to Be Jailbreaker", 2023-11,

  • [paper]

    "A Wolf in Sheep’s Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily", 2023-11, NAACL 24,

  • [paper]

    "AutoDAN: Automatic and Interpretable Adversarial Attacks on Large Language Models", 2023-10,

  • [paper]

    "Language Model Inversion", 2023-11, ICLR(poster) 24,

  • [paper]

    "An LLM can Fool Itself: A Prompt-Based Adversarial Attack", 2023-10, ICLR(poster) 24,

  • [paper]

    "GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts", 2023-09,

  • [paper]

    "Many-shot Jailbreaking", 2024-04,

  • [paper]

    "Rethinking How to Evaluate Language Model Jailbreak", 2024-04,

Papers / Backdoor attack

  • [paper]

    "BITE: Textual Backdoor Attacks with Iterative Trigger Injection", 2022-05, ACL 23,

  • [paper]

    "Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models", 2023-05, EMNLP 23,

  • [paper]

    "Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection", 2023-07, NAACL 24,

Papers / Fingerprinting

  • [paper]

    "Instructional Fingerprinting of Large Language Models", 2024-01, NAACL 24

  • [paper]

    "TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification", 2024-02, ACL 24 (findings)

  • [paper]

    "LLMmap: Fingerprinting For Large Language Models", 2024-07,

Papers / Defense

  • [paper]

    "Baseline Defenses for Adversarial Attacks Against Aligned Language Models", 2023-09,

  • [paper]

    "LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked", 2023-08, ICLR 24 Tiny Paper, ,

  • [paper]

    "Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM", 2023-09, ,

  • [paper]

    "Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models", 2023-12,

  • [paper]

    "AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks", 2024-03,

  • [paper]

    "Protecting Your LLMs with Information Bottleneck", 2024-04,

  • [paper]

    "PARDEN, Can You Repeat That? Defending against Jailbreaks via Repetition", 2024-05, ICML 24,

  • [paper]

    “Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs”, 2024-06,

  • [paper]

    "Improving Alignment and Robustness with Circuit Breakers", 2024-06, NeurIPS 24, ,

Papers / Platform Security

  • [paper]

    "LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI’s ChatGPT Plugins", 2023-09,

Papers / Survey

  • [paper]

    "Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks", 2023-10, ACL 24,

  • [paper]

    "Security and Privacy Challenges of Large Language Models: A Survey", 2024-02,

  • [paper]

    "Breaking Down the Defenses: A Comparative Survey of Attacks on Large Language Models", 2024-03,

  • [paper]

    "Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs)", 2024-07,

Benchmark

  • [paper]

    "JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models", 2024-03,

  • [paper]

    "AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents", 2024-06, NeurIPS 24,

  • [paper]

    "AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents", 2024-10,

Tools

  • Plexiglass

    : a security toolbox for testing and safeguarding LLMs

  • PurpleLlama

    : set of tools to assess and improve LLM security

  • Rebuff

    : a self-hardening prompt injection detector

  • Garak

    : a LLM vulnerability scanner

  • LLMFuzzer

    : a fuzzing framework for LLMs

  • LLM Guard

    : a security toolkit for LLM Interactions

  • Vigil

    : a LLM prompt injection detection toolkit

  • jailbreak-evaluation

    : an easy-to-use Python package for language model jailbreak evaluation

  • Prompt Fuzzer

    : the open-source tool to help you harden your GenAI applications

  • WhistleBlower

    : open-source tool designed to infer the system prompt of an AI agent based on its generated text outputs

Articles

Other Awesome Projects

Other Useful Resources

More related projects

Add a GitHub project

Missing a project or an awesome list? Paste its GitHub URL and we fetch it right away.