Awesome Lists

llm-sp

by chawins

awesome listPythonpushed almost 2 years ago

Papers and resources related to the security and privacy of LLMs 🤖

AI summary

LLM Security Resources

A collection of papers and resources on the security and privacy of Large Language Models (LLMs), providing insights into vulnerabilities and potential attack vectors.

stars
454
forks
34
watching
17
awesome list
1
entries
24
View on GitHubchawins.github.io/llm-sp

Embed the badge

Show how many awesome lists link to your project. The count updates automatically.

Awesome Lists badge
Markdown
[![Awesome Lists Badge](https://awesome.facts.dev/shield/chawins/llm-sp/links.svg)](https://awesome.facts.dev/awesome/chawins/llm-sp)
HTML
<a href="https://awesome.facts.dev/awesome/chawins/llm-sp"><img src="https://awesome.facts.dev/shield/chawins/llm-sp/links.svg" alt="Awesome Lists Badge" /></a>
Image URL
https://awesome.facts.dev/shield/chawins/llm-sp/links.svg

What's in the list

24 links in 7 sections, with live GitHub stats.activeno commit in 2y

Vulnerabilities / Jailbreak

  • red-teaming LLM with LLM paper

    On a high level, the idea is similar to the . They train an LLM (called AdvPrompter) to automatically jailbreak a target LLM. AdvPrompter is trained on rewards (logprob of "Sure, here is...") of the target model. The result is good but maybe not as good as the at the time. However, there are a lot of interesting technical contributions

Vulnerabilities / Privacy

  • Ippolito et al. (2023)

    The authors also advocate for approximate memorization instead of verbatim, similar to

  • GitHub - iamgroot42/mimir: Python package for measuring memorization in LLMs.

    Library of MIAs on LLMs, including Min-k%, zlib, reference-based attack (Ref), neighborhood

  • n-gram overlap

    Temporal shift in member vs non-member test samples contributes to an overestimated MIA success rate. The authors measure this distribution shift with

  • the neighborhood attack

    Ask target LLM to select a verbatim text from a copyrighted book/ArXiv paper in a multiple-choice format (four choices). The other three options are close LLM-paraphrased texts. The core idea is similar to , but using MCQA instead of loss computation. The authors also debias/normalize for effects of the answer ordering, which LLMs are known to have trouble with

  • MemFree

    Propose SHIELD defense which works by (1) detecting copyrighted content in model’s output, (2) verifying it with internet search, and (3) few-shot prompting to let the model refuse or answer as appropriate (summary and QA are ok, but not verbatim). Defense seems very effective and is better than

Defenses / Against Jailbreak & Prompt Injection

Other resources / People/Orgs/Blog to Follow

  • Blog

    ChatGPT Plugin Exploit Explained: From Prompt Injection to Accessing Private Data [ ]

  • Blog

    Advanced Data Exfiltration Techniques with ChatGPT [ ]

  • Blog

    Hacking Google Bard - From Prompt Injection to Data Exfiltration [ ]

  • Blog

    Securing LLM Systems Against Prompt Injection [ ]

  • X

    Meme [ ]

Other resources / Resource Compilation

Other resources / Open-Source Projects

Logistics / Prompt Injection vs Jailbreak vs Adversarial Attacks

  • jailbreakchat.com

    is a method for bypassing safety filters, system instructions, or preferences. Sometimes asking the model directly (like prompt injection) does not work so more complex prompts (e.g., ) are used to trick the model

More related projects

Add a GitHub project

Missing a project or an awesome list? Paste its GitHub URL and we fetch it right away.