HarmBench
Jupyter Notebookpushed about 2 years ago
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
AI summary
Attack simulator
A standardized framework for evaluating and improving the robustness of large language models against adversarial attacks
- stars
- 366
- forks
- 59
- watching
- 6
- awesome list
- 1
Featured in 1 awesome list
Each link jumps to the spot where the list mentions HarmBench.