HarmBench: A standardized evaluation framework for automated red teaming and robust refusal

Systems online

This research is currently only available at its source.

You can find the research at the below link.
Feel free to contact Gray Swan with any questions or comments.