Safety

Emoji Attack

Emoji-Attack / EasyJailbreak results: per-(target LLM, harmful behavior, attacker) binary jailbroken verdicts. 10 attack families x 10 target LLMs over AdvBench-style harmful behaviors.

3,376items
10subjects
unknownlicense
safetydomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

Emoji Attack response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1Mistral-7B-Instruct0.5288
  2. 2Vicuna-13B0.4098
  3. 3GPT-3.5-Turbo0.3868
  4. 4Qwen-7B-Chat0.368
  5. 5Vicuna-7B0.3594
  6. 6InternLM-7B0.3234
  7. 7ChatGLM3-6B0.2027
  8. 8GPT-40.1676
  9. 9Llama-2-7B-Chat0.0913
  10. 10Llama-2-13B-Chat0.047