Safety

XSTest

XSTest checks whether models refuse harmless prompts that merely sound dangerous, pairing them with genuinely unsafe lookalikes so that both excessive caution and unsafe compliance show up.

450items
81subjects
Apache-2.0license
safetydomain
textmodality
item-level responses released
Saturation status: Yes

Response matrix

Responses come from Stanford CRFM's HELM Safety leaderboard (v1.17.0), run without a system prompt; the original evaluation of Rottger et al. labels the same prompts on a 4-level ordinal scale from two human annotators and a GPT-4 judge. Complements the standalone xstest build.

Rasch analysis p = σ(θ − z)

36,450 responses, 80/20 split over cells · 81 subjects · 450 items

AUC train
0.924
AUC test
0.899

Loading session strips…

Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect