Safety
XSTest
XSTest checks whether models refuse harmless prompts that merely sound dangerous, pairing them with genuinely unsafe lookalikes so that both excessive caution and unsafe compliance show up.
450items
81subjects
Apache-2.0license
safetydomain
textmodality
item-level responses released
Saturation status: Yes
Response matrix
Responses come from Stanford CRFM's HELM Safety leaderboard (v1.17.0), run without a system prompt; the original evaluation of Rottger et al. labels the same prompts on a 4-level ordinal scale from two human annotators and a GPT-4 judge. Complements the standalone xstest build.
Rasch analysis p = σ(θ − z)
36,450 responses, 80/20 split over cells · 81 subjects · 450 items
AUC train
0.924
AUC test
0.899
Loading session strips…
Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect