Safety
SimpleSafetyTests
SimpleSafetyTests checks whether a model responds safely to blunt requests touching severe harms such as self harm, violence and illegal activity, with a classifier marking each reply safe or unsafe.
100items
81subjects
Apache-2.0license
safetydomain
textmodality
item-level responses released
Saturation status: Yes
Response matrix
Responses come from Stanford CRFM's HELM Safety leaderboard (v1.17.0), run without a system prompt; the original evaluation of Vidgen et al. tested each model under two system-prompt conditions.
Rasch analysis p = σ(θ − z)
8,100 responses, 80/20 split over cells · 81 subjects · 100 items
24 of 81 subjects answered every item alike, so their θ is unbounded and sits pegged to the column edge.
AUC train
0.958
AUC test
0.880
Loading session strips…
Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect