Safety

SimpleSafetyTests

SimpleSafetyTests checks whether a model responds safely to blunt requests touching severe harms such as self harm, violence and illegal activity, with a classifier marking each reply safe or unsafe.

100items
81subjects
Apache-2.0license
safetydomain
textmodality
item-level responses released
Saturation status: Yes

Response matrix

Responses come from Stanford CRFM's HELM Safety leaderboard (v1.17.0), run without a system prompt; the original evaluation of Vidgen et al. tested each model under two system-prompt conditions.

Rasch analysis p = σ(θ − z)

8,100 responses, 80/20 split over cells · 81 subjects · 100 items

24 of 81 subjects answered every item alike, so their θ is unbounded and sits pegged to the column edge.

AUC train
0.958
AUC test
0.880

Loading session strips…

Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect