Safety
Alignment Faking (RL)
Alignment Faking (RL) records whether a model pretends to go along with a request it would otherwise refuse, and whether that changes with how far its training has progressed and which kind of user it is told it is talking to.
244items
155subjects
CC-BY-4.0license
safetydomain
textmodality
item-level responses released
Saturation status: Yes
Response matrix
Rasch analysis p = σ(θ − z + c)
2,138,969 responses, 80/20 split over cells · 155 subjects · 244 items · 64 conditions
AUC train
0.927
AUC test
0.927
Loading session strips…
Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect