Safety

Alignment Faking (RL)

Alignment Faking (RL) records whether a model pretends to go along with a request it would otherwise refuse, and whether that changes with how far its training has progressed and which kind of user it is told it is talking to.

244items
155subjects
CC-BY-4.0license
safetydomain
textmodality
item-level responses released
Saturation status: Yes

Response matrix

Rasch analysis p = σ(θ − z + c)

2,138,969 responses, 80/20 split over cells · 155 subjects · 244 items · 64 conditions

AUC train
0.927
AUC test
0.927

Loading session strips…

Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect