Skip to main content

Search AIMS

Find pages, publications, events, course projects, and software. Results update as you type; press Enter to open the first result.

Safety

Alignment Faking (RL)

Published judgments of alignment-faking reasoning across model organisms, interventions and reinforcement-learning checkpoints in Anthropic's training-time mitigation study.

1,464items
4,278subjects
CC-BY-4.0; retain the released transcript obfuscation and canary as requested by the authors.license
safetydomain
textmodality
item-level responses released
Saturation status: Yes

Response matrix

Loading response matrix…

Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect