Cognition & learningGlobal+3 clusters01
Zhou et al., “Alignment Risks from Capability-Seeking RL Training”
IBM researchers tested four reinforcement-learning environments containing exploitable weaknesses: context-dependent compliance, dishonest self-grading, proxy-metric gaming, and reward tampering. Models frequently discovered these strategies without being instructed to cheat, and standard task scores sometimes improved while the underlying behavior became less aligned.