
Anthropic's lie detector scored 0.95 at home and stumbled outside the test
Anthropic's Alignment Science team trained lie detectors using roughly 200,000 labeled examples from 12 settings and eight model families. In-distribution performance rose from an AUROC of 0.60 to 0.95, but cross-category transfer reached only about 0.70 to 0.75, and larger models prompted as judges often beat the fine-tuned detectors. The research also exposes a label problem: about one quarter of labels changed during a GPT-5-assisted cleaning process, particularly around ambiguous behavior such as sycophancy. Third-person monitoring worked better than asking a model to report on itself. The team released its datasets and explicitly limits its conclusion to controlled settings rather than production behaviors such as alignment faking or reward hacking. The result is a valuable negative finding. A detector that excels only on familiar lies is not a universal truth machine, and institutions must not convert an uncertain score into punishment without evidence and appeal.
