The detector learned the neighborhood
Training on examples from the same category produced a dramatic gain. But moving to a different category of deceptive behavior exposed a much smaller and inconsistent advantage. Larger models generally performed better, although the relationship was not monotonic.
That gap matters because real deployments encounter unfamiliar motives, contexts, and behaviors. A model may learn surface regularities in a dataset without discovering a transferable representation of deception.
Ground truth was part of the failure
Deception labels are not simple sensor readings. The researchers found substantial instability when cleaning judgments and highlighted sycophancy as especially ambiguous. A detector can appear precise while inheriting disagreement about what counts as a lie.
The responsible use is therefore investigative and limited: surface cases for review, compare independent evidence, retain uncertainty, and prohibit automatic sanctions. The team's release of data enables other researchers to test those limits rather than treating one score as a product claim.
Go to the source
Read the evidence behind this analysis. External links open in a new tab.
Anthropic Alignment Science — Fine-tuned lie detectors failed to generalize


