Search the evidence

Find the signal.

Search titles, impact clusters, countries, organizations and the full text of every analysis.

2 stories found

Reasoning tokens travel along unequal pathways around stereotype symbols before the paths feed into two consequential decision gates.
Technical failuresGlobal+4 clusters01

Reasoning models work harder against stereotypes, and the difference predicts biased outputs

A study in Nature Machine Intelligence proposes a new way to detect bias before it becomes a final answer. The Reasoning Model Implicit Association Test uses the number of reasoning tokens a model spends as a proxy for computational effort, adapting a human test that looks for slower responses when an association conflicts with a learned stereotype. Across o3-mini, DeepSeek-R1, gpt-oss-20b, and Qwen3-8B, models generally used more reasoning tokens for association-incompatible pairings than for compatible ones. Claude 3.7 Sonnet showed a reversed pattern that the researchers linked to explicit internal attention to bias and stereotypes. The important result is not only the token difference. Those patterns predicted bias in two downstream word-association and decision-making tasks, giving the measure convergent validity. The interpretation still needs restraint. Reasoning tokens are a proxy for computational effort, not a window into humanlike implicit attitudes, consciousness, or motive. Model traces can also reflect training style and explicit safety behavior. The study nevertheless shows why final-answer audits are incomplete. When AI influences hiring, health, education, credit, or public services, evaluators should test internal process signals alongside outcomes, verify that the signal predicts real decisions, compare demographic contexts, and disclose where the proxy stops being reliable.

6 min
A luminous forensic scanner assigns conflicting human, AI, and mixed labels to the same edited manuscript while a locked penalty stamp waits behind an evidence folder.
Technical failuresGlobal+4 clusters02

AI detectors improve sharply, but mixed human-machine writing still breaks the verdict

Nature reports that a new generation of commercial AI-text detectors performs far better than earlier systems on clearly human or clearly machine-generated passages. Pangram advertises 99.98 percent accuracy and GPTZero advertises 99 percent, while independent tests found very low false-positive rates on selected human-written datasets. Adoption is spreading through publishing, conferences, preprint tools, and universities. The hard case is mixed authorship. Style imitation and humanizer tools increase false negatives, passages under 50 words reduce performance, different detectors can disagree, and a score can change when a sentence is moved into a larger segment. A label near 100 percent AI does not mean every word was generated, and vendor claims for the newest models inevitably arrive before independent validation. One technical study reported that substantially AI-modified human student essays were still labeled fully human 41 percent of the time. Detectors can prioritize review and expose undisclosed use. They cannot establish intent, contribution, or misconduct on their own. Any consequential decision needs declared rules, original evidence, human investigation, and appeal.

5 min