Search the evidence

Find the signal.

Search titles, impact clusters, countries, organizations and the full text of every analysis.

4 stories found

A housing-court appeal reveals unstable fabricated citations under forensic light beside apartment keys and an eviction notice.
Law & informationUnited States+3 clusters01

AI did not cause the eviction loss. It made a weak appeal look legally real

WKRN reports that a Nashville renter representing himself lost an appeal of his eviction after submitting a filing with AI-fabricated legal support. The opinion said the appeal used real case names but attached wrong dates, fabricated quotations, invented citations, and a false rendering of Tennessee landlord law. The court described the material as having hallmarks of artificial intelligence and affirmed the landlord's judgment. AI was not the sole cause of the loss. The tenant was behind on rent, failed to provide a transcript or statement of evidence, and relied heavily on a national uniform landlord-tenant act that Tennessee never adopted. That nuance makes the case more instructive. A model can turn an already weak position into a confident, finished-looking argument without fixing the underlying facts or procedure. The access-to-justice gap also matters: renters who cannot obtain counsel may choose between navigating the system alone and trusting a tool that can manufacture authority.

5 min
An empty oversight chair sits beside automated congressional workflows processing speeches, legislative summaries, and constituent mail.
Law & informationUnited States+3 clusters02

Congress is handing daily work to chatbots faster than it writes the rules

The Washington Post reports that AI chatbots are spreading through Congress for work including speeches, legislative summaries, and sorting constituent mail while oversight remains limited. The adoption matters because these systems can influence what lawmakers read, say, and send under the authority of public office. A useful governance framework must cover more than whether a staff member used an approved tool. It should define which information can enter a model, who checks factual claims and citations, how constituents are told when automation materially shaped a response, how records are retained, and who corrects an error. Public reporting does not establish that every office uses the same tools or practices, and Congress is not one uniform organization. The signal is institutional: deployment can become routine office work before rules make responsibility visible. A chatbot can draft a sentence, but it cannot accept electoral, ethical, or legal accountability for it.

5 min
A rising AI capability graph is balanced against a warning signal for confident uncertainty and factual hallucinations.
Cognition & learningGlobal+4 clusters03

Claude Opus 5 is more capable—and slightly more prone to factual hallucinations

Anthropic’s system card reports broad gains for Claude Opus 5 in agentic coding, computer use, long-horizon knowledge work, and scientific reasoning. It also documents a reliability tension: on one closed-book factuality benchmark, accuracy was 11% higher than Opus 4.8 while the hallucination rate was 6% higher. Anthropic found cases where the model confidently answered despite internal uncertainty, even as its automated alignment scores and prompt-injection robustness improved.

4 min
A warped molecular structure resolving into a physically constrained chemical lattice.
Work & marketsGlobal+3 clusters04

Liu et al., “Integrating chemical priors and physical laws to mitigate hallucinations in structure-based drug design”

The NUS/Harbin-led team identifies a domain-specific form of generative-AI hallucination: molecular candidates can receive strong predicted binding scores while violating basic chemistry or producing physically impossible atomic arrangements. Its DrugRPG framework incorporates chemical-foundation-model priors and differentiable physical constraints during molecule generation, reducing severe steric clashes by 65.4% relative to the reported state-of-the-art baseline and increasing by 28.6% the share of generated candidates meeting combined potency, stability, and synthetic-feasibility criteria.

2 min