How we read the signal

Analysis frame

Evidence level

Primary-source evidence

Analytical lens

The failure arose at the boundary between fluent summarization and evidentiary provenance: professional responsibility remained human, but the tool and workflow made unsupported factual reconstruction easy to mistake for the record.

Affected groups
  • Defendants whose liberty depends on accurate appellate records and effective counsel
  • Lawyers and public defenders using generative tools under time and cost pressure
  • Judges and court staff responsible for correcting contaminated filings
  • Clients and the public who rely on the integrity of legal records
What remains unknown
  • Which model settings, prompts, transcript format, or context limitations produced the fabricated testimony
  • Whether source-linked legal tools would have prevented or merely exposed the error sooner
  • How frequently AI-assisted factual hallucinations enter filings without being detected
  • Which disclosure and verification rules improve accuracy without creating empty compliance language
Second-order effects to watch
  • Courts may require certificates describing AI use and factual verification in high-stakes filings
  • Legal software vendors may be expected to preserve record-level provenance for every generated factual assertion
  • Insurers and law firms could restrict general chatbots for transcript summarization after professional-liability claims
  • Overbroad bans may deny smaller practices useful tools while leaving ordinary human errors unchecked

The hallucination became part of a criminal record

The court's order identifies wholly fabricated witnesses, false testimony attributed to other witnesses, and misrepresented authority. The lawyer acknowledged using ChatGPT and failing to verify the material before filing.

Because the filing concerned a murder conviction, correction required more than deleting a bad sentence. The court struck the briefs and ordered replacement counsel so the appeal could proceed on a clean record.

A fluent summary is not an evidentiary map

General-purpose models optimize for plausible continuation. A legal reader, by contrast, needs every factual claim and quotation anchored to a specific page, speaker, and source document.

Without that chain, a polished brief can make invention harder to notice because professional tone is mistaken for evidentiary reliability.

Verification should be a product constraint

Responsibility remains with the signing lawyer, but legal tools can make responsible review easier or harder. Source spans, unsupported-claim alerts, quotation locks, and prefiling audits turn the instruction to verify into an observable process.

Courts should test whether those controls reduce errors rather than relying on broad disclosure language that says little about what was actually checked.

  • Link each generated factual claim to the authoritative record.
  • Block direct quotations when no exact source span exists.
  • Require a prefiling report of unsupported claims and unresolved conflicts.
  • Preserve the human reviewer and the verification timestamp for audit.
Primary trail

Go to the source

Read the evidence behind this analysis. External links open in a new tab.

Reuters — ChatGPT invented fake police testimony in murder appeal Supreme Court of New Mexico — Dispositional order of direct contempt