The editorial argues the problem isn't AI quality but the category of risk: AI produces convincingly wrong text with fabricated quotes and invented timeline details that diverge from officers' original notes. When the downstream artifact is a sworn document, this drift ends prosecutions and triggers misconduct inquiries, making the pipeline structurally unsafe regardless of review.
The FT reports the College of Policing's guidance citing concrete failure modes — fabricated witness quotes, invented details, and inconsistencies with body-worn-video footage. The guidance targets specifically the pipeline where machine-generated text gets sworn to as an officer's account, while permitting AI for triage, transcription, and translation.
The editorial warns that officers who reviewed AI-generated statements still demonstrably missed hallucinations, exposing a well-documented cognitive trap. This challenges the broader industry assumption that human review constitutes a defensible audit trail, particularly when the output becomes a legally binding artifact.
By surfacing this story to the HN front page (122 points), the submitter implicitly frames it as a cautionary case study about AI governance failures, signaling that policing's reckoning with LLM review limitations is relevant to the broader developer audience grappling with the same oversight problem.
The editorial emphasizes that the College of Policing carved out a narrow restriction — only court-bound sworn statements are off-limits, while internal triage, transcription, and translation remain permitted. This precision distinguishes the guidance from blanket AI prohibitions and reflects a mature understanding of where AI risk concentrates in the evidentiary chain.
The College of Policing — the professional standards body for officers in England and Wales — has told forces to stop using generative AI tools to draft statements and case material that will end up in court. The Financial Times reported the guidance this week after multiple forces quietly confirmed that officers had been pasting case notes into ChatGPT and similar tools to produce cleaner prose for statements, MG11 witness forms, and disclosure summaries.
The problem wasn't that the AI wrote badly. It's that it wrote convincingly wrong. Officials cite statements that contained fabricated quotes attributed to witnesses, invented timeline details, and small but legally fatal inconsistencies with body-worn-video footage. In at least one referenced case, an AI-polished statement materially diverged from the officer's original notes — the kind of drift that, in a Crown Court, ends a prosecution and starts a misconduct inquiry.
The guidance is not a blanket ban on AI in policing. It targets the specific pipeline where machine-generated text gets sworn to as an officer's own account. Internal triage tools, transcription, and translation are still on the table. What's off the table is the bit where a model writes the words an officer signs.
UK policing is the first major Western institutional user of LLMs to put a hard line in writing about court-bound output. That matters less for what it says about policing and more for what it says about the *category* of risk these tools create when they cross into evidentiary territory.
Every other industry currently rationalising 'human in the loop' as sufficient AI governance is about to discover that 'I reviewed it' is not a defensible audit trail when the downstream artifact is a sworn document. Officers reviewing AI-generated statements were, by all accounts, reviewing them. They were also, demonstrably, missing hallucinations. The cognitive trap is well-documented in the radiology and aviation literature: humans rubber-stamp fluent output far more readily than they rubber-stamp raw data. Policing just hit that wall first, in public, with prosecutions on the line.
The technical question this surfaces is provenance. A statement drafted by a person, edited by a person, and signed by a person has a single chain of authorship. A statement drafted by an LLM, edited by an officer, and signed by that officer has at least two — and disclosure rules in England and Wales (under the Criminal Procedure and Investigations Act) increasingly treat the model's contribution as material the defence is entitled to see. Which means: prompt, model version, temperature, system instructions, and any retrieval context become disclosable. Few forces have any of that logged. Hence the guidance.
Compare the U.S. trajectory. American federal courts have spent two years issuing standing orders requiring lawyers to disclose AI use in filings — driven by the Avianca brief and a steady drip of fabricated-citation sanctions. The UK is now extending that principle one rung up the pipeline, from advocacy to investigation. That's the more consequential move. Lawyers' filings are argument; police statements are evidence. The bar for evidence is — and has to be — higher.
Community reaction has been notably split. Defence solicitors on UK legal Twitter have called the guidance 'overdue by 18 months.' Serving officers in r/policeuk are split between relief (statement-writing is the grimmest part of the job, and AI was helping) and frustration (the guidance offers no replacement for the workload pressure that drove adoption). The Police Federation has asked for clarity on whether transcription tools that produce a draft for officer review count as 'generative' under the rule. The honest answer is: it depends on the model, and nobody has built the regulatory vocabulary to distinguish yet.
If you build, sell, or operate AI tooling that touches regulated workflows — legal, medical, financial advice, anything with a sworn or attested artifact at the end — the UK guidance is a preview of your compliance roadmap. Three concrete implications:
Provenance metadata is no longer optional. Every generation needs to carry, at minimum, the model identifier, the prompt, the retrieval context, and a hash of the output, stored alongside the artifact for the artifact's retention period. If your product can't produce that on request, you're building a future disclosure nightmare for your customers. Tools like Anthropic's and OpenAI's logging APIs cover part of this; the gap is usually in the application layer, where prompts are templated on the fly and never persisted.
'Human review' needs measurable teeth. A checkbox is not review. Forces that adopted AI statement drafting almost universally required officer sign-off, and it didn't catch the hallucinations. If you're shipping a product where review is the safety story, you need diff visualisation, edit tracking, and ideally adversarial highlighting of model-introduced claims — the kind of UI work that legal-tech vendors like Harvey and Spellbook have spent the last year building, and that generic ChatGPT wrappers conspicuously haven't.
The model choice matters legally now, not just performance-wise. Models with documented training data, stable version pinning, and documented evaluation benchmarks are going to be defensible in court. Models accessed through a consumer chat interface, where the underlying weights might silently change between Tuesday and Thursday, are not. Expect 'we used GPT-5 via the consumer app' to become the regulated-industry equivalent of 'we stored the passwords in plaintext.'
The UK guidance is the first domino, not the last. Expect similar moves from Scottish prosecutors, the Australian federal police, and — more slowly — U.S. state attorneys general within twelve months. The deeper shift is that the legal system is starting to treat LLM output as a distinct category of artifact, neither pure machine output nor pure human authorship, requiring its own evidentiary rules. Builders who get ahead of that — with logging, provenance, and review tooling that produces defensible records — will own the regulated verticals. Builders who don't will spend 2027 explaining to their customers why their product can't be used for anything that matters.
If you listen closely to UK cabinet ministers you can intuit that they are being horse whispered into handing over vast sums of taxpayer money to firms for AI who are promising solutions to the productivity gap (chasm?) that the UK is plagued by.I can say with certainty lots of money will be spent,
We can get ambitious and try to head toward a form of statement more probative than even an officer personally typing a report: Have them narrate the facts of the event and the reasons for their decisions as soon as possible after the incident, as a video. Additions and corrections made later would
At nearly £500 a year is an FT subscription worth it? Am I going to get invaluable stock tips that will cover the sub?!
I feel like this is where AI like -Are we thinking about how we’re using it, or???It seems like; there’s two kinds of data that might go into this, boilerplate and subjective information. Subjective information should be input by the police, because I would assert the specific wording matters. It ma
Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.
> [...] he had intervened at forces that were deploying commercially available AI tools before they had been properly assessed [...] “All forces have got a good policy on the use of Copilot,” Murray said. “All forces will have a policy that says, ‘Check everything that it produces’.”Not only are