On 12 June 2026 the Solicitors Regulation Authority added a passage to its effective supervision guidance covering work done with AI. It is three sentences long and reads nothing like a technology rule. Firms must ensure that “outputs produced with the assistance of AI are subject to appropriate human review, scrutiny and professional judgement”, and an authorised individual has to retain ultimate responsibility for anything delivered with AI assistance. For two years the profession has argued about which tools are accurate enough for legal work. The regulator has answered a different question entirely, and its answer is that accuracy was never the thing under regulation. Supervision was.
Mia Leslie, who advises the Law Society on policy covering technology law and digital transformation, set out the mechanism behind that shift in a July explainer written for solicitors. It is the clearest short account of hallucination I have seen from a professional body, and almost none of what makes it alarming is specific to law.
Why the tool cannot be fixed into reliability
Leslie is blunt about what sits under the interface. The words you type are broken into tokens, mathematical stand-ins for fragments of language. The system weighs the statistical relationships between those tokens against the enormous corpus it was trained on, then returns the continuation the numbers favour. She compares it to a very advanced autocomplete, which sounds dismissive until you notice what the comparison rules out. There is no appraisal step anywhere in the sequence. Nothing pauses to ask whether a claim holds, because nothing in the process carries a concept of truth against which to check it.
Two design choices compound this. Models are generally built to sound encouraging and to produce a guess rather than concede they have nothing. Faced with a prompt they cannot serve from anything solid, they reach into adjacent material and assemble something helpful-sounding instead of flagging the gap.
Strategic Reality: A hallucination is not the system failing. Leslie’s formulation is that “a hallucination isn’t a malfunction because the technology is not designed to be accurate”. Every control you build on the assumption that this is a defect awaiting a patch is aimed at the wrong thing.
The consequence for anyone budgeting AI programmes is unwelcome. You cannot procure your way out of this by waiting for the reliable version, because unreliability is not a stage the technology is passing through.
Format is the part that fools you
The genuinely dangerous property is not that models invent things. It is that they invent things in exactly the right shape.
Solicitors work with a narrow set of source types, each with its own conventions: reported cases, statutory instruments, regulatory material, commentary. A citation might follow the OSCOLA referencing style. A model trained on quantities of legal text has absorbed that pattern thoroughly, so a fabricated authority arrives dressed identically to a real one, in a format whose familiarity is precisely what signals reliability to a reader glancing down a page.
That is why fake citations keep reaching court documents rather than being caught at the desk. As Leslie puts it, there are no obvious red flags to watch for, because a hallucination will not look wrong. Dame Victoria Sharp, who presides over the King’s Bench Division, described AI-generated text, in a speech about how lawyers use it before the courts, as material that “can be astonishingly and beguilingly fluent, but false”. Fluency is the failure mode, not a mitigating feature of it.
| What changed | Detail | Source |
|---|---|---|
| Supervision duty extended | Guidance updated 12 June 2026 to cover supervision of AI assisted and AI generated work | SRA effective supervision guidance |
| Conduct rules unchanged | ”The rules and principles of professional conduct are not disapplied as a result of using an AI tool” | Law Society, 21 July 2026 |
| Sanction range | Public admonishment, referral to the regulator, or a finding of contempt of court | Law Society, 21 July 2026 |
| Systemic exposure | Fictitious authorities reaching a judgment could shape how the law develops, then feed the next round of training data | Law Society, 21 July 2026 |
That last row is the one worth sitting with. Because England and Wales run on evolving precedent, an invented case that survives into a judgment does not simply harm one client. It enters the body of law that later work is built from, and potentially the corpus that later models learn from. It is a contamination risk with a feedback loop attached, and it is why the legal sector’s version of this problem is a leading indicator for every profession that runs on citable authority: medicine, audit, engineering, regulatory affairs.
The duty was never delegable
The regulatory logic here is simpler than most AI governance debates allow.
A solicitor owes duties to the client, the court and third parties, and those duties include not misleading anyone or passing off unreliable material as sound. None of that is suspended because a machine drafted the paragraph. Leslie’s framing is that responsibility for confirming accuracy stays with the solicitor, full stop, which means submitting a fictional authority is indefensible regardless of how it got into the document.
What the SRA update does is convert that principle into something a regulator can inspect. Before June, a firm asked about AI could reasonably discuss tool selection and staff guidance. Now the question is what supervision of AI-assisted work looks like in practice inside that firm, and the honest answer for many organisations is that it looks like nothing in particular. Individual fee earners verify when they remember to.
Critical Context: Disclosure policies do not satisfy this. Logging that a tool was used records provenance, not scrutiny. The supervisory question is whether a competent human examined the output and can be shown to have done so.
What each review layer can actually see
| Layer | What it catches | What passes straight through |
|---|---|---|
| Automated citation checker | Authorities that do not exist in the database | Real cases cited for propositions they do not support |
| Supervising partner review | Weak argument, wrong strategy, poor drafting | A plausible authority the partner also has no time to pull |
| Client sign-off | Commercial misalignment | Everything technical, by definition |
| Post-matter file audit | Patterns, after the exposure has already occurred | Nothing useful in time to prevent it |
Read down the right-hand column and the shape of the gap is clear. The one control that reliably works, opening the source and reading it, is the one that consumes exactly the time the tool was bought to save.
Reality Check: If nobody in your organisation can say how many AI-assisted outputs were verified at source last month, you do not have a supervision framework. You have a policy describing one.
Building the supervision layer
Treat this as an operating change rather than a policy document. Three tiers, by where an organisation currently sits.
If AI use is informal or ungoverned. Establish what is actually happening before writing rules about it. Shadow use is consistently higher than leadership believes, a gap our reporting on law firm leaders and shadow AI has already traced. Then set one non-negotiable: any external-facing citation, figure or authority produced with AI assistance is verified at source by a named person before it leaves the building. Record the verification, not the tool use.
If AI use is governed but unmeasured. Distinguish tool categories in policy. A platform built for legal services has been trained with jurisdictional specificity in mind and can draw on curated material. A general-purpose assistant serving everything from code to recipes has no such constraint. Both feel the same to type into, which is why staff treat them as interchangeable. Separate them explicitly, and separate again between a public consumer version and a contracted deployment where you can configure behaviour and control what happens to inputs.
If supervision is established. Move to sampling and evidence. Audit a defined proportion of AI-assisted output against sources each month, track the verification failure rate as a metric with a baseline, and report it upward like any other risk indicator. The UK Jurisdiction Taskforce has argued that not using AI could itself become negligent, which means the destination is not avoidance. It is demonstrable control.
Implementation Note: Budget the verification hours explicitly in the business case. A programme that books the drafting time saved and leaves checking as absorbed overhead will show a return that the supervision requirement then quietly deletes.
Four things this breaks that nobody costs
The efficiency case inverts on high-stakes work. Verification effort scales with the consequence of being wrong, not with the length of the output. On routine material the time saved is real. On the matters that carry the firm’s exposure, thorough checking can cost more than drafting from scratch would have, because you are now reading unfamiliar material sceptically rather than building familiar material confidently. Mitigation: segment work by consequence and set different AI policies per segment rather than one firm-wide rule.
Better tools produce worse vigilance. A model that is right nine times out of ten trains its user to stop checking by the fiftieth output. The tenth then travels further into the process than it would have from a visibly unreliable system. This is ordinary automation complacency, and it means each accuracy improvement partly consumes itself. Mitigation: mandate verification by document class, never by reviewer confidence, and rotate who checks.
Procurement reads “legal AI” as “safe”. Domain training narrows the failure distribution. It does not remove the generative mechanism underneath, and a legally-trained model still predicts plausible continuations. A specialist tool that hallucinates rarely, in convincingly domain-appropriate language, is harder to catch than a general one that does so obviously. Mitigation: require vendors to describe their failure modes and grounding approach, not just their accuracy claims.
The reasoning risk sits beside the citation risk and has no equivalent control. A fabricated case is wrong and findable. A well-structured analysis that the professional recognised and approved rather than actually authored is defensible on its face and leaves no trace to audit, a problem we examined in AI sycophancy and professional judgement. Citation verification does nothing about it. Mitigation: require the reasoning to be articulated independently before the output is reviewed, not after.
The takeaway
The June guidance change is small in wording and large in what it settles. Responsibility for AI-assisted work has not moved anywhere, and the regulator now expects firms to show the mechanism by which it is discharged. That is a management question about who checks what, with what time, evidenced how, and it is one that public sector bodies are already answering the hard way, as the CPS apology for hallucinated material in court documents showed.
Three things separate organisations that will handle this well from those already sleepwalking into it:
- They fund verification as a line item. Not as goodwill, not as something senior people absorb, but as costed capacity attached to the work that generates the exposure.
- They evidence scrutiny rather than disclosure. The artefact a regulator or an insurer wants is a record of what was checked and by whom, not a log of which tool was opened.
- They keep professional judgement as the product. Leslie’s closing point to solicitors is that judgement and expertise are ultimately what the client is paying for. The same holds in any advisory business, and it is the thing an efficiency programme is most likely to optimise away without noticing.
Take Action: Pick the single document class where a fabricated citation would cost you most, and put a named verifier against it this week. One class done properly beats a firm-wide policy nobody has time to follow.
Your next steps, in order:
- Map where AI-assisted output currently reaches clients, courts or regulators without a named verifier
- Define a verification standard by document class, with the evidence format specified
- Set a baseline verification failure rate this quarter so improvement is measurable
- Separate legal-specific from general-purpose tooling in policy, and public from contracted deployments
- Review the supervision guidance against your actual practice, not your written one
Sources and further reading
Original analysis by Resultsense, drawing on How AI tools hallucinate - and why it matters in law by Mia Leslie, published by the Law Society on 21 July 2026, and the Solicitors Regulation Authority’s effective supervision guidance as updated on 12 June 2026. Related coverage: a litigant hiding instructions in court filings and our earlier analysis of why AI hallucinations persist.