Some 94% of UK undergraduates now say they use generative AI to help with assessed work. Fewer than half think their teaching staff are helping them build the skills to do it well. That gap, not the detector arms race, is the real state of academic integrity in British higher education, and closing it is a design problem rather than an enforcement one. The institutions still treating it as enforcement are about to discover they have retained all the cost of surveillance and none of its cover.

The policy went one way, everyone else went the other

Writing in Times Higher Education on 28 July, Doug Specht and Gunter Saunders of the University of Westminster identify the gap at the centre of most institutional responses to generative AI. On one side, integrity rules argue about whether using ChatGPT counts as plagiarism. On the other, students and colleagues had built working relationships with the tools long before any guidance arrived. “The policy process ran in one direction; practice ran in another,” they write.

Their conclusion is blunter than the sector’s usual register. “Universities need not a tighter grip on AI use, but a deeper reckoning with how they live with it,” they argue, in assessments, feedback and day-to-day teaching. That is a shift from regulation to integration, and they are honest that it is the harder of the two.

Strategic Reality: The detection-and-punish era did not end because universities won the argument about pedagogy. It ended because the tools failed, the ombudsman started ruling against institutions, and the students most likely to be wrongly flagged turned out to be the ones paying the bills.

The evidence base against detection is now settled. Stanford researchers, publishing in Patterns, tested seven widely used AI detectors and found that more than 61% of essays written by non-native English speakers were misclassified as machine-generated, whilst the same tools were near-perfect on native writers. A separate large-scale evaluation across 805 samples found average accuracy of 39.5% on unmodified AI text, falling to 17.4% once students applied simple evasion techniques. Weber-Wulff and colleagues, in what remains the most comprehensive independent evaluation, concluded that none of the tools tested met the standard required for high-stakes decisions.

The numbers that closed the argument

FigureWhat it measuresSource
61%Essays by non-native English speakers misclassified as AI-generated across seven detectorsStanford researchers, Patterns
39.5% → 17.4%Detector accuracy on unmodified AI text, then after simple evasion techniquesLarge-scale evaluation, 805 samples
3 of 4Office of the Independent Adjudicator cases published July 2025 upheld or partly upheld against the universityOIA
24% / 51%International students as a share of all UK higher education students, and of all postgraduatesHEPI
43%English universities forecasting deficitsHEPI
94%Undergraduates using generative AI to help with assessed workHEPI / Kortext Survey 2026
48%Undergraduates who feel teaching staff are helping them develop AI skillsHEPI / Kortext Survey 2026

The financial arithmetic is what makes this untenable rather than merely unfair. International students are 24% of UK higher education students and 51% of postgraduates at a point when 43% of English universities are forecasting deficits. Institutions built detection systems that flag their most commercially important cohort most often. Waterloo dropped Turnitin’s AI detection in September 2025, Curtin followed in January 2026, and UCLA and UC San Diego had already moved in 2024. UK institutions have since begun restricting or abandoning the tools too, and the HEPI analysis of who those tools actually catch explains why that was inevitable.

What replaces detection, and what quietly does not

Here is where the sector’s self-congratulation gets ahead of the evidence. Sam Illingworth of Edinburgh Napier University tested the claim that universities are helping students use AI, across 96 institutions, in HEPI Policy Note 71 published in May 2026. His summary of the result: policies “promise critical thinking but deliver audit trails”. They “name support yet deliver surveillance”.

Start with what could not be tested. Illingworth began with 163 UK institutions holding degree-awarding powers and scraped for their AI policies on a single day in February 2026, testing what a student, parent or regulator could find without logging in. Only 96 had a publicly accessible policy. The other 67 had no discoverable policy, or one behind an authentication wall, a broken link or dynamically loaded text invisible to search. That is 41% of the sector with no public AI policy at all. Among specialist institutions — conservatoires, arts universities, smaller providers — only 18% had one.

For the 96 that did, keyword analysis looked reassuring: 83 policies, or 86%, came out education-dominant, with only four dominated by detection language. Then Illingworth read 19 of them properly. Five were genuinely educational. Four used educational vocabulary as a veneer. Two were balanced. Eight were detection instruments. The computational classification was misleading for 12 of the 19, and it failed in a consistent direction: policies using the language of learning for non-educational purposes were systematically scored as educational.

Critical Context: A policy can say “learning” and mean “compliance”. Illingworth’s finding is that word-counting cannot tell the difference, which means every sector-level audit of AI policy conducted by keyword analysis has probably overstated how educational the sector is.

Three mechanisms produced the gap. The first is presentational: aspirational openings followed by regulatory bodies, so the critical-literacy ambition does no work past the first paragraph. Loughborough University is Illingworth’s most revealing case. Its four-step requirement — save outputs, write an acknowledgement, describe use, reference AI — scores as education-dominant on vocabulary whilst functioning as an evidence-retention protocol for future misconduct investigations. “The vocabulary of learning is pressed into the service of surveillance,” he writes.

The second is conditional trust. No policy in his sample of 19 states outright that the institution trusts its students. The dominant model grants trust provided students declare their use, retain evidence of their process and submit to verification on request. A student saving drafts because policy tells them to demonstrate their process is building a defence file. The University of Liverpool makes the inference explicit: “A student that refuses to declare how they have used the technology … may be attempting to hide the fact that the work is not their own.” Non-declaration is read as concealment. Students are required to be transparent about their AI use; institutions are not required to be transparent about their detection practices.

The third mechanism is the one with the clearest management lesson. Where a policy sits on an institution’s website predicts its framing better than the words inside it. Of seven policies located within academic misconduct frameworks, six read as detection-oriented. Policies housed in teaching and learning or study skills architectures, as at Stirling and Canterbury Christ Church, were more likely to be genuinely educational. Durham was the exception that proves the mechanism: it acknowledges its location inside a misconduct policy and then works deliberately to exceed that frame. Illingworth is careful that the gap “may well be unintentional” — a structural consequence of writing educational guidance inside regulatory architecture that predated it. AI did not build these assumptions about students. It inherited them.

The University of East Anglia case shows what the inheritance costs. Its green light / red light framework addresses staff and students in the same publicly accessible document, but the staff guidance is rich in pedagogical reasoning and assessment redesign strategy whilst the student section is directive and rule-bound. Nothing is hidden. The difference in depth simply reveals who the policy was written for. Staff are addressed as professionals trusted with judgement; students are given instructions.

The capability question underneath the compliance question

Specht and Saunders arrive at the same place from the other direction. Their FUTURES framework, published by HEPI in March 2026, names seven domains of human capability that higher education should deliberately cultivate: fluency in AI and digital systems; understanding self and well-being; technology ethics and responsibility; social intelligence; resilience and adaptability; emerging technology awareness; and professional engagement. The framework’s stated relationship to what already exists is complementary rather than competitive. Where the Jisc AI maturity toolkit asks whether an institution is ready to adopt AI responsibly, FUTURES asks whether students and colleagues are developing the human capabilities to make that adoption mean anything.

On assessment they are direct. Blanket prohibitions are hard to enforce and signal the wrong thing about what learning is for. If the goal is graduates who think critically and reason ethically, assessment should test those capabilities rather than the ability to produce text unaided. “This is a reframe, not a capitulation,” they write. The question moves from whether a student can produce a piece of work to whether they can evaluate, critique and responsibly integrate AI-assisted work into their own practice.

StakeholderWhat changes when detection goesWhat they need instead
StudentsThe threat of a probabilistic accusation recedes; ambiguity about acceptable use does notExplicit permitted-use definitions, and assessment that rewards demonstrated reasoning
Academic staffDetection scores stop substituting for academic judgementTime, training in assessment design, and workload models that survive oral components
Registrars and integrity teamsEvidence base shifts from tool output to process and dialogueProcedural safeguards, differential-impact monitoring, disclosure of detection practice
International studentsRemoval of the systematic bias that flagged them mostLanguage support that is not later read as evidence of dishonesty
EmployersGraduate signal changes meaning: AI-assisted work is the baselineRecruitment that assesses judgement about AI use, not its absence

Hidden Cost: Removing the detector without redesigning assessment is the worst available position. The institution loses the fig leaf, keeps the uncertainty, and still has an integrity process it cannot evidence.

What to do about it, by where you are starting from

The sequencing depends less on institutional ambition than on where AI policy currently lives.

If you have no public AI policy — 41% of the sector, on Illingworth’s count — the first move is publication, not drafting perfection. A policy behind an authentication wall or invisible to search is not a policy from the point of view of an applicant, a parent or an ombudsman reviewing a misconduct case. Publish, then improve.

If your policy sits inside a misconduct framework, relocation is the cheapest change available and among the most consequential. Move the guidance into teaching and learning architecture, and take Durham’s approach in the interim: name the location explicitly and state what the policy does beyond enforcement.

If your policy is already educational in structure, the work is procedural. Prohibit findings based solely or predominantly on detection scores. Show students the full evidence before any misconduct meeting. Define permitted use of generative AI, grammar checkers and language-support tools rather than leaving students to infer it. Monitor misconduct outcomes for differential effects on international, disabled and second-language students, because that is the exposure the OIA cases established.

If assessment redesign is the goal, the evidence-based alternatives are consistent across the literature: process-based assessment, staged submissions, oral defences and AI literacy embedded in the curriculum rather than bolted on. None of these is cheap, and pretending otherwise is how transformation programmes fail.

Resource Reality: Oral defences and staged submissions verify learning better than any detector. They also do not scale to a 400-student cohort without either new staff time or fewer, better-chosen verification points. The honest version of this plan names which it is.

For business readers, the read-across is uncomfortably close. Most corporate AI policies were written into existing information-security and disciplinary architecture, use the vocabulary of enablement, and operationalise as declaration requirements and audit logging. That is precisely the structure Illingworth found in universities. If your organisation’s AI policy sits under acceptable-use enforcement rather than under capability development, you have the same performative gap, the same conditional trust, and the same predictable result: people conceal what they are actually doing with the tools, so you lose the visibility the policy was written to gain.

Four problems the redesign consensus is not solving

Access asymmetry inside the same cohort. Specht and Saunders flag that paying students reach far more capable models than free tiers offer, and that capacity to adopt AI differs by discipline and between the Global North and South. Assessment that assumes competent AI use quietly assumes competent AI access. Mitigation: institutionally provisioned tool access, specified in the assessment brief, so the model is a controlled variable rather than a proxy for who can pay.

The trust asymmetry is structural, not attitudinal. The UEA pattern — pedagogical reasoning for staff, instructions for students — is not a drafting slip that better editing fixes. It reflects who the document was designed for. Mitigation: student involvement in policy development, recorded visibly rather than buried in page source, and a single register across both audiences.

Integrity pressure is migrating upstream and downstream simultaneously. The same authorship question is now live in schools, where Ofqual has warned of far tougher scrutiny of coursework, and in research, where arXiv has threatened one-year submission bans for unchecked AI content whilst major funders have agreed to permit AI in grant assessment with human final decisions. A university that solves undergraduate assessment whilst its research integrity and admissions processes run on detection logic has moved the problem, not fixed it.

Process failure is now the live legal risk, not tool failure. When Qualifications Scotland reinstated a pupil’s grade because it could not be sure a fair process had been followed, the concession was procedural. The OIA cases turned on the same thing: evidence not shown, mitigating circumstances not heard, reliability of the tool never considered. Institutions replacing detection with academic judgement inherit that standard, and academic judgement is only defensible when it is documented. Mitigation: build the evidence trail around dialogue and staged work now, before it is needed in a hearing.

The indispensability argument is the actual strategy

Beneath the procedural detail sits the question Specht and Saunders put at the end of their piece: in a world where AI can generate, automate and synthesise at scale, “what makes a human’s education indispensable?” Their answer is not the production of outputs. It is the critical judgement to recognise when a well-formed answer is wrong, the ethical reasoning to handle decisions with human consequences, the social intelligence to build trust, and the resilience to adapt when no model saw the change coming.

That is not a soft-skills addendum to a crowded curriculum. It is the argument for why the degree is worth its cost, made newly visible by tools that handle everything else competently. And it is testable, which is the point that gets lost. An institution can assess whether a student recognises a plausible wrong answer. It cannot assess whether they wrote unaided.

Success Factor: Three things separate institutions that will come out of this well. AI policy housed in teaching and learning rather than misconduct. Permitted use defined explicitly rather than inferred from prohibitions. Assessment that verifies reasoning at a proportionate number of points rather than policing outputs at all of them.

The uncomfortable finding in the HEPI survey data is not that students are using AI. It is that 68% think AI skills are essential to thrive whilst only 48% think their teachers are helping them build those skills, and that they split almost evenly on whether their institution encourages AI use at all — 37% agree, 36% disagree. Meanwhile 65% say assessment has already changed significantly. The change is happening; the explanation is not arriving with it. That is the same failure mode that produced the detection era, running one loop later. Pearson and AWS research puts the wider position starkly: only 14% of UK learners consider themselves highly AI-ready, and that is the pipeline employers will be recruiting from for the rest of the decade.

Take Action: Four questions worth putting to any UK institution this term. Where does your AI policy live on your website? Can a prospective student find it without logging in? Does it state that you trust your students? And if a misconduct finding were challenged tomorrow, what would the evidence be, other than a score?

Sources

Analysis by Resultsense. We cover UK AI policy, adoption and regulation for business and technology decision-makers.