
AI Detection Is Dead. Universities Are Switching to Process-Based Evaluation.
By Thandi Mokoena 5 min read
Turnitin flagged 61% of human-written essays as AI. Universities are pulling the plug. Here's what's replacing it and what it means for your thesis.
In 2024, Australian Catholic University recorded nearly 6,000 alleged academic misconduct cases. About 90 percent were AI-related.
After investigation, roughly a quarter were dismissed. Any case where Turnitin's AI detection tool was the sole evidence? Thrown out immediately.
One student waited six months before the accusations were dropped. Her transcript was marked “results withheld” the entire time. It affected her chances of securing graduate positions.
ACU eventually abandoned the Turnitin AI detector entirely.
And they're far from alone.
Curtin University disabled Turnitin's AI writing detection across all campuses from January 2026. Vanderbilt, Johns Hopkins, Yale and the University of Waterloo have all disabled or restricted Turnitin's AI detection feature.
The reason is the same in every case: these tools don't work well enough to stake a student's academic career on.
Why Are AI Detection Tools Flagging 61% of Human-Written Essays?
AI detection tools promise to identify machine-generated text. In practice, they're flagging real students for work they actually wrote.
Stanford ran the numbers.
Researchers tested seven widely used GPT detectors on 91 TOEFL essays written by non-native English speakers and 88 essays by US-born eighth-graders.
The detectors performed “near-perfectly” on the American students' essays. But they classified 61.22 percent of the TOEFL essays as AI-generated. Every single one of those essays was written by a real person.
All seven detectors unanimously flagged 19.8 percent of those human-written essays as AI-authored. Not one detector got it right.
Why? AI detectors measure two properties: perplexity (how surprising the word choices are) and burstiness (how much sentence structure varies).
Non-native English speakers naturally use simpler vocabulary, shorter sentences and more formulaic structures. Those patterns look a lot like AI-generated text. The detectors can't tell the difference.
But it's not just international students getting hit.
A Common Sense Media study found that 20 percent of Black teens reported being falsely accused of using AI to complete an assignment. Compare that to 10 percent of Latino teens and 7 percent of white teens.
Among those incorrectly flagged, 79 percent said their work had been submitted to AI detection software.
These aren't just stats. There are real people behind them.
NPR reported in December 2025 on Ailsa Ostovitz, a high school junior at Eleanor Roosevelt High School in Maryland. Accused of using AI three times across two different classes.
In one case, her teacher showed her a screenshot: about a 30 percent probability she had used AI. The assignment? A description of the music she listens to. She lost points.
After a meeting weeks later, the teacher said they no longer believed she had used AI. The school district later stated they advise educators not to rely on such tools, citing “potential inaccuracies and inconsistencies.”
And then there's Turnitin's own numbers. They claim a false positive rate under 1 percent. A 2023 Washington Post investigation by Geoffrey Fowler found a misclassification rate closer to 50 percent. Smaller sample, but a massive gap.
That disconnect between the company line and independent testing is a big part of why universities are walking away.
What Universities Are Doing Instead
So if they're ditching the detectors, are universities just giving up on academic integrity?
No. They're changing how they evaluate it.
It's called process-based evaluation. Instead of running a finished paper through detection software, they're looking at how the work was developed.
In practice, this looks like three things:
Draft trails. Students submit outlines, rough drafts and revision history. Shows they engaged with the material over time, not just pasted a finished product.
Source accountability. Can you justify your citations? Can you explain why you picked those sources and how they back up your argument? Way harder to fake than a polished paragraph.
Oral checks. For theses and dissertations, some programs now include short oral components. Can you explain the choices you made? Can you discuss your methodology without reading from the paper?
Think about it. If the goal is making sure students understand what they submitted, checking their process makes way more sense than running their output through a tool that flags non-native English speakers at six times the rate of native speakers.
Curtin University's Academic Board said as much. They pointed to reliability concerns, equity issues with false positives for certain populations and a push to focus on education over surveillance.
Process-friendly AI writing
Litrevu generates drafts from your own uploaded papers. You choose the sources and steer every section. Your process is baked in from the start.
Try Litrevu freeHow Does This Affect Your Thesis?
Most students haven't caught on to this yet.
The old game was simple: write your paper, submit it, hope the detector doesn't flag it.
That game is changing. Universities are moving toward a model where using AI isn't the violation. Hiding it is.
Disclosure is becoming the baseline. Oxford now lets students use generative AI for studies and research as long as they declare it and the course allows it. Same story across most institutions now: be transparent or get in trouble.
So what does this change for you?
If you use ChatGPT to generate a literature review from scratch, you've got no process trail. No uploaded sources. No sign you engaged with the material. Nothing that shows your thinking.
Under process-based evaluation, that's a problem. Not because a detector caught you, but because you can't show you understood any of it.
Compare that to a different approach. You upload your own research papers to a tool that pulls from those sources. You steer the focus. You revise the output. Your process is there from start to finish.
You can point to the papers you uploaded. You can explain why you picked those sources. You can talk through the revisions you made.
That's what process-based evaluation actually rewards.
The Bottom Line
The “paste into ChatGPT and pray” era is ending. Not because detectors got better, but because universities figured out they were asking the wrong question.
The new question isn't “did a machine write this?” It's “does this student understand what they submitted?”
If you actually engage with your research, that's a better system for you.
How Litrevu Compares to Other Research Tools
These tools mostly solve different problems. Three of the four below help you find and evaluate papers. Litrevu starts after that, when you have the papers and have to write. Elicit is the one with real overlap.
| Tool | Best for | Works from | What you get | Where it beats Litrevu |
|---|---|---|---|---|
| Litrevu | Drafting a cited chapter from papers you have already chosen | PDFs you upload | Sectioned prose draft with every citation traceable to an uploaded passage | No discovery, no screening at scale, no view on whether a source is contested |
| Elicit | Finding papers and extracting structured data from them | 138M paper corpus, and PDFs you upload | Reports, tables, summaries and structured comparisons | Far larger discovery corpus, systematic review screening into the thousands, purpose-built data extraction tables |
| Consensus | Answering an empirical yes or no question across the literature | 200M+ peer-reviewed papers | Synthesised answer with a meter showing how much of the literature agrees | Answers questions across all published work in seconds; nothing in Litrevu does this |
| Scite | Checking whether a paper has been supported or contradicted since | 1.6B+ classified citation statements | Citation contexts labelled supporting, contrasting or mentioning | Tells you if a source is contested. Litrevu has no view on this at all |
| Research Rabbit | Exploring outward from one paper to find related work | Citation graph | Visual networks of related papers, authors and topics | Visual citation-graph discovery; Litrevu offers nothing comparable |
The honest split: if you do not yet have your papers, Elicit, Consensus and Research Rabbit will get you there faster than Litrevu can, because Litrevu does not search for papers at all. If you want to know whether a study you already cite has since been contradicted, Scite answers that and Litrevu does not.
Litrevu is an AI literature review assistant that turns the papers a researcher has already gathered into a cited first draft, with every citation traceable to the uploaded source. Elicit also works from uploaded PDFs, so the real difference there is the output: Elicit produces reports, tables and structured comparisons, while Litrevu produces a sectioned prose draft.
Comparison verified as of August 2026 against each vendor's own documentation. These products change quickly, so check the current feature list before deciding.
Working on a Thesis?
Litrevu generates drafts from your own uploaded papers. You choose the sources, steer each section and revise the output. Your process is baked in from the start. Try it free , 800 words, no credit card.
Litrevu is an AI literature review assistant that turns the papers a researcher has already gathered into a cited first draft, with every citation traceable to the uploaded source.
Start writing for free