Preventing AI misuse in Pakistani schools: why detection software alone fails

The tools sold to catch AI-written homework are less reliable than they look, and they can wrongly accuse the students who write in a second language.

A student hands in an unusually polished English essay, and a teacher's first instinct is to run it through an AI detector. It feels like a fair, objective check. The trouble is that the evidence behind these tools is much weaker than the confident percentage on the screen suggests, and in a country where most students write English as a second language, that weakness lands on exactly the wrong students. Preventing misuse is a real job. Detection alone is a poor way to do it.

Start with what the research says. A Stanford-led study published in the journal Patterns, GPT detectors are biased against non-native English writers, found that these detectors consistently misclassify non-native English writing samples as AI-generated, while native writing samples are identified accurately. The likely reason is uncomfortable but understandable: detectors tend to treat simple, predictable phrasing as machine-like, and a student writing carefully in a second language often uses exactly that kind of phrasing. For a Pakistani classroom, where English is the language of the textbook and the exam but rarely the language of the home, the implication is direct: a detector can flag your most careful students as cheaters.

The companies that build the underlying models have run into the same wall. When OpenAI retired its own AI-written-text classifier in July 2023, TechCrunch reported that it was pulled over a low rate of accuracy, with the company saying it was researching more effective provenance techniques for text. TechCrunch's own quick test found the classifier identified only one of seven generated snippets correctly. If a lab with access to its own model couldn't build a dependable detector, a school should be sceptical of a subscription that promises one.

None of this means misuse is imaginary. Students do use AI to write assignments, and teachers are right to care. It means the response should rest on things a school can actually control. The first is a clear written rule, stated in plain words at the start of each term. Many schools find a three-tier version works: uses that are allowed (checking grammar, explaining a concept), uses that must be declared (brainstorming, outlining), and uses that are not allowed (submitting generated text as your own). A student can't follow a rule nobody stated, and an honest student can't tell whether asking an AI to explain photosynthesis counts as cheating unless you say.

The second lever is assessment design, which we treat in its own post on making assignments harder to outsource. In short: ask for the thinking, not just the product. Work done in class, drafts with dated stages, a short oral explanation of what was written, personal or local examples, and marks for process all make it harder to hand in something you can't explain. These methods work regardless of which tool the student used, and they don't need detection software or a discussion about who wrote what.

The third lever is conversation. Before treating a suspicious piece of work as misconduct, ask the student to walk you through it: how did you get this argument, why this example, what would you change if the question were different? A student who wrote it can usually answer; one who didn't usually can't, and the conversation gives you something firmer than a percentage to act on. It also treats the student as a person, which matters when the evidence is uncertain.

If a school does use detection software at all, use it as a prompt for that conversation and never as proof. Write into your policy that no student is penalised on a detector result alone, keep a record of what you asked and how they answered, and be especially careful with students who write in their second language. A wrong accusation costs more than a missed one: it damages trust, and it teaches students that careful writing is dangerous.

A note on where a tool like ours sits. Muallim is built by DIGIT Pakistan for teachers and school admins, and students don't use it, so it isn't an AI-writing tool in your students' hands. We also don't claim to detect AI use, and we don't think a planner or quiz builder should. What Quiz & Assessment and the Worksheet Builder can do is help you produce varied questions faster, including short-answer and fill-in-the-blank items that ask students to apply a concept, so that assessment design isn't limited by how much time you have on Sunday night. Every item opens in an editor for your review first, so you can add the local, specific question that a generic one can't.

To act on this week, do three small things. Write your three-tier AI rule for students and read it aloud in class. Choose one assessment this term and add a process element, such as an in-class draft or a two-minute oral follow-up. And agree, as a staff, that a detector score is never enough on its own. None of those needs a budget, and together they cover more of the problem than software can.

Ways to respond to AI misuse, and how far each can be trusted
ApproachWhy it worksIts limit
AI detection softwareQuick, and feels objective.Research finds detectors misclassify non-native English writing; OpenAI retired its own classifier over low accuracy.
A clear three-tier student ruleStudents know what is allowed, declared, or not allowed.Needs to be repeated each term and enforced consistently.
Process-based assessmentDrafts, in-class writing, and oral follow-ups show the student's own thinking.Takes planning time and some redesign of tasks.
A conversation with the studentLets the student explain the work in their own words.Requires time and a no-blame tone.

โ† Back to Journal ยท See what's live