Almost every AI writing tool treats non-English output as an afterthought — write the content in English, then run it through a translation pass at the end. That approach falls apart the moment you actually watch a Pakistani classroom in session: the textbook and syllabus documents are in English, but the teaching itself — the explanations, the questions, the parent messages — happens in Urdu.
In conversations with teachers across different boards and school types, the same pattern kept coming up: they wanted content generated in Urdu directly, not an English draft they'd have to mentally translate while teaching, and not a stiff, obviously-machine-translated version either. Roman Urdu came up just as often, for teachers who message parents in Urdu written in Latin script.
So Urdu and Roman Urdu are equal, first-class output modes across every generator — the Lesson Planner, the Worksheet Builder, and Quiz & Assessment — chosen per generation, not fixed once for your whole account. It's a design decision, not a feature checkbox: Urdu isn't a layer added on top of an English template, it's one of the outputs each of those modules was built to produce from the start.
The gap between translation and native generation shows up most clearly in ordinary classroom instructions, not technical vocabulary. Take something as simple as "Read the passage and answer the following questions." A literal translation renders every word correctly and still sounds like an instruction manual, not something a teacher would actually say to a class. Ask a Pakistani teacher to phrase the same instruction in Urdu from scratch, and it comes out shorter and closer to how they'd actually say it out loud — "یہ عبارت پڑھیں اور سوالات کے جواب دیں" — the kind of natural phrasing native generation is built to produce, and a translation pass, however good the underlying model, has to work harder to arrive at.
Subject-specific vocabulary is its own separate risk. A science or math term translated literally from an English textbook can come out technically accurate and still be the wrong word for how that subject is actually taught in Urdu in a Pakistani classroom — the kind of small mismatch a teacher catches instantly and a generic translation pass usually doesn't, because it's translating words, not teaching a subject.
Roman Urdu isn't a lesser fallback bolted on for completeness, either — it's how a large share of real school communication already happens, especially over WhatsApp and SMS, where typing the Urdu script is slower for some teachers or a recipient's phone doesn't render it cleanly. A teacher who needs a homework reminder in Roman Urdu isn't settling for a compromise; they're asking for the format they'd have typed by hand anyway.
This also shaped a concrete product decision, not just a wording choice: language is chosen per generation, not set once in account settings and forgotten. A teacher covering three sections of the same subject might want an English draft for their own planning notes, a Roman Urdu version for a parent's WhatsApp message, and a full Urdu version for the classroom handout — same underlying content, three different outputs, chosen fresh each time rather than locked to a single setting that fits none of those three well.
None of this is a claim that machine translation is inherently bad — it's a claim that translation and native generation are different tasks with different failure modes, and a tool built specifically for a Pakistani classroom should be judged on the second, not graded on a curve for having at least attempted the first. It's also, plainly, why this matters only for the people actually producing that classroom content — teachers and school admins — rather than as a general-purpose translation tool aimed at a broader audience. Narrower scope, on this specific problem, is the point.
A fair question: if Urdu-first output matters this much, why not just require Urdu everywhere? Because a teacher's own planning notes, or a report meant for a parent who's more comfortable in English, are real use cases too. The point isn't forcing Urdu into every output — it's making it a genuine first-class choice everywhere it's needed, not an inferior fallback wherever it happens to be chosen.
This is also why "Urdu-first" shows up as a decision worth an entire post, rather than a single bullet on a features page — it touches how every one of Muallim's generators was actually built, from the first line of the product, not just a language toggle bolted on at the end.
This is still early. We don't have large-scale usage data to point to yet — what we have is a beta product built directly from what teachers told us they actually needed, and we'd rather say that plainly than dress it up with numbers we don't have.