Read
See the script
OCR for printed and handwritten Pashto, Urdu, and Latin text, from a single line to a full document.
Zirak AI · Peshawar
Zirak Technologies builds machines that read, write, listen, speak, and understand low-resource languages of South Asia and the Middle East — Pashto, Dari, Urdu, Persian, and Arabic — with the same care the industry gives to English.
01 — Mission
Cutting-edge models still fail where the script is cursive, the data is scarce, and the speakers number in the tens of millions. Zirak AI closes that gap with optical character recognition, speech, and language systems designed around local linguistic needs — preserving diversity while opening the digital world.
02 — Capabilities
Read
OCR for printed and handwritten Pashto, Urdu, and Latin text, from a single line to a full document.
Write
Segmentation, spelling, part-of-speech, and translation models that respect how these languages are actually written.
Listen
Pashto speech-to-text for transcription, dictation, and voice interfaces, including native accents.
Speak
Text-to-speech tuned for clarity, tone, and regional Pashto dialects, ready for apps and platforms.
Understand
Monolingual Pashto BERT, offensive-language detection, and benchmarks the research community can build on.
An open toolkit for spelling, segmentation, tagging, and moderation in a language spoken by over 50 million people.
03 — Flagship
One million synthetic images, a thousand font families, and a public 10,000-image test set. We used it to measure how today’s large multimodal models — Gemini, GPT-4o, Claude, Grok, Qwen, and others — actually read Pashto.
Contact
Governments, universities, and product teams use Zirak AI when the language on the page is not English.