Same marks does not mean same reading
Persian and Arabic share fatha, kasra, damma, shadda, sukun, and many base letters. Persian commonly calls the three short-vowel marks zabar, zir, and pish and reads them as a, e, and o in modern Iranian Persian. Arabic conventionally reads the corresponding marks as a, i, and u. Vocabulary, grammar, homographs, and Persian Ezafe also differ, so Arabic tashkeel software is not a substitute for Persian diacritization.
The script and several mark characters are genuinely shared
Persian is written in a form of the Arabic script. Most base letters are shared, and the core short-vowel marks are the same encoded combining characters. This is why a font, keyboard, or text processor can often display Persian and Arabic marks without knowing which language the sentence uses.
Unicode gives these characters Arabic-derived names because it encodes the script and its shared character history. The name “Arabic fatha” in a character picker does not make the surrounding sentence Arabic. The same code point can serve Persian, Arabic, Urdu, and other languages that use the script.
Persian uses the script for Persian sounds and words
The visible overlap can hide three separate differences: the sound attached to a mark, the word being marked, and the grammar connecting that word to the sentence.
1. The short-vowel values are not identical
Modern Iranian Persian
ـَ ـِ ـُ are commonly taught as a, e, o.
گُل is gol, flower. دِل is del, heart.
Standard Arabic
ـَ ـِ ـُ are conventionally taught as a, i, u.
Arabic dialects and phonetic environments add variation, but the basic teaching values remain different from Iranian Persian e and o.
2. Persian has letters and vocabulary Arabic does not
Persian adds پ، چ، ژ، گ to the basic Arabic inventory. More importantly, a Persian sentence is built from Persian vocabulary and morphology. A model trained to predict vowels inside Arabic roots and patterns has no automatic knowledge of Persian words such as پنجره panjare, window, or Persian verb forms such as میخواند mikhānad, reads.
3. Arabic-origin spelling does not require Arabic pronunciation
Persian contains a large body of Arabic-origin vocabulary. Many loans preserve their Arabic spelling, including letters that no longer represent distinct consonants in ordinary Iranian Persian. Once borrowed, the words participate in Persian pronunciation and grammar. Treating the spelling as an instruction to reconstruct an Arabic inflection can produce the wrong Persian reading.
Etymology asks where a word came from. Reading asks how this sentence pronounces and uses the word now. Those answers can be related without being identical.
Persian Ezafe is a major difference for reading tools
Persian uses Ezafe to link a noun with an adjective, possessor, name, or related element. Inکتاب جدید من, the reading isketāb-e jadid-e man, my new book. The connecting e sounds are normally absent after consonants in everyday spelling.
Arabic expresses noun-adjective agreement and possession through different grammatical structures. An Arabic tashkeel system can add legal Unicode marks and still miss the Persian Ezafe chain that makes a phrase sound like Persian.
Persian ambiguity also depends on sentence context
این گل زیباست.
in gol zibāst
This flower is beautiful.
کفش پر از گل بود.
kafsh por az gel bud
The shoe was full of mud.
The unmarked spelling گل becomesگُل gol in the first context andگِل gel in the second. A language-specific system needs Persian meanings and sentence patterns, not only a list of allowed marks.
Why an Arabic diacritizer can look successful while being wrong for Persian
A visual inspection asks, “Did it add marks?” The real evaluation asks several harder questions.
- Did it preserve every source letter?Diacritization should add cues, not rewrite or translate the submitted Persian.
- Did it choose Iranian Persian vowel values?A kasra-shaped mark that represents an Arabic i prediction may not express the intended Persian e.
- Did it recognize Persian words and morphology?Persian prefixes, suffixes, compound verbs, and colloquial forms need Persian analysis.
- Did it infer Ezafe?The connector depends on Persian noun-phrase structure.
- Did context resolve homographs?The same consonants can require different readings in different sentences.
- Did a Persian reader review representative output?Valid characters and a clean interface are not language-quality evidence.
VowelMarks is built around Persian context: the diacritization tool keeps the submitted Persian letters and adds short-vowel marks and Ezafe for the reading selected by the sentence. You can compare that result with Pinglish and Persian speech around the same source. Poetry, uncommon names, dialect writing, and highly specialized text still deserve an extra human check.
Which search term should you use?
A quick identification rule
If the text contains پ، چ، ژ، گ, it cannot be ordinary Standard Arabic text, though a Persian sentence may contain none of those letters. Language identification should also consider vocabulary, function words, and morphology. The safest workflow is to tell a tool the language explicitly and test it on complete Persian sentences.
Related Persian reading guides
- Persian vowels: short and long vowels explained
- Zabar, zir, and pish in Persian
- How to type Persian vowel marks
- Why Persian context changes an unmarked reading
Frequently asked questions
Are Persian and Arabic diacritics the same?
They share several written combining marks and Unicode characters, including fatha, kasra, damma, shadda, and sukun. Persian speakers often call the first three zabar, zir, and pish. The languages assign those marks inside different sound systems, vocabularies, and grammars, so an Arabic-vowelled result is not automatically correct Persian.
What is tashkeel called in Persian?
English sources use diacritics, vowel marks, vocalization, or diacritization. Arabic commonly uses harakat and tashkeel. Persian commonly uses terms such as zabar, zir, pish, harakat, and e'rab or اعراب. Search terminology varies, but the intended Persian pronunciations still need Persian context.
What are zabar, zir, and pish in Arabic terminology?
Zabar uses the same written mark as Arabic fatha, zir corresponds to kasra, and pish corresponds to damma. In modern Iranian Persian they typically cue a, e, and o; in Standard Arabic the conventional values are a, i, and u.
Can an Arabic diacritizer add vowels to Persian text?
It may preserve the shared script and add valid mark characters, but that does not establish Persian correctness. Persian words, Persian pronunciations of Arabic loans, Persian morphology, homographs, and Ezafe require Persian training data or Persian-specific rules. Test actual Persian sentences instead of judging only whether marks appear.
Does Persian have Ezafe while Arabic does not?
Persian Ezafe is a productive connector between nouns and modifiers or possessors. Arabic expresses comparable relationships through different grammatical systems. An Arabic diacritizer is not designed to infer Persian Ezafe simply because both languages use related script.
Sources and review notes
Sources are listed for the claims they support. Original practice examples are identified in the article. Product capabilities and access terms were checked on the review date and can change.
- University of Texas at Austin, The Writing System. Used for the Persian script inventory, the omission of short vowels, and the treatment of Arabic loanword spelling in Persian.
- University of Texas at Austin, Vowels. Used for the modern Iranian Persian values a, e, o and the long-vowel system described in this guide.
- Unicode Arabic names list. Used for the official character names and code points U+064E fatha, U+064F damma, U+0650 kasra, U+0651 shadda, and U+0652 sukun.
- The Unicode Standard, Chapter 9: Middle Eastern Scripts. Used for combining-mark behavior and the distinction between encoded characters and their rendered placement.
- Encyclopaedia Iranica, Arabic Elements in Persian. Used for the explanation that many Arabic loanwords retain Arabic spelling while entering Persian phonology and grammar.
- Doostmohammadi, Nassajian, and Rahimi, Persian Ezafe Recognition Using Transformers. Used for Ezafe as Persian grammatical information that is ordinarily absent from the written form.
- Ayyoubzadeh and Shahnazari, Persian Homograph Disambiguation. Used for the importance of Persian sentence context when one unmarked spelling has multiple readings.