Accuracy report

Publish the evidence, including the disagreement.

These measurements explain what VowelMarks has tested, what the numbers mean, and where the product remains weakest.

337,107

sentences stress-tested

Zero base-text preservation failures in that test.

96.5% / 93.9%

Ezafe precision / recall

Measured on held-out gold data.

90.8%

homograph disambiguation

Measured on 491 usable held-out rows.

These are separate measurements on different datasets. They should not be combined into a single product-accuracy percentage.

Source preservation

The original page text is guarded separately from model output.

A stress corpus contained 337,107 unique Persian sentences across historical news research material, held-out ambiguity cases, current-news feeds, mixed-script text, punctuation, and zero-width-joiner variants. The source-preservation checks recorded zero failures.

This means the tested renderer did not replace the original base letters, punctuation, links, or mixed-script content. It does not mean every added pronunciation cue was correct.

Ezafe

96.5% precision and 93.9% recall on held-out gold data.

Precision asks how often an emitted Ezafe cue belonged there. Recall asks how many labelled Ezafe links the system found. Guard rules cover compounds, time expressions, existing marks, and unsafe word endings.

Homographs

90.8% on 491 usable held-out rows.

A trained local disambiguator and narrow verified context rules choose among readings for ambiguous spellings. This is held-out evidence, not a statement about every ambiguous word in Persian.

Overall reading quality

Two AI reviews of the same 80 outputs disagreed sharply.

87.5%AI reviewer A
57.5%AI reviewer B
~80%adjudicating AI review
87.5%modern-prose subset

These are AI-reader labels, not human or native-reader verdicts. Publishing the disagreement is part of the evidence: broad quality still needs independent human review.

What remains uncertain

No native-speaker review has been completed yet.

The weakest areas are classical poetry, difficult literature, and highly colloquial writing. Broader device testing and independent Persian-reader review are still needed.

How to report a problem

Include the original sentence, the marked result, and what looked wrong. Do not include private or identifying page content.

Report a reading issue