Model accuracy and limitations
dengjen-tashkeel restores diacritics with an ONNX neural model (from Hareef, trained mainly on Modern Standard Arabic data) — it does not run independent grammatical analysis. Where the correct diacritic depends on context the model wasn’t trained to resolve, it can guess wrong. That’s a model-accuracy limitation, not a bug in this library’s code (see Architecture for what this library is and isn’t responsible for).
Known case: proper nouns without resolvable context
Given the input:
عن أمير المؤمنين أبي حفص عمر بن الخطاب
the model diacritizes the name حفص as حِفَصِ ("hifsi") instead of the grammatically correct
حَفْصٍ ("hafsin"). Per the project’s own README, the correct case ending for a proper noun like
this depends on context the model wasn’t trained to resolve. See
issue #28 for the full discrepancy and any
follow-up.
Mitigating uncertain case endings: taskeen
Every binding exposes a taskeen option (taskeen_threshold in Rust and Python, the second
Optional<Float> argument in Java, --taskeen/--prob in the CLI — see
Per-binding usage). When enabled, the model substitutes a sukoon for a
case-ending diacritic it isn’t confident about, instead of guessing — trading a possibly-wrong
diacritic for an explicit "unresolved" marker.
Taskeen only affects the final case-ending diacritic. It would not fully correct the حفص example
above: the discrepancy there isn’t limited to the case ending (ِ vs. ٍ) — the internal vowel
also differs (فَ vs. فْ), so taskeen doesn’t fully hide an error like this one.
Scope
Per the project’s non-scope note: diacritization only, not a general-purpose Arabic NLP toolkit. For proper nouns, genitive chains, or any other case-ending decision that depends on context spanning multiple words, verify output against domain-relevant text before relying on it in a downstream product.
Want to help? Learn how to contribute to the ZirekHQ docs ›