dengjen-tts
A cross-platform Rust engine for neural TTS models.
Features
-
Phonemization: eSpeak-ng (100+ languages, IPA output) and Arabic diacritization via
libtashkeel -
Multi-speaker voices: select by
speaker_id -
Streaming synthesis: chunked output (
chunk_size/chunk_padding) and a realtime gRPC stream -
Prosody control: rate, pitch, and volume via
dengjen-sonic-sys(libsonic) -
Synthesis modes: lazy, parallel, and batched, selectable per request; realtime for the voices that support it
-
Bindings: native Rust, C-API (
libdengjen), Python (pydengjen), Go and Java bindings over the C API (seebindings/goandbindings/java), gRPC (any language over the wire), and a CLI
RHVoice-style formant/statistical synthesis is a different synthesis paradigm from this engine’s neural-ONNX pipeline and isn’t planned.
To start using dengjen-tts, follow Get started with the gRPC server or Get started with Python.
What each model supports
| Piper | Kokoro | MeloTTS | |
|---|---|---|---|
Lazy, parallel, batched modes |
Yes |
Yes |
Yes |
Realtime streaming |
Voices whose manifest sets |
Yes, by chunking an already synthesized sentence |
No |
Speakers |
Manifest speaker map |
Presets from the manifest’s |
Manifest speaker map |
Rate, pitch, volume, silence |
Yes |
Yes |
Yes |
Tunable inference knobs |
|
None |
|
Phonemizers |
eSpeak-ng, raw text, Hebrew, pinyin; Arabic diacritization via |
eSpeak-ng, fixed to |
eSpeak-ng ( |
The Hebrew and pinyin phonemizers are opt-in Cargo features (hebrew, pinyin) that no frontend in
this repository enables, so the prebuilt dengjen-tts-grpc server and the pydengjen wheels do not
include them; embedding the Rust crates lets you enable them. Arabic diacritization (the tashkeel
feature) is enabled by default in every frontend this repository publishes. See
Choosing and tuning a model backend and
Streaming synthesis and the gRPC frontend for detail.
Crates
| Crate | Purpose |
|---|---|
|
Core error types, synthesis configuration, and shared primitives |
|
Model backends: loading and inference for each model family, using |
|
Wraps a model backend and adds synthesized speech post-processing, including changing prosody. Also provides different modes of parallelism. |
|
Converts text to IPA phonemes using a patched version of eSpeak-ng |
|
Hebrew (diacritization plus IPA) and Mandarin pinyin grapheme-to-phoneme conversion |
|
Audio post-processing primitives: sample buffers, Hann windowing, and WAV encoding |
|
Rust FFI bindings to Sonic: a C library for controlling rate, volume, and pitch of generated speech |
|
gRPC server frontend |
|
C API bindings (static and shared library) |
|
Python bindings, published as |
|
Command-line interface (builds the |
License
Licensed under the GNU General Public License v3.0 or later (GPL-3.0-or-later). dengjen originated
as a fork of Sonata by Musharraf Omer, originally MIT-licensed;
see the repo’s NOTICE file for project history and third-party attributions.
Want to help? Learn how to contribute to the ZirekHQ docs ›