dengjen-tts
A cross-platform Rust engine for neural TTS models.
Features
-
Models: Piper, Kokoro, and MeloTTS ONNX voices, including Kokoro per-voice style embeddings
-
Phonemization: eSpeak-ng (100+ languages, IPA output) and Arabic diacritization via
libtashkeel -
Multi-speaker voices: select by
speaker_id -
Streaming synthesis: chunked output (
chunk_size/chunk_padding) and a realtime gRPC stream -
Prosody control: rate, pitch, and volume via
dengjen-sonic-sys(libsonic) -
Synthesis modes: lazy, parallel, and batched, selectable per request
-
Bindings: native Rust, C-API (
libdengjen), Python (pyo3), gRPC (any language over the wire), and a CLI
Not yet supported, tracked on the issue tracker: native Go/Java/Kotlin bindings and a Matcha-TTS model loader. RHVoice-style formant/statistical synthesis is a different synthesis paradigm from this engine’s neural-ONNX pipeline and isn’t planned.
Crates
| Crate | Purpose |
|---|---|
|
Converts text to IPA phonemes using a patched version of eSpeak-ng |
|
Handles model loading and inference using |
|
Wraps |
|
GRPC frontend for dengjen |
|
C-API binding to dengjen |
|
Python bindings to |
|
Rust FFI bindings to Sonic: a C library for controlling rate, volume, and pitch of generated speech |
License
Licensed under the GNU General Public License v3.0 or later (GPL-3.0-or-later). dengjen began as
a fork of Sonata by Musharraf Omer, originally MIT-licensed;
retained attribution is in the repo’s NOTICE file.
Want to help? Learn how to contribute to the ZirekHQ docs ›