dengjen-tts

A cross-platform Rust engine for neural TTS models.

Features

  • Models: Piper, Kokoro, and MeloTTS ONNX voices

  • Phonemization: eSpeak-ng (100+ languages, IPA output) and Arabic diacritization via libtashkeel

  • Multi-speaker voices: select by speaker_id

  • Streaming synthesis: chunked output (chunk_size/chunk_padding) and a realtime gRPC stream

  • Prosody control: rate, pitch, and volume via dengjen-sonic-sys (libsonic)

  • Synthesis modes: lazy, parallel, and batched, selectable per request; realtime for the voices that support it

  • Bindings: native Rust, C-API (libdengjen), Python (pydengjen), Go and Java bindings over the C API (see bindings/go and bindings/java), gRPC (any language over the wire), and a CLI

RHVoice-style formant/statistical synthesis is a different synthesis paradigm from this engine’s neural-ONNX pipeline and isn’t planned.

To start using dengjen-tts, follow Get started with the gRPC server or Get started with Python.

What each model supports

Piper Kokoro MeloTTS

Lazy, parallel, batched modes

Yes

Yes

Yes

Realtime streaming

Voices whose manifest sets "streaming": true

Yes, by chunking an already synthesized sentence

No

Speakers

Manifest speaker map

Presets from the manifest’s voices list

Manifest speaker map

Rate, pitch, volume, silence

Yes

Yes

Yes

Tunable inference knobs

noise_scale, length_scale, noise_w

None

noise_scale, length_scale, noise_scale_w

Phonemizers

eSpeak-ng, raw text, Hebrew, pinyin; Arabic diacritization via libtashkeel

eSpeak-ng, fixed to en-US

eSpeak-ng (en, es, fr, ja, ko) or pinyin (zh)

The Hebrew and pinyin phonemizers are opt-in Cargo features (hebrew, pinyin) that no frontend in this repository enables, so the prebuilt dengjen-tts-grpc server and the pydengjen wheels do not include them; embedding the Rust crates lets you enable them. Arabic diacritization (the tashkeel feature) is enabled by default in every frontend this repository publishes. See Choosing and tuning a model backend and Streaming synthesis and the gRPC frontend for detail.

Crates

Crate Purpose

dengjen-tts-core

Core error types, synthesis configuration, and shared primitives

dengjen-tts-piper, dengjen-tts-kokoro, dengjen-tts-melotts

Model backends: loading and inference for each model family, using onnxruntime via ort

dengjen-tts

Wraps a model backend and adds synthesized speech post-processing, including changing prosody. Also provides different modes of parallelism.

dengjen-espeak-phonemizer

Converts text to IPA phonemes using a patched version of eSpeak-ng

dengjen-hebrew-phonemizer, dengjen-pinyin-phonemizer

Hebrew (diacritization plus IPA) and Mandarin pinyin grapheme-to-phoneme conversion

dengjen-audio-ops

Audio post-processing primitives: sample buffers, Hann windowing, and WAV encoding

dengjen-sonic-sys

Rust FFI bindings to Sonic: a C library for controlling rate, volume, and pitch of generated speech

dengjen-tts-grpc

gRPC server frontend

libdengjen

C API bindings (static and shared library)

dengjen-tts-python

Python bindings, published as pydengjen, using pyo3

dengjen-tts-cli

Command-line interface (builds the dengjen binary)

License

Licensed under the GNU General Public License v3.0 or later (GPL-3.0-or-later). dengjen originated as a fork of Sonata by Musharraf Omer, originally MIT-licensed; see the repo’s NOTICE file for project history and third-party attributions.