Choosing and tuning a model backend
dengjen-tts’s three model backends (Piper, Kokoro, MeloTTS) all implement the same
DengjenModel trait, but their voice manifests expose different tuning knobs and support
different features. See Usage for the required manifest fields per backend;
this page covers what the numeric knobs actually change, and which backend to reach for.
Backend comparison
| Backend | Tunable inference knobs | Phonemizer options | Realtime streaming |
|---|---|---|---|
Piper |
|
|
Only voices with |
Kokoro |
None — speed is fixed at |
Hardcoded to English ( |
Yes, but by chunking an already-fully-synthesized sentence, not true incremental decoding |
MeloTTS |
|
|
Not supported — returns |
All three backends also accept the same post-synthesis prosody controls
(rate/pitch/volume/silence on the CLI, ProsodyControls over gRPC), applied uniformly
by DengjenSpeechSynthesizer regardless of which model produced the audio.
Piper
Piper’s inference block controls the VITS decoder’s stochastic duration/pitch predictor:
-
noise_scale— variability injected into pitch/energy. Higher values sound more expressive and less monotone, but too high introduces artifacts. -
length_scale— a global speaking-rate multiplier applied to predicted phoneme durations; values above1.0slow speech down, below1.0speed it up. -
noise_w— variability injected into the duration predictor itself, changing pacing/rhythm between repeated runs rather than pitch.
Whether a Piper voice can run in realtime mode is a property of the voice, not a runtime flag:
only manifests with "streaming": true ship the split encoder.onnx/decoder.onnx pair that
VitsStreamingModel needs (see crates/dengjen/models/piper/src/lib.rs); other Piper voices
load as the standard single-file VitsModel and only support the lazy, parallel and batched modes.
Kokoro
Kokoro voices have no inference block at all — crates/dengjen/models/kokoro/src/config.rs
only accepts model_path, voices_dir, vocab_path, sample_rate, and the list of voices
names. Per-voice character comes entirely from a precomputed style vector looked up by voice
name and token length (VoiceStyles::style_for); there’s no noise_scale/length_scale
equivalent to adjust. speed is hardcoded to 1.0 and the phonemizer’s target language is
hardcoded to en-US (crates/dengjen/models/kokoro/src/inference.rs) — a Kokoro voice is
English-only regardless of manifest content.
Kokoro does implement stream_synthesis, so realtime mode is available, but
KokoroAudioStreamer slices an already-fully-synthesized sentence into chunks rather than
generating incrementally. It lowers perceived latency for playback (first chunk arrives before
the whole WAV is written out) but not actual synthesis latency the way a streaming Piper voice
does.
MeloTTS
MeloTTS’s inference block mirrors Piper’s shape but with noise_scale_w instead of noise_w;
the perceptual effect of each knob is the same as described above. The phonemizer is chosen per
voice manifest ({"type": "espeak", "voice": "…"} or {"type": "pinyin", "model_dir": "…"})
rather than being a separate phoneme_type field like Piper’s.
MeloTTS models don’t override stream_synthesis, so they fall back to the DengjenModel trait’s
default, which returns DengjenError::UnsupportedOperation (unimplemented over gRPC) for
realtime mode. Use lazy, parallel or batched mode with MeloTTS voices.
Choosing a backend
-
Piper for the broadest voice availability (any VITS-family ONNX export using the standard 3/4-input tensor convention) and the widest language reach, since its
espeakphonemizer works with any configured espeak-ng voice. -
Kokoro when a fixed set of English voices with no tuning surface is acceptable, and true incremental low-latency generation isn’t required.
-
MeloTTS for tone-aware Mandarin output (via its dedicated pinyin phonemizer with real tone extraction) or its bundled Spanish/French/Japanese/Korean voices, and when realtime streaming isn’t needed.
Want to help? Learn how to contribute to the ZirekHQ docs ›