Synthesizing speech
Dengjen synthesizes speech from a Piper voice. Download a
voice’s .onnx model and matching .onnx.json config from the
Piper voices repository and keep both files
together, e.g. voices/en_US-lessac-medium.onnx and voices/en_US-lessac-medium.onnx.json.
A voice doesn’t have to come from the official Piper voices repository — any VITS-family ONNX
export using the same 3/4-input tensor convention (phoneme ids, lengths, scales, optional
speaker id) can be loaded by writing a matching .onnx.json manifest with
"model_type": "vits". The minimal required fields are audio (sample rate), inference
(noise_scale/length_scale/noise_w), and phoneme_id_map (the symbol vocabulary);
phoneme_type selects the phonemizer (espeak is the default if omitted, and needs an
espeak.voice entry — text, hebrew, and pinyin don’t).
A MeloTTS voice — its own VITS-derived ONNX export, with
a tones tensor alongside phone ids — is loaded with "model_type": "melotts". The required
fields are audio (sample rate), inference
(noise_scale/length_scale/noise_scale_w), phone_id_map and tone_id_map (the phone and
tone symbol vocabularies), and model_path; phonemizer selects the phonemization backend and
must be one of {"type": "espeak", "voice": "<espeak-ng voice name>"} (covers English, Spanish,
French, Japanese, Korean) or {"type": "pinyin", "model_dir": "<g2pW model directory>"}
(Chinese, with real tone extraction — requires building with the pinyin feature).
CLI
Synthesize text from a file to a WAV file, using the dengjen-cli frontend:
cargo run --release -p dengjen-tts-cli -- voices/en_US-lessac-medium.onnx.json \
-f input.txt \
-o output.wav
Or send a single request as JSON on stdin and capture the WAV bytes written to stdout:
echo '{"text": "Hello world"}' | cargo run --release -p dengjen-tts-cli -- voices/en_US-lessac-medium.onnx.json > output.wav
Run dengjen --help (or cargo run --release -p dengjen-tts-cli — --help) for the full list
of options. The main ones:
-
--mode lazy|parallel|batched|realtimeselects the synthesis mode (defaultlazy).--batch-sizesets how many sentencesbatchedsynthesizes per concurrent group, and--chunk-sizeand--chunk-paddingtunerealtimechunking. -
--speaker-idpicks a speaker for multi-speaker voices; when omitted, the voice’s own default speaker from its manifest is used (Kokoro’s is0). -
--rate,--pitchand--volumetake 0-100. When omitted, each stays at 1.0 (no change). The value maps linearly onto a speed of 0.5x-5.5x (so 10 is 1x and 50 is 3x), a pitch of 0.5-1.5 and a volume gain of 0-1 (a volume of 0 mutes the audio).--silenceappends that many milliseconds of silence after each sentence;--ratealso changes how long that silence lasts, because the silence is processed like the rest of the audio. -
--length-scale,--noise-scaleand--noise-woverride inference knobs:--noise-wis Piper’s, while--length-scaleand--noise-scalealso apply to MeloTTS.--param KEY=VALUE(repeatable) sets a named backend parameter to a number, such as MeloTTS’snoise_scale_w(an unrecognized key may be silently ignored); a named flag wins over a conflicting--param.
Want to help? Learn how to contribute to the ZirekHQ docs ›