Get started with Python
pydengjen is the Python binding for dengjen-tts. It loads Piper, Kokoro and MeloTTS voices
through one class, PiperModel, and synthesizes in every mode the engine offers.
What is published
-
PyPI package
pydengjen(2.0.3 at the time of writing). -
Wheels for CPython 3.9 and newer, one
abi3wheel per platform: Linux x86_64 and aarch64 (manylinux 2.28), Windows x64 and macOS arm64. -
ONNX Runtime is linked into the wheel, so you install nothing else.
-
Wheels built after 2.0.3 include the eSpeak-ng data that Piper’s default phonemizer needs. See the note under Prerequisites for 2.0.3 and earlier.
-
On a platform with no wheel, pip falls back to the source distribution. It needs a Rust toolchain to build and does not contain the eSpeak-ng data.
Prerequisites: a voice
The commands in this section are POSIX shell: Linux, macOS, or Git Bash on Windows.
Download a voice, a .onnx model and its .onnx.json manifest. This one is 63 MB:
BASE=https://huggingface.co/rhasspy/piper-voices/resolve/v1.0.0/en/en_US/lessac/low
curl -LO $BASE/en_US-lessac-low.onnx
curl -LO $BASE/en_US-lessac-low.onnx.json
Wheels released after 2.0.3 find their bundled eSpeak-ng data on their own. To use a different copy, set
DENGJEN_ESPEAKNG_DATA_DIRECTORY to the directory that contains a folder named espeak-ng-data
before you import pydengjen. Those wheels never overwrite a value you set.
|
|
Run the example
The example is examples/python/synthesize.py in the
repository. Clone the repository, then run the example with the
manifest path:
python examples/python/synthesize.py en_US-lessac-low.onnx.json
It writes output.wav and prints the size. On one run the line was wrote output.wav (113172 bytes).
Play output.wav to check the result. The file is 16000 Hz, mono, 16-bit, and between 3.5 and 3.9 seconds long
across runs. Sizes and durations change a little from run to run.
Synthesis modes
--mode picks how the text turns into audio. Each mode except file returns an iterator of chunks,
one per sentence, or one per window for streamed.
-
filewrites the whole text to one WAV (synthesize_to_file). -
lazysynthesizes one sentence per step, on demand. It has the lowest latency to the first chunk. -
parallelsynthesizes every sentence at once and yields them in order when all are done. -
batchedsynthesizes groups of--batch-sizesentences at a time and yields each group as it finishes. -
streamedyields audio in fixed windows (--chunk-size,--chunk-padding) as the model generates it.
def synthesize_chunks(synth, mode, text, config, args):
if mode == "lazy":
return synth.synthesize_lazy(text, config)
if mode == "parallel":
return synth.synthesize_parallel(text, config)
if mode == "batched":
return synth.synthesize_batched(text, config, args.batch_size)
return synth.synthesize_streamed(text, config, args.chunk_size, args.chunk_padding)
A lazy run prints one line per sentence. For example:
chunk 0: 1086 ms, real-time factor 0.030390238389372826 chunk 1: 2649 ms, real-time factor 0.026805097237229347
streamed needs a voice whose manifest sets "streaming": true. On any other voice, such as
en_US-lessac-low, it raises DengjenException: Streaming synthesis is not supported for this model.
Chunks from streamed are bytes: raw 16-bit little-endian PCM with no header. WaveSamples.get_wave_bytes()
is raw PCM too, despite its name; save_to_file writes a real WAV. See
Streaming synthesis and the gRPC frontend for how the modes differ.
Prosody
AudioOutputConfig sets speed, volume, pitch and trailing silence:
def output_config(args):
return pydengjen.AudioOutputConfig(
rate=args.rate,
volume=args.volume,
pitch=args.pitch,
appended_silence_ms=args.silence,
)
After 2.0.3, every argument is optional, and None leaves it unset. The positional order is rate, volume, pitch
(the CLI and gRPC list pitch before volume), so use the keyword names. Rate, pitch and volume take 0 to
100 and map as described under Usage. appended_silence_ms adds silence after each
sentence. On one run, --rate 60 --pitch 50 --volume 80 --silence 200 produced a 1.2 second file from
the same text that took 3.5 seconds with no prosody options.
Speakers and model knobs
def load_synthesizer(manifest, speaker, parameters):
model = pydengjen.PiperModel(manifest)
if speaker:
model.speaker = speaker
if parameters:
model.set_parameters(parameters)
return pydengjen.Dengjen.with_piper(model)
model.speaker = "name" selects a speaker by name. An unknown name raises DengjenException.
--list-speakers prints the speaker map the manifest declares, one id: name line per speaker. It
prints nothing for a single-speaker voice such as en_US-lessac-low.
def print_speakers(synth):
for speaker_id, name in sorted((synth.speakers or {}).items()):
print(f"{speaker_id}: {name}")
set_parameters (--param KEY=VALUE, repeatable) sets named inference knobs. Unknown names are
silently ignored. What each model supports lists the keys:
Piper reads noise_scale, length_scale and noise_w, MeloTTS reads noise_scale, length_scale
and noise_scale_w, and Kokoro has none.
Errors
Every failure raises pydengjen.DengjenException. The example prints it as error: <message> and
exits 1. Messages from the example:
-
Missing manifest:
Failed to load resource: Failed to read model config: … -
Unknown speaker:
no speaker named 'nope' -
eSpeak data not found, which happens when
DENGJEN_ESPEAKNG_DATA_DIRECTORYpoints at a directory withoutespeak-ng-data, or on 2.0.3 without the workaround:Failed to phonemize given text using espeak-ng … Failed to initialize eSpeak-ng
Next steps
Read Choosing and tuning a model backend and Streaming synthesis and the gRPC frontend. To build from source, see Build from source.
Want to help? Learn how to contribute to the ZirekHQ docs ›