Streaming synthesis and the gRPC frontend
dengjen-tts can hand back audio four different ways — sequentially, eagerly-parallel,
bounded-concurrent, or as low-latency realtime chunks — and exposes the same choice through
both the CLI and the dengjen-tts-grpc frontend.
The four synthesis modes
StreamMode (crates/dengjen/synth/src/lib.rs) splits text into sentences and then produces
audio one of four ways:
-
lazy(the default) — synthesizes one sentence per call to the stream’snext(), on a single thread, on demand. Lowest latency to the first chunk, no parallelism. -
parallel— synthesizes every sentence concurrently across a rayon thread pool, then yields them in original order once all of them are done. Higher throughput for long, multi-sentence text, but nothing is emitted until the whole request finishes. -
batched(N)— synthesizes sentences in concurrent groups of up toN(parallel within a group, sequential across groups), yielding each group’s audio before starting the next. A middle ground betweenlazyandparallel: partial output streams out sooner thanparallel’s "wait for everything," with less per-request thread/memory pressure than firing every sentence at once. Because every backend serializes the actual inference call behind a `Mutex, `batched(N)’s benefit is bounded resource pressure and earlier partial output, not raw inference throughput. Delivery is bursty by design: the first chunk of each group waits for the whole group to finish, then the remaining chunks in that group arrive immediately from a buffer, before the next group’s stall. -
realtime— requests audio from the model in fixed-size windows (chunk_size, in a backend-specific unit — e.g. mel frames for Piper) as it’s generated, withchunk_paddingextra frames of context trimmed from each window’s edges to avoid decoder boundary artifacts. Runs on a background thread and supports cancellation mid-stream. Chunk size grows for later sentences in the same request (up to 4x the base size) to trade a little latency for fewer, larger chunks as the stream goes on.
Not every backend supports every mode — see Choosing and tuning a model
backend for which Piper/Kokoro/MeloTTS voices support realtime. Requesting realtime on a
model that doesn’t support it returns DengjenError::UnsupportedOperation (unimplemented
over gRPC).
On the CLI (dengjen-tts-cli), select a mode with --mode lazy|parallel|batched|realtime, tune
batched grouping with --batch-size, and tune realtime chunking with
--chunk-size/--chunk-padding. Prosody flags (--rate, --pitch, --volume, --silence)
apply uniformly regardless of mode.
The gRPC frontend
dengjen-tts-grpc serves DengjenGrpc (crates/frontends/grpc/proto/dengjen_grpc.proto) over
tonic. Start it with:
cargo run --release -p dengjen-tts-grpc
It binds 127.0.0.1 on the port from DENGJEN_GRPC_SERVER_PORT (default 49314; set it to
0 to let the OS assign a free port), then prints DENGJEN_GRPC_LISTENING port=<port> to
stdout — useful for a test harness or embedder that needs to discover an OS-assigned port
without parsing log output.
Call sequence
-
LoadVoice(VoiceConfigLocation{path})loads a voice manifest and returns aVoiceDescriptorkeyed by avoice_key— a 16-hex-character xxh3_64 hash of the canonicalized config path. The key is deterministic andLoadVoiceis idempotent: calling it again with the same path returns the already-loaded voice’s descriptor instead of reloading. -
GetVoiceInfo(VoiceRef{voice_key})/GetSynthesisOptions/SetSynthesisOptionsinspect or adjust a loaded voice’s default synthesis settings (speaker,length_scale,noise_scale,noise_w, plus a free-formparametersmap for backend-specific knobs). -
SynthesizeUtterance(SynthesisRequest)returns astream SynthesisChunkof WAV-encoded audio (audio_bytesplus a computedreal_time_factor). -
SynthesizeUtteranceRealtime(SynthesisRequest)returns astream RealtimeAudioChunkof raw PCM (audio_bytesonly, no WAV header) — this is the RPCdengjen-nvdavendors for live screen-reader playback.
|
|
Trying it with grpcurl
With the server running, grpcurl can call it directly against the checked-in .proto (the
server doesn’t register gRPC reflection, so -proto/-import-path are required):
cd crates/frontends/grpc/proto
grpcurl -plaintext -import-path . -proto dengjen_grpc.proto \
-d '{}' 127.0.0.1:49314 dengjen_grpc.DengjenGrpc/GetDengjenVersion
which returns:
{
"version": "<workspace version>"
}
Calling any RPC that needs a loaded voice before one exists returns a clear NotFound, e.g.:
grpcurl -plaintext -import-path . -proto dengjen_grpc.proto \
-d '{"voice_key": "missing"}' 127.0.0.1:49314 dengjen_grpc.DengjenGrpc/GetVoiceInfo
ERROR:
Code: NotFound
Message: A voice with the key `missing` has not been loaded
Loading a real voice and synthesizing needs a Piper/Kokoro/MeloTTS voice manifest on disk (see Usage for where to get one), then follows the same shape:
grpcurl -plaintext -import-path . -proto dengjen_grpc.proto \
-d '{"path": "/path/to/en_US-lessac-medium.onnx.json"}' \
127.0.0.1:49314 dengjen_grpc.DengjenGrpc/LoadVoice
grpcurl -plaintext -import-path . -proto dengjen_grpc.proto \
-d '{"voice_key": "<voice_key from LoadVoice>", "text": "Hello world"}' \
127.0.0.1:49314 dengjen_grpc.DengjenGrpc/SynthesizeUtterance
Want to help? Learn how to contribute to the ZirekHQ docs ›