You are viewing the documentation for a prerelease version. View Latest

Streaming synthesis and the gRPC frontend

dengjen-tts can hand back audio four different ways — sequentially, eagerly-parallel, bounded-concurrent, or as low-latency realtime chunks — and exposes the same choice through both the CLI and the dengjen-tts-grpc frontend.

The four synthesis modes

StreamMode (crates/dengjen/synth/src/lib.rs) splits text into sentences and then produces audio one of four ways:

  • lazy (the default) — synthesizes one sentence per call to the stream’s next(), on a single thread, on demand. Lowest latency to the first chunk, no parallelism.

  • parallel — synthesizes every sentence concurrently across a rayon thread pool, then yields them in original order once all of them are done. Higher throughput for long, multi-sentence text, but nothing is emitted until the whole request finishes.

  • batched(N) — synthesizes sentences in concurrent groups of up to N (parallel within a group, sequential across groups), yielding each group’s audio before starting the next. A middle ground between lazy and parallel: partial output streams out sooner than parallel’s "wait for everything," with less per-request thread/memory pressure than firing every sentence at once. Because every backend serializes the actual inference call behind a `Mutex, `batched(N)’s benefit is bounded resource pressure and earlier partial output, not raw inference throughput. Delivery is bursty by design: the first chunk of each group waits for the whole group to finish, then the remaining chunks in that group arrive immediately from a buffer, before the next group’s stall.

  • realtime — requests audio from the model in fixed-size windows (chunk_size, in a backend-specific unit — e.g. mel frames for Piper) as it’s generated, with chunk_padding extra frames of context trimmed from each window’s edges to avoid decoder boundary artifacts. Runs on a background thread and supports cancellation mid-stream. Chunk size grows for later sentences in the same request (up to 4x the base size) to trade a little latency for fewer, larger chunks as the stream goes on.

Not every backend supports every mode — see Choosing and tuning a model backend for which Piper/Kokoro/MeloTTS voices support realtime. Requesting realtime on a model that doesn’t support it returns DengjenError::UnsupportedOperation (unimplemented over gRPC).

On the CLI (dengjen-tts-cli), select a mode with --mode lazy|parallel|batched|realtime, tune batched grouping with --batch-size, and tune realtime chunking with --chunk-size/--chunk-padding. Prosody flags (--rate, --pitch, --volume, --silence) apply uniformly regardless of mode.

The gRPC frontend

dengjen-tts-grpc serves DengjenGrpc (crates/frontends/grpc/proto/dengjen_grpc.proto) over tonic. Start it with:

cargo run --release -p dengjen-tts-grpc

It binds 127.0.0.1 on the port from DENGJEN_GRPC_SERVER_PORT (default 49314; set it to 0 to let the OS assign a free port), then prints DENGJEN_GRPC_LISTENING port=<port> to stdout — useful for a test harness or embedder that needs to discover an OS-assigned port without parsing log output.

Call sequence

  1. LoadVoice(VoiceConfigLocation{path}) loads a voice manifest and returns a VoiceDescriptor keyed by a voice_key — a 16-hex-character xxh3_64 hash of the canonicalized config path. The key is deterministic and LoadVoice is idempotent: calling it again with the same path returns the already-loaded voice’s descriptor instead of reloading.

  2. GetVoiceInfo(VoiceRef{voice_key}) / GetSynthesisOptions / SetSynthesisOptions inspect or adjust a loaded voice’s default synthesis settings (speaker, length_scale, noise_scale, noise_w, plus a free-form parameters map for backend-specific knobs).

  3. SynthesizeUtterance(SynthesisRequest) returns a stream SynthesisChunk of raw 16-bit little-endian PCM at the voice’s sample rate, mono, with no header (audio_bytes plus a computed real_time_factor).

  4. SynthesizeUtteranceRealtime(SynthesisRequest) returns a stream RealtimeAudioChunk of the same raw PCM (audio_bytes only) — this is the RPC dengjen-nvda vendors for live screen-reader playback.

SynthesisRequest.synthesis_mode (MODE_LAZY/MODE_PARALLEL/MODE_BATCHED) selects between lazy, parallel, and batched(N) for SynthesizeUtterance, matching the CLI’s --mode flag. batch_size (SynthesisRequest.batch_size) only applies to MODE_BATCHED; when unset it defaults to min(available CPU cores, 4), and an explicit 0 is rejected as invalid_argument. MODE_BATCHED does not mean tensor-level ONNX batching — a real trained voice’s duration predictor is stochastic, so there is no reliable way to recover each sentence’s audio length from a single padded multi-row inference call; see the design notes on issue #248 for the underlying experiment.

Trying it with grpcurl

With the server running, grpcurl can call it directly against the checked-in .proto (the server doesn’t register gRPC reflection, so -proto/-import-path are required):

cd crates/frontends/grpc/proto
grpcurl -plaintext -import-path . -proto dengjen_grpc.proto \
    -d '{}' 127.0.0.1:49314 dengjen_grpc.DengjenGrpc/GetDengjenVersion

which returns:

{
  "version": "<workspace version>"
}

Calling any RPC that needs a loaded voice before one exists returns a clear NotFound, e.g.:

grpcurl -plaintext -import-path . -proto dengjen_grpc.proto \
    -d '{"voice_key": "missing"}' 127.0.0.1:49314 dengjen_grpc.DengjenGrpc/GetVoiceInfo
ERROR:
  Code: NotFound
  Message: A voice with the key `missing` has not been loaded

Loading a real voice and synthesizing needs a Piper/Kokoro/MeloTTS voice manifest on disk (see Usage for where to get one), then follows the same shape:

grpcurl -plaintext -import-path . -proto dengjen_grpc.proto \
    -d '{"path": "/path/to/en_US-lessac-medium.onnx.json"}' \
    127.0.0.1:49314 dengjen_grpc.DengjenGrpc/LoadVoice

grpcurl -plaintext -import-path . -proto dengjen_grpc.proto \
    -d '{"voice_key": "<voice_key from LoadVoice>", "text": "Hello world"}' \
    127.0.0.1:49314 dengjen_grpc.DengjenGrpc/SynthesizeUtterance