Get started with the gRPC server
dengjen-tts-grpc is a local gRPC server that loads dengjen-tts voices and streams synthesized audio
to any client that speaks gRPC. The dengjen-nvda screen-reader
add-on uses it this way.
What is published
-
GitHub release
grpc-release-v2.0.3, assetdengjen-tts-grpc-v2.0.3-windows-x64.zip(13.8 MB). It holdsdengjen-tts-grpc.exeandNOTICE. ONNX Runtime is linked into the executable. -
The release is Windows x64 only. On other platforms, build from source with
cargo build --release -p dengjen-tts-grpc. See Build from source. -
The executable needs the Microsoft Visual C++ runtime (it imports
VCRUNTIME140.dll,VCRUNTIME140_1.dllandMSVCP140.dll). Install the Visual C++ Redistributable if the server does not start.
Install
In PowerShell:
curl.exe -LO https://github.com/ZirekHQ/dengjen-tts/releases/download/grpc-release-v2.0.3/dengjen-tts-grpc-v2.0.3-windows-x64.zip
Expand-Archive dengjen-tts-grpc-v2.0.3-windows-x64.zip -DestinationPath dengjen-tts-grpc
Prerequisites
-
A voice. Download
en_US-lessac-low.onnxand its manifest as shown in the Python page. -
The eSpeak-ng data. The zip does not contain it. The server looks for an
espeak-ng-datafolder next to the executable, so extract it there and set no environment variable:curl.exe -LO https://github.com/ZirekHQ/dengjen-tts/archive/refs/tags/v2.0.3.tar.gz tar -xzf v2.0.3.tar.gz --strip-components=3 -C dengjen-tts-grpc dengjen-tts-2.0.3/deps/dev/espeak-ng-data
Start the server
.\dengjen-tts-grpc\dengjen-tts-grpc.exe
The server binds 127.0.0.1 only. Version 2.0.3 prints DENGJEN_GRPC_LISTENING port=<port> on stdout
after it binds its port, and writes its log to stderr. Three environment variables change its behavior:
| Variable | Default | Effect |
|---|---|---|
|
|
TCP port. |
|
|
Log filter, in the |
|
the executable’s directory |
Directory that contains the |
For example, $env:DENGJEN_GRPC_SERVER_PORT = 0 before the command above picks a free port. The examples
below connect to 127.0.0.1:49314. call_sequence.sh and client.py read DENGJEN_GRPC_ADDRESS (another host:port) to change that.
Check that it works
The client steps below (the curl download, bash call_sequence.sh, VAR=value command, ffmpeg)
use a POSIX shell: Linux, macOS, or Git Bash on Windows. The server binds 127.0.0.1 only, so a client
in WSL2 cannot reach a server running on the Windows host. Run the client on the same host as the server.
GetDengjenVersion needs no voice. The server does not register gRPC reflection, so grpcurl needs the
.proto file. Install grpcurl, then run:
curl -LO https://raw.githubusercontent.com/ZirekHQ/dengjen-tts/v2.0.3/crates/frontends/grpc/proto/dengjen_grpc.proto
grpcurl -plaintext -import-path . -proto dengjen_grpc.proto -d '{}' \
127.0.0.1:49314 dengjen_grpc.DengjenGrpc/GetDengjenVersion
If the server runs on another port, for example after DENGJEN_GRPC_SERVER_PORT=0, replace 49314 in
that command.
examples/grpc/call_sequence.sh in the repository makes the same
call through a call function that adds the grpcurl flags and the address:
call dengjen_grpc.DengjenGrpc/GetDengjenVersion -d '{}'
The reply carries the server version, which equals the release version:
{
"version": "2.0.3"
}
Call it with grpcurl
Install jq as well, then copy the script
from the repository next to dengjen_grpc.proto (or set DENGJEN_PROTO_DIR to the directory that holds
the .proto) and run it. Point it at the manifest with an absolute path that
the server can read:
DENGJEN_EXAMPLE_VOICE=/path/to/en_US-lessac-low.onnx.json bash call_sequence.sh
LoadVoice returns a voice_key, a hash of the manifest path, and the voice’s audio format:
voice_key=$(jq -n --arg path "$VOICE" '{path: $path}' \
| call dengjen_grpc.DengjenGrpc/LoadVoice -d @ | jq -r '.voiceKey')
SynthesizeUtterance streams one chunk per sentence. grpcurl prints each chunk’s audioBytes as base64,
and the script decodes them into one file:
jq -n --arg key "$voice_key" --arg text "Hello from the gRPC server." \
'{voice_key: $key, text: $text, synthesis_mode: "MODE_LAZY"}' \
| call dengjen_grpc.DengjenGrpc/SynthesizeUtterance -d @ \
| jq -r '.audioBytes // empty' | while read -r chunk; do printf '%s' "$chunk" | base64 --decode; done > "$OUT"
The script prints the version JSON from the first call, then the voice key and the output size. The last two lines look like this, for example:
voice_key=aa3cceca7d5a09d3 wrote output.pcm (67072 bytes)
The audio is raw 16-bit little-endian PCM with no WAV header, mono, at the voice’s sample rate
(audio.sampleRate in the LoadVoice reply, 16000 for en_US-lessac-low). The byte count changes from
run to run. To get a playable file:
ffmpeg -f s16le -ar 16000 -ac 1 -i output.pcm output.wav
Call it with Python
Install the client libraries and generate the stubs next to a copy of examples/grpc/client.py from the repository:
pip install grpcio grpcio-tools
python -m grpc_tools.protoc -I. --python_out=. --grpc_python_out=. dengjen_grpc.proto
python client.py /path/to/en_US-lessac-low.onnx.json
The script imports the generated modules, loads the voice and collects the chunks:
import dengjen_grpc_pb2 as pb
import dengjen_grpc_pb2_grpc as rpc
def load_voice(stub, manifest):
return stub.LoadVoice(pb.VoiceConfigLocation(path=manifest))
def synthesize(stub, voice_key, text):
request = pb.SynthesisRequest(
voice_key=voice_key, text=text, synthesis_mode=pb.MODE_LAZY
)
return [chunk.audio_bytes for chunk in stub.SynthesizeUtterance(request)]
It prints the server version, the voice key, the chunk count and the output file name, and writes a real
WAV, because write_pcm adds the header from the voice’s audio format. The lines after the version look
like this, for example:
voice_key: aa3cceca7d5a09d3 received 1 PCM chunk(s) wrote output.wav
SynthesizeUtteranceRealtime returns PCM windows as the model produces them. The script calls it only
when voice.supports_streaming_output is true:
def synthesize_realtime(stub, voice_key, text):
request = pb.SynthesisRequest(voice_key=voice_key, text=text)
return b"".join(chunk.audio_bytes for chunk in stub.SynthesizeUtteranceRealtime(request))
The realtime RPC works only for voices that support streaming. en_US-lessac-low does not, and the call
fails with UNIMPLEMENTED: Streaming synthesis is not supported for this model.
The RPCs
| RPC | What it does |
|---|---|
|
Returns the server version. |
|
Loads a voice manifest and returns its |
|
Returns the |
|
Returns a loaded voice’s default synthesis settings. |
|
Changes those defaults: |
|
Streams |
|
Streams |
SynthesisRequest.synthesis_mode picks MODE_LAZY, MODE_PARALLEL or MODE_BATCHED. See
Streaming synthesis and the gRPC frontend for how the modes differ.
Errors
An RPC that names a voice that is not loaded returns NotFound, and so does LoadVoice with a manifest
path that does not exist. For the first case, see the example in
Streaming synthesis and the gRPC frontend.
Want to help? Learn how to contribute to the ZirekHQ docs ›