Architecture
This page shows dengjen-nvda’s own structure as a C4 Container, Component, and Code diagram. For how it fits into the wider ZirekHQ speech stack, see the org-wide System Context diagram.
Container
dengjen-nvda is not one process but two: NVDA loads the add-on itself
in-process, and the add-on launches a vendored, separately-built
dengjen-tts-grpc.exe binary (a dengjen-tts frontend) as its own subprocess,
talking to it over gRPC.
C4Container
System_Ext(nvda, "NVDA", "The screen reader host process")
System_Boundary(addon, "dengjen-nvda") {
Container(synthdriver, "dengjen_neural_voices", "Python 3.13 add-on", "Loaded in-process by NVDA as a synthesizer driver")
Container(grpcbin, "dengjen-tts-grpc.exe", "Fetched Rust binary (pinned in dengjen-tts.lock)", "Launched via subprocess.Popen; a gRPC-frontend build of dengjen-tts")
}
System_Ext(voices, "Voice catalogue", "External download source for voices")
Rel(nvda, synthdriver, "Loads / calls SynthDriver API")
Rel(synthdriver, grpcbin, "Launches, then calls over gRPC")
Rel(synthdriver, voices, "Downloads voices from")
Component
The add-on itself is a small hexagonal (ports-and-adapters) design: domain
logic that knows nothing about NVDA or gRPC, one seam (TTSBackend) that any
synthesis backend must satisfy, and two adapters either side of that seam.
C4Component
Container_Boundary(addon, "dengjen_neural_voices") {
Component(domain, "domain/tts_system.py", "Python module", "Pure domain logic: voices, speech options. No gRPC, no NVDA imports.")
Component(port, "ports/tts_backend.py", "Python Protocol", "TTSBackend -- the one interface a TTS engine adapter must satisfy")
Component(grpcadapter, "adapters/sonata_grpc", "Python package", "Implements TTSBackend by launching and talking to dengjen-tts-grpc.exe")
Component(nvdaadapter, "adapters/nvda/synth_driver.py", "Python module", "Exposes the domain layer to NVDA's own SynthDriver API")
}
Rel(nvdaadapter, domain, "Creates/manages DengjenVoice")
Rel(domain, port, "Calls self.backend, typed as")
Rel(grpcadapter, port, "Implements")
Adding a model backend
dengjen-tts (the vendored gRPC engine) can support more model types than
Piper — it already ships Kokoro and MeloTTS. Wiring a new one into this
add-on needs:
-
A
ModelCatalogimplementation (addon/globalPlugins/dengjen_tts_global_plugin/model_catalog.py’s `Protocol) withmodel_type,is_installed(), andinstall(), followingkokoro_download.py’s `KokoroCatalogas a template. -
That catalog’s installer writes two files per voice directory: the model backend’s own
config.json(whatever shapedengjen-tts’s loader for that `model_typeexpects — checkcrates/dengjen/models/<type>/src/config.rsinzirekhq/dengjen-tts), and avoice.jsonsidecar viadomain/voice_metadata.write()carrying{model_type, name, language, description}— the only fieldsDengjenVoice.from_path()reads. -
If the new model type has per-utterance parameters the existing gRPC
ProsodyControlsmessage doesn’t carry, that needs a change inzirekhq/dengjen-ttsfirst (a different repo) — not an addon-side concern. -
If the new model type’s tuning parameters map onto the
noise_scale/length_scale/noise_wsliders, add aModelProfileentry to_PROFILESindomain/model_profiles.py, listing which sliders apply and, where the engine’s key differs from the slider name, theparameterskey (MeloTTS mapsnoise_wtonoise_scale_w). -
Add a
DENGJEN_<TYPE>_VOICES_DIRconstant toconst.pyand register it inDengjenTextToSpeechSystem.load_all_voices_from_nvda_config_dir()(domain/tts_system.py) — that method iterates a fixed tuple of directories, not a dynamic scan ofvoices/*, so a new backend’s directory has to be added there explicitly or its installed voices are invisible to NVDA.
No install directory is shared between model types: each gets its own
DENGJEN_VOICES_BASE_DIR/voices/<model_type>/ (see const.py), scanned
only once its constant is registered per step 5 above. The one exception is
MeloTTS: voices installed from a local archive currently land in
voices/piper/, because local install passes a single directory and no
MeloTTS directory constant is registered.
Code
TTSBackend is a Protocol, not a base class — the sonata_grpc adapter
satisfies it structurally. Its methods are split deliberately: everything NVDA
calls synchronously from its main thread has to stay blocking; only
synthesize() is a true async generator.
classDiagram
class TTSBackend {
<<Protocol>>
+initialize()
+check_version()
+shutdown()
+load_voice()
+get_synth_options()
+set_synth_options()
+synthesize() AsyncIterator~bytes~
}
class SynthOptions {
<<frozen dataclass>>
}
class LoadedVoice {
<<frozen dataclass>>
}
class SonataGrpcBackend
TTSBackend <|.. SonataGrpcBackend
TTSBackend ..> SynthOptions : passes
TTSBackend ..> LoadedVoice : returns
note for SonataGrpcBackend "Implements TTSBackend by launching and calling dengjen-tts-grpc.exe"
Want to help? Learn how to contribute to the ZirekHQ docs ›