You are viewing the documentation for a prerelease version. View Latest

Architecture

This page shows dengjen-nvda’s own structure as a C4 Container, Component, and Code diagram. For how it fits into the wider ZirekHQ speech stack, see the org-wide System Context diagram.

Container

dengjen-nvda is not one process but two: NVDA loads the add-on itself in-process, and the add-on launches a vendored, separately-built dengjen-tts-grpc.exe binary (a dengjen-tts frontend) as its own subprocess, talking to it over gRPC.

C4Container
    System_Ext(nvda, "NVDA", "The screen reader host process")

    System_Boundary(addon, "dengjen-nvda") {
        Container(synthdriver, "dengjen_neural_voices", "Python 3.13 add-on", "Loaded in-process by NVDA as a synthesizer driver")
        Container(grpcbin, "dengjen-tts-grpc.exe", "Fetched Rust binary (pinned in dengjen-tts.lock)", "Launched via subprocess.Popen; a gRPC-frontend build of dengjen-tts")
    }

    System_Ext(voices, "Voice catalogue", "External download source for voices")

    Rel(nvda, synthdriver, "Loads / calls SynthDriver API")
    Rel(synthdriver, grpcbin, "Launches, then calls over gRPC")
    Rel(synthdriver, voices, "Downloads voices from")

Component

The add-on itself is a small hexagonal (ports-and-adapters) design: domain logic that knows nothing about NVDA or gRPC, one seam (TTSBackend) that any synthesis backend must satisfy, and two adapters either side of that seam.

C4Component
    Container_Boundary(addon, "dengjen_neural_voices") {
        Component(domain, "domain/tts_system.py", "Python module", "Pure domain logic: voices, speech options. No gRPC, no NVDA imports.")
        Component(port, "ports/tts_backend.py", "Python Protocol", "TTSBackend -- the one interface a TTS engine adapter must satisfy")
        Component(grpcadapter, "adapters/sonata_grpc", "Python package", "Implements TTSBackend by launching and talking to dengjen-tts-grpc.exe")
        Component(nvdaadapter, "adapters/nvda/synth_driver.py", "Python module", "Exposes the domain layer to NVDA's own SynthDriver API")
    }

    Rel(nvdaadapter, domain, "Creates/manages DengjenVoice")
    Rel(domain, port, "Calls self.backend, typed as")
    Rel(grpcadapter, port, "Implements")

Adding a model backend

dengjen-tts (the vendored gRPC engine) can support more model types than Piper — it already ships Kokoro and MeloTTS. Wiring a new one into this add-on needs:

  1. A ModelCatalog implementation (addon/globalPlugins/dengjen_tts_global_plugin/model_catalog.py’s `Protocol) with model_type, is_installed(), and install(), following kokoro_download.py’s `KokoroCatalog as a template.

  2. That catalog’s installer writes two files per voice directory: the model backend’s own config.json (whatever shape dengjen-tts’s loader for that `model_type expects — check crates/dengjen/models/<type>/src/config.rs in zirekhq/dengjen-tts), and a voice.json sidecar via domain/voice_metadata.write() carrying {model_type, name, language, description} — the only fields DengjenVoice.from_path() reads.

  3. If the new model type has per-utterance parameters the existing gRPC ProsodyControls message doesn’t carry, that needs a change in zirekhq/dengjen-tts first (a different repo) — not an addon-side concern.

  4. If the new model type’s tuning parameters map onto the noise_scale/length_scale/noise_w sliders, add a ModelProfile entry to _PROFILES in domain/model_profiles.py, listing which sliders apply and, where the engine’s key differs from the slider name, the parameters key (MeloTTS maps noise_w to noise_scale_w).

  5. Add a DENGJEN_<TYPE>_VOICES_DIR constant to const.py and register it in DengjenTextToSpeechSystem.load_all_voices_from_nvda_config_dir() (domain/tts_system.py) — that method iterates a fixed tuple of directories, not a dynamic scan of voices/*, so a new backend’s directory has to be added there explicitly or its installed voices are invisible to NVDA.

No install directory is shared between model types: each gets its own DENGJEN_VOICES_BASE_DIR/voices/<model_type>/ (see const.py), scanned only once its constant is registered per step 5 above. The one exception is MeloTTS: voices installed from a local archive currently land in voices/piper/, because local install passes a single directory and no MeloTTS directory constant is registered.

Code

TTSBackend is a Protocol, not a base class — the sonata_grpc adapter satisfies it structurally. Its methods are split deliberately: everything NVDA calls synchronously from its main thread has to stay blocking; only synthesize() is a true async generator.

classDiagram
    class TTSBackend {
        <<Protocol>>
        +initialize()
        +check_version()
        +shutdown()
        +load_voice()
        +get_synth_options()
        +set_synth_options()
        +synthesize() AsyncIterator~bytes~
    }
    class SynthOptions {
        <<frozen dataclass>>
    }
    class LoadedVoice {
        <<frozen dataclass>>
    }
    class SonataGrpcBackend

    TTSBackend <|.. SonataGrpcBackend
    TTSBackend ..> SynthOptions : passes
    TTSBackend ..> LoadedVoice : returns
    note for SonataGrpcBackend "Implements TTSBackend by launching and calling dengjen-tts-grpc.exe"