You are viewing the documentation for a prerelease version. View Latest

Architecture

This page shows dengjen-tts’s own structure as a C4 Container, Component, and Code diagram. For how it fits into the wider ZirekHQ speech stack, see the org-wide System Context diagram.

Container

dengjen-tts is a single Rust workspace exposing the same synthesis engine through four different frontends. Go and Java bindings call the C ABI frontend directly rather than being separate frontends of their own.

C4Container
    Person(caller, "Caller", "CLI user, bound language, or another process")

    System_Boundary(dengjen, "dengjen-tts") {
        Container(cli, "dengjen", "Rust binary")
        Container(capi, "libdengjen", "C ABI", "Consumed by the Go and Java bindings")
        Container(pyext, "pydengjen", "Python extension")
        Container(grpc, "dengjen-tts-grpc", "gRPC server", "Vendored by dengjen-nvda")
        Container(engine, "Synthesis engine", "Rust library", "See Component diagram")
    }

    Rel(caller, cli, "Invokes")
    Rel(caller, capi, "Calls, via Go/Java bindings")
    Rel(caller, pyext, "Imports")
    Rel(caller, grpc, "Calls over gRPC")
    Rel(cli, engine, "Uses")
    Rel(capi, engine, "Uses")
    Rel(pyext, engine, "Uses")
    Rel(grpc, engine, "Uses")

Component

The engine’s model backends (Piper, Kokoro, MeloTTS) implement a shared DengjenModel trait; the engine itself wraps a chosen model in a decorator that adds prosody post-processing and parallel/batched dispatch.

C4Component
    Container_Boundary(engine, "Synthesis engine") {
        Component(core, "dengjen-tts-core", "Rust crate", "Defines the DengjenModel trait and shared types (Audio, SynthesisConfig, DengjenError)")
        Component(synth, "DengjenSpeechSynthesizer", "Rust struct (dengjen-tts / synth crate)", "Wraps an inner DengjenModel; applies prosody and dispatches via a rayon thread pool")
        Component(piper, "dengjen-tts-piper", "Rust crate", "VitsModel / VitsStreamingModel")
        Component(kokoro, "dengjen-tts-kokoro", "Rust crate", "KokoroModel")
        Component(melotts, "dengjen-tts-melotts", "Rust crate", "MeloTTSModel")
        Component(sonic, "dengjen-sonic-sys", "FFI binding", "libsonic, used for prosody (pitch/speed) adjustment")
    }

    Rel(synth, core, "Implements")
    Rel(piper, core, "Implements")
    Rel(kokoro, core, "Implements")
    Rel(melotts, core, "Implements")
    Rel(synth, piper, "Wraps one model, e.g.")
    Rel(synth, sonic, "Post-processes via")

Code

The interesting shape here is that DengjenSpeechSynthesizer is a decorator: it implements the exact same DengjenModel trait it wraps, rather than exposing a separate API for "a model with prosody applied."

classDiagram
    class DengjenModel {
        <<trait>>
        +phonemize_text()
        +speak_batch()
        +speak_one_sentence()
        +audio_output_info()
    }
    class VitsModel
    class VitsStreamingModel
    class KokoroModel
    class MeloTTSModel
    class DengjenSpeechSynthesizer {
        -backend: Arc~DengjenModel~
        +speak_batch()
    }

    DengjenModel <|.. VitsModel
    DengjenModel <|.. VitsStreamingModel
    DengjenModel <|.. KokoroModel
    DengjenModel <|.. MeloTTSModel
    DengjenModel <|.. DengjenSpeechSynthesizer
    DengjenSpeechSynthesizer o-- DengjenModel : wraps

Versioning

All 14 published crates share a single lockstep version via [workspace.package].version in the root Cargo.toml, rather than letting each crate version independently. A one-line fix in a single low-level crate (e.g. dengjen-audio-ops) still forces a version bump and republish of every crate, including unrelated ones, but the tradeoff is deliberate: any two crate versions that match are known to be compatible, with no compatibility matrix to consult.

Sibling repo dengjen-piper-rs piloted a hybrid split (lockstep core, independent versions for espeak-rs-sys/espeak-rs) and reverted it back to a single shared version, judging it simpler to reason about. dengjen-tts follows the same precedent and keeps lockstep versioning for all crates. scripts/next-version.sh and scripts/bump-version.sh continue to encode this assumption.

Phonemization

phonemize_text lives on DengjenModel itself rather than behind a standalone Phonemizer trait. The crate layout under crates/text/ (dengjen-espeak-phonemizer, dengjen-hebrew-phonemizer, dengjen-pinyin-phonemizer) already provides the reuse this issue asks about: Piper, Kokoro, and MeloTTS each pull in the phonemizer crates they need as optional Cargo dependencies and call them directly, without re-implementing espeak/Hebrew/pinyin phonemization per backend.

What differs per backend, and what a shared Phonemizer trait would have to paper over, is the adaptation from a phonemizer crate’s output to that model’s own phoneme representation:

  • Piper (crates/dengjen/models/piper/src/lib.rs) dispatches on the voice’s configured phoneme type, optionally diacritizes via libtashkeel first, and produces IPA strings shaped for its phoneme_id_map.

  • MeloTTS (crates/dengjen/models/melotts/src/phonemize.rs) dispatches through its own PhonemizerBackend enum and reshapes the engine output into (phone, tone) pairs, a representation Piper and Kokoro don’t use.

  • Kokoro (crates/dengjen/models/kokoro/src/inference.rs) hardcodes its target language and calls its own text_to_kokoro_phonemes wrapper.

None of the three share an output shape beyond the sentence-list Phonemes wrapper DengjenModel already returns. A Phonemizer trait sitting between DengjenModel and the phonemizer crates would either flatten to a type none of the backends actually want, or grow per-model associated types that just rename today’s per-model dispatch functions. The coupling is deliberate: resolved as issue #210.