Architecture
This page shows dengjen-tts’s own structure as a C4 Container, Component, and Code diagram. For how it fits into the wider ZirekHQ speech stack, see the org-wide System Context diagram.
Container
dengjen-tts is a single Rust workspace exposing the same synthesis engine through four different frontends. Go and Java bindings call the C ABI frontend directly rather than being separate frontends of their own.
C4Container
Person(caller, "Caller", "CLI user, bound language, or another process")
System_Boundary(dengjen, "dengjen-tts") {
Container(cli, "dengjen", "Rust binary")
Container(capi, "libdengjen", "C ABI", "Consumed by the Go and Java bindings")
Container(pyext, "pydengjen", "Python extension")
Container(grpc, "dengjen-tts-grpc", "gRPC server", "Vendored by dengjen-nvda")
Container(engine, "Synthesis engine", "Rust library", "See Component diagram")
}
Rel(caller, cli, "Invokes")
Rel(caller, capi, "Calls, via Go/Java bindings")
Rel(caller, pyext, "Imports")
Rel(caller, grpc, "Calls over gRPC")
Rel(cli, engine, "Uses")
Rel(capi, engine, "Uses")
Rel(pyext, engine, "Uses")
Rel(grpc, engine, "Uses")
Component
The engine’s model backends (Piper, Kokoro, MeloTTS) implement a shared
DengjenModel trait; the engine itself wraps a chosen model in a decorator
that adds prosody post-processing and parallel/batched dispatch.
C4Component
Container_Boundary(engine, "Synthesis engine") {
Component(core, "dengjen-tts-core", "Rust crate", "Defines the DengjenModel trait and shared types (Audio, SynthesisConfig, DengjenError)")
Component(synth, "DengjenSpeechSynthesizer", "Rust struct (dengjen-tts / synth crate)", "Wraps an inner DengjenModel; applies prosody and dispatches via a rayon thread pool")
Component(piper, "dengjen-tts-piper", "Rust crate", "VitsModel / VitsStreamingModel")
Component(kokoro, "dengjen-tts-kokoro", "Rust crate", "KokoroModel")
Component(melotts, "dengjen-tts-melotts", "Rust crate", "MeloTTSModel")
Component(sonic, "dengjen-sonic-sys", "FFI binding", "libsonic, used for prosody (pitch/speed) adjustment")
}
Rel(synth, core, "Implements")
Rel(piper, core, "Implements")
Rel(kokoro, core, "Implements")
Rel(melotts, core, "Implements")
Rel(synth, piper, "Wraps one model, e.g.")
Rel(synth, sonic, "Post-processes via")
Code
The interesting shape here is that DengjenSpeechSynthesizer is a decorator:
it implements the exact same DengjenModel trait it wraps, rather than
exposing a separate API for "a model with prosody applied."
classDiagram
class DengjenModel {
<<trait>>
+phonemize_text()
+speak_batch()
+speak_one_sentence()
+audio_output_info()
}
class VitsModel
class VitsStreamingModel
class KokoroModel
class MeloTTSModel
class DengjenSpeechSynthesizer {
-backend: Arc~DengjenModel~
+speak_batch()
}
DengjenModel <|.. VitsModel
DengjenModel <|.. VitsStreamingModel
DengjenModel <|.. KokoroModel
DengjenModel <|.. MeloTTSModel
DengjenModel <|.. DengjenSpeechSynthesizer
DengjenSpeechSynthesizer o-- DengjenModel : wraps
Versioning
All 14 published crates share a single lockstep version via
[workspace.package].version in the root Cargo.toml, rather than letting
each crate version independently. A one-line fix in a single low-level crate
(e.g. dengjen-audio-ops) still forces a version bump and republish of every
crate, including unrelated ones, but the tradeoff is deliberate: any two
crate versions that match are known to be compatible, with no compatibility
matrix to consult.
Sibling repo dengjen-piper-rs piloted a hybrid split (lockstep core,
independent versions for espeak-rs-sys/espeak-rs) and reverted it back to
a single shared version, judging it simpler to reason about. dengjen-tts
follows the same precedent and keeps lockstep versioning for all crates.
scripts/next-version.sh and scripts/bump-version.sh continue to encode
this assumption.
Phonemization
phonemize_text lives on DengjenModel itself rather than behind a
standalone Phonemizer trait. The crate layout under crates/text/
(dengjen-espeak-phonemizer, dengjen-hebrew-phonemizer,
dengjen-pinyin-phonemizer) already provides the reuse this issue asks
about: Piper, Kokoro, and MeloTTS each pull in the phonemizer crates they
need as optional Cargo dependencies and call them directly, without
re-implementing espeak/Hebrew/pinyin phonemization per backend.
What differs per backend, and what a shared Phonemizer trait would have to
paper over, is the adaptation from a phonemizer crate’s output to that
model’s own phoneme representation:
-
Piper (
crates/dengjen/models/piper/src/lib.rs) dispatches on the voice’s configured phoneme type, optionally diacritizes via libtashkeel first, and produces IPA strings shaped for itsphoneme_id_map. -
MeloTTS (
crates/dengjen/models/melotts/src/phonemize.rs) dispatches through its ownPhonemizerBackendenum and reshapes the engine output into(phone, tone)pairs, a representation Piper and Kokoro don’t use. -
Kokoro (
crates/dengjen/models/kokoro/src/inference.rs) hardcodes its target language and calls its owntext_to_kokoro_phonemeswrapper.
None of the three share an output shape beyond the sentence-list Phonemes
wrapper DengjenModel already returns. A Phonemizer trait sitting between
DengjenModel and the phonemizer crates would either flatten to a type none
of the backends actually want, or grow per-model associated types that just
rename today’s per-model dispatch functions. The coupling is deliberate:
resolved as issue #210.
Want to help? Learn how to contribute to the ZirekHQ docs ›