Using the voice manager

Open the voice manager from NVDA’s main menu, under Dengjen voice manager…​. It has two tabs: Download and Installed.

Download tab

Choose a language from the Language dropdown to filter the Available voices list, then select a voice to act on it.

  • Preview plays a short sample of the selected voice so you can hear it before downloading. The sample is streamed from the internet and nothing is installed. While it is playing, the same button becomes Stop preview.

  • Speaker, beside the preview button, is only enabled for voices trained with more than one speaker. It selects which speaker the preview uses.

  • Download standard variant and Download fast variant fetch the voice. Each button is disabled when that variant is already installed, and the fast button is also disabled for voices that have no fast variant.

  • Refresh voices list fetches the catalogue again instead of reusing the copy cached for this session.

Installed tab

The Installed voices list shows each installed voice with its variant, quality and language.

  • Voice model card…​ displays the MODEL_CARD file shipped with the voice, which records where its training data came from and how it is licensed. Not every voice includes one.

  • Remove voice…​ deletes the selected voice after asking you to confirm. It stays disabled unless you have at least two voices installed, and it will not remove the voice that is currently in use.

  • Install from local file installs a voice from a .tar.gz or .tgz archive you already have.

  • Import voices from Sonata copies voices you downloaded with the older Sonata Neural Voices add-on. It only appears while that add-on’s voices are still present and have not been imported yet. The originals are left in place, so Sonata keeps working.

After installing from a local archive or removing a voice, the add-on reloads the synthesizer for you, so the change applies immediately. After a download the new voice shows up in the voice manager right away; if NVDA’s own voice list has not picked it up, restart NVDA.

Voice settings

With Dengjen Neural Voices selected as your synthesizer, the following appear in NVDA’s speech settings (NVDA menu > Preferences > Settings > Speech).

Voice lists your installed voices as name (language) - quality.

Variant switches between the Standard and Fast build of the current voice. Only the variants you actually have installed are listed.

Speaker applies to voices trained with several speakers; on a single-speaker voice it has no effect. It is also available in the synth settings ring.

Rate, Volume and Pitch behave as they do for any NVDA synthesizer. With Rate boost turned off, the rate slider only covers the lower part of the engine’s speed range; turning it on spreads the slider across the whole range, which allows much faster speech.

Fine-tuning how a voice sounds

Length scale, Noise scale and Noise w expose the Piper model’s own inference parameters. All three work the same way: the slider runs from 0 to 100, and 50 means the voice’s trained default, so returning a slider to 50 undoes your changes to it. Of the three, only Length scale is offered in the synth settings ring.

  • Length scale sets how long each speech sound is held. Higher values draw speech out, lower values compress it. This is a separate mechanism from Rate and the two combine, so it is usually easiest to set your speed with Rate and only reach for this if a voice’s natural pacing bothers you.

  • Noise scale sets how much variation the model puts into tone and inflection. Higher values sound more expressive but less predictable.

  • Noise w sets how much the duration of individual speech sounds varies, which comes across as rhythm. Higher values sound less mechanical but can blur articulation.

Above 50, the sliders scale up to twice the voice’s default for Length scale and three times the default for Noise scale and Noise w. Because 50 always means "this voice’s default", a given slider position keeps its meaning when you switch to a different voice.

A note on voice quality

The currently available voices are trained using freely available TTS datasets, which are generally of low quality (mostly public domain audio books or research-quality recordings).

Additionally, these datasets are not comprehensive, hence some voices may exhibit incorrect or weird pronunciation. Both issues could be resolved by using better datasets for training.

Luckily, the Piper developer and some developers from the blind and vision-impaired community are working on training better voices.