VoiceNet dimension LoRAs — scale sweep (page 4/4)

For each VoiceNet perceptual dimension, two rank-32 LoRAs on MOSS-TTS-4.55B — one trained on the HIGH end (score ≥5), one on the LOW end (≤1). Each adapter is merged at four intensities (50 / 100 / 150 / 200%) and generates the same neutral sentence three times; the three takes are ordered with the longest playtime first (most likely to contain full speech). The number is that dimension's own score of the take (0–1). No reference audio.

Pages: 1 2 3 4  ·  🌐 Deutsche Version
Warmthvn_WARM
▲ HIGH
50%
0.89 · 7.84s
0.54 · 6.88s
0.8 · 6.32s
100%
0.88 · 6.32s
0.89 · 6.16s
0.93 · 6.0s
150%
0.93 · 7.68s
0.95 · 7.28s
0.86 · 5.84s
200%
0.89 · 6.72s
0.85 · 6.4s
0.81 · 6.32s
▼ LOW
50%
0.73 · 10.8s
0.7 · 7.76s
0.1 · 1.04s
100%
0.07 · 2.88s
0.12 · 1.68s
0.11 · 1.52s
150%
0.07 · 2.32s
0.19 · 0.56s
0.17 · 0.56s
200%
0.06 · 2.16s
0.06 · 1.04s
0.05 · 0.56s
Whispery / breathyvn_S_WHIS
▲ HIGH
50%
0.65 · 7.36s
0.75 · 7.2s
0.93 · 6.56s
100%
0.62 · 9.04s
0.85 · 6.88s
0.79 · 6.0s
150%
0.76 · 5.92s
0.82 · 4.0s
0.81 · 3.76s
200%
0.78 · 7.2s
0.6 · 5.12s
0.79 · 4.32s
▼ LOW
50%
0.84 · 7.04s
0.64 · 5.68s
0.09 · 5.2s
100%
0.0 · 6.32s
0.58 · 6.16s
0.32 · 3.84s
150%
0.0 · 5.12s
0.22 · 4.72s
0.38 · 3.84s
200%
0.0 · 5.76s
0.03 · 5.52s
0.36 · 5.44s
Pages: 1 2 3 4  ·  🌐 Deutsche Version