What is this? We took the fine-tuned 8-billion-parameter MOSS-TTS-v1.5 voice-acting text-to-speech model and systematically swept its two most important sampling knobs to find the settings that produce the best speech. The model reads a target line of text plus a natural-language performance instruction (age, timbre, emotion, pacing…) and outputs 24 kHz audio built from 32 discrete acoustic codebooks.
Temperature controls how random the model's choices are while it generates each audio token. Low temperature (0.5) = safe, repetitive, "flat" but reliable; high temperature (1.2) = more expressive and varied, but more prone to mistakes, artifacts, or drifting off the script. We tested 8 values: 0.5 → 1.2.
Audio repetition penalty discourages the model from getting stuck repeating the same acoustic token (a common failure that sounds like buzzing, stuttering, or a stuck note). Higher = stronger push away from repeats. We tested 3 values: 1.1, 1.2, 1.3. That gives an 8 × 3 = 24-config grid.
With reference audio: the model is also given a short recording of the target voice to clone (voice + style are anchored). Without reference audio: the model gets only the text and the written instruction and must invent a voice that fits the description. Comparing the two tells us how much the settings should change when you do or don't provide a reference clip.
For every config we generated 5 samples per prompt across a held-out set of English and German prompts, then scored each clip. Total: 2296 clips scored.
WER (Word Error Rate) — we transcribe the generated speech with an ASR model (Parakeet-TDT-0.6B, multilingual) and compare against the intended text. Lower WER = the words are clearer / more intelligible. (invWER = 1−WER, so higher is better — used inside the balanced score.)
Blend score (0–10) — a learned quality head (on VoiceCLAP audio embeddings) that rates overall naturalness/production quality of the take. Higher = better.
Genuineness — how real / authentically-human (vs synthetic-sounding) the voice is, predicted by MLP probes on two independent VoiceCLAP embedding spaces: VoiceCLAP-large-v2 (3584-d, our headline metric) and VoiceCLAP-commercial (768-d, reported for corroboration). Higher = more genuine.
Each table is the 8 temperatures (rows) × 3 repetition penalties (columns). Colour encodes the metric value on a shared scale across both modes so the two panels are directly comparable; the ★ marks the single best cell in each panel.
| rp 1.1 | rp 1.2 | rp 1.3 | |
|---|---|---|---|
| t 0.5 | 0.186 | 0.163 | 0.182 |
| t 0.6 | 0.214 | 0.100 | 0.198 |
| t 0.7 | 0.209 | 0.168 | 0.226 |
| t 0.8 | 0.166 | 0.146 | 0.163 |
| t 0.9 | 0.106 | 0.140 | 0.285 |
| t 1.0 | 0.076 ★ | 0.204 | 0.226 |
| t 1.1 | 0.133 | 0.160 | 0.236 |
| t 1.2 | 0.102 | 0.258 | 0.263 |
| rp 1.1 | rp 1.2 | rp 1.3 | |
|---|---|---|---|
| t 0.5 | 0.109 | 0.058 | 0.125 |
| t 0.6 | 0.108 | 0.135 | 0.066 |
| t 0.7 | 0.117 | 0.128 | 0.112 |
| t 0.8 | 0.056 | 0.099 | 0.124 |
| t 0.9 | 0.052 | 0.046 | 0.091 |
| t 1.0 | 0.068 | 0.125 | 0.097 |
| t 1.1 | 0.034 ★ | 0.142 | 0.135 |
| t 1.2 | 0.119 | 0.110 | 0.121 |
| rp 1.1 | rp 1.2 | rp 1.3 | |
|---|---|---|---|
| t 0.5 | 3.60 | 2.50 | 2.11 |
| t 0.6 | 3.61 ★ | 2.47 | 2.11 |
| t 0.7 | 3.00 | 2.18 | 2.26 |
| t 0.8 | 3.51 | 2.18 | 2.10 |
| t 0.9 | 3.19 | 2.19 | 2.27 |
| t 1.0 | 3.33 | 2.22 | 2.00 |
| t 1.1 | 3.14 | 2.22 | 2.34 |
| t 1.2 | 3.12 | 2.39 | 2.27 |
| rp 1.1 | rp 1.2 | rp 1.3 | |
|---|---|---|---|
| t 0.5 | 3.01 | 2.57 | 2.68 |
| t 0.6 | 3.29 | 2.68 | 2.62 |
| t 0.7 | 2.78 | 2.51 | 2.43 |
| t 0.8 | 3.35 ★ | 2.66 | 2.35 |
| t 0.9 | 2.93 | 1.87 | 2.27 |
| t 1.0 | 2.87 | 2.48 | 2.03 |
| t 1.1 | 2.84 | 2.53 | 2.27 |
| t 1.2 | 2.88 | 2.73 | 2.69 |
| rp 1.1 | rp 1.2 | rp 1.3 | |
|---|---|---|---|
| t 0.5 | 0.89 | 0.91 | 1.08 |
| t 0.6 | 0.94 | 0.87 | 1.00 |
| t 0.7 | 0.86 | 0.98 | 1.12 |
| t 0.8 | 1.06 | 1.07 | 1.24 ★ |
| t 0.9 | 0.91 | 1.05 | 1.15 |
| t 1.0 | 1.07 | 1.09 | 1.13 |
| t 1.1 | 1.08 | 1.17 | 1.24 |
| t 1.2 | 1.07 | 1.14 | 1.21 |
| rp 1.1 | rp 1.2 | rp 1.3 | |
|---|---|---|---|
| t 0.5 | 0.59 | 0.82 | 0.74 |
| t 0.6 | 0.79 | 0.84 | 0.77 |
| t 0.7 | 0.76 | 0.85 | 0.75 |
| t 0.8 | 0.94 | 0.76 | 1.00 |
| t 0.9 | 0.80 | 0.73 | 0.93 |
| t 1.0 | 0.93 | 1.00 | 0.90 |
| t 1.1 | 0.92 | 0.99 | 0.95 |
| t 1.2 | 0.88 | 1.03 | 1.11 ★ |
| rp 1.1 | rp 1.2 | rp 1.3 | |
|---|---|---|---|
| t 0.5 | 1.23 | 1.37 | 1.44 |
| t 0.6 | 1.29 | 1.37 | 1.47 |
| t 0.7 | 1.28 | 1.32 | 1.46 |
| t 0.8 | 1.43 | 1.38 | 1.47 |
| t 0.9 | 1.29 | 1.43 | 1.45 |
| t 1.0 | 1.48 | 1.54 | 1.58 |
| t 1.1 | 1.50 | 1.66 | 1.58 |
| t 1.2 | 1.46 | 1.67 | 1.75 ★ |
| rp 1.1 | rp 1.2 | rp 1.3 | |
|---|---|---|---|
| t 0.5 | 0.95 | 1.18 | 1.10 |
| t 0.6 | 1.12 | 1.21 | 1.15 |
| t 0.7 | 1.01 | 1.22 | 1.15 |
| t 0.8 | 1.29 | 1.22 | 1.23 |
| t 0.9 | 1.16 | 1.12 | 1.13 |
| t 1.0 | 1.15 | 1.36 ★ | 1.19 |
| t 1.1 | 1.23 | 1.26 | 1.27 |
| t 1.2 | 1.26 | 1.34 | 1.28 |
| Language | Mode | mean WER | mean blend | mean genu (large) | mean genu (comm) | best genuineness config |
|---|---|---|---|---|---|---|
| English | With reference audio | 0.192 | 2.94 | 1.297 | 1.651 | temp 0.8, rep-pen 1.3 |
| English | Without reference audio | 0.093 | 2.95 | 1.068 | 1.316 | temp 1.2, rep-pen 1.3 |
| German | With reference audio | 0.133 | 1.31 | 0.132 | 0.7 | temp 1.2, rep-pen 1.3 |
| German | Without reference audio | 0.122 | 1.49 | 0.129 | 0.738 | temp 1.2, rep-pen 1.2 |
Overall means. With-ref: WER 0.18, blend 2.6, genuineness(large) 1.056, genuineness(comm) 1.455. Without-ref: WER 0.099, blend 2.64, genuineness(large) 0.869, genuineness(comm) 1.194.
Difference (without − with). WER -0.081, blend +0.04, genuineness(large) -0.187, genuineness(comm) -0.261.
| Metric | Best — With reference | Best — Without reference |
|---|---|---|
| WER (word error rate) | temp 1.0, rep-pen 1.1 0.0761 | temp 1.1, rep-pen 1.1 0.0339 |
| Blend score (0-10) | temp 0.6, rep-pen 1.1 3.6129 | temp 0.8, rep-pen 1.1 3.355 |
| Genuineness — VoiceCLAP-large-v2 (headline) | temp 0.8, rep-pen 1.3 1.243 | temp 1.2, rep-pen 1.3 1.111 |
| Genuineness — VoiceCLAP-commercial | temp 1.2, rep-pen 1.3 1.7517 | temp 1.0, rep-pen 1.2 1.3619 |
| Balanced (rank-avg of invWER+blend+genuineness) | temp 1.0, rep-pen 1.1 | temp 0.8, rep-pen 1.1 |
With a reference clip: temperature 1.0, repetition penalty 1.1.
Without a reference clip: temperature 0.8, repetition penalty 1.1.
These are the balanced optima (best average rank across intelligibility, blend quality, and genuineness). If you care most about one axis, use the per-metric best-config table above. In general, push temperature higher for genuineness/expressiveness and lower for intelligibility (WER); the repetition penalty mainly guards against stuck-token artifacts.