Best-of-100 voice-acting selection across speaker identities
Study: best-of-100 voice-acting selection across speaker identities.
Model: an 8B autoregressive expressive TTS model, generated WITH reference-voice conditioning
(audio temperature 1.0, repetition penalty 1.1, text temperature 0.7, top-p 0.95, top-k 25, bf16).
Reference source: a large open crowd-sourced corpus of expressive human speech (one distinct
speaker clip per group, diverse emotions). Reward: combined = invWER x (norm_genu + norm_blend),
where genu is a commercial-voice genuineness score (0-6) and blend an expressive-quality score (0-10),
each min-max normalized across all generated clips; invWER = max(0, 1-WER) from multilingual ASR.
Achieved throughput: 2.87 clips/s on a single GPU.
What this tests
For 50 distinct reference voices (diverse emotions), we author one
3-sentence English voice-acting script + an emotion-matched delivery instruction, and a natural German
translation of the script. Each script is generated 100 times under three reference conditions:
the original reference voice; the same content voice-converted to a random speaker A; and
voice-converted to a random speaker B. Before conversion the source is randomly sped up by
+0-50%. We ask: does the best-of-100 uplift over the reference hold across different speaker identities,
across English vs the German translation, and how does the random speed-up affect things?
Conclusions
- Best-of-100 lifts quality well above the reference in every condition. Mean genuineness ceiling uplift = +1.46 (0-6 scale), blend ceiling uplift = +3.33 (0-10 scale) across all conditions.
- The best take beats the reference in genuineness for 100% of group-cells and in blend for 100% on average.
- Speaker identity: original voice genu uplift +1.36 vs converted-A +1.99 vs converted-B +1.02; blend uplift +3.93 / +3.38 / +2.70. The uplift holds across all three identities.
- Language (EN vs DE-translation): genu uplift EN +1.08 vs DE +1.83; blend uplift EN +1.54 vs DE +5.13.
- Diminishing returns with N: mean genu uplift grows from -0.05 (N=1) to +0.97 (N=10), +1.20 (N=32), +1.40 (N=100) — the N=100 ceiling sits above the earlier N=10/N=32 ceilings but gains taper.
- Source speed-up (+0-50%) before conversion: correlation with genu uplift r=-0.44, with picked-clip invWER r=-0.16 — a notable effect.
Uplift by condition x language (means over groups)
| Condition x Language | ref genu | best genu | genu ceiling uplift | ref blend | best blend | blend ceiling uplift | % beat ref (genu) | % beat ref (blend) |
|---|
| Original reference voice / English | 0.53 | 1.60 | +1.07 | 0.44 | 2.52 | +2.08 | 100% | 100% |
| Original reference voice / German | 0.53 | 2.19 | +1.66 | 0.44 | 6.21 | +5.78 | 100% | 100% |
| Voice-converted to a random speaker A / English | 1.37 | 2.46 | +1.09 | 1.18 | 1.90 | +0.72 | 100% | 100% |
| Voice-converted to a random speaker A / German | 1.37 | 4.25 | +2.88 | 1.18 | 7.20 | +6.03 | 100% | 100% |
| Voice-converted to a random speaker B / English | 1.35 | 2.45 | +1.09 | 1.13 | 2.94 | +1.81 | 100% | 100% |
| Voice-converted to a random speaker B / German | 1.59 | 2.54 | +0.95 | 0.00 | 3.58 | +3.58 | 100% | 100% |
Best-of-N ceiling curve (N=1/10/32/100, from the same 100 samples)
| best-of-N | mean genu | genu uplift vs ref | mean blend | blend uplift vs ref |
|---|
| N=1 | 1.10 | -0.05 | 0.22 | -0.56 |
| N=10 | 2.13 | +0.97 | 1.79 | +1.01 |
| N=32 | 2.36 | +1.20 | 2.30 | +1.52 |
| N=100 | 2.56 | +1.40 | 3.90 | +3.12 |
Per-group pages
| Group | Emotion | mean genu ceiling uplift | mean blend ceiling uplift | link |
| g00 | contempt disdain loathing and detestation | +1.50 | +3.24 | open |
| g01 | sadness sorrow grief melancholy and heartache | +0.85 | +2.40 | open |