Best-of-100 voice-acting selection across speaker identities

Study: best-of-100 voice-acting selection across speaker identities. Model: an 8B autoregressive expressive TTS model, generated WITH reference-voice conditioning (audio temperature 1.0, repetition penalty 1.1, text temperature 0.7, top-p 0.95, top-k 25, bf16). Reference source: a large open crowd-sourced corpus of expressive human speech (one distinct speaker clip per group, diverse emotions). Reward: combined = invWER x (norm_genu + norm_blend), where genu is a commercial-voice genuineness score (0-6) and blend an expressive-quality score (0-10), each min-max normalized across all generated clips; invWER = max(0, 1-WER) from multilingual ASR. Achieved throughput: 2.87 clips/s on a single GPU.

What this tests

For 50 distinct reference voices (diverse emotions), we author one 3-sentence English voice-acting script + an emotion-matched delivery instruction, and a natural German translation of the script. Each script is generated 100 times under three reference conditions: the original reference voice; the same content voice-converted to a random speaker A; and voice-converted to a random speaker B. Before conversion the source is randomly sped up by +0-50%. We ask: does the best-of-100 uplift over the reference hold across different speaker identities, across English vs the German translation, and how does the random speed-up affect things?

Conclusions

Uplift by condition x language (means over groups)

Condition x Languageref genubest genugenu ceiling upliftref blendbest blendblend ceiling uplift% beat ref (genu)% beat ref (blend)
Original reference voice / English0.531.60+1.070.442.52+2.08100%100%
Original reference voice / German0.532.19+1.660.446.21+5.78100%100%
Voice-converted to a random speaker A / English1.372.46+1.091.181.90+0.72100%100%
Voice-converted to a random speaker A / German1.374.25+2.881.187.20+6.03100%100%
Voice-converted to a random speaker B / English1.352.45+1.091.132.94+1.81100%100%
Voice-converted to a random speaker B / German1.592.54+0.950.003.58+3.58100%100%

Best-of-N ceiling curve (N=1/10/32/100, from the same 100 samples)

best-of-Nmean genugenu uplift vs refmean blendblend uplift vs ref
N=11.10-0.050.22-0.56
N=102.13+0.971.79+1.01
N=322.36+1.202.30+1.52
N=1002.56+1.403.90+3.12

Per-group pages