MOSS-8B Voice-Acting — Full Scaling Pilot

🎧 Top-3 audio samples for every group → (all 43 groups, ranked by combined reward, with prompts & scores)
Batched generation + reward scoring (invWER × normalized genuineness + voice-blend) with best-of-N selection. One expressive TTS model, three reference voices from held-out clips plus a diverse in-the-wild set. Scored 18862 clips (2968 Study-C takes, 12891 Study-D takes, 2960 restoration takes).

Conclusions

  1. Pipeline bottleneck: MOSS generation (0.54 s/clip at batch 100) is the slowest per-clip stage and dominates wall time; Parakeet ASR (0.32 s/clip) is the scoring bottleneck.
  2. Throughput: generation scales strongly with batch — 0.088→1.846 clips/s from batch 1→100 (36.38x realtime). Scoring runs at 2.7 clips/s (ASR-bound). Speech restoration is cheap (4.93 clips/s).
  3. Best-of-N plateau (Study D, n=1000, English): the combined-reward selection gain plateaus around k=600. Going 100→1000 adds only +0.177 reward (34% more than the N=100 gain); 500→1000 adds +0.039. Most of the win is captured by the first few hundred takes; 1000 does NOT meaningfully beat 300–500.
  4. Test-ref performance (Study C, fair text-independent): best-of-100 beats the held-out reference voice on genuineness by +1.679 (100% of groups) and voice-blend by +5.279 (100%). Combined-reward selection gain +0.648.
  5. Speech-restoration effect (Study B): Sidon is essentially neutral on the combined reward (+0.003 avg; Δgenu +0.040, Δ1−WER -0.000) — mixed by study (C +0.018, D -0.032). It does NOT systematically improve the reward, but it reshuffles near-tied top takes: the best-of-N pick changed in 48% of groups.

1 · Per-stage throughput (bottleneck: generation)

Single GPU. xRT = seconds of audio produced per second of wall time (higher is faster than realtime). Batching is the only generation speed lever.
Stagesamples/ssec/samplexRT
MOSS generation — batch 10.08811.4171.76x
MOSS generation — batch 321.5470.64630.8x
MOSS generation — batch 100 (used)1.8460.54236.38x
Voice conversion (per conversion)0.561.7771
Speech restoration (per clip)4.930.2028~80x
WER / ASR (per clip)3.10.3226
Voice-embed encode (per clip)21.450.0466
Blend regressor MLP (per clip)1404.80.00071
Genuineness MLP (per clip)1901.30.00053
SCORING pipeline total (per clip)2.70.3704

2 · Study C — test-ref performance (best-of-100)

5 held-out reference voices × 3 performance texts × {EN, DE} = 30 groups, 100 takes each; reference = the held-out clip. Δ = best-of-100 minus the reference clip's own score (positive = generation beats the reference).

Fair comparison to the reference voice (text-independent): best-of-100 beats the reference clip on genuineness by +1.679 (100% of groups) and on voice-blend by +5.279 (100% of groups). Best-of-100 selection gain on the combined reward (vs the mean take) = +0.648.

Note: the reference clips do not speak the performance text, so their invWER≈0 and the raw combined-reward gap vs reference is an artifact; genuineness/voice-blend (audio-only) are the fair reference comparison.

By language

langn grpΔ genuΔ blend%beat(combined)
English15+1.636+4.70095%
German15+1.722+5.85794%

By reference voice

refnΔ genuΔ blend
voice_a6+1.930+2.890
voice_d6+3.244+7.191
voice_c6+1.475+5.608
voice_b6+0.960+6.675
voice_e6+0.786+4.029

By performance text

textnΔ genuΔ blend
funny_ecstasy10+1.870+4.882
horror_basement10+1.732+5.729
love_vulnerable10+1.435+5.224

All 30 groups (ref × text × language), sorted by genuineness uplift

reftextlangbest-of-100 rewardsel-gainΔ genuΔ blend
voice_dfunny_ecstasyGE1.276+0.840+3.596+7.727
voice_dfunny_ecstasyEN0.922+0.560+3.526+6.590
voice_dhorror_basementGE1.228+0.782+3.511+7.277
voice_dhorror_basementEN1.119+0.631+3.206+7.499
voice_dlove_vulnerableEN0.837+0.428+2.915+7.404
voice_chorror_basementEN0.620+0.335+2.843+4.493
voice_dlove_vulnerableGE1.237+0.780+2.709+6.650
voice_afunny_ecstasyEN1.002+0.621+2.473+1.209
voice_alove_vulnerableEN0.872+0.548+2.097+3.460
voice_alove_vulnerableGE1.255+0.613+2.084+2.597
voice_afunny_ecstasyGE1.080+0.604+2.014+3.647
voice_bfunny_ecstasyGE1.392+0.906+1.898+7.276
voice_ahorror_basementGE1.026+0.632+1.792+3.628
voice_cfunny_ecstasyEN0.727+0.455+1.519+5.460
voice_cfunny_ecstasyGE1.012+0.745+1.348+5.409
voice_bhorror_basementGE1.405+0.870+1.342+7.405
voice_chorror_basementGE0.636+0.372+1.258+7.248
voice_ehorror_basementGE1.245+0.966+1.238+5.182
voice_clove_vulnerableGE0.986+0.737+1.147+6.535
voice_ahorror_basementEN1.066+0.648+1.122+2.800
voice_bfunny_ecstasyEN0.817+0.449+0.944+5.493
voice_elove_vulnerableGE0.774+0.464+0.781+5.738
voice_efunny_ecstasyEN0.674+0.441+0.773+2.256
voice_clove_vulnerableEN0.744+0.513+0.736+4.502
voice_elove_vulnerableEN0.845+0.588+0.707+1.806
voice_blove_vulnerableEN1.212+0.820+0.673+5.772
voice_efunny_ecstasyGE1.051+0.769+0.611+3.758
voice_ehorror_basementEN0.791+0.520+0.606+5.435
voice_blove_vulnerableGE1.555+1.053+0.504+7.780
voice_bhorror_basementEN1.264+0.737+0.401+6.323

3 · Study D — best-of-N plateau (n=1000, English)

13 groups (10 diverse in-the-wild voices + 3 held-out voices), 1000 takes each. For each k we draw 20 random size-k subsets, take the best combined reward in each, and average. Curve = mean best-of-k uplift over the reference vs k.
0.290.460.640.810.99501002003004005006007008009001000plateau k=600Best-of-N selection gain (combined reward − mean single take) vs N — yellow=avg, blue=per groupN (takes drawn, best-of-N)

Combined-reward selection gain: N=100 → +0.514, N=500 → +0.651, N=1000 → +0.691. Going 100→1000 adds only +0.177 (34% over the N=100 gain); 500→1000 adds +0.039. Plateau at k=600.

Avg best-of-N selection gain vs N (combined reward)

k→501002003004005006007008009001000
avg+0.436+0.514+0.562+0.599+0.640+0.651+0.657+0.671+0.675+0.688+0.691

Avg uplift over reference — genuineness & voice-blend (text-independent, fair) vs N

k→501002003004005006007008009001000
Δgenu+0.706+0.852+0.980+1.097+1.154+1.216+1.279+1.315+1.354+1.404+1.441
Δblend+2.953+3.586+4.241+4.463+4.709+4.933+4.940+5.072+5.140+5.223+5.327

Per-group selection gain vs N (combined reward)

group501002003004005006007008009001000
ref_g06+0.429+0.554+0.681+0.722+0.798+0.782+0.820+0.890+0.876+0.954+0.954
voice_c+0.364+0.641+0.626+0.696+0.853+0.888+0.870+0.887+0.901+0.918+0.920
ref_g13+0.473+0.526+0.570+0.624+0.679+0.698+0.687+0.725+0.707+0.749+0.761
ref_g09+0.503+0.566+0.601+0.636+0.673+0.634+0.687+0.694+0.694+0.717+0.723
voice_b+0.493+0.543+0.637+0.661+0.677+0.695+0.700+0.709+0.710+0.717+0.720
ref_g08+0.460+0.472+0.549+0.582+0.645+0.652+0.671+0.670+0.676+0.677+0.679
voice_a+0.495+0.599+0.609+0.646+0.666+0.671+0.673+0.674+0.674+0.674+0.674
ref_g00+0.389+0.461+0.489+0.556+0.575+0.621+0.606+0.627+0.641+0.643+0.644
ref_g11+0.442+0.499+0.536+0.589+0.602+0.629+0.619+0.633+0.639+0.641+0.642
ref_g01+0.514+0.532+0.572+0.605+0.606+0.609+0.614+0.607+0.624+0.625+0.628
ref_g03+0.460+0.519+0.569+0.566+0.581+0.605+0.604+0.605+0.611+0.612+0.615
ref_g10+0.328+0.366+0.424+0.439+0.482+0.496+0.499+0.509+0.513+0.511+0.513
ref_g05+0.320+0.399+0.446+0.464+0.480+0.486+0.495+0.498+0.503+0.502+0.504

4 · Study B — speech-restoration (Sidon) effect on scores

A deterministic ~50% of all groups (language-balanced) were passed through single-speaker speech restoration before scoring; both raw and restored versions were scored on the same takes.

Overall

metricraw μSidon μΔ (Sidon−raw)
inv_wer0.7730.772-0.000
blend1.8381.825-0.013
genu1.2371.277+0.040
combined0.3900.392+0.003

Study C

metricraw μSidon μΔ (Sidon−raw)
inv_wer0.7180.717-0.001
blend1.8471.920+0.073
genu1.2921.366+0.074
combined0.3700.388+0.018

Study D

metricraw μSidon μΔ (Sidon−raw)
inv_wer0.8970.897+0.000
blend1.8201.608-0.212
genu1.1101.073-0.037
combined0.4340.402-0.032

English

metricraw μSidon μΔ (Sidon−raw)
inv_wer0.8060.806-0.001
blend1.6891.645-0.044
genu1.1871.204+0.017
combined0.3840.379-0.005

German

metricraw μSidon μΔ (Sidon−raw)
inv_wer0.7100.709-0.000
blend2.1182.162+0.044
genu1.3301.414+0.084
combined0.4010.417+0.016

Best-of-N pick changed by Sidon in 48% of groups (does restoration change which take wins).

5 · Audio examples

External MP3 players. Top / median / bottom take per study by combined reward, plus raw-vs-restored pairs.

Study C — reward-ranked examples

TOP · voice_b · love_vulnerable · GE
combined 1.555 · 1−WER 1.0 · blend 7.9 · genu 2.51
“Ich hätte nie gedacht, dass ich das jemals laut aussprechen würde. Aber du gibst jedem zerbrochenen Teil das Gefühl, das”
MEDIAN · voice_a · love_vulnerable · GE
combined 0.341 · 1−WER 1.0 · blend 0.36 · genu 1.07
“Ich hätte nie gedacht, dass ich das jemals laut aussprechen würde, aber du gibst jedem zerbrochenen Teil das Gefühl, das”
BOTTOM · voice_d · love_vulnerable · EN
combined 0.0 · 1−WER 0.0 · blend 0.0 · genu 0.74
“”

Study D — reward-ranked examples

TOP · ref_g09 · fear_dread_apprehension_and_horror · EN
combined 1.381 · 1−WER 1.0 · blend 8.03 · genu 1.78
“Do not move. Do not make a single sound. It is right outside the door. If it hears us breathing in here, then we are fin”
MEDIAN · ref_g05 · malevolence_spite_sadism_malice_and_schadenfreude · EN
combined 0.412 · 1−WER 0.778 · blend 2.32 · genu 0.95
“Oh, how absolutely delicious. Everything you built is crumbling right before your eyes. I have waited so very long for t”
BOTTOM · ref_g11 · intoxication_stupor_altered_perception_and_being_drunk · EN
combined 0.0 · 1−WER 0.0 · blend 5.17 · genu 1.79
“”

Study B — same take: raw vs Sidon-restored

c_voice_a_funny_ecstasy_de · seed 19
raw
Sidon
best raw 1.080 → Sidon 1.773 (Δ+0.693)
c_voice_a_love_vulnerable_de · seed 75
raw
Sidon
best raw 1.255 → Sidon 1.627 (Δ+0.372)
Provenance kept generic. Reward = invWER × (norm genuineness + norm voice-blend), min-max normalized per study.