Batched generation + reward scoring (invWER × normalized genuineness + voice-blend) with best-of-N selection. One expressive TTS model, three reference voices from held-out clips plus a diverse in-the-wild set. Scored 18862 clips (2968 Study-C takes, 12891 Study-D takes, 2960 restoration takes).
Conclusions
Pipeline bottleneck: MOSS generation (0.54 s/clip at batch 100) is the slowest per-clip stage and dominates wall time; Parakeet ASR (0.32 s/clip) is the scoring bottleneck.
Throughput: generation scales strongly with batch — 0.088→1.846 clips/s from batch 1→100 (36.38x realtime). Scoring runs at 2.7 clips/s (ASR-bound). Speech restoration is cheap (4.93 clips/s).
Best-of-N plateau (Study D, n=1000, English): the combined-reward selection gain plateaus around k=600. Going 100→1000 adds only +0.177 reward (34% more than the N=100 gain); 500→1000 adds +0.039. Most of the win is captured by the first few hundred takes; 1000 does NOT meaningfully beat 300–500.
Test-ref performance (Study C, fair text-independent): best-of-100 beats the held-out reference voice on genuineness by +1.679 (100% of groups) and voice-blend by +5.279 (100%). Combined-reward selection gain +0.648.
Speech-restoration effect (Study B): Sidon is essentially neutral on the combined reward (+0.003 avg; Δgenu +0.040, Δ1−WER -0.000) — mixed by study (C +0.018, D -0.032). It does NOT systematically improve the reward, but it reshuffles near-tied top takes: the best-of-N pick changed in 48% of groups.
1 · Per-stage throughput (bottleneck: generation)
Single GPU. xRT = seconds of audio produced per second of wall time (higher is faster than realtime). Batching is the only generation speed lever.
Stage
samples/s
sec/sample
xRT
MOSS generation — batch 1
0.088
11.417
1.76x
MOSS generation — batch 32
1.547
0.646
30.8x
MOSS generation — batch 100 (used)
1.846
0.542
36.38x
Voice conversion (per conversion)
0.56
1.7771
—
Speech restoration (per clip)
4.93
0.2028
~80x
WER / ASR (per clip)
3.1
0.3226
—
Voice-embed encode (per clip)
21.45
0.0466
—
Blend regressor MLP (per clip)
1404.8
0.00071
—
Genuineness MLP (per clip)
1901.3
0.00053
—
SCORING pipeline total (per clip)
2.7
0.3704
—
2 · Study C — test-ref performance (best-of-100)
5 held-out reference voices × 3 performance texts × {EN, DE} = 30 groups, 100 takes each; reference = the held-out clip. Δ = best-of-100 minus the reference clip's own score (positive = generation beats the reference).
Fair comparison to the reference voice (text-independent): best-of-100 beats the reference clip on genuineness by +1.679 (100% of groups) and on voice-blend by +5.279 (100% of groups). Best-of-100 selection gain on the combined reward (vs the mean take) = +0.648.
Note: the reference clips do not speak the performance text, so their invWER≈0 and the raw combined-reward gap vs reference is an artifact; genuineness/voice-blend (audio-only) are the fair reference comparison.
By language
lang
n grp
Δ genu
Δ blend
%beat(combined)
English
15
+1.636
+4.700
95%
German
15
+1.722
+5.857
94%
By reference voice
ref
n
Δ genu
Δ blend
voice_a
6
+1.930
+2.890
voice_d
6
+3.244
+7.191
voice_c
6
+1.475
+5.608
voice_b
6
+0.960
+6.675
voice_e
6
+0.786
+4.029
By performance text
text
n
Δ genu
Δ blend
funny_ecstasy
10
+1.870
+4.882
horror_basement
10
+1.732
+5.729
love_vulnerable
10
+1.435
+5.224
All 30 groups (ref × text × language), sorted by genuineness uplift
ref
text
lang
best-of-100 reward
sel-gain
Δ genu
Δ blend
voice_d
funny_ecstasy
GE
1.276
+0.840
+3.596
+7.727
voice_d
funny_ecstasy
EN
0.922
+0.560
+3.526
+6.590
voice_d
horror_basement
GE
1.228
+0.782
+3.511
+7.277
voice_d
horror_basement
EN
1.119
+0.631
+3.206
+7.499
voice_d
love_vulnerable
EN
0.837
+0.428
+2.915
+7.404
voice_c
horror_basement
EN
0.620
+0.335
+2.843
+4.493
voice_d
love_vulnerable
GE
1.237
+0.780
+2.709
+6.650
voice_a
funny_ecstasy
EN
1.002
+0.621
+2.473
+1.209
voice_a
love_vulnerable
EN
0.872
+0.548
+2.097
+3.460
voice_a
love_vulnerable
GE
1.255
+0.613
+2.084
+2.597
voice_a
funny_ecstasy
GE
1.080
+0.604
+2.014
+3.647
voice_b
funny_ecstasy
GE
1.392
+0.906
+1.898
+7.276
voice_a
horror_basement
GE
1.026
+0.632
+1.792
+3.628
voice_c
funny_ecstasy
EN
0.727
+0.455
+1.519
+5.460
voice_c
funny_ecstasy
GE
1.012
+0.745
+1.348
+5.409
voice_b
horror_basement
GE
1.405
+0.870
+1.342
+7.405
voice_c
horror_basement
GE
0.636
+0.372
+1.258
+7.248
voice_e
horror_basement
GE
1.245
+0.966
+1.238
+5.182
voice_c
love_vulnerable
GE
0.986
+0.737
+1.147
+6.535
voice_a
horror_basement
EN
1.066
+0.648
+1.122
+2.800
voice_b
funny_ecstasy
EN
0.817
+0.449
+0.944
+5.493
voice_e
love_vulnerable
GE
0.774
+0.464
+0.781
+5.738
voice_e
funny_ecstasy
EN
0.674
+0.441
+0.773
+2.256
voice_c
love_vulnerable
EN
0.744
+0.513
+0.736
+4.502
voice_e
love_vulnerable
EN
0.845
+0.588
+0.707
+1.806
voice_b
love_vulnerable
EN
1.212
+0.820
+0.673
+5.772
voice_e
funny_ecstasy
GE
1.051
+0.769
+0.611
+3.758
voice_e
horror_basement
EN
0.791
+0.520
+0.606
+5.435
voice_b
love_vulnerable
GE
1.555
+1.053
+0.504
+7.780
voice_b
horror_basement
EN
1.264
+0.737
+0.401
+6.323
3 · Study D — best-of-N plateau (n=1000, English)
13 groups (10 diverse in-the-wild voices + 3 held-out voices), 1000 takes each. For each k we draw 20 random size-k subsets, take the best combined reward in each, and average. Curve = mean best-of-k uplift over the reference vs k.
Combined-reward selection gain: N=100 → +0.514, N=500 → +0.651, N=1000 → +0.691. Going 100→1000 adds only +0.177 (34% over the N=100 gain); 500→1000 adds +0.039. Plateau at k=600.
Avg best-of-N selection gain vs N (combined reward)
k→
50
100
200
300
400
500
600
700
800
900
1000
avg
+0.436
+0.514
+0.562
+0.599
+0.640
+0.651
+0.657
+0.671
+0.675
+0.688
+0.691
Avg uplift over reference — genuineness & voice-blend (text-independent, fair) vs N
k→
50
100
200
300
400
500
600
700
800
900
1000
Δgenu
+0.706
+0.852
+0.980
+1.097
+1.154
+1.216
+1.279
+1.315
+1.354
+1.404
+1.441
Δblend
+2.953
+3.586
+4.241
+4.463
+4.709
+4.933
+4.940
+5.072
+5.140
+5.223
+5.327
Per-group selection gain vs N (combined reward)
group
50
100
200
300
400
500
600
700
800
900
1000
ref_g06
+0.429
+0.554
+0.681
+0.722
+0.798
+0.782
+0.820
+0.890
+0.876
+0.954
+0.954
voice_c
+0.364
+0.641
+0.626
+0.696
+0.853
+0.888
+0.870
+0.887
+0.901
+0.918
+0.920
ref_g13
+0.473
+0.526
+0.570
+0.624
+0.679
+0.698
+0.687
+0.725
+0.707
+0.749
+0.761
ref_g09
+0.503
+0.566
+0.601
+0.636
+0.673
+0.634
+0.687
+0.694
+0.694
+0.717
+0.723
voice_b
+0.493
+0.543
+0.637
+0.661
+0.677
+0.695
+0.700
+0.709
+0.710
+0.717
+0.720
ref_g08
+0.460
+0.472
+0.549
+0.582
+0.645
+0.652
+0.671
+0.670
+0.676
+0.677
+0.679
voice_a
+0.495
+0.599
+0.609
+0.646
+0.666
+0.671
+0.673
+0.674
+0.674
+0.674
+0.674
ref_g00
+0.389
+0.461
+0.489
+0.556
+0.575
+0.621
+0.606
+0.627
+0.641
+0.643
+0.644
ref_g11
+0.442
+0.499
+0.536
+0.589
+0.602
+0.629
+0.619
+0.633
+0.639
+0.641
+0.642
ref_g01
+0.514
+0.532
+0.572
+0.605
+0.606
+0.609
+0.614
+0.607
+0.624
+0.625
+0.628
ref_g03
+0.460
+0.519
+0.569
+0.566
+0.581
+0.605
+0.604
+0.605
+0.611
+0.612
+0.615
ref_g10
+0.328
+0.366
+0.424
+0.439
+0.482
+0.496
+0.499
+0.509
+0.513
+0.511
+0.513
ref_g05
+0.320
+0.399
+0.446
+0.464
+0.480
+0.486
+0.495
+0.498
+0.503
+0.502
+0.504
4 · Study B — speech-restoration (Sidon) effect on scores
A deterministic ~50% of all groups (language-balanced) were passed through single-speaker speech restoration before scoring; both raw and restored versions were scored on the same takes.
Overall
metric
raw μ
Sidon μ
Δ (Sidon−raw)
inv_wer
0.773
0.772
-0.000
blend
1.838
1.825
-0.013
genu
1.237
1.277
+0.040
combined
0.390
0.392
+0.003
Study C
metric
raw μ
Sidon μ
Δ (Sidon−raw)
inv_wer
0.718
0.717
-0.001
blend
1.847
1.920
+0.073
genu
1.292
1.366
+0.074
combined
0.370
0.388
+0.018
Study D
metric
raw μ
Sidon μ
Δ (Sidon−raw)
inv_wer
0.897
0.897
+0.000
blend
1.820
1.608
-0.212
genu
1.110
1.073
-0.037
combined
0.434
0.402
-0.032
English
metric
raw μ
Sidon μ
Δ (Sidon−raw)
inv_wer
0.806
0.806
-0.001
blend
1.689
1.645
-0.044
genu
1.187
1.204
+0.017
combined
0.384
0.379
-0.005
German
metric
raw μ
Sidon μ
Δ (Sidon−raw)
inv_wer
0.710
0.709
-0.000
blend
2.118
2.162
+0.044
genu
1.330
1.414
+0.084
combined
0.401
0.417
+0.016
Best-of-N pick changed by Sidon in 48% of groups (does restoration change which take wins).
5 · Audio examples
External MP3 players. Top / median / bottom take per study by combined reward, plus raw-vs-restored pairs.