LAION · Voice-Acting Data Pipeline · reproducibility run

MOSS-8B LGT best-of-N — scaling & Sidon study

50 LAION's-Got-Talent voices × 3 reference conditions × EN/DE, plus named-reference fair-text groups and an n=1000 scaling sweep. Every clip scored with invWER (Parakeet-TDT-0.6b-v3) and VoiceCLAP blend + genuineness heads; ranked by reward = invWER · (norm·genu + norm·blend).

modelmoss-tts-v1.5-8b-voice-acting
codecMOSS-Audio-Tokenizer · 24kHz mono
clips scored70,893
hardware8 × A100-80GB
0.826
base best-of-100 reward
+0.808
uplift vs cloned reference
100% of groups ↑
-0.001
Sidon Δ reward
≈ neutral
1000
plateau k (95% of gain)
still rising at 1000
0.51
gen clips/s/GPU
batched ×100
01

Best-of-N reward & uplift

Across all 300 base cells, the best-of-100 take beats the reference it was cloned from in 100% of cases (uplift +0.808 reward). The cloned reference speaks different words than the eval script, so its invWER is low by construction — best-of-N reliably recovers an on-script, high-quality take.

conditionnbest rewardmean rewardbest invWERuplift
orig (raw LGT ref)1000.8210.3960.866+0.803100%
emolia_vc1000.9020.3320.895+0.886100%
testref_vc1000.7560.3140.878+0.736100%

By language

langnbest rewardbest invWERuplift
EN1500.7450.936+0.714
DE1500.9080.823+0.902

German scores higher on reward (stronger blend/genuineness) while English scores higher on invWER — the ASR transcribes English more cleanly, but the VoiceCLAP heads rate the German takes as more genuine.

02

Sidon-before-scoring ablation

Sidon speech restoration (w2v-BERT 2.0 + DAC, 16→48 kHz) was applied to a deterministic 50% of groups before scoring — 12,923 clips restored and scored paired against their raw versions. The verdict: restoration does not help best-of-N selection.

metricrawSidonΔ (Sidon − raw)% clips improved
invWER0.7340.733-0.0005.7%
blend2.2192.177-0.04242.1%
genuineness1.2791.313+0.03555.5%
reward0.3430.343-0.00144.4%

On the best-of-N pick per cell, Sidon reward is -0.019 vs raw (Sidon wins in only 44.7% of 132 cells). Restoration slightly lowers VoiceCLAP blend similarity while marginally raising genuineness — the two nearly cancel, leaving reward flat. For this pipeline, Sidon is a cost with no scoring benefit; its value would be in output audio quality, not selection.

03

Per-sample throughput

Measured on the real run, per stage. ASR dominates end-to-end cost — the VoiceCLAP scoring heads are effectively free next to Parakeet decoding.

MOSS-8B generation
1.969s
per clip · batched ×100
WER / ASR — Parakeet-TDT
1.304s
per clip · dominant cost
Sidon restoration
0.609s
per clip
Scoring — blend + genuineness
0.033s
per clip · VoiceCLAP
Chatterbox voice conversion
10.075s
per group (2 refs built)
04

n=1000 scaling — does best-of-N plateau?

Expected best-of-k reward as a function of how many of the 1000 takes you sample (bootstrap over random subsets). Reward keeps climbing all the way to 1000 — with clear diminishing returns, but no plateau before the ceiling.

0.710.770.830.890.95 100 = 0.788best-of-1000 = 0.926502004006008001000 number of takes sampled (k)

Going 100 → 1000 takes lifts best-of-k reward by +0.138 (+17.5%). The marginal gain shrinks from +0.049 (50→100) to +0.010 (900→1000), but 95% of the 100→1000 gain is not reached until k = 1000 — so within this budget, more samples still pay off. The knee is around k ≈ 300–500, where you capture most of the gain at a quarter of the compute.

05

Named-reference fair-text groups

Three DramaBox emotional prompts — fear/terror, bliss, vulnerable love — 3 sentences each, EN + DE, cloned onto each named voice. Best-of-100 per (voice × emotion × language).

By reference voice

voicenbest rewardbest invWERuplift
samantha61.0500.947+1.050
dragon60.9170.928+0.905
chris60.7930.814+0.780
orc60.7920.930+0.792
goblin60.7140.900+0.712

By emotion prompt

emotionnbest rewardbest invWER
vulnerable love100.8930.906
fear / terror100.8510.936
bliss100.8150.870

Samantha clones best (reward 1.05, invWER 0.95); goblin is hardest (0.71). Across voices the tender-love prompt scores highest and bliss lowest — the model lands intimate delivery more reliably than sustained elation.