MOSS-TTS-v1.5 8B — Voice-Acting

Expressive, instruction-controlled text-to-speech. A Qwen3-8B backbone with a 32-codebook delay audio head at 24 kHz. Live demo grids and evaluation dashboards for the fine-tuned model. Play the clips directly in your browser.

Production emotion grid

Production best-of-64 — 8B emotion grid (40 EmoNet emotions)
The production best-of-64 emotion grid (recipe 1b) regenerated with the 8B MOSS model — it replaces the 4B local transformer of the original reference run. For each of the 40 EmoNet emotions, 64 diverse no-reference takes (audio-temp 0.8) from laion/voice-acting-prompts, ranked by reward = norm(target emotion) + norm(blend) + norm(genuineness) (min-max within each 64-group; target emotion = the Empathic-Insight-Voice-Plus head on BUD-E-Whisper). Per-emotion top-3 of 64 + summary table.

Samples

Voice-Cloning + Paraphrase
Original real clip vs. a generated clone that keeps the voice & style but speaks a freshly reworded line — data-augmentation demo across several source styles and two languages.

Rankings

Genuineness ranking (best-of-10)
Ten temperature-1.0 takes per prompt, ranked by two independent VoiceCLAP genuineness probes (large-v2 3584-d & commercial 768-d). Top-3 / bottom-2 per probe.
Combined-reward ranking (best-of-10)
The same takes ranked by a combined reward: invWER × (norm(blend) + norm(genuineness)), folding in perceptual quality and transcription fidelity.

Benchmarks

Generation-settings grid search
Temperature × repetition-penalty sweep, scored on WER, blend quality, and genuineness — with and without a reference clip. Source of the recommended defaults.
Best-of-10 uplift
Voice-clone 20 reference clips (EN + DE), 10 takes each. Can the best-of-10 take out-score the real human reference on quality & genuineness? Uplift analysis included.
Best-of-32 uplift
50 references × 2 languages, 32 batched takes each. Larger best-of-N pool — how much extra uplift ceiling does drawing 32 seeds buy over 10?
Best-of-100 across speaker identities (pilot)
Best-of-100 selection, and whether the uplift over the reference holds when the same performance is voice-converted (via Chatterbox VC, +0–50% random speed-up) to two different speaker identities — EN + DE translation. Pilot subset with conclusions & per-group audio.

Reproduce study

Reproduce-and-improve, ranked two ways
25 expressive reference performances × 3 temperatures × best-of-8. Clone each voice, then rank the takes with two independent rewards: a voice-identity combined reward (ECAPA speaker-sim + invWER + genuineness + blend + duration) and an emotion-profile reward — cosine of a 42-dim EmoNet emotion vector (Empathic-Insight-Voice-Plus on BUD-E-Whisper) × invWER. Per temperature × group: reference clip + prompt + best-of-8 top-3 under both rankings side by side. They disagree on ~48% of takes.

Scaling study

Scaling pilot — throughput, best-of-N plateau & Sidon
Per-stage throughput profiling (generation is the bottleneck), the best-of-N reward plateau (~600 samples), and the effect of Sidon speech-restoration on scores.
🎧 Top-3 audio samples per group
The best 3 takes of every group across the scaling study, ranked by the combined reward, each with its spoken text, performance instruction, and scores.

Prompting ablation

Prompting style × sample size
Does keeping the quoted lines inside a structured two-part performance prompt help or hurt vs. passing them only as the text to speak? 25 prompts × 2 instruction variants × 16 no-reference takes, scored on combined reward, with a best-of-k (4/6/8/16) scaling curve — split by character-consistent vs. scenario-driven pathways. 🎧 top-3 of 16 · best-of-8.

LGT scaling & Sidon study (n up to 1000)

Study report — reward/uplift, Sidon ablation, throughput, n=1000 plateau
50 LGT voices × 3 reference conditions × EN/DE, up to 1000 takes. Best-of-N beats the cloned reference in 100% of groups; Sidon-before-scoring is neutral; reward still rising at k=1000.
🎧 Top-3 samples per group (78 groups)
The 3 best takes of every group with all scores, prompts, and a plain-language explainer.

DramaBox prompting ablation (dialogue in instruction? reference audio? rewards)

Analysis & plots — style × reference × best-of-k + prompt adherence
with_dialogue vs directions-only × no-ref vs random emotion ref, 25 prompts × 60 takes. Includes VoiceCLAP prompt-adherence (near-orthogonal to reward, r≈0).
🎧 Top-3 grids · best-of-16/8/6/4 · prompt-reward ranking
Every variant's best takes with full scores on each player.
🎧 Three rewards side-by-side — best-of-16 / best-of-60 · 16 vs 128 gain
Same take pool ranked by all / original / adherence rewards; expanding 16→128 lifts the best take ~14–16%.

Emotion prompting (Empathic-Insight-Plus scored)

How to prompt extreme emotion — 6 emotions × 20 instruction styles
Escalation-arc and loudness instructions win; the emotional text itself carries most of the emotion; empty instruction is best for speaker similarity; genuineness can't be asked for.
🎧 All insights + top-3 takes per group (456 players)
Every emotion × style group's best takes with target/arousal/valence/volume scores and instruction texts.

voice-acting-prompts best-of-64 (procedural, 40 emotions)

🎧 Pilot: 5 emotions × 2 groups × 64 takes, weighted-reward top-3
The 40-emotion dataset maps 1:1 to EI-Plus heads — fully procedural pipeline; reward = invWER × (1.5·target + genu + blend). Full run: 256k clips ≈ 8–9 h on 8×A100.
🎧 moss-local (48 kHz) same pilot — head-to-head vs 8B
MOSS-TTS-Local + voice-acting LoRA on the identical 10 groups: better word accuracy (invWER 0.97 vs 0.83) and higher target-emotion scores.
Fast batched inference — every measured lever
Batch scaling, the broken FA2 path, 134× batched ASR, fused gen→sidon→score workers, token-budget pacing. Reference numbers for production runs.
🎧 Production best-of-64 — MOSS-TTS-Local 4.55B (48 kHz), top-3 raw per group
From the 40-emotion × 100-prompt × 64-take production run (local-transformer model, raw audio), reward = invWER × (target + blend + genu); grows as emotions complete. Links to the 4 dataset repos.

Checkpoint eval

Held-out emotional prompts
The trained 8B checkpoint rendering a set of held-out emotional prompts (no reference audio) — a quick qualitative read on expressive range.
MOSS-TTS-v1.5 8B voice-acting · Apache-2.0 · reference clips in the eval grids come from a public emotion-attribute reference set (provenance kept generic). Audio is served as external mp3 files alongside each page.