LAION Voice Acting Pipeline

Open-weights voice-acting data pipeline: structured taxonomy sampling, DramaBox TTS synthesis, and best-of-N reward ranking. Live benchmark pages below.

Documentation

How the DramaBox Prompt Dataset is Sampled & Generated

Plain-English, reproduce-it-yourself walkthrough of the laion/dramabox-cutscene-prompts dataset: the 5 sampling pathways, every taxonomy it draws from (VoiceNet, 40 EmoNet emotions, archetypes, situations, 180 vocal bursts), the exact prompts the model receives, and real generated examples with their sampled metadata.

Benchmark

Vanilla DramaBox TTS — Generate + Reward-Rank

Raw two-scene CUT TO: DramaBox prompts fed directly to the 8B MOSS voice-acting TTS, 4 seeds per prompt, scored and reward-ranked. Listenable takes, best-of-k quality/compute trade-off, and the full k=1..32 seed-scaling walltime table.

Throughput study

Local-LLM DramaBox Prompt-Generation

Throughput and token stats for generating DramaBox two-scene prompts with a local LLM on a single A100, per pathway and language, with a 1M-prompt estimate — plus the DramaBox TTS seed-scaling table for downstream best-of-k planning.

Demo

Sidon + ChatterboxVC Sample Groups

20 groups × 25 candidates with LLM-guided CUT TO: splitting, Whisper-turbo ASR, and Gemma re-annotation.

Code Apache-2.0 · docs & sample pages CC-BY-4.0 · GitHub repo

🎭 DramaBox reinterpretations — reference-guided best-of-8 (emotion-similarity × inverse-WER reward)

🎚️ DramaBox reinterpretation — dynamic emotion+VoiceNet LoRA merge vs. baseline (best-of-8)