Open-weights voice-acting data pipeline: structured taxonomy sampling, DramaBox TTS synthesis, and best-of-N reward ranking. Live benchmark pages below.
DocumentationPlain-English, reproduce-it-yourself walkthrough of the laion/dramabox-cutscene-prompts dataset: the 5 sampling pathways, every taxonomy it draws from (VoiceNet, 40 EmoNet emotions, archetypes, situations, 180 vocal bursts), the exact prompts the model receives, and real generated examples with their sampled metadata.
Raw two-scene CUT TO: DramaBox prompts fed directly to the 8B MOSS voice-acting TTS, 4 seeds per prompt, scored and reward-ranked. Listenable takes, best-of-k quality/compute trade-off, and the full k=1..32 seed-scaling walltime table.
Throughput and token stats for generating DramaBox two-scene prompts with a local LLM on a single A100, per pathway and language, with a 1M-prompt estimate — plus the DramaBox TTS seed-scaling table for downstream best-of-k planning.
Demo20 groups × 25 candidates with LLM-guided CUT TO: splitting, Whisper-turbo ASR, and Gemma re-annotation.
Code Apache-2.0 · docs & sample pages CC-BY-4.0 · GitHub repo
🎭 DramaBox reinterpretations — reference-guided best-of-8 (emotion-similarity × inverse-WER reward)
🎚️ DramaBox reinterpretation — dynamic emotion+VoiceNet LoRA merge vs. baseline (best-of-8)