Expressive, instruction-controlled text-to-speech. A Qwen3-8B backbone with a 32-codebook delay audio head at 24 kHz. Live demo grids and evaluation dashboards for the fine-tuned model. Play the clips directly in your browser.
laion/voice-acting-prompts, ranked by reward = norm(target emotion) +
norm(blend) + norm(genuineness) (min-max within each 64-group; target emotion = the
Empathic-Insight-Voice-Plus head on BUD-E-Whisper). Per-emotion top-3 of 64 + summary table.invWER × (norm(blend) + norm(genuineness)), folding in perceptual quality and transcription fidelity.k (4/6/8/16) scaling curve — split by
character-consistent vs. scenario-driven pathways.
🎧 top-3 of 16 ·
best-of-8.