g01: sadness sorrow grief melancholy and heartache
Study: best-of-100 voice-acting selection across speaker identities.
Model: an 8B autoregressive expressive TTS model, generated WITH reference-voice conditioning
(audio temperature 1.0, repetition penalty 1.1, text temperature 0.7, top-p 0.95, top-k 25, bf16).
Reference source: a large open crowd-sourced corpus of expressive human speech (one distinct
speaker clip per group, diverse emotions). Reward: combined = invWER x (norm_genu + norm_blend),
where genu is a commercial-voice genuineness score (0-6) and blend an expressive-quality score (0-10),
each min-max normalized across all generated clips; invWER = max(0, 1-WER) from multilingual ASR.
Achieved throughput: 2.87 clips/s on a single GPU.
English script: We built a whole life together, brick by brick, dream by dream. And now I walk through these empty rooms and all I can hear is the echo of everything we used to be. I miss you so much that some mornings I can barely get out of bed.
German script (translation): Wir haben ein ganzes Leben zusammen aufgebaut, Stein fuer Stein, Traum fuer Traum. Und jetzt gehe ich durch diese leeren Raeume, und alles, was ich houere, ist das Echo von allem, was wir einmal waren. Ich vermisse dich so sehr, dass ich an manchen Morgen kaum aus dem Bett komme.
Delivery instruction:Quiet sorrow and aching heartache, the voice trembling with grief, thick with unshed tears. High quality recording.
Original reference voice
REFgenu=- blend=-
Voice-converted to a random speaker A — source sped +2%, target=Source-A_de_09
REFgenu=- blend=-
Voice-converted to a random speaker B — source sped +16%, target=Source-B