An evolutionary search for the natural-language performance instruction that best makes a 4.55-billion-parameter text-to-speech model become a character — zero-shot, no reference audio — refined across five rounds and scored by a VoiceNet + EmoNet reward.
Most voice cloning needs a reference recording of the target voice. This study does the opposite: it writes a director's instruction in plain English — the way you'd brief a voice actor — and the model performs it from scratch. The challenge is finding the exact wording that lands each character. We searched for it automatically, one evolutionary round at a time.
For every prompt the model generates many takes (N=8 while searching, N=64 for the final pick). The single best take is far better than the average — randomness is a feature you harvest.
Each round starts from the previous champion, spawns ~10 mutated prompts per generation, scores them, and breeds the winners forward — an evolutionary search over language, not model weights.
Every take is scored by VoiceNet (57 acoustic dimensions) and Empathic-Insight EmoNet (emotion intensities), plus naturalness, vocal-burst blend and word-error rate. That number drives the search.
Every best-of-N call is scored, then min-max normalized within that call (so takes only compete against their own siblings). The reward blends four signals, then multiplies by a word-error-rate factor so unintelligible takes can't win.
The round-5 champion take for each character — its in-character acoustic score (char), its emotion score (char_emo), and the single strongest emotion the EmoNet panel detected.
| Character | char | char_emo | Top firing emotion | reward | blend | genu | invWER |
|---|---|---|---|---|---|---|---|
| 🧟 Zombie | 0.80 | 0.55 | Helplessness · 3.16 | 3.05 | 7.21 | 1.73 | 0.89 |
| 👹 Ork | 0.67 | 0.93 | Malevolence/Malice · 2.40 | 3.08 | 2.72 | 0.97 | 1.00 |
| 👺 Goblin | 0.56 | 0.57 | Malevolence/Malice · 1.45 | 2.45 | 3.25 | 1.89 | 0.85 |
| 🐉 Dragon | 0.72 | 0.88 | Anger · 3.71 | 3.21 | 5.25 | 0.81 | 1.00 |
| 👻 Evil Ghost | 0.68 | 0.75 | Malevolence/Malice · 2.01 | 3.93 | 6.81 | 1.31 | 1.00 |
| 🐭 Cartoon Mouse | 0.81 | 0.89 | Interest · 3.17 | 3.40 | 2.66 | 1.38 | 1.00 |
| 🧚 Fairy | 0.72 | 0.89 | Affection · 3.30 | 3.23 | 3.90 | 1.01 | 1.00 |
| 🗯️ Ranting Fury | 0.93 | 0.85 | Anger · 5.11 | 5.07 | 0.66 | 1.06 | 1.00 |
| 😱 Pain Scream | 0.69 | 0.92 | Distress · 4.07 | 4.13 | 1.23 | 1.55 | 0.67 |
| 🌙 ASMR Man | 0.62 | 0.88 | Contentment · 3.26 | 3.31 | 6.94 | 1.13 | 1.00 |
| 🌙 ASMR Woman | 0.82 | 0.56 | Affection · 4.05 | 2.98 | 8.73 | 1.27 | 0.92 |
| 😢 Grieving Woman | 0.67 | 0.96 | Distress · 4.80 | 4.41 | 7.42 | 2.12 | 0.94 |
| 😔 Grieving Man | 0.73 | 0.99 | Sadness · 4.53 | 4.69 | 6.30 | 2.28 | 1.00 |
Each card carries the top three best-of-64 takes (champion, runner-up and third), the exact prompt to reuse with its strategy dissected word-by-word, per-take scores, the signature VoiceNet dimensions and firing emotions, and a fresh procedural caption generated by running all 57 VoiceNet heads on each clip. Press play.
For each character we take its final champion prompt — the winning instruction plus the spoken text — and generate it 64 times, each with a different random seed. That yields 64 slightly different recordings of the exact same prompt: same words, same instruction, only the random seed changes.
All 64 takes are then scored by the reward and ranked best-to-worst. The three players on every card are simply the top three of that ranking:
Why bother generating 64? Because the takes differ only by seed, and a lucky seed often lands a noticeably better take than the average one. Some effects surface only occasionally — e.g. the goblin's rare, clean “malevolence spike” that most seeds miss entirely — so generating many and keeping the best (best-of-N) is how you reliably catch them.
The most valuable output of the whole study: everything five rounds taught us about writing a director instruction that lands. Read the principles, steal the template, use the lookup table.
“cartoon mouse voice” loses to “extremely high, thin, pinched, buzzing nasal, bright piercing head voice.” The model listens to the description — spell out how it sounds, never just what it is.
A perfect zombie rasp that sounds bored is a bad zombie. Put the feeling in words — helpless, murderous glee, seething contempt, giddy delight, drowning in pain. Adding an EmoNet emotion reward was round 2's single biggest lever.
Not just “deep” or “bright” — name the cavity: chest (hollow/warm), throat (guttural/raspy), nose (nasal/buzzing), head (glassy/light). Resonances stack: a zombie can have a hollow chest groan and a wet throat rasp. Naming the cavity as a place (“a vast cavern of a chest”) beats the adjective “deep.”
Register = how high/low it sits (“subterranean bass” vs “sparkling high register”). Range = how far it moves (“flat and unwavering” vs “swooping wildly”). Wide swings read as mania/fury/excitement; flat narrow pitch reads as menace/authority/deadness.
“Commanding, immovable, planting each word with authority” pushes authority up (ork, dragon, rant). “Feeble, powerless, meek, no authority whatsoever” pushes it down (zombie, goblin, fairy, ASMR). Meekness is a real, rewardable direction — negating dominance is as useful as asserting it.
Naming authority alone (“an orc warlord, commanding”) made a clean commander, not a monster — character collapsed to ~0.30. Keep the guttural/nasal/whisper timbre and append the new layer. Over-describing resonance (a second “immense chest” clause) actively craters emotion and blend.
Squeaky interjections (“Eee… Ooh… Wow!”) lift the mouse's Amusement/Astonishment; screamed “Aaargh!… it hurts” drives pain-scream's Distress to ~4; a stammer (“I-I should have been there…”) induces audible cracking in the grieving man; broken “I can't— I can't” keeps the grieving woman's blend high. The text carries affect the instruction cannot.
A single emotion-matched burst — (agonized moan), (delighted giggle) — reliably helps. A second usually hurts, except for the highest-arousal voices: ranting uses (enraged shout)(seething snarl) and pain-scream's paired (agonized scream)(sobbing gasp) is exactly what broke its char-vs-emotion tradeoff.
Extra “weightless / gossamer / crystalline” on the fairy, a second “immense chest resonance” on the ork/dragon, or over-stacked hollow/numbness on the zombie all added variance without adding reward. The leanest prompt that keeps the core intact wins the same-call playoff.
Pushing one dimension to the extreme (“ear-splitting volume”, “impossibly high”) garbles words or thins the voice. Keep a safety clause — “roaring but the words still land” / “whispered but every word clear” — to protect the WER-gated archetypes. It structurally raises the reward ceiling (asmr_woman: dropping a mis-heard “Shhh” lifted invWER 0.79→1.00).
The goblin's Malevolence fires as a rare per-take spike (most takes ~0.05, ~1-in-8 jumps to 1–2.5). best-of-8 mean is a noisy selector; only best-of-64 reliably catches a clean malicious jackpot. Concentrate on one emotion — diluting malice with “contempt/bitter” killed the Malevolence head entirely.
The reward is min-max normalized within each best-of-N run, so a 2.4 in one run ≠ a 2.4 in another. Compare finalists by putting them in one shared run; search at best-of-8, decide the final champion at best-of-64.
Round 4–5 added vn_w (per-dimension) and emo_w (per-emotion) weights plus per-character blend/genu weights, so the search chases the 2–3 signals that define a character — the mouse's pitch (RANG ×2.5, REGS ×2.0), the goblin's Malevolence (×2.5), the grieving man's Sadness (×3.0), evil-ghost's whisper (S_WHIS ×2.5).
Bursts go in the instruction, in round brackets — never inside the spoken text. Search at best-of-8, decide the champion at best-of-64, and compare finalists in one shared run.
| Desired effect | Words to push | Dimensions / emotions | Used by |
|---|---|---|---|
| Menace / authority | low flat narrow pitch · commanding, immovable, planting each word | S_AUTH, STNC ↑ · RANG ↓ | ork, dragon |
| Meek / pitiful | feeble, powerless, meek, no authority whatsoever, yielding | S_AUTH ↓ · VULN ↑ | zombie, goblin, ASMR, fairy |
| Mania / fury / excitement | pitch swinging violently / darting giddily up and down | RANG, VOLT, AROU ↑ | ranting, mouse, goblin |
| Deep body | booming from a vast cavern of a chest; deep guttural throat rumble | R_CHST, R_THRT, FULL ↑ | ork, dragon |
| Thin / piercing | extremely high thin pinched buzzing nasal head voice; no chest | R_NASL, R_HEAD, BRGT ↑ · R_CHST ↓ | mouse, goblin, fairy |
| Intimacy | soft breathy intimate close-mic whisper, warm chest underneath | S_WHIS, S_ASMR, WARM ↑ | ASMR man/woman, evil-ghost |
| Grief / crying | breaking down, gasping and hitching, hard sobs, rising into a wail | Sadness, AROU ↑ | grieving man/woman |
| Evil / malice | truly evil, sadistic, cruel, remorseless, pure malice to the core | Malevolence, Contempt ↑ | goblin, evil-ghost |
| Joy / warmth | brimming with warm loving affection and giddy joy, radiant delight | Affection, Elation ↑ | fairy, ASMR |
| Keep it clean | but every word lands / whispered but every word clear | invWER ↑ (protects blend/genu) | all WER-gated voices |
Each round added a new lever, and the scores climbed. Round 2 introduced the EmoNet emotion score (there was none in R1). Round 3 added resonance, pitch and dominance. Rounds 4–5 added per-dimension / per-emotion weights and pushed each voice to its extreme — the ork's hatred (char_emo 0.85→0.93), the grieving man's overt sobbing (0.89→0.99), the goblin caught cleaner and more malevolent, asmr_man's serenity (0.67→0.88), the mouse and asmr_woman climbing hard on character.
| Character | char R1 | char R3 | char R4 | char R5 | Δ char | emo R3 | emo R4 | emo R5 |
|---|---|---|---|---|---|---|---|---|
| 🧟 Zombie | 0.693 | 0.809 | · | 0.799 | ▲ +0.11 | 0.546 | · | 0.552 |
| 👹 Ork | 0.336 | 0.531 | · | 0.670 | ▲ +0.33 | 0.847 | · | 0.928 |
| 👺 Goblin | 0.584 | 0.737 | 0.476 | 0.565 | -0.02 | 0.275 | 0.665 | 0.575 |
| 🐉 Dragon | 0.569 | 0.724 | · | 0.648 | ▲ +0.08 | 0.705 | · | 0.858 |
| 👻 Evil Ghost | · | · | 0.592 | 0.676 | ▲ +0.08 | · | 0.583 | 0.748 |
| 🐭 Cartoon Mouse | 0.752 | 0.757 | 0.869 | 0.805 | ▲ +0.05 | 0.660 | 0.642 | 0.885 |
| 🧚 Fairy | 0.464 | 0.711 | · | 0.724 | ▲ +0.26 | 0.782 | · | 0.888 |
| 🗯️ Ranting Fury | 0.648 | 0.802 | 0.798 | 0.925 | ▲ +0.28 | 0.962 | 0.742 | 0.853 |
| 😱 Pain Scream | 0.846 | 0.714 | 0.737 | 0.712 | -0.13 | 0.814 | 0.938 | 0.882 |
| 🌙 ASMR Man | 0.666 | 0.615 | · | 0.625 | -0.04 | 0.669 | · | 0.884 |
| 🌙 ASMR Woman | 0.491 | 0.624 | · | 0.824 | ▲ +0.33 | 0.865 | · | 0.555 |
| 😢 Grieving Woman | · | · | 0.706 | 0.668 | -0.04 | · | 0.917 | 0.956 |
| 😔 Grieving Man | · | · | 0.623 | 0.732 | ▲ +0.11 | · | 0.894 | 0.990 |
Blank = that round was not run for this character (voices were added at different stages; R1 had no emotion score at all). char / char_emo are absolute VoiceNet / EmoNet dimension scores and are directly comparable across rounds; the round-normalized reward is not.
Every voice on this page was produced by laion/moss-tts-local-transformer-4.55b-voice-acting — a 4.55B MOSS-TTS-Local-Transformer at native 48 kHz — from a natural-language instruction and target text, with no reference audio. Each round ran an evolutionary prompt search (best-of-8 while searching, best-of-64 for the final), scored by VoiceNet's 57 acoustic heads, Empathic-Insight EmoNet emotions, a VoiceCLAP burst-blend head, a genuineness head and a Whisper WER gate. The captions and dimension profiles shown here were freshly recomputed by re-embedding each champion clip with VoiceCLAP-commercial and running all 57 heads.