Evolving director instructions for character voices
A prompt-evolution study over 10 character/style archetypes with MOSS-TTS-Local 4.55B — no reference audio. We search for the plain-English performance instruction that makes the model best embody each character, scored by a VoiceNet-based reward.
MOSS-TTS-Local is a 4.55-billion-parameter text-to-speech model that you can steer with a natural-language performance instruction — a bit like a film director's note ("a rotting corpse's wet gurgling throat voice… slurred and groaning"). Crucially, in this study we give it no reference audio at all: the character voice has to emerge from the words of the instruction alone.
The question: which instruction makes the model best become the target character? We answer it automatically with a search loop:
• Best-of-N. For any single instruction we sample the model N times (N = 8 during the search, N = 64 for the final round) and keep the best-scoring take. TTS is random per seed, so more draws = a better shot at a great take.
• Evolutionary search. We start a population of ~10 candidate instructions, score each with best-of-8, keep the winners, mutate them (reword, add/remove one acoustic cue or one director-burst), and repeat for ~5–7 generations. The final champion is then re-drawn with best-of-64.
• The reward is what tells "good" from "bad" — a learned voice-analysis model (VoiceNet) scores every take, described next.
The reward, in plain language
Each take is embedded with VoiceCLAP-commercial and passed through 57 VoiceNet regression heads that rate voice properties (roughness, chest resonance, brightness, whisper, arousal, …). From those we build, per best-of-N call:
reward = norm(blend) + norm(genu) + 1.5 · char [ then × an invWER factor ]
blend = how well the take matches the character's overall VoiceCLAP fingerprint (an esthetic/coherence term). genu = genuineness/naturalness. char = the mean of the archetype's characteristic VoiceNet dimensions — the positive dims pushed high, the penalized dims pushed low. char carries a 1.5× weight; blend and genu are ~1× each.
invWER factor (intelligibility). Some characters must stay intelligible, some need not. The factor is OFF for zombie & pain-scream (groans and screams are free), soft = 0.4 + 0.6·invWER for orc/mouse/goblin/ranting, and full = ×invWER for fairy/dragon/asmr-man/asmr-woman.
Important methodology caveat. The min-max normalization (norm()) is computed per gen_score call — i.e. within one best-of-N batch. So absolute reward numbers are only comparable inside a single call, never across generations or across archetypes. Throughout the search we compared candidates by putting them into the same scoring call and reading the head-to-head "playoff" ranking. Read the tables below with that in mind.
High-level results
One row per archetype. Metrics are for the top take of the final best-of-64 draw (best-of-8 champion where best-of-64 wasn't separately logged). Click a name to jump to its section.
Wailing and shrieking in raw physical torment, voice torn, ragged and guttural, harsh grating rasp shredding t…
2.55
6.16
1.59
0.91
0.85
4
768
0.513
Per-archetype detail
For each archetype: the champion instruction, the reward's positive/penalized dims and intelligibility gate, the reward trajectory, and the top-3 takes of best-of-8 vs best-of-64 side by side. Every take has an audio player, its reward breakdown, the spoken text it produced, and a procedural voice profile — the top-5 dimensions (out of 57) on which that specific take deviates most from the average voice, computed by re-embedding the clip and running all 57 heads.
Zombie rotting undead, deep in the throat
Reward dims:ROUGR_THRTCOGLDFLUCLRTESTH· invWER gate:OFF (no intelligibility gate)
A rotting, diseased zombie whose voice sounds SICK and comes from deep in the THROAT: throaty, raspy, guttural, wet, laboured, decayed. Barely articulate: slurred, groaning, gurgling, struggling to form words; not loud, strained and ill. Reward favors roughness/throat-resonance/cognitive-load/disfluency and does NOT gate on intelligibility.
Champion instruction (no reference audio — this text alone steers the 4.55B model):
A rotting corpse's wet gurgling throat voice, thick phlegmy guttural rasp, hoarse and raw. Slurred groaning, laboured, mumbling and trailing off, decayed. (wet gurgle)
Reward trajectory (best mean per generation, 5 gens): 2.00 → 2.02 → 2.10 → 2.04 → 1.99 Absolute values are min-max normalized per gen_score call, so they are only comparable within a call — champions were picked by same-call head-to-head playoffs.
A monstrous, brutal orc warrior — deep, guttural, growling, snarling, aggressive, rough, heavy chest+throat resonance, brutal menacing force and rage. Words growled but present.
Champion instruction (no reference audio — this text alone steers the 4.55B model):
A huge vicious orc spitting his words with mounting fury. Very very deep, extremely guttural bass voice, enormous chest and throat resonance, extremely rough grinding rasp. Extremely aggressive, brutal, tense and forceful, seething and ranting with rage. Dark and heavy, zero brightness. High quality recording.
Reward trajectory (best mean per generation, 5 gens): 1.76 → 1.74 → 1.71 → 1.75 → 1.65 Absolute values are min-max normalized per gen_score call, so they are only comparable within a call — champions were picked by same-call head-to-head playoffs.
A tiny delicate fairy: very soft, light, airy, high-pitched, sweet, gentle, melodic almost singsong, twinkling, bright and warm, high head resonance, no roughness; pretty and pleasant.
Champion instruction (no reference audio — this text alone steers the 4.55B model):
Delicate high fairy voice, feather-light and airy, warm and very bright, sweet and gentle, lightly melodic. Smooth pure glassy head resonance, tender and pleasant, crystal clear, no roughness at all. (light giggle) High quality recording.
Reward trajectory (best mean per generation, 5 gens): 1.59 → 1.97 → 1.67 → 1.85 → 1.85 Absolute values are min-max normalized per gen_score call, so they are only comparable within a call — champions were picked by same-call head-to-head playoffs.
A cute tiny squeaking mouse: small, high-pitched, quick, squeaky/pippy, bright, young, chipper and adorable; perky and nasal-bright, tiny and energetic (NOT the soft dreamy singsong of a fairy). Reward favors brightness/head+nasal-resonance/cartoonish-style/tension; penalizes voice-age/chest-resonance/fullness; soft intelligibility gate.
Champion instruction (no reference audio — this text alone steers the 4.55B model):
An extremely high squeaky cartoon-mouse voice, thin and bright with intense nasal buzz and piercing head resonance. Pinched, tense, quick and chipper, exaggerated and cute. Very small and young, weightless, no chest resonance or fullness. High quality recording.
Reward trajectory (best mean per generation, 4 gens): 1.89 → 1.79 → 1.84 → 1.72 Absolute values are min-max normalized per gen_score call, so they are only comparable within a call — champions were picked by same-call head-to-head playoffs.
A colossal ancient dragon - extremely deep, booming, resonant bass from a vast chest cavity; slow thunderous rumble; full, dark, authoritative, commanding, menacing, powerful. Reward favors chest-resonance/fullness/authoritative+dramatic-style/voice-age/throat-resonance (dims R_CHST,FULL,S_AUTH,S_DRAM,AGEV,R_THRT), penalizes brightness (BRGT), full invWER gate on intelligibility.
Champion instruction (no reference audio — this text alone steers the 4.55B model):
A grand, theatrical, dramatic proclamation in an extremely deep commanding bass. Massive chest resonance, guttural throat rumble, slow and thunderous. Dark, full, domineering, ancient and menacing. (deep rumble) High quality recording.
Reward trajectory (best mean per generation, 6 gens): 1.52 → 1.76 → 1.66 → 1.69 → 1.58 → 1.62 Absolute values are min-max normalized per gen_score call, so they are only comparable within a call — champions were picked by same-call head-to-head playoffs.
A small, sneaky, nasty goblin: high-ish nasal, raspy, whiny, scheming, sniveling, sharp, quick, gleeful-cruel and mean; thin nasal resonance with a rasp. Reward favors nasal-resonance/roughness/cartoonish-style/tension/arousal-shift, penalizes fullness/chest-resonance/warmth; soft intelligibility gate.
Champion instruction (no reference audio — this text alone steers the 4.55B model):
A sneaky little goblin, thin, pinched and nasal, buzzing high and reedy through the nose with a rough scratchy grating rasp. The voice keeps cracking higher, sharper, shriller and whinier with gleeful cruel malice, tightening and quickening as it goes, giddy with spite. Mean and scheming. No chest, no warmth, no fullness. High quality recording.
Reward trajectory (best mean per generation, 7 gens): 1.43 → 1.49 → 1.83 → 1.43 → 1.47 → 1.59 → 1.51 Absolute values are min-max normalized per gen_score call, so they are only comparable within a call — champions were picked by same-call head-to-head playoffs.
A man doing gentle ASMR: very soft, breathy, intimate close-mic whisper, slow, soothing, warm, tender, calm, low arousal, saying something kind/comforting. Reward favors ASMR/whisper/warmth/esthetics/vulnerability/respiration and MALE gender + low arousal, and IS gated on intelligibility — keep words clear while maximally soft/whispered.
Champion instruction (no reference audio — this text alone steers the 4.55B model):
A deep, gentle man in a soft, breathy, intimate close-mic whisper, very gentle and slow. Warm, tender, soothing, deeply calm, low arousal. Soft audible breaths, clear whispered words. (soft breath) High quality recording.
Reward trajectory (best mean per generation, 5 gens): 1.60 → 2.03 → 1.94 → 2.00 → 2.07 Absolute values are min-max normalized per gen_score call, so they are only comparable within a call — champions were picked by same-call head-to-head playoffs.
A woman doing gentle ASMR — very soft, breathy, intimate close-mic whisper, slow, soothing, warm, tender, calm, low arousal, saying something kind. Female, low arousal, words clear.
Champion instruction (no reference audio — this text alone steers the 4.55B model):
A gentle woman doing ASMR. (soft breath) Extremely soft, breathy, intimate close-mic whisper, slow and soothing, hushed and quiet. Warm, tender, calm, feminine and delicate, very low arousal, kind and comforting, tenderly caring, soothing you to rest. Whispered but clear. High quality recording.
Reward trajectory (best mean per generation, 5 gens): 1.80 → 1.74 → 2.06 → 1.90 → 1.85 Absolute values are min-max normalized per gen_score call, so they are only comparable within a call — champions were picked by same-call head-to-head playoffs.
Someone ranting furiously: scolding at maximum volume, explosive rage, shouting, harsh, strained, aggressive, veins-popping intensity; words still present but yelled.
Champion instruction (no reference audio — this text alone steers the 4.55B model):
A person screaming a furious tirade at maximum volume, seething and spitting with rage. Voice cracking and strained, harsh and tense, sharp percussive attack on the consonants, pounding emphasis, explosive and volatile, surging louder and louder. Enraged but the words survive. (furious snarl) High quality recording.
Reward trajectory (best mean per generation, 5 gens): 1.52 → 1.63 → 1.76 → 1.73 → 1.48 Absolute values are min-max normalized per gen_score call, so they are only comparable within a call — champions were picked by same-call head-to-head playoffs.
"Get out Get out right now and never, ever come back here."
Procedural voice profile (top-5 deviations vs the average voice):
Extremely distressed and negative. [VALN z=-3.5]
Extremely denasal and clear. [R_NASL z=-3.2]
Extremely darkening in mood. [VALS z=-2.2]
Extremely loud and non-intimate. [S_ASMR z=-2.1]
Extremely emotionally volatile. [VOLT z=+2.1]
Pain Scream raw agonized screaming
Reward dims:AROUTENSVOLTS_RANTROUGARSH· invWER gate:OFF (no intelligibility gate)
Someone SCREAMING in great physical pain: raw, agonized, wild, strained, desperate, tortured wails and cries; mostly non-verbal anguish. Reward favors arousal/tension/volatility/ranting-style/roughness/harshness; NO intelligibility gate (wer=none) - lean fully into raw screaming.
Champion instruction (no reference audio — this text alone steers the 4.55B model):
Wailing and shrieking in raw physical torment, voice torn, ragged and guttural, harsh grating rasp shredding through, breaking on every word. Extreme arousal, extreme tension, violently volatile, uncontrolled and frantic. (agonized scream) (wail) High quality recording.
Reward trajectory (best mean per generation, 4 gens): 1.85 → 1.88 → 1.81 → 2.04 Absolute values are min-max normalized per gen_score call, so they are only comparable within a call — champions were picked by same-call head-to-head playoffs.
"No It hurts It hurts so much Please make it stop"
Procedural voice profile (top-5 deviations vs the average voice):
Extremely non-instructional in style. [S_TECH z=-2.4]
Extremely cold and sterile. [WARM z=-2.4]
Extremely strongly emphasized. [EMPH z=+2.4]
Extremely one-dimensional in resonance. [R_MIXD z=-2.3]
Extremely emotionally volatile. [VOLT z=+2.2]
Compute & timing
4,845
clips generated
2.6 h
audio GPU-seconds (9,486s)
~0.51
mean clips / s
10
archetypes
Each archetype ran ~5–7 generations of a ~10-instruction population × best-of-8, plus a final best-of-64 draw and (most) a variety-text check. That is the "10 per population" scene-generation cost the study was sized around: one generation = 10 candidate instructions × 8 samples = 80 clips ≈ 160 s of generation at ~0.5 clips/s (~2.7 min of GPU per generation), so a full ~6-generation evolution + best-of-64 + variety lands around 460–770 clips per archetype. Mouse and pain-scream were generated in one combined run (768 clips shared), counted once above.
Archetype
clips
seconds
clips/s
gens
note
Zombie
462
924
0.5
5
Orc
472
946
0.499
5
Fairy
472
895
0.527
5
Mouse
768
1497
0.513
4
combined across both archetypes on GPU3; per-archetype ~half
Dragon
616
1232
0.5
6
Goblin
648
1284
0.505
7
ASMR Man
463
837
0.553
5
ASMR Woman
472
902
0.524
5
Ranting
472
970
0.487
5
Pain Scream
768
1497
0.513
4
combined across both archetypes on GPU3; per-archetype ~half
Detailed lessons — build on this next time
The point of the study is a reusable playbook. First the strong convergent findings that held across archetypes, then the raw per-archetype notes.
Convergent findings
1. Describe concrete acoustics, not the archetype label.
Naming what the voice physically does (wet phlegmy throat, nasal buzz, massive chest resonance) beats naming the character. The literal word “cartoon” scored dead-last for goblin despite S_CART being a rewarded dim; “raspy gravelly cartoon goblin” = 1.03, worst of its generation. Describe cartoonish acoustics implicitly (exaggerated, shrill, gleeful) instead.
2. Balance beats char-maxing — because char is only 1.5×.
Reward = norm(blend) + norm(genu) + 1.5·char. Pushing roughness / whisper / pitch to extremes spikes the character dim on a single take but collapses the VoiceCLAP blend (esthetics) term, which drops overall reward. Zombie roughness-only hit char 0.77 but blend 3.27 → lowest reward; over-whispering crushed ASMR blend toward 0–1.5; ranting/scream pure-volume garbled the words. Every archetype rediscovered this.
3. One ()-burst maximum; extra bursts and roars hurt.
A single round-bracket director note — (wet gurgle), (light giggle), (deep rumble), (soft breath), (furious snarl) — reliably added character with no WER cost. Two or three stacked bursts, and especially escalation-into-a-ROAR, lowered blend and intelligibility for no char gain (dragon double-burst dropped invWER to ~0.79; goblin burst-heavy prompt was last). Bursts always go in the instruction, never in the spoken text (that ~5×’s WER).
4. Continuous-escalation phrasing is a big lever.
“The voice keeps cracking higher, sharper, shriller … tightening and quickening as it goes” jumped goblin from ~1.43 to ~1.83 mean by driving the arousal-shift (ARSH) and tension (TENS) dims. Ranting’s “surging louder and louder” and orc’s “mounting fury” work the same way. (Exception: for the static-timbre archetypes zombie/dragon, escalation was neutral-to-slightly-negative.)
5. Explicitly negate the penalized dims.
Spelling out “no chest, no warmth, no fullness” (goblin), “no roughness at all, no chest” (fairy), “dark, zero brightness” (orc/dragon) directly suppressed the negative reward dims and was harmless to intelligibility. Cheap, reliable points.
6. The ONE place a style label won: ASMR-woman.
“A gentle woman doing ASMR” as a literal frame anchored the female GEND dim + S_ASMR/S_WHIS better than a pure acoustic description. This is the single exception to the describe-acoustics-not-labels rule — because “ASMR” names a very specific, well-represented recording style AND fixes perceived gender at once.
7. Per-call normalization is a real methodology caveat.
The min-max normalization inside gen_score runs per call, so absolute rewards are only comparable within one best-of-N call. Cross-generation trajectories are noisy (~±0.2). Every agent worked around this by running same-call “playoffs” — putting the champions of different generations into one shared scoring call and reading the head-to-head ranking.
8. Lean prompts beat kitchen-sink prompts.
Stacked “very very / extremely” intensifiers and long flowery descriptions plateaued or hurt. Zombie’s leanest single-burst finalist produced the highest single-take rewards (max 2.94); “max-everything” prompts made the model clean up rather than commit to the effect.
9. Per-archetype specifics worth reusing.
Pain-scream: short 5–7 s non-verbal bursts beat long ~12 s takes, which drift into calmer articulate speech. Ranting: the safety clause “enraged but the words survive” keeps invWER ~1.0 under the soft gate. Orc/dragon: name chest AND throat resonance together. Mouse vs fairy: mouse wants nasal-buzz + tense + cartoon-perky, fairy wants glassy head-resonance + warmth — describing the wrong one collapses the target.
What worked (per archetype)
Zombie
Framing 'a rotting corpse's voice' + 'all resonance packed into a wet constricted throat' simultaneously maximized throat-resonance/roughness (char) AND kept VoiceCLAP blend high — the gen3 champ g3z3 hit char 0.738 with blend 6.97.
Naming a concrete wet/phlegmy/gurgling throat plus 1-2 (wet gurgle)/(raspy wheeze) director bursts drove ROUG+R_THRT+DFLU; adding 'mind rotted and vacant'/'brain-dead' cheaply lifted cognitive-load (COGL).
Leaner prompts (finalist g5z8, single burst) produced the highest single-take rewards in best-of-N (max 2.94 in gen5, 2.74 in the best-of-64) — fewer competing descriptors gave more consistent extreme decay.
The WER gate is OFF for zombie, so slurred/groaning delivery is free; leaning all the way into decay never costs reward.
Orc
Concrete stacked acoustics won: very very deep + extremely guttural bass + enormous chest and throat resonance + extremely rough grinding rasp. Naming chest AND throat resonance explicitly (both are positive dims R_CHST/R_THRT) is key.
A ranting action frame — spitting his words with mounting fury / seething and ranting with rage — reliably lifted mean reward by engaging the S_RANT/AROU/TENS dims without hurting intelligibility (this framing won the strong gen4 playoff).
zero brightness explicitly suppresses the BRGT negative dim; dark and heavy reinforce it.
Words stayed fully intelligible (invWER 1.0 on every top take) even at max roughness, so the soft WER gate never bit — roughness and clarity coexisted.
Fairy
Compact, dense acoustic descriptions ('delicate high, feather-light and airy, warm and very bright, glassy head resonance, crystal clear') beat long flowery ones.
A single trailing (light giggle) director-note reliably lifted the characteristic (brightness/head-resonance) score and mean_reward without hurting WER.
Explicit negation of the penalized dims — 'no roughness at all', 'no chest' — directly raised reward (ROUG/R_CHST are penalized dims).
Intensifier 'very bright' on the top positive dim (BRGT) gave the cleanest separation in the final playoff.
'feather-light and airy' + 'glassy/crystal clear' pushed head-resonance & esthetics while keeping full intelligibility (WER-gated archetype).
Mouse
Naming concrete acoustics drove the characteristic dims: 'strong nasal buzz' (R_NASL), 'bright/piercing head resonance' (R_HEAD/BRGT), 'pinched and tense throat' (TENS), 'cartoon-style, exaggerated' (S_CART).
Explicit smallness/youth negations ('tiny and young', 'weightless, no chest resonance or fullness') suppressed the penalized dims AGEV/R_CHST/FULL and lifted char.
The balanced full-acoustic description (m4c7) that pairs high char with high genuineness beat pure char-maxers in the fair shootout; best-of-64 top take reached reward 2.73 (char 0.75, blend 5.4).
'Very high squeaky voice, thin and bright' + fast/chipper pacing reliably produced the small perky timbre; soft WER gate meant near-perfect intelligibility (invwer ~0.93-1.0) so it never cost reward.
Dragon
'A grand, theatrical, dramatic proclamation' framing directly targets the S_DRAM+S_AUTH reward dims and is the single strongest opener - both finalists use it.
'Dark' / 'menacing dark bass' suppresses the penalized BRGT dim; pairs well with 'deep/booming/bass'.
Age words 'ancient, aeons-old' cheaply raise AGEV with no intelligibility cost.
Exactly ONE vocal burst - either (deep rumble) or (low growl) - as a round-bracket director note; keeps delivery clean (invwer often 1.0).
A concrete acoustic metaphor ('like distant thunder rolling through a vast cavern') helped the runner-up family; evocative-but-clean imagery adds character without hurting WER.
Trailing 'High quality recording.' tag, and audio_temp 1.1 for good best-of-N breadth.
'full and heavy' / 'full, thick and heavy' nudges the FULL dim upward.
Goblin
Continuous-escalation phrasing is the single biggest lever: 'the voice keeps cracking higher, sharper, shriller and whinier ... tightening and quickening as it goes' drove the arousal-shift (ARSH) and tension (TENS) dims and jumped best mean from ~1.43 (gen2) to ~1.83 (gen3).
'giddy with spite' / 'gleeful cruel malice' emotional-affect words compounded the escalation gain (gen4->gen5 champion g5c2).
Fusing three concrete acoustic axes wins: thin/pinched NASAL buzz ('buzzing high and reedy through the nose') + rough scratchy grating RASP + escalation. 'rough scratchy grating rasp' beat plain 'dry rasp' in the gen7 playoff (g7c10).
Explicit negation of the penalized dims ('no chest, no warmth, no fullness') in every prompt reliably suppressed FULL/R_CHST/WARM and was harmless to intelligibility.
Soft WER gate stayed satisfied (invWER 0.77-1.0) because all bursts/stage-notes were kept OUT of the spoken text; best-of-64 top take hit invWER 0.923 while keeping 'Heh heh!'.
ASMR Man
'A deep, gentle man' prefix reliably pushed the (negative) GEND dim toward male and lifted char without hurting intelligibility.
'clear whispered words' + a SINGLE (soft breath) burst was the key balance: it protected VoiceCLAP blend/ESTH (7-9) while keeping whisper/respiration high — finalist g5a1 hit char 0.60 / blend 7.27 / genu 1.86, and the best-of-64 top-3 all had char ~0.67 with blend ~7.8.
Stacking 'soft, breathy, intimate close-mic whisper, very gentle and slow, low arousal, warm/tender/soothing' hit S_ASMR+S_WHIS+WARM+VULN+low-AROU together.
invWER stayed ~1.0 through EVERY generation — a soft male whisper did not hurt Parakeet intelligibility, so the WER gate never actually bound; whispering freely was safe.
ASMR Woman
The literal frame A gentle woman doing ASMR anchors the female (GEND) + S_ASMR/S_WHIS dims better than a pure acoustic description without the ASMR label — the one place naming the style beat describing only properties.
Paired director bursts (soft breath) and (gentle whisper) inside the instruction raised the RESP/whisper dims with zero intelligibility cost (every top take invWER 1.0).
Stacking stillness/quietness words — hushed and quiet, hushed quiet and still — was the single biggest jump (gen3), suppressing the AROU negative dim; adding a soft narrative tail soothing you to rest lifted warmth/vulnerability and won the gen5 playoff.
Explicit very low arousal plus warm/tender/kind/comforting kept AROU low while boosting WARM/VULN/ESTH; final best-of-64 top take hit blend 9.69 with char 0.5, all perfectly clear.
Ranting
'screaming a furious tirade... voice cracking and strained' captured arousal/tension far better than generic 'yells angrily'.
Naming the reward dims as concrete acoustics — 'sharp percussive attack on the consonants' (ATCK), 'pounding emphasis' (EMPH), 'explosive and volatile / surging louder and louder' (VOLT) — steadily raised the characteristic score.
(furious snarl) as a director-note was the best single burst; it added rasp/aggression while the words survived.
The safety clause 'Enraged but the words survive' / 'still intelligible' kept invWER ~1.0 (soft gate still rewards intelligibility), letting high-arousal takes keep their reward.
'seething and spitting with rage' phrasing gave the best char+intelligibility balance of any variant.
Pain Scream
Action verbs 'Wailing and shrieking ... shredding through' plus texture words 'torn, ragged and guttural, harsh grating rasp' drove ROUG/ARSH; 'Extreme arousal, extreme tension' drove AROU/TENS; 'violently volatile, uncontrolled and frantic' drove VOLT/S_RANT.
Adding 'guttural' and 'uncontrolled and frantic' to the winning line (p4c2) gave the highest AND most consistent char (0.84, mean 2.039 with tight spread), beating higher-blend but lower-char rivals.
No intelligibility gate (wer=none) let fully non-verbal, shredded delivery score freely; best-of-64 top take hit reward 2.55 (char 0.85, blend 6.2).
(agonized scream) and (wail) director bursts inside the instruction helped push the raw-scream character without hurting anything (no WER cost).
What failed (per archetype)
Zombie
Max-everything prompts ('very very extremely' + three bursts, g1z10) COLLAPSED char to 0.41 — piling on conflicting cues made the model clean up rather than decay.
Pushing roughness alone (g1z6 char 0.767) tanked blend to 3.27 and gave the LOWEST reward: reward = norm(blend)+norm(genu)+1.5*char rewards BALANCE, not pure roughness.
Heavy stammering / 'half-formed words dissolving' disfluency (g2z4) crushed blend to 2.56.
Orc
Pushing roughness to extremes (voice like grinding boulders, thick vocal fry, snarling-only) spiked the char dim but collapsed the VoiceCLAP blend score to 2-4, tanking overall reward. char alone is only 1.5x weight vs blend+genu.
Loud/roar bursts and (roar) director notes did not help mean reward and sometimes lowered blend/consistency.
Fairy
Inline vocal bursts in the text were never used (per protocol); the giggle only works as a (round-bracket) note in the instruction.
Doubling the giggle ((light giggle)(light giggle)) or adding (airy sigh) did NOT beat a single giggle.
Naming the label ('a little flower fairy', 'enchanting fairy voice') without dense acoustics scored mid-pack.
Mouse
Pure char-maximizing phrasing (e.g. double bursts, extreme pinch) spiked char to ~0.77-0.86 on single takes but collapsed blend/genuineness, lowering mean_reward.
Exotic pitch metaphors ('piccolo-high', 'featherlight', 'helium') did not outperform plain 'very high squeaky' and sometimes reduced consistency.
Dragon
Escalation into a ROAR ('swells into a thunderous commanding roar', g1c5) was the worst gen1 prompt - it trades fullness/intelligibility for loudness the reward doesn't want.
Two stacked bursts '(deep rumble) (low growl)' (g1c9, g3c7) lowered intelligibility (invwer dropped to ~0.79) with no char gain.
'so deep the very stone trembles' (g5c2) and other purple-prose intensity phrases hurt WER (invwer 0.79) for no reward gain.
Putting bursts/stage directions in the TEXT was never done (protocol) - would ~5x WER.
Goblin
Vocal-burst director notes like '(nasal snicker)' / '(snickering cackle)' inside the instruction did NOT help and usually hurt: burst-heavy gen1 g1c9 was last (1.22), and every A/B with vs without a burst on the same base (g4c5, g5c6, g6c10) ranked below its burst-free twin.
The literal word 'cartoon' tanked despite S_CART being a positive reward dim: gen2 g2c5 'raspy gravelly cartoon goblin' scored 1.03, dead last. Describe cartoonish acoustics implicitly (exaggerated, shrill, gleeful) rather than naming the style.
Over-whispering ('extremely soft' + heavy airy/breath emphasis, g1a5/g2a6/g4a7) collapsed blend toward 0-1.5 — ESTH tanks when the voice becomes pure air. This was the dominant failure mode.
Two bursts (soft breath)+(gentle whisper) or 'softening/escalation' phrasing raised char but added variance and lowered genu/blend on average.
ASMR Woman
Over-breathy framings (speaking almost entirely in breath, breath audible on every word) pushed the char/whisper dim up but crushed the VoiceCLAP blend to ~4 and lost mean reward — the full WER gate and blend term punish an unclear whisper.
Adding soft as a feather / soothing ASMR as the ONLY framing (gen1 a1c9/a1c10) under-scored vs the concrete ASMR + very-low-arousal + hushed formula.
Ranting
Pushing pure volume ('so hard the voice strains', 'bellowing at ear-splitting volume') garbled words (invWER dropped to ~0.73) and cost reward despite the soft gate.
(shout) as the primary burst was noisier and less reliable than (furious snarl); stacking (shout)(furious snarl) did not help.
Extra body cues ('veins popping','chest heaving','jaw clenched') added nothing measurable.
Pain Scream
Long takes (~12s) consistently scored worst - they drift into calmer, more articulate speech that dilutes scream arousal/roughness. Winning takes were short (~5-7s) bursts of pure anguish.
Char-maxing lines (e.g. p2c2, p2c1) hit char up to 0.96 on single takes but with weak blend, so mean_reward lagged the balanced phrasings.
Mediocre / within-noise (per archetype)
Zombie
(low guttural groan) vs (raspy wheeze) vs (wet gurgle) bursts were roughly interchangeable; two bursts was the sweet spot, three slightly worse.
Escalation phrasing ('starts as a groan and escalates') was neutral-to-slightly-negative for this archetype.
mean_reward is min-max normalized WITHIN each gen_score call, so cross-generation absolute values are noisy (~+/-0.2); a same-call champions head-to-head (gen5) was needed to pick the finalist fairly.
Orc
Single (guttural growl)/(snarl) bursts were roughly neutral vs no bursts — occasionally a small max_reward boost but no mean gain; the acoustic description carries the delivery.
Escalation framing (starts low, escalates into a roar) was decent but slightly below the sustained ranting-fury frame.
Extra intensifier words beyond the core set gave diminishing returns; reward plateaued around 1.6-1.75 mean from gen1 onward.
Fairy
'tinkling bells' / 'silvery' imagery was decent but no better than plain 'bright/glassy' acoustics.
'extremely high-pitched' produced high-variance takes (great max, weaker mean) — pitch too extreme sometimes thinned the voice.
Mouse
Adding (tiny squeak)/(excited squeak) director bursts was roughly neutral for mouse - occasionally helped a single take but did not raise mean over the burst-free description.
Stacking more intensifiers ('very very', 'extremely', 'intense') past a point gave diminishing returns; reward plateaued around gen1's level for 4 generations.
Dragon
Adding 'authoritative' on top of 'imperious and commanding' (g4c2/g6f5): redundant, no measurable lift.
'sub-bass' vs 'bass' (g5c1) and 'gravelly guttural' vs 'guttural' (g5c4): within noise.
'pitch-black ... no brightness' explicit BRGT-suppression (g3c4/g5c7): about even with plain 'dark' - 'dark' already does the job.
Goblin
Pure static acoustic descriptions (nasal+rasp, no escalation) plateaued around the pack average; they are solid but never won once the escalation phrasing existed.
Adding 'strangled/tight in the throat' tension words was roughly neutral (g6c7, g4c9) - the escalation already supplies TENS.
'honking' vs 'buzzing' for nasal onset was a wash across playoffs.
ASMR Man
'very very gentle'/'unhurried'/'hushed' intensifiers gave higher single-take ceilings (g4a8 max 2.78) but noisier means; 'extremely soft' traded blend for char roughly evenly.
Explicit vulnerability wording ('gently vulnerable') did little for the VULN dim and sometimes hurt blend.
Because reward is normalized within each call, a same-call gen5 head-to-head of champions was used to select the finalist.
ASMR Woman
Leading with (gentle whisper) instead of (soft breath) was roughly neutral to slightly worse.
feathery and perfectly clear helped char but did not consistently beat the plainer hushed and quiet, whispered but clear phrasing.
soft gentle woman vs gentle woman made little difference; very very low arousal did not beat very low arousal.
Ranting
'red-faced and raging drill sergeant' framings were intelligible but landed low on the characteristic dims.
Ranting reward was noisy across best-of-8 (high per-take variance): high-arousal takes occasionally slur, so best-of-N matters more here than for fairy.
Pain Scream
'howling' as a third verb ('shrieking, wailing and howling') was roughly neutral vs the two-verb form.
Triple bursts (agonized scream)(pained cry)(wail) sometimes raised blend but did not reliably raise char/mean over the two-burst (agonized scream)(wail) winner.