Evolving prompt+sampling to push VoiceNet dimension S_CASU in the high direction with the vn_S_CASU__high LoRA. Steering phrase: a voice with strongly pronounced casual style, casual style-forward and unmistakable.
Best in-direction score 0.841 at 100% vs neutral 0.301 (shift +0.540).
Method & reproducibility — how the evolution ran
Base model:laion/moss-tts-local-transformer-4.55b-voice-acting-v2 · LoRA: the new VoiceNet-dimension LoRAvn_<CODE>__<high|low> (retrained with ≤25% game-character data, rank 32), applied via PEFT adapter scaling at strength λ. Algorithm: genetic search — population 12, 10 generations, elitism (top 5). During evolution every genome is generated best-of-6with the LoRA at λ=100%; its fitness = the mean over those 6 candidates of the target VoiceNet dimension score in the intended direction (for a __high LoRA the raw score; for __low, 1−score), minus 0.2·max(0, WER−0.6). Genome (what evolves): the GENERAL voice description, the SCRIPT cue, the carrier sentence, and the sampling params (temperature, top_p, top_k). λ is fixed at 100% during the search. Mutation operators: 40% regenerate description, 25% swap carrier, 17% jitter temperature ±0.1, 18% change top_k. Offspring = 60% crossover + mutation, 40% mutation. Dose measurement (below): the winning recipe is then re-generated at λ=0 (LoRA OFF), λ=50% and λ=100% (each best-of-6) to show exactly what 50% vs 100% of the LoRA brings over the no-LoRA baseline.
Dose — what 50% vs 100% brings (top recipe)
λ=0 (no LoRA) — dir 0.512 (raw 0.51)
WER 0.00 · genu 0.15
λ=50% — dir 0.616 (raw 0.62)
WER 0.33 · genu 0.19
λ=100% — dir 0.841 (raw 0.84)
WER 0.90 · genu 0.23
GENERAL: A voice that is extremely a voice with strongly pronounced casual style, casual style-forward and unmistakable.
SCRIPT: (markedly a voice with strongly pronounced casual style, casual style-forward and unmistakable) "Please remember to lock the door before you head out." · temp 1.0 top_p 0.9 top_k 30
Other evolved recipes (100%)
— dir 0.773 (raw 0.77)
WER 0.52 · genu 0.22
A voice that is strongly a voice with strongly pronounced casual style, casual style-forward and unmistakable.
— dir 0.770 (raw 0.77)
WER 0.42 · genu 0.19
A voice that is distinctly a voice with strongly pronounced casual style, casual style-forward and unmistakable.
— dir 0.763 (raw 0.76)
WER 0.36 · genu 0.25
A voice that is distinctly a voice with strongly pronounced casual style, casual style-forward and unmistakable.
— dir 0.759 (raw 0.76)
WER 0.41 · genu 0.25
A voice that is extremely a voice with strongly pronounced casual style, casual style-forward and unmistakable.
— dir 0.745 (raw 0.74)
WER 0.68 · genu 0.19
A voice that is extremely a voice with strongly pronounced casual style, casual style-forward and unmistakable.