← all VoiceNet LoRAs

Cartoonish Style — low

Evolving prompt+sampling to push VoiceNet dimension S_CART in the low direction with the vn_S_CART__low LoRA. Steering phrase: natural realistic human voice.

Best in-direction score 0.989 at 100% vs neutral 0.788 (shift +0.200).
Method & reproducibility — how the evolution ran
Base model: laion/moss-tts-local-transformer-4.55b-voice-acting-v2 · LoRA: the new VoiceNet-dimension LoRA vn_<CODE>__<high|low> (retrained with ≤25% game-character data, rank 32), applied via PEFT adapter scaling at strength λ.
Algorithm: genetic search — population 12, 10 generations, elitism (top 5). During evolution every genome is generated best-of-6 with the LoRA at λ=100%; its fitness = the mean over those 6 candidates of the target VoiceNet dimension score in the intended direction (for a __high LoRA the raw score; for __low, 1−score), minus 0.2·max(0, WER−0.6).
Genome (what evolves): the GENERAL voice description, the SCRIPT cue, the carrier sentence, and the sampling params (temperature, top_p, top_k). λ is fixed at 100% during the search.
Mutation operators: 40% regenerate description, 25% swap carrier, 17% jitter temperature ±0.1, 18% change top_k. Offspring = 60% crossover + mutation, 40% mutation.
Dose measurement (below): the winning recipe is then re-generated at λ=0 (LoRA OFF), λ=50% and λ=100% (each best-of-6) to show exactly what 50% vs 100% of the LoRA brings over the no-LoRA baseline.

Dose — what 50% vs 100% brings (top recipe)

λ=0 (no LoRA) — dir 0.773 (raw 0.23)
WER 0.26 · genu 0.06
λ=50% — dir 0.939 (raw 0.06)
WER 0.09 · genu 0.04
λ=100% — dir 0.989 (raw 0.01)
WER 0.09 · genu 0.06
GENERAL: A voice that is strongly natural realistic human voice. SCRIPT: (markedly natural realistic human voice) "The meeting is scheduled for three o'clock on Tuesday afternoon." · temp 1.1 top_p 0.9 top_k 25

Other evolved recipes (100%)

— dir 0.984 (raw 0.02)
WER 0.09 · genu 0.05
A voice that is markedly natural realistic human voice.
— dir 0.982 (raw 0.02)
WER 0.15 · genu 0.06
A voice that is extremely natural realistic human voice.
— dir 0.979 (raw 0.02)
WER 0.11 · genu 0.04
A voice that is strongly natural realistic human voice.
— dir 0.978 (raw 0.02)
WER 0.11 · genu 0.06
A voice that is markedly natural realistic human voice.
— dir 0.968 (raw 0.03)
WER 0.09 · genu 0.06
A voice that is markedly natural realistic human voice.