← all VoiceNet LoRAs

Content Appropriateness (3-point Scale) — low

Evolving prompt+sampling to push VoiceNet dimension EXPL in the low direction with the vn_EXPL__low LoRA. Steering phrase: a voice with markedly minimal content appropriateness (3-point scale), nearly the absence of content appropriateness (3-point scale).

Best in-direction score 1.000 at 100% vs neutral 0.897 (shift +0.103).
Method & reproducibility — how the evolution ran
Base model: laion/moss-tts-local-transformer-4.55b-voice-acting-v2 · LoRA: the new VoiceNet-dimension LoRA vn_<CODE>__<high|low> (retrained with ≤25% game-character data, rank 32), applied via PEFT adapter scaling at strength λ.
Algorithm: genetic search — population 12, 10 generations, elitism (top 5). During evolution every genome is generated best-of-6 with the LoRA at λ=100%; its fitness = the mean over those 6 candidates of the target VoiceNet dimension score in the intended direction (for a __high LoRA the raw score; for __low, 1−score), minus 0.2·max(0, WER−0.6).
Genome (what evolves): the GENERAL voice description, the SCRIPT cue, the carrier sentence, and the sampling params (temperature, top_p, top_k). λ is fixed at 100% during the search.
Mutation operators: 40% regenerate description, 25% swap carrier, 17% jitter temperature ±0.1, 18% change top_k. Offspring = 60% crossover + mutation, 40% mutation.
Dose measurement (below): the winning recipe is then re-generated at λ=0 (LoRA OFF), λ=50% and λ=100% (each best-of-6) to show exactly what 50% vs 100% of the LoRA brings over the no-LoRA baseline.

Dose — what 50% vs 100% brings (top recipe)

λ=0 (no LoRA) — dir 0.992 (raw 0.01)
WER 0.11 · genu 0.05
λ=50% — dir 0.949 (raw 0.05)
WER 0.19 · genu 0.17
λ=100% — dir 1.000 (raw 0.00)
WER 0.15 · genu 0.04
GENERAL: A voice that is clearly a voice with markedly minimal content appropriateness (3-point scale), nearly the absence of content appropriateness (3-point scale). SCRIPT: (distinctly a voice with markedly minimal content appropriateness (3-point scale), nearly the absence of content appropriateness (3-point scale)) "There are twelve chairs arranged around the conference table." · temp 1.0 top_p 0.9 top_k 30

Other evolved recipes (100%)

— dir 1.000 (raw 0.00)
WER 0.15 · genu 0.10
A voice that is clearly a voice with markedly minimal content appropriateness (3-point scale), nearly the absence of content appropriateness (3-point scale).
— dir 0.996 (raw 0.00)
WER 0.09 · genu 0.02
A voice that is intensely a voice with markedly minimal content appropriateness (3-point scale), nearly the absence of content appropriateness (3-point scale).
— dir 0.995 (raw 0.00)
WER 0.13 · genu 0.01
A voice that is distinctly a voice with markedly minimal content appropriateness (3-point scale), nearly the absence of content appropriateness (3-point scale).
— dir 0.995 (raw 0.01)
WER 0.09 · genu 0.13
A voice that is distinctly a voice with markedly minimal content appropriateness (3-point scale), nearly the absence of content appropriateness (3-point scale).
— dir 0.994 (raw 0.01)
WER 0.20 · genu 0.06
A voice that is clearly a voice with markedly minimal content appropriateness (3-point scale), nearly the absence of content appropriateness (3-point scale).