What the VoiceNet LoRAs do — evolutionary study (50% vs 100% + optimal prompting)

For each new VoiceNet-dimension LoRA (retrained with ≤25% game-character data), a 10-generation evolutionary search finds the prompt+sampling that best pushes its dimension, then measures the effect at 0% / 50% / 100% strength. Shows what each LoRA brings, how 50% compares to 100%, and the best way to prompt it. Fitness = mean over 6 candidates of the dimension in its intended direction. Every per-LoRA page has a Method & reproducibility box; dose players are marked λ=0 OFF / λ=50% / λ=100%.

Method & reproducibility — how the evolution ran
Base model: laion/moss-tts-local-transformer-4.55b-voice-acting-v2 · LoRA: the new VoiceNet-dimension LoRA vn_<CODE>__<high|low> (retrained with ≤25% game-character data, rank 32), applied via PEFT adapter scaling at strength λ.
Algorithm: genetic search — population 12, 10 generations, elitism (top 5). During evolution every genome is generated best-of-6 with the LoRA at λ=100%; its fitness = the mean over those 6 candidates of the target VoiceNet dimension score in the intended direction (for a __high LoRA the raw score; for __low, 1−score), minus 0.2·max(0, WER−0.6).
Genome (what evolves): the GENERAL voice description, the SCRIPT cue, the carrier sentence, and the sampling params (temperature, top_p, top_k). λ is fixed at 100% during the search.
Mutation operators: 40% regenerate description, 25% swap carrier, 17% jitter temperature ±0.1, 18% change top_k. Offspring = 60% crossover + mutation, 40% mutation.
Dose measurement (below): the winning recipe is then re-generated at λ=0 (LoRA OFF), λ=50% and λ=100% (each best-of-6) to show exactly what 50% vs 100% of the LoRA brings over the no-LoRA baseline.
Averaged over all LoRAs, 50% already delivers 36% of the full-strength shift.

Conclusions

dimensiondirneutral50%100%Δ100
Ranting/Angry Stylehigh0.000.560.78+0.78
Tensionhigh0.220.610.88+0.67
Dynamic Archigh0.160.340.81+0.65
Respirationhigh0.080.280.73+0.65
Monologue Stylelow0.360.620.99+0.63
Cartoonish Stylehigh0.210.240.83+0.61
Cognitive Loadhigh0.120.450.73+0.61
Attackhigh0.260.460.87+0.61
Pitch Rangehigh0.250.510.85+0.60
Dramatic Stylehigh0.390.720.96+0.58
Warmthlow0.390.660.95+0.56
Whisper-Talk Stylehigh0.400.590.96+0.56
Roughnesshigh0.220.390.77+0.55
Casual Stylehigh0.300.620.84+0.54
Registerhigh0.340.690.87+0.53
Articulation Claritylow0.410.760.94+0.53
Emphasishigh0.420.650.94+0.52
Focushigh0.320.540.83+0.51
Volatilityhigh0.290.580.80+0.50
Formal Stylelow0.410.780.91+0.50
Estheticslow0.500.581.00+0.50
Recording Qualitylow0.300.330.79+0.49
Perceived Genderlow0.470.910.96+0.49
Arousalhigh0.430.770.92+0.48
Arousal Shiftlow0.520.821.00+0.48
Velocity Fluxhigh0.460.600.94+0.47
Tempohigh0.260.400.73+0.47
Vulnerabilityhigh0.280.520.75+0.47
Playful Stylehigh0.270.450.74+0.46
Throat Resonancehigh0.330.560.79+0.46
Chest Resonancehigh0.400.740.85+0.45
ASMR Stylehigh0.160.390.60+0.44
Disfluencyhigh0.090.330.53+0.44
Stancehigh0.500.770.94+0.44
Fullnesshigh0.530.650.97+0.43
Perceived Genderhigh0.530.750.95+0.43
Arousal Shifthigh0.480.460.90+0.42
Conversational Stylehigh0.230.540.65+0.42
Valencehigh0.260.450.68+0.42
Metallic Characterhigh0.200.440.61+0.41
Whisper-Talk Stylelow0.600.711.00+0.40
Nasal Resonancehigh0.290.410.69+0.39
Narrator Stylehigh0.530.620.92+0.39
Stancelow0.500.740.89+0.39
Chest Resonancelow0.600.850.98+0.38
Storytelling Stylehigh0.570.830.94+0.37
Authoritative Stylehigh0.280.540.65+0.37
Voice Agehigh0.360.530.72+0.36
Dramatic Stylelow0.610.810.97+0.36
Narrator Stylelow0.470.470.82+0.35
Harmonicityhigh0.460.580.81+0.35
Arousallow0.570.630.92+0.35
Registerlow0.660.991.00+0.34
Harmonicitylow0.540.690.88+0.34
Content Appropriateness (3-point Scale)high0.100.220.44+0.34
Throat Resonancelow0.670.911.00+0.33
Head Resonancehigh0.240.290.56+0.32
Oral Resonancelow0.500.590.81+0.32
Valence Shiftlow0.640.730.94+0.30
Background Noiselow0.430.540.74+0.30
Teacher/Didactic Stylehigh0.280.340.57+0.29
Fullnesslow0.470.550.76+0.29
Velocity Fluxlow0.540.720.82+0.29
Brightnesslow0.560.800.84+0.28
Vulnerabilitylow0.720.661.00+0.28
Warmthhigh0.610.640.89+0.27
Valence Shifthigh0.360.480.63+0.27
Casual Stylelow0.700.720.92+0.22
Playful Stylelow0.730.810.94+0.21
Pitch Rangelow0.750.860.96+0.21
Estheticshigh0.500.600.71+0.21
Monologue Stylehigh0.640.770.85+0.21
Head Resonancelow0.760.870.97+0.21
Mask Resonancehigh0.510.590.72+0.21
Chunkinghigh0.400.450.60+0.21
Structurehigh0.540.630.74+0.20
Metallic Characterlow0.800.981.00+0.20
Cartoonish Stylelow0.790.940.99+0.20
Teacher/Didactic Stylelow0.720.930.92+0.20
Brightnesshigh0.440.540.64+0.19
Nasal Resonancelow0.710.780.90+0.19
Formal Stylehigh0.590.590.77+0.19
Valencelow0.740.800.91+0.17
Newsreader Stylehigh0.260.280.42+0.16
Tensionlow0.780.810.94+0.16
Smoothnesshigh0.480.530.63+0.16
ASMR Stylelow0.840.991.00+0.16
Recording Qualityhigh0.700.740.83+0.13
Oral Resonancehigh0.500.400.63+0.13
Articulation Clarityhigh0.590.560.72+0.13
Dynamic Arclow0.840.940.97+0.13
Background Noisehigh0.570.600.69+0.12
Cognitive Loadlow0.880.931.00+0.12
Content Appropriateness (3-point Scale)low0.900.951.00+0.10
Volatilitylow0.710.740.81+0.10
Disfluencylow0.910.981.00+0.09
Respirationlow0.920.991.00+0.08
Mixed Resonancehigh0.470.490.53+0.06
Roughnesslow0.780.840.83+0.06
Mixed Resonancelow0.530.550.57+0.04