TL;DR. Best reward is WITHOUT the LoRA (BASE 0.485; BASE_P 0.465, LoRA50 0.468). This target is effectively unlearnable in this rig — emo sits at or below zero at baseline (-0.004) and even LoRA150 only reaches 0.084, so use a plain neutral prompt and don't expect audible intoxication. Pushing the LoRA does lift genuineness (0.110 to 0.265 at 150%) but collapses blend (0.453 to 0.254), quality (0.625 to 0.342) and WER (up to 0.579), a net reward loss. Worse, the merged model doesn't render intoxication at all — it shifts up Triumph, Amusement and Malevolence and correlates with Confusion/Fear, i.e. you get slurred, degraded audio miscolored as other high-arousal states.
Without a LoRA
Best: neutral prompt, no LoRA → reward 0.49
GENERAL: A voice overwhelmingly expressing intoxication altered states of consciousness, a woozy, slurring voice, intoxicated and untethered, in an altered haze, impossible to hide.
SCRIPT: (woozy, slurring, intoxicated) "<neutral sentence>" · temp 1.0 top_p 0.9 top_k 30
− anti-correlates: Focus (-0.64), Authoritative Style (-0.58), Tempo (-0.54)
Conclusion
Best reward is WITHOUT the LoRA (BASE 0.485; BASE_P 0.465, LoRA50 0.468). This target is effectively unlearnable in this rig — emo sits at or below zero at baseline (-0.004) and even LoRA150 only reaches 0.084, so use a plain neutral prompt and don't expect audible intoxication. Pushing the LoRA does lift genuineness (0.110 to 0.265 at 150%) but collapses blend (0.453 to 0.254), quality (0.625 to 0.342) and WER (up to 0.579), a net reward loss. Worse, the merged model doesn't render intoxication at all — it shifts up Triumph, Amusement and Malevolence and correlates with Confusion/Fear, i.e. you get slurred, degraded audio miscolored as other high-arousal states.