DramaBox Reinterpretation — dynamic emotion+VoiceNet LoRA merge vs. baseline
The long DramaBox clips, reinterpreted two ways and reward-ranked (emotion-similarity × inverse-WER), best-of-8 each. Left = the plain v2 baseline (no LoRA). Right = v2 with a dynamic mix of LoRAs merged in RAM, whose strengths are read off the input clip itself: the top-3 EmoNet emotions (raw 0–6 → 1→0%, 2→50%, 3→100%, 4→200%) and the top-3 VoiceNet dimensions by deviation from the neutral middle (|v−0.5|·2 → 100% at the extremes, high/low adapter by side). Same reference audio, same script text (now correctly aligned), generous length so nothing is cut off.
reward -0.000 · emo-sim -0.135 · WER 1.167 · 7.92s
#3
reward 0.000 · emo-sim 0.046 · WER 1.000 · 4.88s
✦ Dynamic emotion+VoiceNet merge
#1
reward 0.073 · emo-sim 0.441 · WER 0.833 · 9.36s
#2
reward 0.052 · emo-sim 0.312 · WER 0.833 · 0.96s
#3
reward 0.050 · emo-sim 0.150 · WER 0.667 · 0.96s
#11 · 21.7s · “A cruel dream, I should I should just walk towards it, even if even if it's not what I think, it's just so far, so so far.”
A the audio delivers a, voice at peak biological vigor. Timbre: perfect; standard human baseline; standard. Delivery: casual, conversational, playful. Emotion: contemplation, sadness, fatigue exhausti
#83 · 20.32s · “and in the chapter that's right here. Oh”
A feminine voice, voice at peak biological vigor. Timbre: the treble frequencies are significantly; highly functional but distinctly neutral-cool; lmost completely obliterated by violent. Delivery: ca