DramaBox Reinterpretation — dynamic emotion+VoiceNet LoRA merge vs. baseline
The long DramaBox clips, reinterpreted two ways and reward-ranked (emotion-similarity × inverse-WER), best-of-8 each. Left = the plain v2 baseline (no LoRA). Right = v2 with a dynamic mix of LoRAs merged in RAM, whose strengths are read off the input clip itself: the top-3 EmoNet emotions (raw 0–6 → 1→0%, 2→50%, 3→100%, 4→200%) and the top-3 VoiceNet dimensions by deviation from the neutral middle (|v−0.5|·2 → 100% at the extremes, high/low adapter by side). Same reference audio, same script text (now correctly aligned), generous length so nothing is cut off.
Best-of-8: baseline vs dynamic merge (mean over 10 long clips)
variant
mean reward (top1)
mean WER
mean emo-sim
Baseline (no merge)
0.139
0.699
0.414
Dynamic LoRA merge
0.042
0.950
0.223
Emotional similarity to the input changes from 0.414 to 0.223 with the dynamic merge; reward 0.139 → 0.042.
#111 · 31.24s · “And look at that! Look at what your good rain has done, cleared the way hasn't it? Cle- we've cleared away there a little boundary”
A the audio delivers a, transitional acoustics of an adolescent. Timbre: perfect; distinctly tinny and highly top-heavy; the acoustic signal is heavily. Delivery: cartoonish, dramatic, casual. Emotion
reward -0.000 · emo-sim -0.138 · WER 1.000 · 42.24s
✦ Dynamic emotion+VoiceNet merge
#1
reward 0.118 · emo-sim 0.511 · WER 0.768 · 17.44s
#2
reward 0.015 · emo-sim 0.417 · WER 0.963 · 1.36s
#3
reward 0.002 · emo-sim 0.178 · WER 0.988 · 0.96s
#37 · 30.23s · “The the application process, quite lengthy, was it. I s- quite a long wait personally. I've spent an incredible amount of time on ”
A the audio delivers a, the acoustic signature reflects a fully. Timbre: the treble frequencies are significantly; standard human baseline; a distinct. Delivery: monologue, cartoonish, narrator. Emoti
reward -0.000 · emo-sim -0.084 · WER 1.000 · 0.24s
#38 · 29.28s · “And then the email arrived. It you always know I had a little impression of shock because I I it stated I know congratulations. It”
A the audio delivers a, the acoustic signature reflects a fully. Timbre: the treble frequencies are significantly; standard human baseline; a distinct. Delivery: casual, conversational, monologue. Emo
#47 · 28.06s · “es, wie still es hier ist, wie ich dein Herz nicht mehr lesen kann. Bitte, sag mir, dass es nicht vorbei ist, dass es noch einen W”
A feminine voice, voice at peak biological vigor. Timbre: perfect; standard human baseline; standard. Delivery: casual, monologue, dramatic. Emotion: distress, helplessness, disappointment, fear, long
reward -0.000 · emo-sim -0.178 · WER 1.000 · 0.32s
#3
reward -0.000 · emo-sim -0.204 · WER 1.000 · 0.32s
#1 · 23.65s · “so it's so big I could barely hear myself all those people just looking. But then the music started and I remembered the high note”
A feminine voice, voice at peak biological vigor. Timbre: perfect; standard human baseline; standard. Delivery: casual, conversational, monologue. Emotion: interest, relief, elation, pride, pleasure e