What this is (for dummies). We asked: which wording of the instruction makes the voice-acting model actually sound furious / terrified / heartbroken / ecstatic / disgusted / astonished? We generated 4 takes for every combination of 6 emotions × 20 instruction styles × 3 emotional texts (1,440 clips, no reference audio), and scored every clip with Empathic-Insight-Voice-Plus (emotion-annotations): the target emotion score, arousal (energy), valence, expressiveness, plus the average volume (rms in dB) of the clip. two extra arms test which instructions maximise speaker similarity to a reference voice and genuineness.
Reading the tables: higher = more of that thing. "target" = the Empathic-Insight score of the emotion we asked for (e.g. Anger head for anger prompts). "ctrl" rows are the neutral-instruction control. rms dB is negative — closer to 0 = louder.
1 · Per emotion: which instruction style works best?
expr_index = mean z-score of target emotion + arousal + volume + expressiveness. What reliably cranks the model up — and what does nothing.
style
what it does
expr index
target
arousal
rms dB
invWER
s11_escalation
instruct an escalation arc
0.86
2.336
2.577
-22.5
0.96
s14_loudness
explicit maximum-volume wording
0.59
2.222
2.513
-21.9
0.93
s16_movie_climax
oscar-movie climax framing
0.54
2.274
2.462
-22.1
0.95
s19_kitchen_sink
everything combined
0.51
2.383
2.532
-21.8
0.95
s04_extreme_adj
extreme adjectives, 'pushed to the limit'
0.46
2.283
2.507
-22.3
0.96
s07_shouting
explicit shouting/screaming cues
0.44
2.253
2.509
-21.9
0.95
s13_ellipses
ellipses + exclamation marks
0.35
2.409
2.521
-22.7
0.95
s02_plain
plain 'a person speaking, X'
0.33
2.183
2.474
-22.3
0.93
s06_physical
physical body cues
0.31
2.291
2.467
-21.9
0.93
s01_minimal
just the emotion word
0.18
2.216
2.492
-22.2
0.94
s15_contrast
negate calm ('not calm, not measured')
0.13
2.295
2.470
-22.6
0.93
s03_dramabox_base
dramabox baseline sentence
0.06
2.097
2.521
-22.4
0.93
s10_actor_coach
method-actor coaching language
-0.20
2.264
2.524
-22.8
0.95
s12_imperative
second-person commands: 'be furious!'
-0.61
2.219
2.451
-23.1
0.91
s09_scenario
ground it in a concrete scenario
-0.61
2.152
2.505
-23.2
0.96
s18_vocal_desc
describe the raw unstable voice
-0.65
2.157
2.386
-22.3
0.96
s17_arousal_words
adrenaline/arousal vocabulary
-0.69
2.184
2.406
-22.6
0.92
s08_parenthetical
stage directions in (parentheses)
-0.80
2.104
2.430
-21.9
0.87
s05_repetition
repeat the emotion word many times
-1.19
2.157
2.387
-23.0
0.91
volume ↔ arousal correlation r = 0.26; volume ↔ target emotion r = 0.09. loudest style: s19_kitchen_sink, quietest: s09_scenario.
3 · Speaker similarity: how to prompt "talk exactly like the reference"
template
spk similarity (mean)
best take
invWER
genu
p1_empty
0.931
0.968
0.86
1.65
p7_neutral_ctrl
0.930
0.974
0.88
1.88
p2_exact
0.929
0.966
0.84
1.77
p4_same_person
0.928
0.961
0.63
1.52
p5_identical
0.927
0.970
0.82
2.10
p8_do_not_change
0.924
0.963
0.91
1.99
p3_clone
0.920
0.954
0.73
1.81
p6_session
0.907
0.959
0.83
1.69
similarity = cosine of Chatterbox VoiceEncoder embeddings vs the reference. best template: p1_empty. per reference voice: chris 0.883 · goblin 0.939 · samantha 0.952
p1_empty (samantha) — best take
spk sim0.968
p1_empty (samantha) — best take
spk sim0.968
p7_neutral_ctrl (samantha) — best take
spk sim0.974
4 · Genuineness: which template sounds most real?
template
genuineness (VoiceCLAP)
authenticity (EI-Plus)
blend
invWER
g1_control
1.956
3.256
1.96
0.94
g2_genuine
1.641
3.288
1.67
0.93
g6_anti_acting
1.551
3.227
2.13
0.97
g8_voice_memo
1.539
3.404
2.84
0.99
g3_hesitations
1.538
3.361
2.58
0.97
g4_documentary
1.490
3.319
2.60
0.99
g7_method
1.396
3.324
2.89
0.93
g5_spontaneous
1.338
3.342
2.92
0.99
two independent judges: the VoiceCLAP genuineness head and the Empathic-Insight Authenticity head (agreement r = -0.03). best template: g1_control.