All four ears share one recipe. A frozen VoiceCLAP encoder turns a clip into a single embedding, and a
tiny MLP head reads it to output one score. We train every head here on the VoiceCLAP-Small (768-d)
embeddings — the genuineness head, the vocal-burst-blend head, and the 57 VoiceNet dimension heads — so that
perception is a single forward pass, cheap enough to run thousands of times inside a search loop. (VoiceCLAP-Large,
3584-d, gives a modestly higher ceiling but is not needed for this speed regime.) The 40-emotion model sits on a
separate Whisper-based encoder.
VoiceNet predictor heads (on VoiceCLAP-Small)
The heads. We take the 57-dimension VoiceNet taxonomy as given. To make it actionable inside a generation
loop, we train one small head per dimension on VoiceCLAP-Small (768-d) embeddings —
Linear(768→H) → GELU → Dropout → Linear — in two flavours: a regression head (a smooth
number, e.g. 3.4) and a classification head (a hard 0–6 bucket with a confidence). 57×2 heads,
each a single forward pass on the frozen encoder.
Labels. Rather than spend the taxonomy’s human benchmark on head-training, we distil the heads from a
strong instruction-following multimodal LLM used as a per-dimension annotator (a Gemini-class model in a strict
non-reasoning configuration: reasoning budget 0, temperature 0), shown the verbatim 0–6 rubric for one
dimension plus the audio. Candidates were first spread across 0–6 by zero-shot VoiceCLAP similarity to each
level’s description, then LLM-rescored: ~196k single-dimension labels per round over EmoLia (~392k across
two rounds). Heads use Huber loss, per-dimension standardization, best-validation checkpoint.
Fidelity. Mean validation MAE ≈ 0.81, Pearson ≈ 0.76 (regression); within-±1 accuracy
0.85 (classification). Fast, imperfect, and — crucially — a usable reward for steering generation.
EmoNet — 40 fine-grained emotions
What it gives the agent. Intensity along the EmoNet taxonomy of 40 fine-grained emotions — well
beyond the usual six — e.g. Amusement, Elation, Affection, Awe, Longing, Contempt, Malevolence/Malice, Distress,
Helplessness, Emotional Numbness, Triumph, Bitterness.
Architecture. A Whisper-based encoder produces a full sequence embedding; one expert MLP head per emotion
maps it to that emotion’s intensity (40 experts), validated against a human-expert emotion benchmark.
Why it matters here. The agent can target feeling directly — reward a grieving voice for Sadness
and Longing, an ork for Anger and Malevolence.
Vocal-Burst Blend — 0–10 (on VoiceCLAP-Small)
What it measures. How naturally a vocal burst — a laugh, giggle, chuckle, sob, gasp, sigh, groan
or scream — blends into the surrounding speech, instead of sounding pasted in. A burst can be fine in
isolation yet ruin a take if it does not belong: wrong emotion, wrong timing, an audible seam between the
spoken words and the laugh. The scale: 0 = disconnected / spliced / robotic / wrong-emotion (also clips with no
burst); 5 = the burst fits but sounds performed — a stagey, acted laugh; 10 = fully organic,
indistinguishable from a spontaneous human reaction that arose in the moment.
Model. A 768→H→1 MLP on VoiceCLAP-Small, Huber loss; labels from a strong multimodal LLM
asked to listen and rate the burst’s integration with a short justification. Validation Pearson ≈ 0.63.
Why it matters here. It keeps the fairy’s giggle, the goblin’s cackle and the grieving voice’s sob
believable — an evolved take that wins on emotion but bolts its laugh on is punished.
Genuineness — 0–6 (on VoiceCLAP-Small)
What it measures. How spontaneous a delivery sounds. If a clip sounds rehearsed, recited, practised,
read off a page, it scores low; if it sounds spontaneous, authentic, arising unplanned from the situation —
with the natural timing, breath, hesitations and micro-imperfections of real speech — it scores high. It is
deliberately not about audio fidelity: a crisp studio read can be very un-genuine, and a rough
spontaneous outburst very genuine. Anchors: 0 = completely rehearsed / script-reading; 3 = conversational
and believable; 6 = indistinguishable from a genuine, unplanned real-life moment.
Model. A 768→50→1 MLP on VoiceCLAP-Small, Huber loss; labels from a strong multimodal LLM
over ~9.7k clips spanning many TTS systems and natural speech. Validation MAE 1.00, Pearson 0.77.
Why it matters here. It is the anti-robotic reward — it stops the search from winning on caricature while
sounding like a machine, and it is an axis our generation work is now actively pushing on.