Giving a Coding Agent Ears

VoiceNet as a reusable sense of hearing — for understanding voices and guiding speech generation. A listening companion (anonymized for review).
We show that VoiceNet is not only a benchmark, but a reusable sense of hearing that can be handed to an autonomous coding agent. Turning the taxonomy into a set of fast VoiceCLAP-Small predictor heads — alongside a fine-grained emotion model and two naturalness gauges (vocal-burst blend and genuineness) — lets an agent both understand voices (score any clip along the 57 dimensions) and guide speech generation (run an evolutionary search over natural-language voice instructions, with the predictors as the fitness signal). This page is a listening companion for the rebuttal: it lets you hear what the agent hears, and what it can make a text-to-speech model produce — from tender ASMR to dragons, goblins and grieving voices — with no reference recordings and no copyrighted training voices.

The four ears

All four ears share one recipe. A frozen VoiceCLAP encoder turns a clip into a single embedding, and a tiny MLP head reads it to output one score. We train every head here on the VoiceCLAP-Small (768-d) embeddings — the genuineness head, the vocal-burst-blend head, and the 57 VoiceNet dimension heads — so that perception is a single forward pass, cheap enough to run thousands of times inside a search loop. (VoiceCLAP-Large, 3584-d, gives a modestly higher ceiling but is not needed for this speed regime.) The 40-emotion model sits on a separate Whisper-based encoder.

VoiceNet predictor heads (on VoiceCLAP-Small)
The heads. We take the 57-dimension VoiceNet taxonomy as given. To make it actionable inside a generation loop, we train one small head per dimension on VoiceCLAP-Small (768-d) embeddings — Linear(768→H) → GELU → Dropout → Linear — in two flavours: a regression head (a smooth number, e.g. 3.4) and a classification head (a hard 0–6 bucket with a confidence). 57×2 heads, each a single forward pass on the frozen encoder.

Labels. Rather than spend the taxonomy’s human benchmark on head-training, we distil the heads from a strong instruction-following multimodal LLM used as a per-dimension annotator (a Gemini-class model in a strict non-reasoning configuration: reasoning budget 0, temperature 0), shown the verbatim 0–6 rubric for one dimension plus the audio. Candidates were first spread across 0–6 by zero-shot VoiceCLAP similarity to each level’s description, then LLM-rescored: ~196k single-dimension labels per round over EmoLia (~392k across two rounds). Heads use Huber loss, per-dimension standardization, best-validation checkpoint.

Fidelity. Mean validation MAE ≈ 0.81, Pearson ≈ 0.76 (regression); within-±1 accuracy 0.85 (classification). Fast, imperfect, and — crucially — a usable reward for steering generation.
EmoNet — 40 fine-grained emotions
What it gives the agent. Intensity along the EmoNet taxonomy of 40 fine-grained emotions — well beyond the usual six — e.g. Amusement, Elation, Affection, Awe, Longing, Contempt, Malevolence/Malice, Distress, Helplessness, Emotional Numbness, Triumph, Bitterness.

Architecture. A Whisper-based encoder produces a full sequence embedding; one expert MLP head per emotion maps it to that emotion’s intensity (40 experts), validated against a human-expert emotion benchmark.

Why it matters here. The agent can target feeling directly — reward a grieving voice for Sadness and Longing, an ork for Anger and Malevolence.
Vocal-Burst Blend — 0–10 (on VoiceCLAP-Small)
What it measures. How naturally a vocal burst — a laugh, giggle, chuckle, sob, gasp, sigh, groan or scream — blends into the surrounding speech, instead of sounding pasted in. A burst can be fine in isolation yet ruin a take if it does not belong: wrong emotion, wrong timing, an audible seam between the spoken words and the laugh. The scale: 0 = disconnected / spliced / robotic / wrong-emotion (also clips with no burst); 5 = the burst fits but sounds performed — a stagey, acted laugh; 10 = fully organic, indistinguishable from a spontaneous human reaction that arose in the moment.

Model. A 768→H→1 MLP on VoiceCLAP-Small, Huber loss; labels from a strong multimodal LLM asked to listen and rate the burst’s integration with a short justification. Validation Pearson ≈ 0.63.

Why it matters here. It keeps the fairy’s giggle, the goblin’s cackle and the grieving voice’s sob believable — an evolved take that wins on emotion but bolts its laugh on is punished.
Genuineness — 0–6 (on VoiceCLAP-Small)
What it measures. How spontaneous a delivery sounds. If a clip sounds rehearsed, recited, practised, read off a page, it scores low; if it sounds spontaneous, authentic, arising unplanned from the situation — with the natural timing, breath, hesitations and micro-imperfections of real speech — it scores high. It is deliberately not about audio fidelity: a crisp studio read can be very un-genuine, and a rough spontaneous outburst very genuine. Anchors: 0 = completely rehearsed / script-reading; 3 = conversational and believable; 6 = indistinguishable from a genuine, unplanned real-life moment.

Model. A 768→50→1 MLP on VoiceCLAP-Small, Huber loss; labels from a strong multimodal LLM over ~9.7k clips spanning many TTS systems and natural speech. Validation MAE 1.00, Pearson 0.77.

Why it matters here. It is the anti-robotic reward — it stops the search from winning on caricature while sounding like a machine, and it is an axis our generation work is now actively pushing on.

Two demonstrations

Perception you can measure (Demo 1) and perception you can optimise (Demo 2). Each opens on its own page — lighter and faster to load.

◆  Understanding
Hear clips scored live along the 57 VoiceNet dimensions — each with its full radar profile, genuineness and vocal-burst blend. This is what the agent hears.
Open the radar profiles →
◆  Generation
13 characters designed by evolutionary search over voice instructions — dragons, goblins, fairies, grieving voices — with the four ears as fitness. Listen to the winners.
Open the character voices →

Outlook

Outlook. The generation results here are an early look at a direction we are actively pursuing: speech generation steered entirely by machine listening, on permissively-licensed data only. Because the voices are designed from language and selected by perceptual models — not cloned from copyrighted recordings — an agent can invent voices that no dataset contains: orcs, ghosts, fairies and other fantasy creatures, as well as finely-controlled human affect. We see the perceptual stack shown here as the enabling ingredient: give a capable coding agent a reliable sense of hearing, and it can both understand voice and take authorship of it. A dedicated treatment of the generative side is the subject of forthcoming work.
Anonymized companion for double-blind review. All voices are synthetic, generated by an open-weight text-to-speech model with no reference recordings; all scores are produced by the perceptual models described here. No author, institution, or repository identifiers are included by design.