What is VoiceNet? VoiceNet listens to a short speech clip and estimates 57 perceptual voice & speech dimensions — human-interpretable qualities such as voice age, arousal, warmth, brightness, tempo, breathiness, speaking style and resonance placement. Each dimension is scored on an ordinal scale (mostly 0–6) whose numbers map to plain-language rubric levels. Every prediction comes from a single VoiceCLAP-commercial embedding per clip, read by small model "heads". Two extra heads are shown: Genuineness (0–6, authentic vs. performed) and Vocal-Burst Blend (0–10, how naturally a laugh / sob / gasp blends into speech).
The caption (new). Each card leads with a procedural caption — the 5 most distinctive dimensions for that clip. We compare every dimension's predicted score to the average of 1000 random Emolia voices, rank by how many standard deviations it is above or below that average (its z-score), and turn the top 5 into short English phrases. The wording reflects the direction (above / below average) and the intensity of the deviation (“somewhat”, “notably”, “very”, “extremely”). It is a fingerprint of what makes this voice stand out — no taxonomy knowledge required.
Regression vs. classification. For every dimension we show two predictions: Regression a smooth continuous score, and Classification a single hard bucket with its plain-language level text. The radar plots the 57 v3 regression values.