model_v2.pt @ 0.50 / merge 0.10 s / min-dur 0.10 s, 30 s windows stitched onto one timeline, named by laion/vocal-burst-detector-v2, no-burst gate, terse tags; emotions and the gender gate from laion/Empathic-Insight-Voice-Plus; captions z-scored against the in-domain baseline_stats.json with reliability weighting. The LLM path rewords the procedural draft keeping burst positions. (The pages above this line predate that switch and show locator v1 output.)
Ready-made dataset: laion/moss-character-voices-top3-captioned.
(cues), inline vocal bursts, [pause] markers).What is this? Each of the 100 speech clips is scored by 99 models — 57 VoiceNet perceptual voice/speech dimensions, a Genuineness score, a Vocal-Burst-Blend score, and 40 EmoNet emotions. The procedural module turns those raw numbers into plain English by describing how far the voice deviates from the average voice.
baseline_stats.json): z = (value − median) / spread, spread = 1.4826·MAD (robust std).The VoiceNet + genuineness + blend predictions are reused from the VoiceNet demo set; the 40 EmoNet emotion scores were computed with BUD-E-Whisper + the EmoNet emotion heads. Open “Show the numbers” under any caption to see the exact dimensions, emotions, z-scores and the genuineness verdict that produced it.
preload="none") — click play to fetch each clip.