Procedural Voice Captions — Live Demo Grid

Every caption below is written automatically by the LAION-AI/procedural-voice-captions module — no hand-writing, no LLM.
Current demos → 25 clips, current stack (new baseline + reliability weighting) · 14 clips: locator v2 + classifier v2 · character voices: procedural vs LLM (bursts + gates) · gender A/B (VoiceNet vs Empathic-Insight) · burst predictions · single-burst · detect→classify combo · burst-captions demo
Current default: vocalburst-locator model_v2.pt @ 0.50 / merge 0.10 s / min-dur 0.10 s, 30 s windows stitched onto one timeline, named by laion/vocal-burst-detector-v2, no-burst gate, terse tags; emotions and the gender gate from laion/Empathic-Insight-Voice-Plus; captions z-scored against the in-domain baseline_stats.json with reliability weighting. The LLM path rewords the procedural draft keeping burst positions. (The pages above this line predate that switch and show locator v1 output.) Ready-made dataset: laion/moss-character-voices-top3-captioned.
🧠 MOSS-Audio-Thinking experiment → — feed these procedural captions + Parakeet timestamps + the LAION taxonomies to the MOSS-Audio-Thinking models (4B & 8B), which listen to the audio and synthesize the final voice-acting format (general caption + per-sentence (cues), inline vocal bursts, [pause] markers).

What is this? Each of the 100 speech clips is scored by 99 models — 57 VoiceNet perceptual voice/speech dimensions, a Genuineness score, a Vocal-Burst-Blend score, and 40 EmoNet emotions. The procedural module turns those raw numbers into plain English by describing how far the voice deviates from the average voice.

1
z-score every dimension against a 99-dim baseline (baseline_stats.json): z = (value − median) / spread, spread = 1.4826·MAD (robust std).
2
Keep the top-5 VoiceNet dimensions and top-3 EmoNet emotions by |z| — the traits where this voice stands out most.
3
Age, Gender, Register and Tempo are always described, even when close to average, so every caption anchors the basic voice profile.
4
A genuineness gate decides how strongly emotion is worded: authentic delivery → full intensity; performed delivery → intensity is capped or emotion is dropped. Synonym clusters rotate wording for variety.
5
Each phrase carries a direction (above/below baseline) and an intensity from |z| (Extremely / Very / Notably / Somewhat).
6
The same phrase content is then arranged by one of 10 surface templates (identity-first, emotion-first, telegraphic, two-sentence, "sounds like …", bulleted, varied-connectors, quality-led, minimal-identity, default), with optional deterministic dim/emotion shuffling — so captions don't all read identically (which would invite overfitting when they are reused as fine-tuning targets). The template badge on each card names the form used; the section just below shows the same clip under several templates. Selection is deterministic (seeded per clip), so a given clip always renders the same way.

The VoiceNet + genuineness + blend predictions are reused from the VoiceNet demo set; the 40 EmoNet emotion scores were computed with BUD-E-Whisper + the EmoNet emotion heads. Open “Show the numbers” under any caption to see the exact dimensions, emotions, z-scores and the genuineness verdict that produced it.

Audio is lazy-loaded (preload="none") — click play to fetch each clip.
English 30 · German 25 · French 12 · Chinese 15 · Korean 10 · Japanese 8
↑ Top