Voice-Acting-Pipeline

Path SIT — Situation-Driven Acting

Situation-driven performance generation. The actor is physically and socially embedded in a specific situation from the Situation Taxonomy — their body posture, physical activity, social context, environment, health, or pain state naturally affects how they speak and perform.

Situation Taxonomy

The Situation Taxonomy extends the core VoiceNet taxonomy (57 voice attribute dimensions) with 11 situation-dependent dimensions containing 289 total situations that describe how a speaker’s physical state and environment alter their vocal output.

Based on the Extended VoiceNet Taxonomy from Schuhmann et al., 2025.

Dimension Code Levels Description
Body Posture & Gravitational Alignment POSE 32 How gravity, skeletal alignment, and thoracic compression alter the vocal tract and diaphragm
Physical Activity & Dynamic Load ACTV 69 How metabolic demand, physical movement, and interaction with objects compete with the speech signal
Speaking Target & Projection TRGT 25 Who or what the speaker is addressing — determines throw of voice, register, and feedback loop
Social Situation & Context SOCT 56 Social dynamics, power relations, and contextual norms that shape vocal behavior
Environment & Acoustic Space ENVI 22 Physical space and its acoustic properties affecting the voice
Health & Physiological Condition HLTH 18 Medical conditions, illnesses, and physiological states that alter vocal production
Face & Head Obstructions GEAR 14 Physical obstructions (masks, helmets, food) that filter or modify the voice
Climate & Atmospheric Conditions CLMT 10 Temperature, humidity, and weather effects on vocal tract and breathing
Substance & Chemical Influence SUBST 12 Chemical substances affecting vocal control, coordination, and quality
Fatigue, Sleep & Energy State FATG 19 Sleep deprivation, exhaustion, and energy levels impacting vocal effort
Pain & Physical Distress PAIN 12 Active pain states and their effect on breathing, tension, and vocal production

Each situation level includes:

Acting Challenge Database (Situation-Inspired)

5,749 acting challenges generated from the situation taxonomy (289 situations x 20 variants each, with a small number of API failures). Each variant samples:

The challenges place the actor genuinely IN the situation — the physical/social context naturally affects the voice, breathing, and emotional delivery.

SIT — Standalone Situation

Sampling Strategy

  1. Situation: Random selection from 289 situations across 11 dimensions
  2. Emotions: 1-3 random emotions from EmoNet (40 categories) with random intensity
  3. Speaker Gender: Random from 7 VoiceNet GEND levels
  4. Speaker Age: Random from 6 AGEV levels
  5. Word Count: 40-80 words of spoken dialogue

Key Characteristics


SIT-CC — Character Consistent

Two-Scene Format

Same actor, same situation, two different emotional moments separated by “CUT TO:”. The speaker’s physical situation stays identical — they’re still lying down, still in the boardroom, still freezing — but the emotional delivery shifts dramatically.

Sampling Strategy

Same as standalone SIT, plus:

Scene Structure

[Speaker description — age, gender, timbre — applies to BOTH scenes]

[Situation context: the actor is IN this physical/social situation]

[Scene 1: emotional moment with situation-appropriate vocal effects]

CUT TO:

[Same situation, but dramatic emotional shift]

[Scene 2: contrasting emotional moment, same physical constraints]

Pre-Generated DramaBox Prompts

dramabox_sit_situation.json contains 5,749 pre-generated SIT-CC DramaBox prompts — one per situation-inspired acting challenge, distributed across four languages:

Language Count
English 1,438
French 1,437
Spanish 1,437
German 1,437

Directions (in parentheses) and speaker descriptions are in English; spoken dialogue (in double quotes) is in the target language.

Audio Processing

  1. DramaBox TTS: Raw audio synthesis (produces one continuous audio file)
  2. RE-USE Enhancement: Chunked (15s chunks, 1s overlap) — SIT-CC audio is longer
  3. Best-of-N Scoring: 3 candidates, select best
  4. Audio Splitting: Qwen3-ASR word-level timestamps find the “CUT TO:” boundary, then the audio is split into Scene 1 and Scene 2 with 100ms fades