Situation-driven performance generation. The actor is physically and socially embedded in a specific situation from the Situation Taxonomy — their body posture, physical activity, social context, environment, health, or pain state naturally affects how they speak and perform.
The Situation Taxonomy extends the core VoiceNet taxonomy (57 voice attribute dimensions) with 11 situation-dependent dimensions containing 289 total situations that describe how a speaker’s physical state and environment alter their vocal output.
Based on the Extended VoiceNet Taxonomy from Schuhmann et al., 2025.
| Dimension | Code | Levels | Description |
|---|---|---|---|
| Body Posture & Gravitational Alignment | POSE | 32 | How gravity, skeletal alignment, and thoracic compression alter the vocal tract and diaphragm |
| Physical Activity & Dynamic Load | ACTV | 69 | How metabolic demand, physical movement, and interaction with objects compete with the speech signal |
| Speaking Target & Projection | TRGT | 25 | Who or what the speaker is addressing — determines throw of voice, register, and feedback loop |
| Social Situation & Context | SOCT | 56 | Social dynamics, power relations, and contextual norms that shape vocal behavior |
| Environment & Acoustic Space | ENVI | 22 | Physical space and its acoustic properties affecting the voice |
| Health & Physiological Condition | HLTH | 18 | Medical conditions, illnesses, and physiological states that alter vocal production |
| Face & Head Obstructions | GEAR | 14 | Physical obstructions (masks, helmets, food) that filter or modify the voice |
| Climate & Atmospheric Conditions | CLMT | 10 | Temperature, humidity, and weather effects on vocal tract and breathing |
| Substance & Chemical Influence | SUBST | 12 | Chemical substances affecting vocal control, coordination, and quality |
| Fatigue, Sleep & Energy State | FATG | 19 | Sleep deprivation, exhaustion, and energy levels impacting vocal effort |
| Pain & Physical Distress | PAIN | 12 | Active pain states and their effect on breathing, tension, and vocal production |
Each situation level includes:
5,749 acting challenges generated from the situation taxonomy (289 situations x 20 variants each, with a small number of API failures). Each variant samples:
The challenges place the actor genuinely IN the situation — the physical/social context naturally affects the voice, breathing, and emotional delivery.
Same actor, same situation, two different emotional moments separated by “CUT TO:”. The speaker’s physical situation stays identical — they’re still lying down, still in the boardroom, still freezing — but the emotional delivery shifts dramatically.
Same as standalone SIT, plus:
[Speaker description — age, gender, timbre — applies to BOTH scenes]
[Situation context: the actor is IN this physical/social situation]
[Scene 1: emotional moment with situation-appropriate vocal effects]
CUT TO:
[Same situation, but dramatic emotional shift]
[Scene 2: contrasting emotional moment, same physical constraints]
dramabox_sit_situation.json contains 5,749 pre-generated SIT-CC DramaBox prompts — one per situation-inspired acting challenge, distributed across four languages:
| Language | Count |
|---|---|
| English | 1,438 |
| French | 1,437 |
| Spanish | 1,437 |
| German | 1,437 |
Directions (in parentheses) and speaker descriptions are in English; spoken dialogue (in double quotes) is in the target language.