MOSS Voice-Acting β€” Manual & Studies

A measured, audio-backed usage guide for driving laion/moss-tts-local-transformer-4.55b-voice-acting-v2 β€” the 4.55B expressive voice-acting TTS model. For any target (an emotion, a voice dimension, a delivery style, a whole scene) it tells you the best prompt, which LoRA at which merge %, the expected reward, and the measured side-effects. Every page here is live; the conditioning manuals additionally have a Markdown mirror (linked MD) so agents can read them without HTML.

40
emotions
57
voice dimensions
114
VoiceNet LoRAs
120
character voices
64
vocal-burst LoRAs
4
LLM brains compared
πŸ’₯ Newest (Aug 2026) β€” bursts, merge doses & evaluation: for a vocal burst inside a sentence, Ξ» = 0.75 is the knee β€” 91 % of full-merge burst presence at a much better blend, and it is where genuineness peaks; burst@1.0 + emotion@0.5 is the best cell of 2,304 generations, while a half-merged burst adapter paired with a half-merged emotion adapter is the worst (they compete). Cross-lingual reference audio (Japanese anime β†’ DE/EN) does not transfer the voice β€” similarity 0.188 over 1,280 generations against a 0.105 unrelated-speaker floor β€” but pitch still tracks the reference (r = 0.591, compressed toward the model's own register), which makes it a cheap source of new character voices rather than a cloning method. And always evaluate against a real-audio ceiling control: a 13-cell sports-commentator sweep looked like a win until real human commentary scored the same as the un-adapted base. Read the page β†’
πŸ§ͺ What's new (Aug 2026): an autonomous search agent now discovers recipes on its own. Key validated learnings, folded into the manuals: the reward is multiplied by (1βˆ’WER) to stop over-acted, unintelligible takes; at high energy, add the stabilizer adapters vn_ARSH_high@0.5 + vn_BRGT_high@0.5 (and vn_S_NARR_high for storytelling) to keep speech clear; and a cheap API brain (gpt-5.6-luna) matches the local model at cents per task. Example β€” a favourite audiobook-narrator recipe: vn_S_STRY_high@0.65 + vn_ARSH_high@0.20 + vn_BRGT_low@0.20 β†’ storytelling 0.998, aesthetics 0.59, WER 0.000, ~11 s. Brain used in the comparison: google/gemma-4-31B-it-qat-w4a16-ct (see MODEL_COMPARISON.md).

πŸš€ Start here β€” conditioning manuals

How to drive the model: for a target (emotion, voice dimension, delivery), what prompt + which LoRA at which merge % β€” with measured side-effects. Each has an agent-readable Markdown mirror.

🎭 Emotion conditioning manual (40 emotions)

Per emotion: the best prompting strategy **with and without** the emotion LoRA, the best merge Ξ», expected reward, and the top-3 Β± correlations on other emotions / VoiceNet dims / genuineness / blend / quality. Start here to make a voice *feel* something.

🎚️ VoiceNet dimension manual (57 dims Γ— high/low)

Per speech dimension (tempo, warmth, arousal, roughness…): the dose-response across 25β†’125% merge, the best dose, the cost to genuineness/blend/quality, and the Β± side-effects. The reference for *how a voice sounds*.

πŸ” Reinterpretation guide (procedural, at scale)

A deterministic rule to pick LoRAs + scales + prompt from a clip's scores **without an LLM call**, so you can reinterpret a huge corpus (e.g. 1M DramaBox clips) β€” including the LoRA hot-swap / batch-bucketing answer.

πŸ”₯ Expressive & NSFW recipe book (agent warm-start)

Validated, ready-to-use LoRA+prompt recipes distilled from autonomous agent runs β€” per emotional target (desire, ecstasy, teasing, dominance, anger, pain…) and the explicitness LoRAs. Written for a fast warm-start (human or agent). 18+, synthetic.

πŸ’₯ Vocal bursts, merge doses & how to evaluate a LoRA

New (Aug 2026). The measured dose curve for vocal-burst adapters inside a sentence (presence rises 27β†’71 % from Ξ» 0.25β†’1.0, but blend falls 5.73β†’3.97 β€” Ξ» 0.75 is the knee), how burst and emotion adapters interact (burst@1.0 + emotion@0.5 is the best cell; a half-merged pair is the worst), and the evaluation checklist that turned an apparent sports-commentator win into a null result. 2,304 generations.

πŸ”Š Edge-case delivery (screams, crying, laughter…)

Reward-maximised recipes (prompt + Ξ» + sampling) for hard non-speech-heavy deliveries β€” fear-scream, pain-scream/groan, sobbing, shivering, genuine laughter β€” evolved with and without the LoRA.

open β†—

πŸ”¬ Interactive studies & audio grids

Listen to the effects. Every clip is labelled with its recipe and scores.

⭐ Emotion Γ— Voice conditions

6 cloned voices Γ— 40 emotions Γ— 4 delivery conditions (intense/moderate Γ— free/contained), best-of-32 with a speaker-similarity floor. The flagship study β€” shows that intensity, not containment, breaks a cloned voice.

open β†—

🧬 40-emotion evolution

Evolutionary search (10 generations, mean-of-8) for the best Ξ» / prompt / sampling per emotion β€” the source of the per-emotion recipes used elsewhere.

open β†—

πŸŽ›οΈ Prompt Γ— LoRA strength study

Neutral vs prompt vs LoRA @50/100/150% for all 40 emotions, with a cross-steering heatmap β€” shows when the LoRA helps and when a good prompt alone wins.

open β†—

πŸ’₯ Edge-case evolution

The audio grid behind the edge-case recipes: screaming, crying, laughter, groaning, shivering, maximised with/without LoRA.

open β†—

🎚️ VoiceNet LoRA evolution (dose 0/50/100%)

Per-dimension controllability + best prompting at three doses β€” 50% already delivers most of the shift for most dimensions.

open β†—

↕️ VoiceNet high/low scaling grid

57 dimensions Γ— high/low Γ— 5 merge doses (25β†’125%), 3 audio samples each β€” hear the full scaling sweep per dimension side by side.

open β†—

πŸ—£οΈ Character-voice clusters

VoiceNet-embedding clusters of the generated character voices with procedural captions β€” browse the space of distinct synthetic voices.

open β†—

♻️ Character LoRA refinement grids

120 character voices refined by SIDON self-distillation (generate β†’ filter β†’ restore β†’ warm-start FT) β€” before/after quality per character.

open β†—

πŸ€– Autonomous search agent

An agent (local or API LLM brain) that *searches* for merge + prompt recipes on its own, guided by the manuals above and scored by the full perceptual stack.

πŸ“– Taxonomies (definitions)

The vocabularies everything is scored against.

πŸ€— Models & adapters (Hugging Face)

The checkpoints and LoRA collections referenced throughout.

Model: laion/moss-tts-local-transformer-4.55b-voice-acting-v2 Β· scoring: EmoNet-40 Β· VoiceNet-57 Β· genuineness Β· vocal-burst-blend Β· WER Β· ECAPA. All pages verified live. Human pages here; agent-readable Markdown mirrors are linked MD where available.