A measured, audio-backed usage guide for driving laion/moss-tts-local-transformer-4.55b-voice-acting-v2 β the 4.55B expressive voice-acting TTS model. For any target (an emotion, a voice dimension, a delivery style, a whole scene) it tells you the best prompt, which LoRA at which merge %, the expected reward, and the measured side-effects. Every page here is live; the conditioning manuals additionally have a Markdown mirror (linked MD) so agents can read them without HTML.
burst@1.0 + emotion@0.5 is the best cell of 2,304 generations, while a half-merged burst adapter paired with a half-merged emotion adapter is the worst (they compete). Cross-lingual reference audio (Japanese anime β DE/EN) does not transfer the voice β similarity 0.188 over 1,280 generations against a 0.105 unrelated-speaker floor β but pitch still tracks the reference (r = 0.591, compressed toward the model's own register), which makes it a cheap source of new character voices rather than a cloning method. And always evaluate against a real-audio ceiling control: a 13-cell sports-commentator sweep looked like a win until real human commentary scored the same as the un-adapted base. Read the page βvn_ARSH_high@0.5 + vn_BRGT_high@0.5 (and vn_S_NARR_high for storytelling) to keep speech clear; and a cheap API brain (gpt-5.6-luna) matches the local model at cents per task. Example β a favourite audiobook-narrator recipe: vn_S_STRY_high@0.65 + vn_ARSH_high@0.20 + vn_BRGT_low@0.20 β storytelling 0.998, aesthetics 0.59, WER 0.000, ~11 s. Brain used in the comparison: google/gemma-4-31B-it-qat-w4a16-ct (see MODEL_COMPARISON.md).Per emotion: the best prompting strategy **with and without** the emotion LoRA, the best merge Ξ», expected reward, and the top-3 Β± correlations on other emotions / VoiceNet dims / genuineness / blend / quality. Start here to make a voice *feel* something.
Per speech dimension (tempo, warmth, arousal, roughnessβ¦): the dose-response across 25β125% merge, the best dose, the cost to genuineness/blend/quality, and the Β± side-effects. The reference for *how a voice sounds*.
A deterministic rule to pick LoRAs + scales + prompt from a clip's scores **without an LLM call**, so you can reinterpret a huge corpus (e.g. 1M DramaBox clips) β including the LoRA hot-swap / batch-bucketing answer.
Validated, ready-to-use LoRA+prompt recipes distilled from autonomous agent runs β per emotional target (desire, ecstasy, teasing, dominance, anger, painβ¦) and the explicitness LoRAs. Written for a fast warm-start (human or agent). 18+, synthetic.
New (Aug 2026). The measured dose curve for vocal-burst adapters inside a sentence (presence rises 27β71 % from Ξ» 0.25β1.0, but blend falls 5.73β3.97 β Ξ» 0.75 is the knee), how burst and emotion adapters interact (burst@1.0 + emotion@0.5 is the best cell; a half-merged pair is the worst), and the evaluation checklist that turned an apparent sports-commentator win into a null result. 2,304 generations.
Reward-maximised recipes (prompt + Ξ» + sampling) for hard non-speech-heavy deliveries β fear-scream, pain-scream/groan, sobbing, shivering, genuine laughter β evolved with and without the LoRA.
6 cloned voices Γ 40 emotions Γ 4 delivery conditions (intense/moderate Γ free/contained), best-of-32 with a speaker-similarity floor. The flagship study β shows that intensity, not containment, breaks a cloned voice.
Evolutionary search (10 generations, mean-of-8) for the best Ξ» / prompt / sampling per emotion β the source of the per-emotion recipes used elsewhere.
Neutral vs prompt vs LoRA @50/100/150% for all 40 emotions, with a cross-steering heatmap β shows when the LoRA helps and when a good prompt alone wins.
The audio grid behind the edge-case recipes: screaming, crying, laughter, groaning, shivering, maximised with/without LoRA.
Per-dimension controllability + best prompting at three doses β 50% already delivers most of the shift for most dimensions.
57 dimensions Γ high/low Γ 5 merge doses (25β125%), 3 audio samples each β hear the full scaling sweep per dimension side by side.
VoiceNet-embedding clusters of the generated character voices with procedural captions β browse the space of distinct synthetic voices.
120 character voices refined by SIDON self-distillation (generate β filter β restore β warm-start FT) β before/after quality per character.
Single-dimension optimisation, the 4-brain comparison (gpt-5.6-luna vs Gemma-4 31B/12B/MoE, side by side with audio), the supervised edge-case swarm, and the corrected per-dimension runs.
Gemini-judged, same-budget: gpt-5.6-luna (API) scored 9/9/10 β tied with the local Gemma-4 31B (9/10/9), ahead of 12B/MoE. Both independently found the stabilizer adapters.
Measured token/$ figures: an API brain runs a search for cents per task (luna: $0.31 for 4 missions, no GPU), and supervision is ~$0.007/verdict β see the 100/1000-task projections.
The plan for scaling to a 32β64-agent self-improving swarm: two-tier hive-mind knowledge (raw experience vs. approved manual), task registry, supervisor loop, offline consolidation.
The distilled, validated learnings folded into every agent's context: winner recipes, the stabilizer adapters, the (1βWER) reward, and the task-difficulty map (what the TTS model can and cannot do).
The 57 perceptual speech dimensions with 7-point anchor descriptions β what each VoiceNet score means.
The extended dimension set used by the predictor heads.
266 hard expressive-delivery situations, each expressed as VoiceNet/EmoNet/vocal-burst tokens β also mirrored (JSON + docs) in the [voice-taxonomies](https://github.com/LAION-AI/voice-taxonomies) repo.
The scene/situation vocabulary behind the DramaBox reinterpretation prompts.
The 4.55B expressive voice-acting TTS model everything conditions on.
One rank-64 LoRA per emotion (β€25% Got-Talent), the adapters behind the emotion manual.
57 dimensions Γ high/low, the adapters behind the VoiceNet manual and the high/low grid.
One adapter per non-verbal burst (chuckle, gasp, sob, scream, moan, whistleβ¦), each confirmed by a listening audit against real recordings and a no-LoRA baseline. See the merge-dose page above before choosing Ξ».
Best 3 cells of a 13-cell sweep, trained on filtered English generations + real German broadcast commentary. Ships with its own negative result: the base model already matches real human commentary on the judge metric.
120 SIDON-refined from-scratch character voices (public mirror; a genuine-voices public repo also exists).
Model: laion/moss-tts-local-transformer-4.55b-voice-acting-v2 Β· scoring: EmoNet-40 Β· VoiceNet-57 Β· genuineness Β· vocal-burst-blend Β· WER Β· ECAPA. All pages verified live. Human pages here; agent-readable Markdown mirrors are linked MD where available.