For each VoiceNet perceptual dimension, two rank-32 LoRAs on MOSS-TTS-4.55B — one trained on the HIGH end (score ≥5), one on the LOW end (≤1). Each adapter is merged at four intensities (50 / 100 / 150 / 200%) and generates the same neutral sentence three times; the three takes are ordered with the longest playtime first (most likely to contain full speech). The number is that dimension's own score of the take (0–1). No reference audio.