For each of the 98 voice-clone prompts (same reference speaker, same performance instruction, same paraphrased text as the Phase A demo), we sampled the fine-tuned MOSS-TTS 8B model 10 times at temperature 1.0 — ten different takes of the same line.
Each take is then scored for “genuineness” (how real / natural-human it sounds, 0–6) by two independently-trained MLP probes on two different VoiceCLAP audio embeddings: VoiceCLAP-commercial (768-d) and VoiceCLAP-large-v2 (3584-d).
Within every group we rank the 10 takes descending by predicted genuineness — separately for each probe — and show the Top-3 (green) and Bottom-2 (amber), each with its predicted score. The group's original reference clip and the spoken paraphrase are shown for context.