Listening along the 57-dimension VoiceNet taxonomy — this is what the agent hears.
Below are clips from EmoLia, each scored live by the four ears. This is the agent’s perception: for
any voice it can name the loudest of the 57 dimensions and read its genuineness and blend. We deliberately include
high- and low-genuineness cases, natural vs. bolted-on vocal bursts, and clips that are striking on a single
dimension (whispered, very bright, rough, high-arousal, nasal…) so you can hear what each number means. It is
what makes the generation loop possible — you cannot optimise what you cannot measure.
Anonymized companion for double-blind review. All voices are synthetic, generated by an open-weight text-to-speech model with no reference recordings; all scores are produced by the perceptual models described here. No author, institution, or repository identifiers are included by design.