one script, three renders: Gemini, ElevenLabs, and Mato Voice. same fifteen-line two-host exchange about a thermostat, identical speaker turns, all loudness-matched to the same target, no timeline edits between them.
a voice demo without disclosed normalization is trivially riggable. push your own clip a couple dB hotter and most listeners will call it warmer and more present without knowing why they preferred it.
next time you're on a voice vendor's comparison page: what did they tell you they normalized? if the answer is nothing, you're judging the mix, not the model.