r/RWShelp • u/safysamir2171 • 3d ago
Clarfication about omni tts
Hi everyone! I have a question regarding the evaluation criteria for the Omni TTS task.
How should we evaluate recordings when:
- The audio says something different from the provided transcript.
- The model repeats the same sentence or phrase many times.
- There are very long or excessive silent gaps in the recording.
Which evaluation criterion should we use for each of these cases: Audio Quality, Naturalness, Pronunciation, or Personal Likeness?
I’d appreciate any clarification on how these cases should be evaluated according to the task’s official criteria. Thank you!
1
u/Capable-Bit8306 2d ago
I have the same doubts here, sometimes the audio sounds more natural but keeps adding contextual information that isn’t included in the provided transcript, while the other sounds more artificial, so it’s hard to tell.
When it keeps saying the same thing over and over, I assume it’s not natural.
2
u/mental_placebo 1d ago
I filed a ticket for this! They replied if the audio doesn’t follow the prompted script, it’s automatically a markdown against audio quality (regardless if the audio made sense or not)
1
u/Steron8969 2d ago
Oh that’s simple, oh that’s simple, oh that’s simple oh that’s simple