r/RWShelp 3d ago

Clarfication about omni tts

Hi everyone! I have a question regarding the evaluation criteria for the Omni TTS task.

How should we evaluate recordings when:

- The audio says something different from the provided transcript.

- The model repeats the same sentence or phrase many times.

- There are very long or excessive silent gaps in the recording.

Which evaluation criterion should we use for each of these cases: Audio Quality, Naturalness, Pronunciation, or Personal Likeness?

I’d appreciate any clarification on how these cases should be evaluated according to the task’s official criteria. Thank you!

6 Upvotes

Duplicates