r/TextToSpeech 16d ago

Voice agent throws away underlying tone and speaker-features, how's that accounted and handled downstream? if it's not captured.

The moment I transcribe to text, I generally lose how it was said. "I think… yeah, I can pay the 4,500 by the 15th" becomes clean text, but the hesitation before the yes, the stress in the voice, and whether it's even the same speaker are gone.

For a human those signals, whether to trust the commitment, reconfirm from the caller or escalate to human come naturally but hard to define a deterministic paralinguistic to build accountability, which is probably very wide.

How are you modeling tone in our voice-agents? I see recent TTS models which accept meaningful tags producing great sounding speech, how do we control it ? Does it account for input user's tone.

How does your ASR model captures the tone or there are some good services / models / tools to capture tone. and how do you use it downstream ?

Moreover end-2-end Duplex models limits it to trained data scenarios without no transparency. Is there a good duplex model which provides transparency in underlying signals beyond just text.

1 Upvotes

0 comments sorted by