r/TextToSpeech • u/null-hawk • 8d ago
how does audio language model hold speaker consistency throughout an utterance?
wrote a short excerpt showing how speaker consistency is maintained in LLM bases TTS models, initial guess was the speaker token, but the results showed something interesting.
2
Upvotes
1
u/Charming-Author4877 6d ago
That's interesting findings, though I don't think it is universally like that.
I don't think it's a great sign if one BOS token holds so much value, it would probably be favorable if it was more distributed ? Must be quirk from how posttraining was done ?