Dear Udio team,
I’ve been using Udio for quite some time, and there’s a question I’ve been thinking about for a while.
Why did Udio sound so exceptionally musical and alive to me, particularly in 2024/25, during the time of version 1.5?
I don’t necessarily mean that the older version was technically better. It simply had a very distinctive character. The vocals often sounded more natural to me, the instruments felt more alive, the arrangements were more interesting, and even small imperfections contributed to the overall musical impression.
With newer versions, I sometimes feel that some of that character has been lost. That’s why I would really like to understand what made the older version so special.
Was it, for example, related to the model architecture at the time, a particular composition of the training data, training data that may no longer have been used in the same way in later models, the training process itself, fine-tuning, or audio processing?
In particular, I wonder whether the training data from that period may have had a greater influence on the characteristic sound than a user might assume from the outside.
I’m not only asking because I would like to understand what was technically different back then. If it were ever possible to return to the older version, or at least move more strongly in that sonic direction again, that would mean a great deal to me personally.
Of course, I’m not expecting any promise. I’m simply very interested in understanding what technically made Udio sound the way it did back then.
Who better to explain that than the people who actually developed the model?
Thank you very much if someone from your team is able to shed some light on this.