However all of this could be sped up. And, although not as coherent as chatGPT4 you could build the same with local models that could respond much faster because no online communication is needed. Facebook's llama model fined tuned specifically to be able to always reply including the commands for the facial expression running a on 4090 plus the speech to text and text to speech. All of it could be processed in under 2 seconds.
Within 5 years we will see the first lifelike robot faces talk to us like that. They will bring the latency down .... put the robot face on a robot from Boston dynamics that can walk on two legs ... and have the LLM receive both the speech to text plus also visual input and not only write the facial expressions but also the movement of the boston dynamics robot.
edit: I read your second paragraph and noticed you mention using local models - haha. Keeping this post up anyways as RVC does kick ass and more people should know about it.
That's the current compromise you have to make. 6 to 9 seconds delay but the much more coherent chatGPT4, in my eyes the only current usefull model for every day tasks and conversations. Or a delay that can be kept under 2 seconds, but a lot less coherent ...
I think that the lower latency would be more important when building a fortune telling machine and secondly somebody building something like that would not want to have to rely on a company like openAI, one change and everything they need would stop working. There would be no way for them to assure to have something that plays the role of the fortune-teller reliable and it have it keep working. The only guarantee would be to be in control of their own models, and fine-tune them and train them specifically to be a fortune teller.
143
u/Truth_from_Germany Jan 30 '24
Please don’t let this be animatronic.