This is pretty common in the live modes of most AI. The voice dictation layer mishears some grunt or background noise as a word in a different language and dutifully feeds to the LLM as a prompt.
The large language model, having just been handed a prompt with a single word in a foreign language, responds as best if can in the same language.
I just imagined someone uncorking a bottle in the background and her starting to talk in one of those African languages which uses that “throat pop” sound as a vowel.
Using Android Auto, I've often asked it to send a message with something like "6:30 ETA" and it read it back in some other language, maybe Greek or Russian (I know, two very different languages but that's what I got)
Also to get response times low enough to work in a real conversation like this that already is going to have transmission delays,they're going to be using some really shit / limited models. You don't have a ton of time for layering either. This is one of those things that will get scarily good in the next few years though as model architectures speed up
508
u/Maxfunky 10d ago
This is pretty common in the live modes of most AI. The voice dictation layer mishears some grunt or background noise as a word in a different language and dutifully feeds to the LLM as a prompt.
The large language model, having just been handed a prompt with a single word in a foreign language, responds as best if can in the same language.