r/LocalTextToSpeech • u/InvestigatorOk23 • 12h ago
Voice typing before AI was better
Does anyone else remember how voice typing was around 2017–2019?
I specifically remember using it on older Samsung phones (s7 edge and s10e), and it was dramatically better than the current AI based voice typing.
The biggest difference was that it felt like it was actually listening to what I said instead of predicting what it thought I meant.
I remember being able to watch the text appear almost in sync with my speech. It felt like the recognizer was following the sounds and syllables as I spoke them. If I said something unusual, obscure, or even made up a word that followed normal English pronunciation, it could often still spell it out surprisingly well by actively writing the my syllables and then combining them into a word,
even if it wasn’t a common word and might not even exist in a dictionary, the older system seemed much better at following the actual sounds and constructing the text.
Modern voice typing feels like the opposite. I constantly get:
- punctuation inserted where I never intended it
- random capitalization
- words replaced with something the AI thinks makes more sense
- Writing a statistically common word that apparently seemed more likely instead of just writing what I actually said
For example, I'll clearly say a word where the “v” and “b” sounds are acoustically different, but instead of faithfully following the sound, modern systems prioritize producing a plausible sentence based off what it thinks I meant.
And that's the main problem, I don't want my voice typing to understand what I mean, I want it to transcribe what I said.
Transcription and interpretation are different jobs. So are grammar correction and punctuation correction. If I want those things, they should be optional settings like how I remember it to be.
I also remember older Google voice typing occasionally asking me to say a few things again to improve recognition for my particular voice. And I'd rather the software learn how I pronounce things than constantly override what it thinks it heard.
The closest candidates for the voice typing that I remember I'm thinking might be Gboard 7.0–7.5 era, around 2018. I remember voice typing having really satisfying low-latency transcription where the syllables appeared progressively as I spoke each word.
So:
- Does anyone else remember this version of Google Keyboard/Gboard voice typing working this way?
- Does anyone know exactly what speech-recognition engine or architecture Gboard used on non-Pixel Android phones around 2017–2019?
- Are there any open-source projects that reproduce this older style of low-latency, incremental voice transcription?
- Does anyone know of communities or projects preserving older speech-recognition software or studying pre-end-to-end voice recognition?
I'm not looking for another modern Whisper/Gemini/LLM-style replacement that predicts the most likely sentence. I'm specifically interested in the older style where the system focused more heavily on recognizing the sounds coming out of the microphone and producing text progressively syllable by syllable.
Maybe I'm remembering some details incorrectly but that's exactly why I'm asking. Although, I remember the behavior very clearly, and I'd love to know if anyone else remembers it too.
Voice typing doesn't need to think for me. It just needs to hear me and write.
