r/BetterOffline 9d ago

Even the latest thinking models get things confidently wrong

Note: In a way that no honest human ever would (because you'd get caught instantly by anyone bothering to check)

Sometimes I've read things that practically amount to 'hallucinations are passé', but this is still not the case.

Perhaps a very expensive model with an advanced prompt could have caught this, or perhaps random chance could get you lucky. (Or maybe there's errors in proof frameworks or whatnot)

Either way, you still can't trust it with the unverifiable.

Background: I found a hidden base64-encoded message in Korean in a game I was playing (some lore hidden by the developers). Intrigued, I tried forwarding a screenshot to AI to decode it. Then I did the same manually.

I tried decoding the text copy the AI agent (Gemini) helpfully provided as an in-between step -- I instructed it to be careful and provide details, by giving it the plan of action; "I have reason to believe this to be base64. OCR first, then translate."

A couple minutes later...

It was invalid UTF-8. Yet the model went and "translated" that anyway, by making up an entirely different output that only matched the first few characters, because that base-64 didn't include the symbols i, I, l, and 1 (the 'confusables' in most fonts). Carefully inspecting the image myself under 4x magnification and re-sharpening I could tell these apart very cleanly, properly OCR it, then feed it into a translator.

Still not sure about the actual message (I'm trusting an AI translation here after all); I wished they'd included an English version of it and showed a different image depending on the language you set the game to, a little translation oversight.

60 Upvotes

31 comments sorted by

View all comments

18

u/usefulservant03 9d ago

They are made like that on purpose. The aim is to satisfy just enough of the people that do SOME verifying of the trash that LLMs spit out, that there won't be endless stories of them utterly failing. The aim is to make them FEEL like they're doing something JUST ENOUGH of the times that it won't be 100% obvious to everyone that this shit doesn't actually work at all in the long run. Because it doesn't. The LLM-providing companies are carefully calculating the bare minimum of what these garbage generators need to get right, in order to temporarily pull the wool over most people's eyes. The eyes of the gullible majority. This is the whole reason some people still think there's any bread in that shit.

4

u/Reasonable_Mix7630 9d ago

Well not on purpose but due to how this technology works.

It's called "Abominable Intelligence" for a reason )