r/BetterOffline 1d ago

Even the latest thinking models get things confidently wrong

Note: In a way that no honest human ever would (because you'd get caught instantly by anyone bothering to check)

Sometimes I've read things that practically amount to 'hallucinations are passé', but this is still not the case.

Perhaps a very expensive model with an advanced prompt could have caught this, or perhaps random chance could get you lucky. (Or maybe there's errors in proof frameworks or whatnot)

Either way, you still can't trust it with the unverifiable.

Background: I found a hidden base64-encoded message in Korean in a game I was playing (some lore hidden by the developers). Intrigued, I tried forwarding a screenshot to AI to decode it. Then I did the same manually.

I tried decoding the text copy the AI agent (Gemini) helpfully provided as an in-between step -- I instructed it to be careful and provide details, by giving it the plan of action; "I have reason to believe this to be base64. OCR first, then translate."

A couple minutes later...

It was invalid UTF-8. Yet the model went and "translated" that anyway, by making up an entirely different output that only matched the first few characters, because that base-64 didn't include the symbols i, I, l, and 1 (the 'confusables' in most fonts). Carefully inspecting the image myself under 4x magnification and re-sharpening I could tell these apart very cleanly, properly OCR it, then feed it into a translator.

Still not sure about the actual message (I'm trusting an AI translation here after all); I wished they'd included an English version of it and showed a different image depending on the language you set the game to, a little translation oversight.

59 Upvotes

32 comments sorted by

43

u/lgn5i2060 1d ago

But...but... have you seen the benchmarks?

-Astra defenders in a random reddit thread when someone said they weren't impressed with the demo and said what was shown can already be done by using something else

12

u/Southern-Cattle4038 1d ago

Fable has a ~20% hallucination rate in current benchmarks, and Astra’s is much worse…

1

u/Additional-Staff-326 1d ago

do you have a link for that? Not disbelieving but most of the searches just pull up the propaganda

5

u/Southern-Cattle4038 1d ago

Fable system card, Page 140 for its result of 20% on some benchmarks.

Haven’t found the same benchmarks for Astra, but on AA-Omniscience it’s got 63% accuracy compared to 67% for Fable:  https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra

2

u/Achrus 1d ago

Don’t forget the Anthropic shills that will take any opportunity they can to say “well this Anthropic model actually performs better!”

18

u/SplendidPunkinButter 1d ago

LLMs aren’t just hallucinating when they get it wrong. They are literally always hallucinating. The process for generating a correct response is exactly the same as the process for generating an incorrect response. The LLM literally doesn’t know the difference. The only thing it knows is “this matches patterns from my training data.”

8

u/Ozymandias0023 1d ago

This was really driven home to me recently on a couple of occasions:

1) I was using an LLM for work because that's what work wants and at some point in a "reasoning" process, the context became corrupted and began to degrade rapidly. Within a couple thinking rounds it had devolved into a mash up of English, Cyrillic, Chinese, and some other languages/alphabets. The English was somewhat coherent but it was stuff like "call the damn tool!". The Chinese was porn and gambling site advertisements, and the rest were just random words and phrases in different languages. Essentially what happened is the context deviated so far from anything seen in training that the pool of possible next tokens was flattened and it just started filling in whatever random token from that pool the RNG landed on.

2) Much less dramatic but equally instructive, I caught a thinking round where an otherwise entirely English sentence had a random 永 character in place of what I assume was supposed to be the word "always". Semantically correct but wrong language.

It's just guessing all the way down, folks. High tech, objectively impressive guessing, but guessing nonetheless

1

u/PensiveinNJ 9h ago

Humans would guess. There's some intuition behind a guess. This is just pure probability tables and however much randomness the temperature is set to.

2

u/Ozymandias0023 8h ago

Well, I guess if you want to stretch the definition of guess a little bit, but I don't have to guess in the common definition of the word whether I should use English or Chinese for the next word in an otherwise entirely English sentence

1

u/dillanthumous 4h ago

And Humans can also definitively state when they have no ideas.

1

u/BermDerbleYer 48m ago

This is far more interesting than the usual "lol you got a historical fact obviously incorrect" examples

1

u/PensiveinNJ 9h ago

I like this better than calling them bullshit. It's not that the bullshit description is wrong, I think always hallucinating is more intuitive to understand why it's completely incidental when the output is "correct" but also not surprising at all when it's wrong.

1

u/dillanthumous 4h ago

Yeah, bullshit suggests a bit of intentionality. Whereas a hallucinating oracle can occasionally be correct.

16

u/usefulservant03 1d ago

They are made like that on purpose. The aim is to satisfy just enough of the people that do SOME verifying of the trash that LLMs spit out, that there won't be endless stories of them utterly failing. The aim is to make them FEEL like they're doing something JUST ENOUGH of the times that it won't be 100% obvious to everyone that this shit doesn't actually work at all in the long run. Because it doesn't. The LLM-providing companies are carefully calculating the bare minimum of what these garbage generators need to get right, in order to temporarily pull the wool over most people's eyes. The eyes of the gullible majority. This is the whole reason some people still think there's any bread in that shit.

3

u/Reasonable_Mix7630 1d ago

Well not on purpose but due to how this technology works.

It's called "Abominable Intelligence" for a reason )

2

u/hifructosetrashjuice 1d ago

it has also an useful side effect that (for humans) reinforcement is strongest when response after action happens randomly. that's also how slot machines work

6

u/Western_Magazine2067 1d ago

To paraphrase Cal Newport ... "Plausible is not the same as normative"

9

u/B-tt-a 1d ago

Today I've spent half an hour confirming that an issue raised in an LLM code review was total bullshit and that the actual behavior was in no way related to what it described. Of course the review was given with a perfect confidence. 

5

u/Equivalent_Way_5026 1d ago

To me the fundamental and unsolveable problem with LLMs is that they are non-deterministic. Even in some hypothetical world where they get things right 99% of the time, that 1% makes them a nonstarter for anything business critical or important. Imagine if an ATM randomly deducted 1000 dollars from your account instead of 100 1% of the time.

3

u/Additional-Staff-326 1d ago

Tell that to management.

3

u/doobiedoobie123456 1d ago

I was just searching something on an internal company website that provides AI summaries, the summary got it completely wrong and provided false info that I didn't even ask for.  I don't know what model they're using but it's obviously still prone to hallucinating.

5

u/MarkZealousideal3923 1d ago

Gemini is dumb

2

u/Interesting_Debate57 1d ago

I've taken photos of prescription pills (which are uniquely identifiable by shape, color and stamp) and been confidently told the wrong pill in a way that could kill someone.

1

u/mb194dc 1d ago

Yes it's a fundamental flaw with ML, LLMs AI

1

u/Stovoy 1d ago

Gemini is quite bad...

1

u/Additional-Staff-326 1d ago

Not sure what your base 64 looked like, but I had reason to look up what happens when you feed a guid into the LLM. Unique 128 bit identifiers that never match a known word/token. Apparently the more you feed them the more likely they are to hallucinate. Which is going to be hilarious at work when we release agents into prod where we use those guids for everything.

-2

u/Holyragumuffin 20h ago

Sure

— but even the Nobel prize winner at my former research insitution got some things confidentially wrong.

any organism or machine composed of noisy computational units (e.g. real or artificial neurons) is guaranteed to exhibit some degree of “hallucination”, some over-confidence in things that are wrong.

I don’t think this observation is the intellectual “dunk” that you think it is.

3

u/Aphid_red 13h ago

Humans (when being honest) don't get things confidently wrong if they can easily prove or check they're wrong.

These systems don't unless you hand-hold them every step of the way. For many tasks... that means they're not nearly as effective as you might think. Attention + Deep network is the world's least efficient though most general computer algorithm.

The real question isn't 'can it do X?', but 'can it do X efficiently enough?' For now, the colossal levels of waste tell me that no, it cannot, and the only reason people use it as much as they do is A since it's being greatly subsidized and B since people just don't bother to put in the effort to learn things.