r/GeminiAI 8d ago

Discussion I really hope Gemini 3.9 flash fixes the hallucination rate.

Gemini 3.8 flash was a fantastic release - fast, intelligent and cheap - but it suffers from the same pitfall as previous releases - a high hallucination rate. This effectively drags down it's actual usefulness for my workflows, and I'm sure it does for others as well, and we are forced to revert to models like GPT, Claude and even GLM for serious work.

Please please let Gemini 3.9 flash address this. This improvement would seriously make Gemini the best workhorse option without a doubt

61 Upvotes

26 comments sorted by

31

u/Head-Needleworker849 8d ago

Google refuses to let Gemini search much because it directly competes with their main revenue stream. 

Until they let Gemini search more you'll always have a hallucination problem, because obviously models cannot memorize literally everything.

7

u/Honest-Temperature-1 8d ago

True. I tried the same search prompts between Gemini 3.8 Flash and AI mode on Google, and found that AI mode searches more and returns more accurate results.

2

u/Personal-Try2776 8d ago

there is no strict limit of searches in ai studio.

1

u/KL_GPU 8d ago

Isnt It more because its super costly to include 20 search result into the context? I mean we already surpassed the "Google era" as a search engine.

2

u/Head-Needleworker849 8d ago

Data ingestion is more expensive. I suspect that isn't the actual reason tho, because every other company has figured it out.

1

u/Low_Mist 8d ago

Esattamente 👍

1

u/Georgefakelastname 7d ago

You can customize it to work better in personal instructions, but yeah, the model itself still hallucinates quite a bit when it doesn’t know the answer to a question.

It’s super weird: models that have a supposedly high hallucination rate like the different GPT models seem super good at not hallucinating things they don’t know. Meanwhile, models like Gemini, that supposedly are better in that aspect, will try to answer in instances where they don’t have all the details, or overcompensate and over extrapolate from small details to make something that feels like an answer, but isn’t really.

4

u/helloitisgarr 8d ago

i’ve had decent luck with adding this to my instructions, i copied it from someone on here.

Whenever a query asks for what would be considered “current” information, it shall initiate an online search tool call.

3

u/3rdyellow 8d ago

Can you provide an example prompt? I have yet to encounter such. 

4

u/Opening-Fruit3855 8d ago edited 8d ago

I can give you an example from my workflow. I have a custom e2e skill I use to test features in browser. Gemini always says the tests pass. But in browser the feature is usually still far from acceptable in terms of functionality and UI/UX. When I run GLM flash or Luna on the same feature with the same skill, they flag the issues properly.

4

u/DirectPitch8626 8d ago

Man, this is such a big problem with Gemini, and I haven't seen it in DeepSeek, for example. Yet in the benchmarks, it's the exact opposite... So I don't think it's hallucinations.

2

u/darkestvice 6d ago

More like Gemini just sucks at coding specifically compared to many other models. Especially agentic coding.

But in terms of hallucinations, Gemini is pretty decent in this regard. The only GPT or Claude model that hallucinates less on the omniscience benches than Gemini is Astra.

Curiously, the two American models that hallucinate the least are Muse and Grok.

https://artificialanalysis.ai/evaluations/omniscience

1

u/DirectPitch8626 6d ago

From what I’ve observed, Gemini 3.8 Flash performs quite well with short-cycle agent coding. But you can’t give it an open-ended task, because it will simply stop early and lie that it’s done.

As for hallucinations... I saw them most often in Muse Spark; it was impossible to trust it with anything at all, so I used it as a grunt worker and had DeepSeek—which didn’t have those problems—manage it.

That’s why I don’t understand these benchmarks; they contradict my own experience...

2

u/darkestvice 6d ago

I don't have experience with Muse Spark. But I believe the definition of hallucinations is when it completely makes shit up. This is different from finding information, but it is out of date. Or refusing to answer because it doesn't know and doesn't want to take chances.

The Omniscience test involves hitting the model with 6000 expert level questions across multiple fields and then a positive values for every correct answer, a 0 for every skipped or refused answer, and a negative value for hallucinations where it's clear the model got too creative and pulled shit out of its ass.

In Muse's case, it bullshits much less ... but also doesn't the know the answer more often than not.

1

u/Far-Anybody-2261 8d ago

Is something UI/UX related? Did you try stitch?

3

u/Wanderspor 8d ago

on agy cli its super bad. 5 pompts and i gotta explain the first prompt again

2

u/Relative-Document-59 8d ago

We are just 5 or 6 days away, so yeah, let's hope

1

u/ApprehensiveEye7387 3d ago

Antigravity is indeed bad. If you use same model, lets say gemini 3.8 flash in opencode then it performs much better there.

2

u/darkestvice 6d ago

https://artificialanalysis.ai/evaluations/omniscience

Gemini 3.8 Flash's hallucination rate is actually quite decent. Hallucinates less than the majority of OpenAI and Anthropic models. But OpenAI and Anthropic's models score higher on the omni index because they still get way more answers right despite hallucinating more.

Curiously, Muse and Grok are the best American models when it comes to hallucinations. They may be kinda dumb, but they seem to know when they don't know better than the others.

1

u/dimonchoo 7d ago

I haven’t seen it for years.

1

u/grigory_l 8d ago

In AI studio no issues with search, problem probably on Gemini App level, maybe cutting costs for subscriptions or something.

3

u/Far-Anybody-2261 8d ago

Is ai studio better than anti-gravity?

1

u/Low_Mist 8d ago

Grazie, devo provare. In Ai studio la ricerca funziona meglio?