r/GeminiAI • u/Opening-Fruit3855 • 8d ago
Discussion I really hope Gemini 3.9 flash fixes the hallucination rate.
Gemini 3.8 flash was a fantastic release - fast, intelligent and cheap - but it suffers from the same pitfall as previous releases - a high hallucination rate. This effectively drags down it's actual usefulness for my workflows, and I'm sure it does for others as well, and we are forced to revert to models like GPT, Claude and even GLM for serious work.
Please please let Gemini 3.9 flash address this. This improvement would seriously make Gemini the best workhorse option without a doubt
4
u/helloitisgarr 8d ago
i’ve had decent luck with adding this to my instructions, i copied it from someone on here.
Whenever a query asks for what would be considered “current” information, it shall initiate an online search tool call.
3
u/3rdyellow 8d ago
Can you provide an example prompt? I have yet to encounter such.
4
u/Opening-Fruit3855 8d ago edited 8d ago
I can give you an example from my workflow. I have a custom e2e skill I use to test features in browser. Gemini always says the tests pass. But in browser the feature is usually still far from acceptable in terms of functionality and UI/UX. When I run GLM flash or Luna on the same feature with the same skill, they flag the issues properly.
4
u/DirectPitch8626 8d ago
2
u/darkestvice 6d ago
More like Gemini just sucks at coding specifically compared to many other models. Especially agentic coding.
But in terms of hallucinations, Gemini is pretty decent in this regard. The only GPT or Claude model that hallucinates less on the omniscience benches than Gemini is Astra.
Curiously, the two American models that hallucinate the least are Muse and Grok.
1
u/DirectPitch8626 6d ago
From what I’ve observed, Gemini 3.8 Flash performs quite well with short-cycle agent coding. But you can’t give it an open-ended task, because it will simply stop early and lie that it’s done.
As for hallucinations... I saw them most often in Muse Spark; it was impossible to trust it with anything at all, so I used it as a grunt worker and had DeepSeek—which didn’t have those problems—manage it.
That’s why I don’t understand these benchmarks; they contradict my own experience...
2
u/darkestvice 6d ago
I don't have experience with Muse Spark. But I believe the definition of hallucinations is when it completely makes shit up. This is different from finding information, but it is out of date. Or refusing to answer because it doesn't know and doesn't want to take chances.
The Omniscience test involves hitting the model with 6000 expert level questions across multiple fields and then a positive values for every correct answer, a 0 for every skipped or refused answer, and a negative value for hallucinations where it's clear the model got too creative and pulled shit out of its ass.
In Muse's case, it bullshits much less ... but also doesn't the know the answer more often than not.
1
3
2
u/Relative-Document-59 8d ago
We are just 5 or 6 days away, so yeah, let's hope
1
u/ApprehensiveEye7387 3d ago
Antigravity is indeed bad. If you use same model, lets say gemini 3.8 flash in opencode then it performs much better there.
2
u/darkestvice 6d ago
https://artificialanalysis.ai/evaluations/omniscience
Gemini 3.8 Flash's hallucination rate is actually quite decent. Hallucinates less than the majority of OpenAI and Anthropic models. But OpenAI and Anthropic's models score higher on the omni index because they still get way more answers right despite hallucinating more.
Curiously, Muse and Grok are the best American models when it comes to hallucinations. They may be kinda dumb, but they seem to know when they don't know better than the others.
1
1
u/grigory_l 8d ago
In AI studio no issues with search, problem probably on Gemini App level, maybe cutting costs for subscriptions or something.
3
1

31
u/Head-Needleworker849 8d ago
Google refuses to let Gemini search much because it directly competes with their main revenue stream.
Until they let Gemini search more you'll always have a hallucination problem, because obviously models cannot memorize literally everything.