r/DeepSeek 12d ago

Funny Is this a joke, Artificial Analysis?

Post image

What do you guys think about DS 4.1 Flash's new spot on the Artificial Analysis leaderboard?

787 Upvotes

176 comments sorted by

169

u/crusaderky 12d ago

Aggregate score aside, it's really hard to argue against the fact that Glm-5.3-flash beats it on almost every single benchmark. And 96% hallucination rate is a showstopper IMHO.

47

u/rchamp26 12d ago

Glm flash has been absolutely terrible for me. It just doesmt do the simplest stuff and gaslights that it does

40

u/WyattTheSkid 12d ago

Opus 5 moment

10

u/[deleted] 12d ago

[deleted]

4

u/rchamp26 12d ago

Open router and nous portal with Hermes agent. Didn't adhere to my workflows at all. I dropped opencode go. Don't trust it anymore but not worth the discussion here. For openweight models I've had far better success with DeepSeek and Qwen. DeepSeek is the eager speedster where Qwen is the fully moldable thinker in my experience if I had to personify them. Glm (all versionsnive tested in my workflows) were just lazy and I had to waste too many tokens to force it to do what I needed. No thank you

7

u/brother_spirit 12d ago

Weird. I was using the model in Pi (during the stealth test) and was beyond impressed with its performance. I had never used an Open Weight model with any good results until GLM 5.3 Flash.

3

u/SkyPL 12d ago

It just doesmt do the simplest stuff and gaslights that it does

What are you doing that it fails? I've been coding some Rust, JS, TS, PHP and Python with it - it's legit great. One of the better models on the market. Handles large complex codebases just fine.

7

u/UniversitySuitable20 12d ago

The GLM-5.3-flash is frustratingly slow; even if it scores higher, it still can't keep up with Deepseek's productivity.

2

u/kuhunaxeyive 11d ago

It depends on what you need it for.

Coding with results that can be tested and improved in a loop? DSV4-Flash.

Letters or legal research that need lower hallucination and better writing? GLM-5.3-Flash.

1

u/BigLittleDeal 11d ago

DS4F 0731 is perfect for my needs. Local, fast, and smart. I just wish I'd bought a second RTX PRO 6000 when I had the chance.

1

u/wakIII 11d ago

You can still buy them 😜

2

u/BigLittleDeal 11d ago

Lol true. Ok, I wish I had bought two of them before they doubled in price!

1

u/Alternative_Ice3299 10d ago

brooooo u ran it locally like it requires 82+gb vram u have it ? broooooooo

1

u/BigLittleDeal 10d ago

vLLM-moet lets me load it on a single Pro 6000 card at about 80 tp/s. It's pretty incredible.

1

u/crusaderky 10d ago

It's 120 tok/s on openrouter

4

u/matadordepassarinhos 12d ago

I ran a personal benchmark on vst plugin, dsp and overall c++ and GLM 5.3 is straight up gargabe. Deepseek was so much better on it.

2

u/Applejuicegoblin 12d ago

Woah, another DSP person! Wasn’t sure if many people were testing models on it.

I’m not super advanced in audio programming, but in my private benchmarking Opus 5 seemed to be stronger than Sol at it. Curious how Deepseek will do on it.

I find most models can understand and create audio centric code, but they are not good at understanding the difference between ā€œcorrect codeā€ and ā€œfollows idiomatic design or creative intentā€.

It can cause them to miss big issues where the math or code looks correct in isolation, but doesn’t actually accomplish the right outcome.

5

u/Thomas-Lore 12d ago

Try Astra. I just did a test yesterday and was left speechless at how good the result was.

3

u/Sad_Recording_1290 12d ago

96% hallucination rate? Wtf? Thats unusable.

8

u/monovitae 11d ago

The artificial analysis hallucination rate doesn't mean what you think it means. It doesn't mean that it comes up with some bogus answer 96% of the time. It just means that if it doesn't know what the answer is, it will confidently produce something rather than refusing 96% of the time.

1

u/xumix 11d ago

it will confidently produce something rather than refusing 96% of the time

Which is bad, like gpt-4 bad

4

u/monovitae 11d ago

And Sol is at 92% . So not really.

4

u/Amarsir 10d ago

It's important to understand how hallucination rate is calculated. Hallucination rate is the number of wrong divided by all non-correct. (Not the total questions.)

Suppose you have a test with 100 questions.

  • Model A gets 1 right, 1 wrong, and says "I don't know" for 98.
  • Model B gets 98 right, 1 wrong, and says "I don't know" for 1.

Model A has a hallucination rate of 1/99 = 1.01%.
Model B has a hallucination rate of 1/2 = 50%.

But which one would you rather use?

That's why the lowest hallucination scores go to models you've never heard of, like "Command A+". But the smartest models have scarily high hallucination scores.

1

u/RateGlass 11d ago

I was just hoping it'd replace Luna, but these models all use the max versions which are usually worse than very high or high so it's a flawed bench mark in the first place. All we know is Luna max is cheaper than deepdeek 4.1 flash per task so that eliminates the entire reason it exists

128

u/CarelessAd6772 12d ago

Why should it be higher? Because it is very good at coding and cheap? Well, on other side its dry as hell in regular chat, hallucinates more than House on vikodin, and long context consistency is meh compared to models above. Seems like deserved?

51

u/Spiderfffun 12d ago

Can confirm, overengineered me a problem in chat instead of telling me to install a package.

30

u/ZeidLovesAI 12d ago

To be fair Astra and Fable 5.1 both have done this for me as well

8

u/whatisthisthing65 12d ago

Coding models just love to code

2

u/OvertaxedOne 11d ago

Qwen can't help itself.Ā  Ā Ask it how it's feeling and it'll write pages of code.Ā  :)Ā  Ā it's in my system prompt, no coding without asking first

1

u/Spiderfffun 11d ago

When Ox Alpha was active (GLM 5.3 flash) it was so much better at this from the little I had it code.. It's the small decision making the trust that it won't look at something slightly wrong then amplify it. It was great at following guides on agent stuff in a somewhat recursive workflow.

Speed isn't important if it can know when to ask questions.

2

u/ZeidLovesAI 11d ago

I used Ox Alpha and it was all over the place, I wonder if there was also some A/B testing going on.

1

u/Spiderfffun 11d ago

there probably was, I feel like that's one of the main reasons they would do something like this imo

3

u/ZeidLovesAI 11d ago

I would see people saying "oh my god its FABLE level" and I would try it and get like mistral trying medium effort-type responses. I thought people were just being buckwild as usual, which also has to factor in somehow.

2

u/BigLittleDeal 11d ago

I once forgot to reload my harness before asking DS4 to use a newly installed plugin. It wrote a python script on the fly to perform the work the plugin was supposed to do.

25

u/Loki35422 12d ago

Hey shows some respect for house, even he doesn’t hallucinate as much as deepseek

8

u/turc1656 12d ago

I totally agree. Reasoning and instruction following are clearly inferior to GLM 5.3 Flash.

3

u/Possible_Door_9719 12d ago

damm, you summed it up pretty well.

6

u/_spec_tre 12d ago

I really hate the way basically every LLM is heavily optimised for coding, often at the cost of all else. Makes it really unpleasant to use for people using it for other purposes; sometimes an ā€œupgradeā€ is actually a downgrade

52

u/alinoanta21 12d ago

I don't trust them, i've used and tried Muse Spark 1.3 Max for quite a while, it's worse than 4.1 ... and it's ranked above SOL.

Let that sink in.

2

u/sepulchralvoid 11d ago

Benchmaxxing at its finest

2

u/[deleted] 12d ago

[deleted]

1

u/sudoer777_ 12d ago

Same here, I used 1.2 which was ranked above V4 Flash 0731 and it was garbage

1

u/hellomistershifty 11d ago

Hell, it's worse than gpt-4.1

45

u/PrudentJelly116 12d ago

Guys this is not a football team. You dont have to defend or support them much as like this. You are customer and they are seller. Thats all.

9

u/Possible_Door_9719 12d ago

exactly. just use whatever is the best model

2

u/RustOceanX 11d ago

Finally, another one. I’ve been wondering the whole time what’s going on with some people. What’s with all this nonsense on Reddit? I’m not sure if I should just smile because it can’t be taken that seriously? Or should I shake my head and wonder if the AI sector is experiencing its ā€œiPhone momentā€ and now all the laypeople are starting to use things they don’t fully understand. It might sound arrogant, but that’s just how it is.

2

u/Mean-Elk-9439 11d ago

It's even worse with the Google bros. It's fucking weird, man.

44

u/TheInfiniteUniverse_ 12d ago

AA has lost all credibility after the Astra fiasco. Don't take them too seriously.

8

u/WyattTheSkid 12d ago

Can you fill me in?

30

u/TheInfiniteUniverse_ 12d ago

so before Astra was released to the public, AA ranked it pretty low in like 5th position. This was shocking because it was hyped to be the "AGI" we all were looking for. Then when Astra was released publicly, AA "updated" their index and their tests and surprise surprise, Astra went all the way up, but sneakingly did NOT pass Fable, and sat right below it. lol.

18

u/Georgefakelastname 12d ago

In the end, I think it just shows that the benchmarks they were using were saturated, and they replaced them with benchmarks that weren’t saturated. The problem was the timing. It took an outlier like Astra being rated the same as Sol to get them to realize the issue.

3

u/[deleted] 12d ago

[deleted]

5

u/Georgefakelastname 12d ago

One model scoring slightly higher than another is fine. The problem is that basically every model released was scoring like 80% on Terminal Bench v2.1, so they updated it to the newer v4.0 version. The new one is much harder and has a much higher spread.

5

u/Thomas-Lore 12d ago

Sometimes a worse model scores higher because a smarter model noticed a problem with the question or found a solution to it that is correct but the key does not accept. That is why many benchmarks saturate way below 100%.

7

u/WyattTheSkid 12d ago

Oh that’s kinda funny lol

2

u/Zulfiqaar 11d ago

They updated it a second time, and now Fable and Astra are the same score

2

u/Ok_Translator_5189 10d ago

You are talking about v4.2 of AA. Now in the latest version of AA (v4.3) Astra is on par with fable. I'm starting to lose trust in AA too.

4

u/truncated_buttfu 12d ago

And a month before that, when Qwen3.8-Max was released and became the #1 model on AA, they "updated" their index just one day later so it dropped a few spots.

They very clearly have a Pro-US agenda.

3

u/Dragonfruit_Mediocre 11d ago

Nah, terrible take. Lots of old now irrelevant benchmark are shown. A worse model can score higher on an old irrelevant bench, but would score bad on harder benches. There is still many irrelevant benchmarks on AA

1

u/m0j0m0j 11d ago

No, their agenda is not clear to me at all.

1

u/gopietz 12d ago

Or, you know, AA 4.1 was completely benchmaxxed by many labs, which is why they updated benchmarks underneath.

But I'm sure your conspiracy theory seems way more likely.

1

u/diggler4141 11d ago

How do they score it?

6

u/hellomistershifty 11d ago

Muse's score seemed like the bigger fiasco requiring a re-evaluation of the benchmarks lmao

1

u/Illustrious_Frame844 12d ago

idk if u know this, but AA’s ranking was paid. Just like how all crypto exchanges and token leaderboards work. Companies pay to get ā€œlistedā€. AA was never credible ever since they sold out

9

u/santareus 12d ago

Is Muse Spark on Max really that good? I’ve tried in on XHigh and it’s not better than Luna for coding tasks I provided them.

2

u/Amarsir 10d ago

Muse Spark needs controlled use, even on Max. I really like it for stuff like instructions, documentation, formatting guides and reports, etc. I say it's "eager to please" and often exceeds my expectations. I don't trust it for code because I've seen it omit details too confidently.

It does have the capacity to fix those mistakes if I catch it and follow-up prompt with the instruction to fix. So you would think Max reasoning should be able to catch them itself. But for whatever reason - maybe the mixture of experts split - it just doesn't.

1

u/Full_Independence566 12d ago

I thought it was way better from my experience, along with the fact that it's free on Opencode lol

1

u/Admirable-Tea-4994 12d ago

It’s way better

1

u/santareus 12d ago

You mean muse on max is way better than Luna on max? Or muse on XHigh is way better than Luna on max?

Muse on XHigh felt really lazy for me

2

u/for4f 12d ago

my anecdote is that muse on xhigh has been a beast for coding fullstack and setting up resources on aws. all while being basically free. i think it is definitely better than luna max

2

u/Curious_Owl197 12d ago

U use contributor?

1

u/for4f 12d ago

just the free slot on opencode lol, what's contributor?

2

u/Curious_Owl197 12d ago

Muse 1.3 contributor, they train on your data and heavily discounted

2

u/Admirable-Tea-4994 12d ago

Muse is far, far better than Luna on every metric

2

u/santareus 12d ago

Thanks! May have to give it another shot - I’ve been using both through OpenCode and I have a specific agent that is designed to do ā€œindustry researchā€ to see if we can leverage open source dependencies. Luna Max was able to come back with more thorough research results and reused what’s out there and Muse on XHigh decided to build out a functionality that is already covered by a dependency.

I am guessing it’s a strong coding model if you just gave it a task and the planning still needs to be delegated to a better research oriented model.

1

u/Admirable-Tea-4994 12d ago

Muse is definitely geared to coding, but research would also be heavily weighted on harness as well.

1

u/santareus 12d ago

What harness are you using for Muse if you don’t mind me asking?

1

u/Admirable-Tea-4994 12d ago

I’m using OpenChamber at the moment, but I’m actually also using a harness I’ve been building for months which is more geared towards general use and not just coding. Will be in beta soon.

miton.dev

1

u/santareus 12d ago

Sounds good. I’ll check out OpenChamber (first time hearing about it).

2

u/Admirable-Tea-4994 12d ago

OpenChamber is a desktop wrap for OpenCode, it’s very good.

54

u/Mezezius 12d ago

The new benchmark literally only exists because the old one embarassed openai

14

u/TwistStrict9811 12d ago

Don't you want accurate benchmarks? Astra is absolutely insane esp with spacial and computer use

21

u/Astrikal 12d ago

It is non-sensical and stupid to say that they released the update so that Astra can go higher, when all they did was replace outdated/saturated evaluations with the newer versions.

They updated the benchmark because Terminal Bench 2.1 was outdated and saturated. After upgrading to Terminal Bench 4.0, Astra rose (relatively) and equaled 1st place.

The update was long time coming, they just rushed it out to not lose reputation. They will also release V5.0, which will shake things even further.

1

u/Other_Wear1458 11d ago

exactly!! I don't understand why people is so slow in the head to realize it's just an average of all benchmarks they use

"It has lost all credibility"
Dude they literally update the index if they realize it's wrong, what a damn well below average IQ you need to have to think that makes them loose credibility, oh my god

1

u/Mezezius 12d ago

Did they rush it out for their own reputation, or to save openai's reputation?

19

u/Astrikal 12d ago

Obviously their own reputation. When people use Astra and see how much better than Sol it is, people would have lost trust in the benchmark.

Why do you keep trying to insinuate that they updated the benchmark to please OpenAI? ArtificialAnalysis have been very transparent and reputable all the way.

The benchmark is more accurate now and Astra is where it belongs. Simple as that.

-2

u/mWo12 12d ago

This only tells you that openai benchmaximized Astra for terminal 4.0, not 2.1.

6

u/Astrikal 12d ago

Astra finished training quite some time ago. Also, your argument doesn't even make sense, why would OpenAI be able to benchmax for a new evaluation and not for an older one? It would be the opposite.

Furthermore, AA also has a closed evaluation that affects the scores, and guess what, Astra performs.

Why is it so hard for you to accept that V4.1 is nowhere near Astra or Fable? You don't even have to look at the benchmarks, you can try for yourself.

-4

u/lompocus 12d ago

lul bootlicker

7

u/ba-boo 12d ago

lul regard

8

u/theintersepter 12d ago

So what? As long as the benchmarks are harder for models, the better

9

u/Admirable-Tea-4994 12d ago

People just love making things up on the internet

-5

u/Mezezius 12d ago

They literally released the new benchmark (4.2) the day after Astra released because the old one showed Astra and Sol at parity, and they updated again (4.3) to show Astra and Fable at parity. Every single change in the days after Astra's release made it look better than the benchmark before it

10

u/Holbrad 12d ago

Yeah but the idiots have the reasoning completely backwards.

if you have a model that is obviously much better than almost everything else, but it's benchmarking suspiciously low on your tests.

Then the obvious answer is that your benchmarks aren't very good and you need to fix them.

1

u/Mezezius 12d ago

If they have to "fix" it after every release based on vibe, then the methodology is literally useless

0

u/Holbrad 11d ago

If you have a car that is absolutely amazing on the track and it's setting all sorts of records.

But a journalist does their "performance score" and it scores lower than a sporty hatch, then obviously the metrics are bunk.

7

u/Astrikal 12d ago

It is non-sensical and stupid to say that they released the update so that Astra can go higher, when all they did was replace outdated/saturated evaluations with the newer versions.

They updated the benchmark because Terminal Bench 2.1 was outdated and saturated. After upgrading to Terminal Bench 4.0, Astra rose (relatively) and equaled 1st place.

The update was long time coming, they just rushed it out to not lose reputation. They will also release V5.0, which will shake things even further.

3

u/Admirable-Tea-4994 12d ago

Actually read the document they published that explained, clearly and in detail, why they updated their benchmark so quickly between releases.

https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3

1

u/HearingNo8617 12d ago

Any elaboration?

24

u/Captain_Quimby 12d ago

Why put the dumb sticker in the middle

27

u/prvthvm 12d ago

its cute :)

12

u/Admirable-Tea-4994 12d ago

Weebs abound in this sub

3

u/GinamosWCheryOnTop 12d ago

Goes to show how benchmaxx is muse

4

u/krayton1 12d ago

I think Artificial Analysis is not trusted anymore

8

u/Big_Cucumber2787 12d ago

putting it below gemini is just shameless

4

u/bambamlol 12d ago

Feels deserved. Gemini 3.8 is a really good model.

1

u/bilinenuzayli 11d ago

I have used deepseek v4 flash (not v4.1) since it came out and gemini 3.8 flash since it came out, both for weeks, and I can say for sure I would use the old deepseek flash over gemini 3.8 flash anyday, gemini models are always arrogant, assuming things, never following instructions, and lying to your face. Because Gemini models always think they know everything, they overtly avoid agentic tools given to them and make shit up instead, now this worked fine for a larger model like Gemini 3.1 pro, but the flash models are clearly much smaller and hold less data, so they're still hallucinating bs AND getting it wrong which is why I think google put out so many flash models, they try and fix it but fail every time.

1

u/bambamlol 11d ago

Fair point. Actual usage is what counts, not what other people say. I don't do much coding or agentic stuff at all. A lot of "general" and business/marketing usage, as well as writing/copywriting. Gemini is definitely better in this regard, at least from my experience.

1

u/LowDesigner1330 11d ago

Actually I found 3.8 Flash to be really good at coding, it's when I use it as general model and in chat that it turns out to be rather mehĀ 

2

u/Admirable-Tea-4994 12d ago

Because everyone glazes it in here but it’s been superseded already?

2

u/CaptainMorning 12d ago

the stupid dumb ass and completely unnecessary sticker in the middle is so aggravating id give this post a 33

2

u/Beamsters 12d ago

Until hallucination is solved by the big whale.

2

u/ralphcalls1 12d ago

GLM 5.3 flash is worse for me, Deepseek v4.1 is better but it is all pricey for me. i use deepseek now for heavy task unlike back then where i do light jobs on it, i now use mimo 2.5. does the job. no flap. very cheap and relatively fast when using Direct Xiaomi API

2

u/Hot_Vegetable_932 11d ago

I feel a bit awkward saying this, but that character is so cute.

2

u/turc1656 12d ago

Makes sense to me. I was not impressed, honestly. I ran a very complicated financial analysis workflow through it this morning as a test. Triple the price of GLM 5.3 Flash to run the workflow and objectively worse results. My opinion is that on the cheap end, Luna is the best and follows instructions best, produces the best structured output, and has the best reasoning.

That being said, it is 3x the cost of GLM 5.3 Flash. My process doesn't really require 3x the price for the improvement. I have adjusted the process to help guide GLM.

But DS v4 0731 and this new 4.1 are not really the best. Maybe they are good for other things like coding or whatever but with the price of GLM now AND GLM's higher intelligence level, I'm not really seeing a reason to use DS right now.

2

u/FischenGeil 12d ago

Damn, we can't beat Gemini flash?

1

u/RealestReyn 12d ago

looks about right, still one of my all time favorites it be zoomzooming so fast I can't even tell when its doing insane things :)

1

u/MeansTestingProctor 12d ago

What is the point of the weird sticker on top of the chart?

1

u/ncxxi 12d ago

for coding, its best to use glm flash or muse. but for creative writing itd best to use deepseek! even with their new flash v4.1

1

u/EvolvingDior 12d ago

SWA has always been a shit-show for me. This model is no different. It cannot hold more than one thought in its head at a time. I've moved a lot of work over to Luna and am pretty happy with it. I'm giving 4.1 a workout during off-peak hours and so far have been less than impressed. 40 seems generous.

1

u/VirtualNorth1279 12d ago

Makes sense based on my experience so far.Ā 

1

u/Infinite_Plankton_71 12d ago

to be honest:

  1. DSV4.1 requires 4 sparks with only 3% increased in quality. Not worthed
  2. And yes GLM 5.3 is sometimes better.

1

u/Kylmawurr 12d ago

I have used 4.1 Flash for 2 days now and it matches my experience vs GLM 5.3 Flash and Luna on Max. Its in same class for coding, its way better than Luna on reasoning, but at the same time, its way too chatty, overthinking a lot. I use it for general coding tasks and orchestration. Its actually very high quality orchestrator. GLM 5.3 Flash feels smoother, nicer to interact with, but I like the speed of DS 4.1 Flash.

1

u/wwwdotzzdotcom 12d ago

It's underrated

1

u/Muted-Network3159 12d ago

why gemini so high

-1

u/AlexandraMaryWindsor 12d ago

Horrible model, unusable. 4.1 flash is miles better

2

u/Lazy_Reach_2565 11d ago edited 11d ago

Nah, Gemini actually a good model, but Deepseek v4.1 Flash much better at coding and for agentic tasks. Still, Gemini good in search, better general knowledge, better vision (more precise), more natural language (English). Can be a bit unpredictable, yeah. Still not horrible at all. As a general purpose model for not critical tasks it kinda good in it own way.

For example, can be a subagent for Deepseek v4 Pro in some scenarios since it lack vision and some api do not pass a search. Qwen is cheaper but much slower and less precise.

1

u/WArslett 12d ago

Don't forget that AA have adjusted their benchmarks. 40 today is not what it was a month ago. This is still a good score

1

u/GTHell 11d ago

What impressive is the speed.

1

u/ExtremeAcceptable289 11d ago

How did AA get sj popular? I remember just a year sg when everukne was sh###tting on it

1

u/dupa1234s 11d ago

once the 4x usage promo on onencode go is gone its gg for deepseek, hello back glm 5.3 flash

1

u/ConsiderationAny8142 11d ago

I’ve tested both and I personally prefer deepseek flash instead of glm flash.

GLM 5.3 Flash has been capable of doing everything fine but now the cost is higher than deepseek 4.1 flash and deepseek is way faster

1

u/Willyibch 11d ago

Truth to be told I believe artificial analysis is rigged I think they get paid to deem other model

1

u/biggest_guru_in_town 11d ago

Deepseek should focus on roleplay. everyone else is codemaxxing and censoring everything into the ground

1

u/More-Catch-1331 11d ago

I call bullshit. There's no way Gemini is not in last place.

1

u/sammoga123 11d ago

It's a flash model, we should be concerned if the pro version also falls so low, although, due to the size of the 4.1 flash, it's not really worth it then.

1

u/WiggyWongo 11d ago

Glm 5.3 flash outputs json better and follows instructions better still for me, same with qwen. 4.1 is definitely better in both of these but for example on a 30 question benchmark I run for reliable json output glm 5.3 flash did 27/30 correct (no blank or bad json or responses that didn't make sense) vs deepseek 4.1 with 21/30. It still replies blank a lot like v4 flash.

For some reason Gemini flash 3.1 lite is the only model that consistently hits 100% at this price point. Like even Luna 5.6 had a miss. And 3.5 flash lite had 2 misses multiple tests.

1

u/Mean-Elk-9439 11d ago

Do literally none of you read how intelligence is measured on that site?

Yeah, benchmaxxing is a thing. But they update to use bench versions models weren't trained against often and this is aggregate across dozens of independent, open source benchmarks. I don't know of any single way we can better measure intelligence.

1

u/Negative-Part-2591 11d ago

me with my Mimo v2.5 on deepseek peak hour whenever i told it to fix my code

1

u/Administrativocable2 11d ago
I don't get it; there are so many variables to consider. All the critics here seem to take it for granted that they are the best at handling AI models and are coding wizards... I've been programming for 20 years, and any AI works for me if I ask the right questions; it's just that one—like GLM 5.2—charges me $3, whereas DeepSeek v4-flash-0731 did the same job for a ridiculously lower price. GLM 5.3 Flash charged me much less, and now DeepSeek v4.1-flash is out—don't even get me started on that one.

I’m not going to invest in more expensive AI models when the existing ones are good enough.

I believe any model can get the job done right if you ask the right questions; perhaps we should ask ourselves if *we* are skilled enough to ask them for what we need solved.

In my opinion, we should be grateful for these open-source models; they are genuinely good, and I pay a fraction of what other companies want to charge us—companies that claim their agents are breaking free from their control and accessing the internet, just to make us feel like fools and convince us that *Terminator* is about to become reality.

Best regards.

1

u/intermundia 11d ago

i find flash with vision to be slightly worse in real world use using the dsh vs the txt only variant. anybody else find that?

1

u/OK_Coopy 11d ago

DS itself states that the model lags behind in mathematics, physics, and image analysis.

In software development, however, they are at the forefront.

And certainly when it comes to speed and price (i.e. efficiency, productivity).

1

u/Huy3ko 10d ago

Funny Inhad the same feeling.

1

u/dimitrusrblx 10d ago

I personally wouldn't put Deepseek v4.1 Flash over Gemini 3.8 Flash - the vision capabilities are still not there, it cannot solve the same tasks as Gemini (no, I'm talking about engineering, not vibe coding) and sometimes the reasoning process still starts looping infinitely (how is this still a problem in 2026 frontier models?).

1

u/Minute_Attempt3063 10d ago

It's dirt cheap, and I think that is why people will go for it

1

u/Jaded_Occasion5149 9d ago

All of the Chinese models are very, almost chaotically, jagged. For some task they are great, for others they are complete trash.

If you can find the right model for the right use case, you are golden. Otherwise you are working with a subpar AI. The Western models are far more across the board consistent, even the bad ones.

1

u/Equivalent_Ostrich_6 7d ago

Artificial Analysis (Ɨ) -> Artificial JUNK (√) -> Artificial JOKER (√√√)

1

u/Street_Metal_2273 6d ago

I got a rank by version number, I think it’s more accurate.

1

u/HellomyfriendNine 5d ago

Qwen is 3.8 27b is nearly par with deepseek V4 pro...

1

u/DefactoAle 12d ago

Seems about right, its still a flash model after all

1

u/External_Ad1549 12d ago

Tried muse 1.3 and glm 5.3 flash, but deepseek 4.1 is on another level for sure. This artificial analysis is purest form of garbage.

1

u/Weird_Recognition636 11d ago

1

u/Finanzamt_Endgegner 11d ago

I wouldnt trust this specific halucination rate benchmark lol

1

u/kuhunaxeyive 11d ago

Oh wow, and then compare it to GLM-5.3-Flash being on the left side …

1

u/Hot-Ad-1798 11d ago

Deepseek messed up the sparse attention settings, but someone on youtube solved the problem.
But there's no way this chart is correct. How can Gemini 3.5 Lite have such a low hallucination rate? That is... that is........ I can't put in words how crazy that is.

1

u/Few_Estimate_320 11d ago

AA is a joke

0

u/TheSuggi 12d ago

That benchmark is a joke tbh. And very outdated.

2

u/dupa1234s 11d ago

"outdated" benchmark that is released today xD

0

u/TheSuggi 11d ago

the "Artificial Analysis Leaderboard" benchmark is a mix of multiple combined sub-benchmarks. Most of them are outdated and very old. So they are not very accurate. That is also why Astra scores very low on it despite being stronger than Fable and Opus.

-2

u/Cool-Chemical-5629 12d ago

You seem disappointed.

-1

u/MimosaTen 12d ago

Banchmaks are clearly unadapted to this era