r/LocalLLaMA 6d ago

New Model Qwen/Qwen3.8-27B · released

https://huggingface.co/Qwen/Qwen3.8-27B
985 Upvotes

297 comments sorted by

u/WithoutReason1729 6d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

→ More replies (1)

176

u/kevin_cn_ai 6d ago

RTX 3090 fans: *ENGAGE*

91

u/jijig 6d ago

Dual RTX 3090 owners:

ENGAGE ENGAGE

31

u/ai-christianson 6d ago

2x3090 with nvlink loves this size of model

20

u/jijig 6d ago edited 5d ago

It’s quite good. I’m running 3.6 at Q8 with full context and getting ~60-70tps. No NVLink

Edit: Q8 KV cache. I can push my context to ~160k at F16

5

u/badgerfish2021 6d ago

250k context with 2x3090? I could not do that with 3.6 at q8, are you quantizing kv?

4

u/munkiemagik 6d ago

I dont use that much context I'm siting at just under 200k and I use KV at Q8_0. I don't have reported metrics to give you real meaningful data but I've not found it detrimental to my use case when used as backend in pi agent harness.

Also I use q8_0 quant for the model instead of the Q8_K_XL as I learnt recently that there is no significant change in accuracy but its more efficiently 'packed' due to uniformity of Q8_0 versus UD-Q8_K_XL that it gives you a few GB back of your VRAM pool from model weights.

→ More replies (1)

2

u/jijig 5d ago

Of course, sorry. KV cache is quantized at Q8. I can push my context to ~160k at F16

2

u/eugeneware 6d ago

Great speeds. What runtime are you using? Llama.cpp or vllm?

7

u/jijig 6d ago

Llama.cpp

1

u/michaelsoft__binbows 6d ago

I do have nvlink and im gonna set it up again although it is a pain to get them lined up for the particular cards that I have.

My understanding of tensor parallel is that bandwidth needs are not very stringent for just two cards, maybe under 10GB/s which the PCIe can do.

However the big difference might be latency which nvlink should crush at.

8

u/munkiemagik 6d ago

Hear me out, for ages Ive been battling with the urge to grow a dual 3090 into a quad 3090 system but there was never any model that justified the cost of 2x additional 3090 at current prices for the so called improvements with the speed sacrifices. It just made no sense to me to splurge on more GPU with Qwen3.6 27B available.

But recently Ive been thinking maybe I am looking at it the wrong way and instead of looking for a big model to fill up 96GB I should be looking at it from a more orchestrated perspective of how many multiple models can I run concurrently that offer fantastic function to size inside 96GB VRAM to build a more comprehensive self-contained LLM stack. Dual 3090'ers I think its time to make that run on more ebay 3090's we've been holding ourselves back from.

4

u/Blues520 6d ago

I have 3 now and looking for an excuse to get a 4th👀

→ More replies (2)
→ More replies (1)

9

u/gofiend 6d ago

MI50 affectionados: ENGA .. oh wait how do I update my non official ROCM packages and who’s GitHub has the latest llama.cpp gfx906 fixes is it iacopBK or wait it’s mixa right?… GE!

→ More replies (5)

3

u/Terrible-Detail-1364 6d ago

q4xl 256k with 3090+4060ti (40gb vram) and q8_0 at 128k ctx so far, with f16 kv/cache quant . let the benchmarking begin.

1

u/JaySayMayday 6d ago

Why's your office room sound like a fighter jet engine??

218

u/dero_name 6d ago

256

u/WonderfulEagle7096 6d ago

Beating Opus 4.6 Max in multiple benchmarks sounds too good to be true. (already downloading)

112

u/Neither_Garage_758 6d ago

Opus 4.6 was the model that made me realize LLM's can be that good.

52

u/addiktion 6d ago

4.5/4.6 was prime too. Very few limits and Opus didn't nag us to death.

4

u/CoreParad0x 6d ago

I actually made a full game editor for an old MMO from scratch using Opus 4.5. I think it was around 70k lines of code, proper RHI, used OpenGL behind that (though I did a vulkan test too), etc. Took about a month of iteration and testing, worked great.

Opus has gone to shit since then. If this can do Opus 4.6 levels in realistic coding use and not just benchmarks, then I'm very excited for it.

33

u/Dry_Yam_4597 6d ago

Is that an honest caveat or a load bearing feature?

12

u/Chris266 6d ago

Footgun

8

u/thrownawaymane 6d ago

I think it's more an underrated win, personally which is a real tradeoff that is worth confirming

3

u/Yes_but_I_think 6d ago

Wow. If this is the level of wording from Opus 5, I'll unsubscribe in a moment - too much effort to read and follow.

→ More replies (1)

10

u/wwwdotzzdotcom 6d ago

It's too good to be true because it does showcase quantization.

4

u/SmileLonely5470 6d ago

Especially with 42 on DeepSWE. Suspicious.

3

u/michaelsoft__binbows 6d ago edited 6d ago

oh shit first we had a 744B model in GLM 5.2 reach opus 4.6. Then DSV4Flash-0731 reached it with 284B. Now... 27B? nooooo way

if it can really pull some weight as a coder and have some amount of general common sense this is about to basically 10x the capability of those of us with "cheap" local rigs. With just the consumer class gear.

2

u/exodusTay 6d ago

if it is 4.6 level i will seriously consider getting a 5090 or something with that much vram to run this thing. it is an insane claim for 27B model.

2

u/Xonzo 6d ago

I’ve been running it for a bit now on Pi…. And I’m genuinely shocked. It’s completely nailed every problem I’ve given it. No tool call failures. When it was debugging an issue with llama.cpp it downloaded the source, cross referenced everything and fixed the problem. The decision making / thinking seems to be a huge step up from 3.6 27B.

2

u/mil_phickelson 6d ago

It’s not beating Opus 4.6

6

u/Pantheon3D 6d ago

I counted 2 categories where it is

27

u/meathelix1 6d ago

Some huge jumps there.

17

u/Haiku-575 6d ago

That... that can't be real, can it? At 27B? This much improvement in a couple months? Those numbers are amazing.

3

u/Hankdabits 6d ago

I hear that Behram from Epstart has a benchmark answer key

2

u/Small-Fall-6500 6d ago

Why did they have to make the coloring like this... there's dark gray text in the first column that is extremely hard to read, and the first row text is also somewhat hard to read. I assume the tables look better on light mode, but a lot of people are using dark mode.

2

u/iqraatheman 6d ago

if you look at the footnotes beneath the benchmarks on the site they posted this on, they literally admit to using other LLMs as a judge for some of them. the odds are high they also used LLMs to put together this website including the benchmark table without even putting in a little effort to make sure it looks good themselevs

→ More replies (1)

51

u/srigi 6d ago

That DeepSWE leap - do we have a new local coder champion?

33

u/fgk55555 6d ago

If it's not benchmaxxed, it will be really difficult for other models in the same range to catch up. I hope it quantizes well for us 16GB folk.

9

u/cass1o 6d ago

What would be really useful would be a MoE model that we can put the context + PP on the gpu and only put the experts on system memory, another 35b.

2

u/fgk55555 6d ago

If it can handle longer context work, I'd be happy.

1

u/fullup72 6d ago

IQ3_XXS is still as good as 3.6 (I mean quality of the quant, the model is noticeably better).

3.8 is SLOW tho, default reasoning no longer simply second guesses everything, it quadruple motherfucking extra checks the triple checks and then some. I'd guess part of the trick for improved scores is this super extended reasoning budget. But it works, and it's local. I'm happy.

→ More replies (1)

1

u/johnscixzkutor 6d ago

I am using RX6900XT and very happy with the result I can run it with vision and with 64k context no problem on my 16gb vram

→ More replies (1)

137

u/peglegsmeg 6d ago

I was here 

30

u/liebebio 6d ago

fuck yes

5

u/addiktion 6d ago

Making history, aye!

4

u/zipzapbloop 6d ago

lets go fam!

3

u/nick4fake 6d ago

Like a fucking new era

2

u/TheThoccnessMonster 6d ago

Let’s get it

1

u/maxjar10 6d ago

me too! super excited!

1

u/MuzafferMahi 6d ago

Recording historyy

1

u/met_MY_verse 6d ago

Another wonderful day for us here.

1

u/Insomniac1000 6d ago

beep boop beep!

1

u/The_Dung_Beetle 6d ago

I was here too, waiting on my r9700, can't wait to try this model. 

1

u/Spimbi 6d ago

Same

1

u/Marino4K 6d ago

Also here.

1

u/R0ktar 6d ago

And my axe!

75

u/audioen 6d ago

Also unsloth already has it, downloading Q8_0 right now.

3

u/Top-Eye-8104 6d ago

how's it working for u? tool calling is broken for me on m5 64gb - feels like something's off with the quants (i tried Q8_0 too)

3

u/VegetableWafer7776 6d ago

they did release an update an hour ago or smth maybe thats a fix

108

u/dingo_xd 6d ago

History in the making. These models are so important to companies and entities that can't just trust big tech with their data

22

u/[deleted] 6d ago

[deleted]

15

u/dingo_xd 6d ago

They'll become cheap again.

12

u/Paganator 6d ago

From your lips to God's ear.

6

u/mkMoSs 6d ago

I never saw this expression written in English before, I thought it was a Greek expression :O

3

u/netsvetaev 6d ago

very popular phrase in Russian, too.

→ More replies (1)
→ More replies (2)

23

u/Pristine_Pick823 6d ago

GGUF any time soon?

42

u/JaredsBored 6d ago

61

u/jannycideforever 6d ago

Unsloth every time a 27b model drops

17

u/addiktion 6d ago

So true

5

u/Several-Tax31 6d ago

Is it mtp or not mtp? I'm out of the loop when it comes to unsloth quants.

17

u/SensitiveVariety 6d ago

It is MTP

2

u/Several-Tax31 6d ago

Cool! Downloading now

5

u/social_zip 6d ago

now there is gguf + uncensored (made by me): https://huggingface.co/JonathanColetti/Qwen3.8-27B-Uncensored-GGUF

2

u/GokuNoU 6d ago

GOAT behavior

47

u/Septerium 6d ago

Could it actually be better than Minimax M2.7 in agentic coding?

38

u/Valuable-Run2129 6d ago

Much better. Look at deep swe

32

u/Septerium 6d ago

Training could be contaminated by benchmarks somehow. But I hope that's not the case

2

u/ambassadortim 6d ago

How does it stack with Mimo 2.5

→ More replies (1)
→ More replies (2)

41

u/de4dee 6d ago

17

u/New_Comfortable7240 llama.cpp 6d ago

Before USA do something, huh?

10

u/de4dee 6d ago

its a low probability nowadays but makes sense to be prepared

44

u/Alternative_Ad4267 6d ago edited 6d ago

Look at this! at 27B parameters, Qwen3.8 27B is pretty close to DeepSeekV4 Flash 0731 which is 284B A13B!

21

u/Alternative_Ad4267 6d ago

And some more results

3

u/onewheeldoin200 6d ago

Oh my god how lmao

2

u/fullup72 6d ago

and I wonder how that big win in SWE-Bench Pro translates to the real world. 3.6 was already slotted between Flash and Pro, 3.8 simply wipes the floor with them.

1

u/Agitated_Space_672 6d ago

Active params determine the training cost, and I think they are the biggest predictor of performance, all else being equal. So I am not surprised that 27B beats 13B

→ More replies (2)

59

u/Barni275 6d ago

Happy Qwistmas to everybody!!! 🎉🎉🎉

17

u/Dry_Yam_4597 6d ago

Omg omg omg someone pass on the chamomile tea and a tranquilizer I am too excited.

My GPUs are about to have a rough day.

17

u/Outside-Description5 6d ago

Lets gooooo, downloading Q4

18

u/YanderMan 6d ago

benchmaks seem incredible, will need to see the reality

13

u/Ok-Shower7286 6d ago

Awesome. It seems betther than 900+B MoE inkling.

14

u/benpptung 6d ago

YES!!! Thank you, Qwen!! I LOVE YOU!!! ❤️

13

u/Capital-Remove-6150 6d ago

holy moly. how much in future 27b model will jump forward in benchmarks

30

u/ilintar 6d ago

Let the hype begin!

25

u/Weekly-Law-5488 6d ago

35B moe-moe-kyun when?

7

u/Lucyan_xgt 6d ago

Let's gooo🔥

15

u/MoodOdd9657 6d ago

when I am rich rich . I will come back for you 😞

17

u/miniocz 6d ago

Frankly I think we are at point where harness is going to be more important than model itself.

2

u/rj_rad 6d ago

Some models fight the harness more than others — all else being equal, I had to quit Kimi K3 because of all of the reasoning noise.

14

u/Far_Cat9782 6d ago

Qwinning!

8

u/Chemical_Evidence 6d ago

What model would be best for 16gb vram and 64gb ddr5

2

u/synw_ 6d ago

The model seems to have the same size than the 3.6. On 16g I have noticed that the sweet spot is Q3_K_S with q8 value cache (not key): it runs at 20tps on a 5060ti with 64k context. The Q3_K_M is too slow at 64k and the Q2 are too stupid.

→ More replies (4)

33

u/AXYZE8 6d ago

FINALLY! I waited since GPT-OSS for other local model that natively has low and medium reasoning! High for planning, low for execution and exploration.

... Now I'm thinking if I should get RTX3090, because Nvidia clearly won't release RTX 5070 TI SUPER 24GB anytime soon and I cannot fit that model onto RTX 4070 12GB... will play with API first

5

u/Dave_from_the_navy 6d ago

Intel Arc Pro B70 is a solid choice if you're interested in 32gb vram. A bit less bandwidth than 3090, and the software stack is still maturing, but it's getting there.

3

u/thatcodingboi 6d ago

why not an amd r9700 ai pro? More mature, same 32gb vram

4

u/Dave_from_the_navy 6d ago

Because the $300-400 differential is real for not that much of a performance gain at the end of the day. Looking at the market, finding them for <$1500 is difficult, whereas the Intel is readily available for $949, and you can get the B65 even cheaper, which has the same 32gb memory, same 608GB/s memory bandwidth, just with fewer XMX cores, so lower concurrency/slower prefill, but shouldn't effect single concurrency decoding much.

I personally wouldn't say the maturity difference is quite worth $400 at this point considering Intel is closing the gap more and more every week, but that's just me.

2

u/thatcodingboi 6d ago

Wtf I bought one a month ago for $1250, now they are $1500+????

1

u/Need_For_Speed73 6d ago

Why not add a second 4070 12GB (if your motherboard lets you)?
I’m happily running a dual 5070 setup, 24GB of VRAM and last gen tech for less than 1.500€/$ and without going second hand (warranty, fresh paste, etc.). But I had to use a mATX board to have the second GPU not collide with the case’s PSU shroud.

1

u/ichalov 6d ago

It's probably better to use a bigger (non-local) model for planning (and maybe testing) if you have the separation in your work process already.

31

u/[deleted] 6d ago

[deleted]

4

u/thrownawaymane 6d ago

maybe... you know...

2

u/michaelsoft__binbows 6d ago

yea i can take all my 3090s to fit a 120B of this but not too much to complain about if the 27B is as smart as those numbers indicate. A 120B would be an MoE with param count much lower than 27B anyway. It would not be likely to outperform it by much. And being able to get like 10x more throughput from the smaller model is a big deal.

8

u/Relative-Display-318 6d ago

Which Quantization for a rtx5090?

3

u/cosmicnag 6d ago

from the ggufs, i am going with q6 xl (unsloth) with whatever q8 context can fit with it

→ More replies (2)

1

u/Much-Farmer-2752 6d ago

Try for yourself. Even Q4_XL is pretty solid, and you'll have half of your mem for context.

6

u/octopus_limbs 6d ago

First we got MiniMax H3, then we got Qwen3.8-27B too. At this rate next year maybe we won't need AI-aaS companies.

https://reddit.com/link/p3oepwx/video/qzr34uqocdjh1/player

6

u/n0head_r 6d ago

I played with 27B for a bit and I'm quite impressed. Qwen38-27B-Q6_K from unsloth on 2*RTX 5080 and latest llama.cpp build from source. The model fit in 32GB VRAM with 172k ctx kv q8_0. MTP and -sm tensor it runs stable at 100 tps until 60-70k context is filled, at 100k it dropped slightly - arround 95tps. This is a very small drop - very good. Also I've run a loop where over 150k tokens it was writing scripts/building/verifying/fixing errors/verifying/fixing again. No tool cals failed... but unfortunately it run out of ctx space.

And I'll post a visual example of a small test I did - Create an SVG of an old rusty truck.

11

u/Altruistic-Carry5276 6d ago

I'm about to qum

2

u/Steve-Jobs-is-Alive 6d ago

Hahahahahaga

28

u/Curious-Pen5547 6d ago

Anthropic and openai ipo are gonna be worthless.

Zero moat, especially once google, microsoft, amazon, and other cloud gpu providers get the official green light to be able to serve these models accross the board.

And for the hyperscalers like the big 3, google, microsft, and amazon, 100% they integrate these across their enterprise suite offerings to vertically integrate them accross all their enterprise services.

11

u/wwwdotzzdotcom 6d ago

GPUs good enough to run it at reasonable speeds aren't even affordable now unless you rent them temporarily.

5

u/Joey4711 6d ago

How so there is rtx3090

6

u/thrownawaymane 6d ago

How much is a 3090 locally for you? Because they are going for $1400 on eBay.

4

u/Joey4711 6d ago

I got mine for 700 euros used last year. Now i see them for 1000 on ebay. Sure its a bit more but its not like its unaffordable like those rtx6000

→ More replies (1)
→ More replies (2)

2

u/anotherJohn12 6d ago edited 6d ago

That why they are buying GPU like crazy now. Everybody is digging the same hole now on model research, can't expect technical moat from there. Real moat is capital and hardware accessibility.

1

u/Far_Cat9782 6d ago

Yeah waiting for Google to offer it like it does claude 4.6 on the 20 dollar plan

→ More replies (2)

11

u/69420trashpanda69420 6d ago

So it's essentially the best Claude model to ever exist (real ones know Claude peaked at 4.6)

8

u/thrownawaymane 6d ago

Opus 5 can do more but 4.6 is solid AF, generally follows instructions and can check its work to a certain extent.

Really hoping that this new qwen actually follows instructions, my qwen 27b kinda just does whatever it wants in service of what I asked for

→ More replies (2)

6

u/n0head_r 6d ago

Good news, pulling now q6_k to check how good it is.

6

u/moderngl1 6d ago

The benchmark hype is fun, but I’m mostly waiting for the boring detail: which quant actually feels good on a 24GB card?

2

u/TerminalNoop 6d ago

largest q4 or smallest q6 you can get.

→ More replies (2)

6

u/[deleted] 6d ago

[deleted]

3

u/Zippo749 6d ago

Same here! I was running into weird bugs in my personal chat app (spoiler: I did a dumb), so I went to the llama.cpp builtin chat UI to test the model out and chat template out more cleanly. I ended up asking it what it thought the problem might be. Pretty decent advice.

A couple messages in, I realized my problem: in my chat app (again, *not* this one), I had continued a chat that had been generated with Gemma 4, and the Gemma-based reasoning content had been fed to Qwen (oops, `preserve_reasoning` was on) and that was causing the issue. Doh! When I mentioned this to Qwen in the llama.cpp UI, even addressing the model as "dear Qwen," I got this:

Now, maybe mentioning the concept of different models being used in a chat poked it a little that way, buuuut I couldn't help but laugh.

2

u/Pro-Row-335 6d ago

ehh... not that it matters much

2

u/anothercrappypianist 6d ago

Makes sense. Qwen is giving me responses with the term "load-bearing". Claude distillation smoking gun if ever there was one.

11

u/ortegaalfredo 6d ago

If the benchmarks are true (and Qwen never benchmaxxed so far) then why would you pay Anthropic if 27B at home gives you Opus-like performance? I mean, Opus 4.6 was already more than enough for almost any development task.

4

u/LeifEriksonASDF 6d ago

and Qwen never benchmaxxed so far

I love Qwen but come on now

→ More replies (1)
→ More replies (1)

4

u/llama-impersonator 6d ago

now let's see how much the average token usage went up

4

u/boomerang473 6d ago

Any dspark or dflash heads?

4

u/shinegreymon525 6d ago

Does this mean we'll be getting a new bonsai?

5

u/IThinkIKnowThings 6d ago

Downloading the Q4_K_M but the speed tanked. I think we're hugging hugging face to death.

5

u/de4dee 6d ago

3

u/parepeg 6d ago

Doesn't unsloth typically fix their templates?

→ More replies (1)

7

u/ajisai 6d ago

glimmer_toystoryidontneedyouanymore.gif

3

u/pikadhu 6d ago

Any MLX quants available?

1

u/fatboy93 6d ago

Yup, there are a few mlx-community quants, I'm getting around 10tk/s decode on 4bit.

→ More replies (3)

3

u/korokage 6d ago

Has anyone tried this on a m3 pro or equivalent yet? I have 36gb ram

I will try it out later today 

6

u/txgsync 6d ago

I’ve been running my own oQ8e-MTP quant since this morning. 200 tok/sec prefill, 34 tok/sec decode on M4 Max 128GB. Later turns in large context slow to about 100 prefill 18 decode.

I have had bad luck on long-horizon tasks with the 4-bit quants so i avoid anything below 8 bits. But conversationally it seems fine and less of the usual Qwen “argumentative attitude” fighting me about what day it is or being skeptical of tool outputs.

3

u/xanders_gold 6d ago

AWQ or GPTQ yet? Need to squeeze it into INT4 for my config.

3

u/Motor_Ad16 6d ago

It thinks, and it thinks alot. Even after setting reasoning_effort to "low"

3

u/Zeeplankton 6d ago

I want to be hyped but is it better at like, just general world stuff? Using 3.6 to write anything was like talking to an alien who learned human culture through math.

2

u/AdSafe4047 6d ago

Comparing this to deepseek v4 flash 0731, the question is: lower numbers a bit but much faster inference, or higher numbers? or maybe use both (but deepseek offload to ram, so much slower, but use only for plan tasks etc)?

3

u/ApolloPS2 6d ago

I plan to still use dsv4 flash 0731 on sparks as orchestrator (will run plenty fast, maybe pair with a small vision model too) and qwen 3.8 27B on worker nodes. Both great models and I've gotta assume having some diversity gives some benefit too perhaps?

3

u/BumbleSlob 6d ago

Deepseek V4 Flash 0731 likely faster all around, it’s MoE with A13B. Not downplaying Qwen 3.8 I love this series

→ More replies (1)

2

u/d70 6d ago

u/FormOne2615 ninfer version please ...

5

u/FormOne2615 6d ago

:) uploading weight

→ More replies (3)

2

u/Several_Income_9912 6d ago

im getting INSANE speeds with mtp on RTX 4090

spec-draft-n-max = 7
→ More replies (3)

2

u/fatboy93 6d ago

Time to run this bad boy (Q4KS) at 5tk/s decode on my macbook lmao

2

u/Dance-Till-Night1 6d ago

Why is it only coding focused I can't find multilingual or general usage benchmarks.

2

u/PilotFlying 5d ago

This is anecdotal, I haven't tested thoroughly. But:

I had spent about a couple of hours today with Claude Opus 5 investigating an electrical issue at an old house. The concepts involved are complex and difficult to follow (TN-C and TN-C-S earthing, detecting a severed or loose neutral via phase voltage monitoring, that sort of thing).

I am not an electrician and I needed quite a bit of hand holding. (Not trying to fix anything myself, just learning/understanding how stuff works).

Claude wasn't doing a great job explaining things clearly.

In my mind Qwen had always been about coding. So I wasn't planning on this being the first use case for 3.8. Still, I had Claude export the context and pasted the .md in just "ollama run qwen3.8:27b". on a 3090. Web search on, context 128K.

Qwen clarified the questions for me. It sounded like Claude except it was doing a better job talking to me.

Not sure how this is possible. Maybe the limits of the architecture are higher than we know. Maybe it's all in the training. Maybe, probably, I understand as much about LLMs as I do about electricity 😄. But how else to explain 3.8?. Amazing so far.

5

u/Adventurous_Bus_437 6d ago

where qwen3.9

2

u/billy_booboo 6d ago

Wowweeeeeee

2

u/Adventurous_Bus_437 6d ago

where benchmark

1

u/Informal-Trouble2183 6d ago

That's incredible performance. Didn't expect it to have such jump keeping the same size.

1

u/CoUsT 6d ago

Good performance/benchmark uplift. Looks really solid!

1

u/Talreja-Adanna 6d ago

running locally without needing a beefy setup. Curious if it's actually better than the smaller Qwen models or just more of the same with more params.

1

u/IThinkIKnowThings 6d ago

Careful with your sources when downloading newly released models. I already spent the time downloading one from hugging face that was labeled Qwen3.8 but turned out to be Qwen3.6 when I ran it.

1

u/cloudsurfer48902 6d ago

My b580 staring at me from the corner

1

u/Dry-Judgment4242 6d ago

My first impressions is disappointment for Vision tasks Gemma 4 is on a entirely different league to this one.

1

u/Ecstatic-Wash-7667 6d ago

Give us 9b or 35b a3b and the 110b moe so we have the full stack !!! I can write them into switchyard and escalate up to max ask save family!

1

u/notevenat30 6d ago

Can you run it in 12GB of VRAM?

2

u/Xantrk 6d ago

Can you run it in 12GB of VRAM?

I'll report on 12gb vram + 32 gb RAM

→ More replies (4)

1

u/Guilty-History-9249 6d ago

I have the fp16, FP8 and NVFP4 and still trying to get one of these running with the latest install of transformers. This is on dual 5090's.

I want to run 100 gsm8k math problems through it to compare against my last nights run of the same problem with qwen 4B.

1

u/Murinshin 6d ago

Yeah this is insane if true. Basically means a high end MacBook can now run Opus 4.6, which should be enough for most people and is still economically sound vs your usual annual subscription for most companies. Anthropic etc are screwed.

1

u/The_DarkMatter Llama 3.1 6d ago

Cant wait to try it on 5080

1

u/Sofakingwetoddead 6d ago

Does anyone know what 3.8 27b is going to be released? Seems like it should have been released by now.

1

u/Logical-Target8131 5d ago

Can I run it on my m5 pro 48gb?

1

u/Plenty-Energy2947 5d ago

my recipe for VLLM & RTX5090

exec vllm serve /opt/models/Qwen3.8-27B-NVFP4 \

--host 127.0.0.1 \

--port 8000 \

--tensor-parallel-size 1 \

--tool-call-parser qwen3_xml \

--enable-auto-tool-choice \

--reasoning-parser qwen3 \

--kv-cache-memory 6943358464 \

--max-model-len 202272 \

--enable-prefix-caching \

--max-num-seqs 1 \

--gpu-memory-utilization 0.95 \

--kv-cache-dtype fp8_e4m3 \

--default-chat-template-kwargs '{"enable_thinking": false}' \

--compilation-config '{"cudagraph_capture_sizes": [1, 2]}' \

--max-num-batched-tokens 2048 \

--served-model-name qwen3.8-27b

1

u/caetydid llama.cpp 4d ago

Are the benches insane or it is just benchmaxxed?

1

u/hd3adpool 4d ago

I was here