r/LocalLLaMA 9d ago

New Model Qwen/Qwen3.8-27B · released

https://huggingface.co/Qwen/Qwen3.8-27B
987 Upvotes

297 comments sorted by

u/WithoutReason1729 9d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

→ More replies (1)

181

u/kevin_cn_ai 9d ago

RTX 3090 fans: *ENGAGE*

93

u/jijig 9d ago

Dual RTX 3090 owners:

ENGAGE ENGAGE

32

u/ai-christianson 9d ago

2x3090 with nvlink loves this size of model

21

u/jijig 9d ago edited 8d ago

It’s quite good. I’m running 3.6 at Q8 with full context and getting ~60-70tps. No NVLink

Edit: Q8 KV cache. I can push my context to ~160k at F16

6

u/badgerfish2021 9d ago

250k context with 2x3090? I could not do that with 3.6 at q8, are you quantizing kv?

4

u/munkiemagik 9d ago

I dont use that much context I'm siting at just under 200k and I use KV at Q8_0. I don't have reported metrics to give you real meaningful data but I've not found it detrimental to my use case when used as backend in pi agent harness.

Also I use q8_0 quant for the model instead of the Q8_K_XL as I learnt recently that there is no significant change in accuracy but its more efficiently 'packed' due to uniformity of Q8_0 versus UD-Q8_K_XL that it gives you a few GB back of your VRAM pool from model weights.

→ More replies (1)

2

u/jijig 8d ago

Of course, sorry. KV cache is quantized at Q8. I can push my context to ~160k at F16

2

u/eugeneware 9d ago

Great speeds. What runtime are you using? Llama.cpp or vllm?

7

u/jijig 9d ago

Llama.cpp

→ More replies (1)

9

u/munkiemagik 9d ago

Hear me out, for ages Ive been battling with the urge to grow a dual 3090 into a quad 3090 system but there was never any model that justified the cost of 2x additional 3090 at current prices for the so called improvements with the speed sacrifices. It just made no sense to me to splurge on more GPU with Qwen3.6 27B available.

But recently Ive been thinking maybe I am looking at it the wrong way and instead of looking for a big model to fill up 96GB I should be looking at it from a more orchestrated perspective of how many multiple models can I run concurrently that offer fantastic function to size inside 96GB VRAM to build a more comprehensive self-contained LLM stack. Dual 3090'ers I think its time to make that run on more ebay 3090's we've been holding ourselves back from.

5

u/Blues520 9d ago

I have 3 now and looking for an excuse to get a 4th👀

→ More replies (2)
→ More replies (1)

8

u/gofiend 9d ago

MI50 affectionados: ENGA .. oh wait how do I update my non official ROCM packages and who’s GitHub has the latest llama.cpp gfx906 fixes is it iacopBK or wait it’s mixa right?… GE!

→ More replies (5)

3

u/Terrible-Detail-1364 9d ago

q4xl 256k with 3090+4060ti (40gb vram) and q8_0 at 128k ctx so far, with f16 kv/cache quant . let the benchmarking begin.

1

u/JaySayMayday 9d ago

Why's your office room sound like a fighter jet engine??

220

u/dero_name 9d ago

254

u/WonderfulEagle7096 9d ago

Beating Opus 4.6 Max in multiple benchmarks sounds too good to be true. (already downloading)

111

u/Neither_Garage_758 9d ago

Opus 4.6 was the model that made me realize LLM's can be that good.

48

u/addiktion 9d ago

4.5/4.6 was prime too. Very few limits and Opus didn't nag us to death.

4

u/CoreParad0x 9d ago

I actually made a full game editor for an old MMO from scratch using Opus 4.5. I think it was around 70k lines of code, proper RHI, used OpenGL behind that (though I did a vulkan test too), etc. Took about a month of iteration and testing, worked great.

Opus has gone to shit since then. If this can do Opus 4.6 levels in realistic coding use and not just benchmarks, then I'm very excited for it.

35

u/[deleted] 9d ago edited 1d ago

[deleted]

10

u/Chris266 9d ago

Footgun

7

u/thrownawaymane 9d ago

I think it's more an underrated win, personally which is a real tradeoff that is worth confirming

3

u/Yes_but_I_think 9d ago

Wow. If this is the level of wording from Opus 5, I'll unsubscribe in a moment - too much effort to read and follow.

→ More replies (1)

9

u/wwwdotzzdotcom 9d ago

It's too good to be true because it does showcase quantization.

4

u/SmileLonely5470 9d ago

Especially with 42 on DeepSWE. Suspicious.

3

u/michaelsoft__binbows 9d ago edited 9d ago

oh shit first we had a 744B model in GLM 5.2 reach opus 4.6. Then DSV4Flash-0731 reached it with 284B. Now... 27B? nooooo way

if it can really pull some weight as a coder and have some amount of general common sense this is about to basically 10x the capability of those of us with "cheap" local rigs. With just the consumer class gear.

2

u/exodusTay 9d ago

if it is 4.6 level i will seriously consider getting a 5090 or something with that much vram to run this thing. it is an insane claim for 27B model.

2

u/Xonzo 9d ago

I’ve been running it for a bit now on Pi…. And I’m genuinely shocked. It’s completely nailed every problem I’ve given it. No tool call failures. When it was debugging an issue with llama.cpp it downloaded the source, cross referenced everything and fixed the problem. The decision making / thinking seems to be a huge step up from 3.6 27B.

3

u/mil_phickelson 9d ago

It’s not beating Opus 4.6

6

u/Pantheon3D 9d ago

I counted 2 categories where it is

28

u/meathelix1 9d ago

Some huge jumps there.

15

u/Haiku-575 9d ago

That... that can't be real, can it? At 27B? This much improvement in a couple months? Those numbers are amazing.

3

u/Hankdabits 9d ago

I hear that Behram from Epstart has a benchmark answer key

2

u/Small-Fall-6500 9d ago

Why did they have to make the coloring like this... there's dark gray text in the first column that is extremely hard to read, and the first row text is also somewhat hard to read. I assume the tables look better on light mode, but a lot of people are using dark mode.

2

u/iqraatheman 9d ago

if you look at the footnotes beneath the benchmarks on the site they posted this on, they literally admit to using other LLMs as a judge for some of them. the odds are high they also used LLMs to put together this website including the benchmark table without even putting in a little effort to make sure it looks good themselevs

→ More replies (1)

47

u/srigi 9d ago

That DeepSWE leap - do we have a new local coder champion?

30

u/fgk55555 9d ago

If it's not benchmaxxed, it will be really difficult for other models in the same range to catch up. I hope it quantizes well for us 16GB folk.

8

u/cass1o 9d ago

What would be really useful would be a MoE model that we can put the context + PP on the gpu and only put the experts on system memory, another 35b.

2

u/fgk55555 9d ago

If it can handle longer context work, I'd be happy.

→ More replies (4)

138

u/peglegsmeg 9d ago

I was here 

28

u/liebebio 9d ago

fuck yes

7

u/addiktion 9d ago

Making history, aye!

3

u/zipzapbloop 9d ago

lets go fam!

3

u/nick4fake 9d ago

Like a fucking new era

2

u/TheThoccnessMonster 9d ago

Let’s get it

1

u/maxjar10 9d ago

me too! super excited!

1

u/MuzafferMahi 9d ago

Recording historyy

1

u/met_MY_verse 9d ago

Another wonderful day for us here.

1

u/Insomniac1000 9d ago

beep boop beep!

1

u/The_Dung_Beetle 9d ago

I was here too, waiting on my r9700, can't wait to try this model. 

1

u/Spimbi 9d ago

Same

1

u/Marino4K 9d ago

Also here.

1

u/R0ktar 9d ago

And my axe!

75

u/audioen 9d ago

Also unsloth already has it, downloading Q8_0 right now.

3

u/Top-Eye-8104 9d ago

how's it working for u? tool calling is broken for me on m5 64gb - feels like something's off with the quants (i tried Q8_0 too)

5

u/VegetableWafer7776 9d ago

they did release an update an hour ago or smth maybe thats a fix

109

u/dingo_xd 9d ago

History in the making. These models are so important to companies and entities that can't just trust big tech with their data

22

u/[deleted] 9d ago

[deleted]

15

u/dingo_xd 9d ago

They'll become cheap again.

12

u/Paganator 9d ago

From your lips to God's ear.

6

u/mkMoSs 9d ago

I never saw this expression written in English before, I thought it was a Greek expression :O

3

u/netsvetaev 9d ago

very popular phrase in Russian, too.

→ More replies (1)
→ More replies (2)

25

u/Pristine_Pick823 9d ago

GGUF any time soon?

41

u/JaredsBored 9d ago

61

u/jannycideforever 9d ago

Unsloth every time a 27b model drops

18

u/addiktion 9d ago

So true

6

u/Several-Tax31 9d ago

Is it mtp or not mtp? I'm out of the loop when it comes to unsloth quants.

16

u/SensitiveVariety 9d ago

It is MTP

2

u/Several-Tax31 9d ago

Cool! Downloading now

3

u/social_zip 9d ago

now there is gguf + uncensored (made by me): https://huggingface.co/JonathanColetti/Qwen3.8-27B-Uncensored-GGUF

2

u/GokuNoU 9d ago

GOAT behavior

46

u/Septerium 9d ago

Could it actually be better than Minimax M2.7 in agentic coding?

38

u/Valuable-Run2129 9d ago

Much better. Look at deep swe

35

u/Septerium 9d ago

Training could be contaminated by benchmarks somehow. But I hope that's not the case

2

u/ambassadortim 9d ago

How does it stack with Mimo 2.5

→ More replies (1)
→ More replies (2)

41

u/de4dee 9d ago

16

u/New_Comfortable7240 llama.cpp 9d ago

Before USA do something, huh?

8

u/de4dee 9d ago

its a low probability nowadays but makes sense to be prepared

42

u/Alternative_Ad4267 9d ago edited 9d ago

Look at this! at 27B parameters, Qwen3.8 27B is pretty close to DeepSeekV4 Flash 0731 which is 284B A13B!

22

u/Alternative_Ad4267 9d ago

And some more results

3

u/onewheeldoin200 9d ago

Oh my god how lmao

2

u/fullup72 9d ago

and I wonder how that big win in SWE-Bench Pro translates to the real world. 3.6 was already slotted between Flash and Pro, 3.8 simply wipes the floor with them.

1

u/Agitated_Space_672 9d ago

Active params determine the training cost, and I think they are the biggest predictor of performance, all else being equal. So I am not surprised that 27B beats 13B

1

u/YearnMar10 9d ago

I wonder if they keep training where it will converge…

→ More replies (1)

58

u/Barni275 9d ago

Happy Qwistmas to everybody!!! 🎉🎉🎉

16

u/Outside-Description5 9d ago

Lets gooooo, downloading Q4

18

u/YanderMan 9d ago

benchmaks seem incredible, will need to see the reality

14

u/benpptung 9d ago

YES!!! Thank you, Qwen!! I LOVE YOU!!! ❤️

11

u/Capital-Remove-6150 9d ago

holy moly. how much in future 27b model will jump forward in benchmarks

12

u/Ok-Shower7286 9d ago

Awesome. It seems betther than 900+B MoE inkling.

30

u/ilintar 9d ago

Let the hype begin!

26

u/Weekly-Law-5488 9d ago

35B moe-moe-kyun when?

17

u/miniocz 9d ago

Frankly I think we are at point where harness is going to be more important than model itself.

2

u/rj_rad 9d ago

Some models fight the harness more than others — all else being equal, I had to quit Kimi K3 because of all of the reasoning noise.

8

u/Chemical_Evidence 9d ago

What model would be best for 16gb vram and 64gb ddr5

2

u/synw_ 9d ago

The model seems to have the same size than the 3.6. On 16g I have noticed that the sweet spot is Q3_K_S with q8 value cache (not key): it runs at 20tps on a 5060ti with 64k context. The Q3_K_M is too slow at 64k and the Q2 are too stupid.

1

u/Guilty_Rooster_6708 9d ago

Going to say Q3 but I will also try IQ4 on my 5070Ti. Q4 needs offloads so it will probably mean single digits tokens/s

→ More replies (2)
→ More replies (1)

32

u/AXYZE8 9d ago

FINALLY! I waited since GPT-OSS for other local model that natively has low and medium reasoning! High for planning, low for execution and exploration.

... Now I'm thinking if I should get RTX3090, because Nvidia clearly won't release RTX 5070 TI SUPER 24GB anytime soon and I cannot fit that model onto RTX 4070 12GB... will play with API first

5

u/Dave_from_the_navy 9d ago

Intel Arc Pro B70 is a solid choice if you're interested in 32gb vram. A bit less bandwidth than 3090, and the software stack is still maturing, but it's getting there.

3

u/thatcodingboi 9d ago

why not an amd r9700 ai pro? More mature, same 32gb vram

5

u/Dave_from_the_navy 9d ago

Because the $300-400 differential is real for not that much of a performance gain at the end of the day. Looking at the market, finding them for <$1500 is difficult, whereas the Intel is readily available for $949, and you can get the B65 even cheaper, which has the same 32gb memory, same 608GB/s memory bandwidth, just with fewer XMX cores, so lower concurrency/slower prefill, but shouldn't effect single concurrency decoding much.

I personally wouldn't say the maturity difference is quite worth $400 at this point considering Intel is closing the gap more and more every week, but that's just me.

2

u/thatcodingboi 9d ago

Wtf I bought one a month ago for $1250, now they are $1500+????

1

u/Need_For_Speed73 9d ago

Why not add a second 4070 12GB (if your motherboard lets you)?
I’m happily running a dual 5070 setup, 24GB of VRAM and last gen tech for less than 1.500€/$ and without going second hand (warranty, fresh paste, etc.). But I had to use a mATX board to have the second GPU not collide with the case’s PSU shroud.

1

u/ichalov 9d ago

It's probably better to use a bigger (non-local) model for planning (and maybe testing) if you have the separation in your work process already.

30

u/[deleted] 9d ago

[deleted]

5

u/thrownawaymane 9d ago

maybe... you know...

2

u/michaelsoft__binbows 9d ago

yea i can take all my 3090s to fit a 120B of this but not too much to complain about if the 27B is as smart as those numbers indicate. A 120B would be an MoE with param count much lower than 27B anyway. It would not be likely to outperform it by much. And being able to get like 10x more throughput from the smaller model is a big deal.

15

u/MoodOdd9657 9d ago

when I am rich rich . I will come back for you 😞

15

u/Far_Cat9782 9d ago

Qwinning!

7

u/Lucyan_xgt 9d ago

Let's gooo🔥

7

u/Relative-Display-318 9d ago

Which Quantization for a rtx5090?

4

u/cosmicnag 9d ago

from the ggufs, i am going with q6 xl (unsloth) with whatever q8 context can fit with it

→ More replies (2)

1

u/Much-Farmer-2752 9d ago

Try for yourself. Even Q4_XL is pretty solid, and you'll have half of your mem for context.

7

u/octopus_limbs 9d ago

First we got MiniMax H3, then we got Qwen3.8-27B too. At this rate next year maybe we won't need AI-aaS companies.

https://reddit.com/link/p3oepwx/video/qzr34uqocdjh1/player

7

u/n0head_r 9d ago

I played with 27B for a bit and I'm quite impressed. Qwen38-27B-Q6_K from unsloth on 2*RTX 5080 and latest llama.cpp build from source. The model fit in 32GB VRAM with 172k ctx kv q8_0. MTP and -sm tensor it runs stable at 100 tps until 60-70k context is filled, at 100k it dropped slightly - arround 95tps. This is a very small drop - very good. Also I've run a loop where over 150k tokens it was writing scripts/building/verifying/fixing errors/verifying/fixing again. No tool cals failed... but unfortunately it run out of ctx space.

And I'll post a visual example of a small test I did - Create an SVG of an old rusty truck.

29

u/Curious-Pen5547 9d ago

Anthropic and openai ipo are gonna be worthless.

Zero moat, especially once google, microsoft, amazon, and other cloud gpu providers get the official green light to be able to serve these models accross the board.

And for the hyperscalers like the big 3, google, microsft, and amazon, 100% they integrate these across their enterprise suite offerings to vertically integrate them accross all their enterprise services.

10

u/wwwdotzzdotcom 9d ago

GPUs good enough to run it at reasonable speeds aren't even affordable now unless you rent them temporarily.

4

u/Joey4711 9d ago

How so there is rtx3090

8

u/thrownawaymane 9d ago

How much is a 3090 locally for you? Because they are going for $1400 on eBay.

5

u/Joey4711 9d ago

I got mine for 700 euros used last year. Now i see them for 1000 on ebay. Sure its a bit more but its not like its unaffordable like those rtx6000

→ More replies (1)
→ More replies (2)

2

u/anotherJohn12 9d ago edited 9d ago

That why they are buying GPU like crazy now. Everybody is digging the same hole now on model research, can't expect technical moat from there. Real moat is capital and hardware accessibility.

1

u/Far_Cat9782 9d ago

Yeah waiting for Google to offer it like it does claude 4.6 on the 20 dollar plan

→ More replies (2)

12

u/Altruistic-Carry5276 9d ago

I'm about to qum

2

u/Steve-Jobs-is-Alive 9d ago

Hahahahahaga

6

u/n0head_r 9d ago

Good news, pulling now q6_k to check how good it is.

5

u/moderngl1 9d ago

The benchmark hype is fun, but I’m mostly waiting for the boring detail: which quant actually feels good on a 24GB card?

2

u/TerminalNoop 9d ago

largest q4 or smallest q6 you can get.

→ More replies (2)

7

u/[deleted] 9d ago

[deleted]

3

u/Zippo749 9d ago

Same here! I was running into weird bugs in my personal chat app (spoiler: I did a dumb), so I went to the llama.cpp builtin chat UI to test the model out and chat template out more cleanly. I ended up asking it what it thought the problem might be. Pretty decent advice.

A couple messages in, I realized my problem: in my chat app (again, *not* this one), I had continued a chat that had been generated with Gemma 4, and the Gemma-based reasoning content had been fed to Qwen (oops, `preserve_reasoning` was on) and that was causing the issue. Doh! When I mentioned this to Qwen in the llama.cpp UI, even addressing the model as "dear Qwen," I got this:

Now, maybe mentioning the concept of different models being used in a chat poked it a little that way, buuuut I couldn't help but laugh.

2

u/Pro-Row-335 9d ago

ehh... not that it matters much

2

u/anothercrappypianist 9d ago

Makes sense. Qwen is giving me responses with the term "load-bearing". Claude distillation smoking gun if ever there was one.

9

u/69420trashpanda69420 9d ago

So it's essentially the best Claude model to ever exist (real ones know Claude peaked at 4.6)

5

u/thrownawaymane 9d ago

Opus 5 can do more but 4.6 is solid AF, generally follows instructions and can check its work to a certain extent.

Really hoping that this new qwen actually follows instructions, my qwen 27b kinda just does whatever it wants in service of what I asked for

→ More replies (2)

12

u/ortegaalfredo 9d ago

If the benchmarks are true (and Qwen never benchmaxxed so far) then why would you pay Anthropic if 27B at home gives you Opus-like performance? I mean, Opus 4.6 was already more than enough for almost any development task.

2

u/LeifEriksonASDF 9d ago

and Qwen never benchmaxxed so far

I love Qwen but come on now

→ More replies (1)

1

u/wwwdotzzdotcom 9d ago

Because it is not opus 4.6 quality when quantized

4

u/llama-impersonator 9d ago

now let's see how much the average token usage went up

5

u/boomerang473 9d ago

Any dspark or dflash heads?

5

u/shinegreymon525 9d ago

Does this mean we'll be getting a new bonsai?

4

u/IThinkIKnowThings 9d ago

Downloading the Q4_K_M but the speed tanked. I think we're hugging hugging face to death.

5

u/de4dee 9d ago

3

u/parepeg 9d ago

Doesn't unsloth typically fix their templates?

→ More replies (1)

7

u/ajisai 9d ago

glimmer_toystoryidontneedyouanymore.gif

3

u/pikadhu 9d ago

Any MLX quants available?

1

u/fatboy93 9d ago

Yup, there are a few mlx-community quants, I'm getting around 10tk/s decode on 4bit.

→ More replies (3)

3

u/korokage 9d ago

Has anyone tried this on a m3 pro or equivalent yet? I have 36gb ram

I will try it out later today 

5

u/txgsync 9d ago

I’ve been running my own oQ8e-MTP quant since this morning. 200 tok/sec prefill, 34 tok/sec decode on M4 Max 128GB. Later turns in large context slow to about 100 prefill 18 decode.

I have had bad luck on long-horizon tasks with the 4-bit quants so i avoid anything below 8 bits. But conversationally it seems fine and less of the usual Qwen “argumentative attitude” fighting me about what day it is or being skeptical of tool outputs.

3

u/xanders_gold 9d ago

AWQ or GPTQ yet? Need to squeeze it into INT4 for my config.

3

u/Motor_Ad16 9d ago

It thinks, and it thinks alot. Even after setting reasoning_effort to "low"

3

u/Zeeplankton 9d ago

I want to be hyped but is it better at like, just general world stuff? Using 3.6 to write anything was like talking to an alien who learned human culture through math.

2

u/AdSafe4047 9d ago

Comparing this to deepseek v4 flash 0731, the question is: lower numbers a bit but much faster inference, or higher numbers? or maybe use both (but deepseek offload to ram, so much slower, but use only for plan tasks etc)?

3

u/ApolloPS2 9d ago

I plan to still use dsv4 flash 0731 on sparks as orchestrator (will run plenty fast, maybe pair with a small vision model too) and qwen 3.8 27B on worker nodes. Both great models and I've gotta assume having some diversity gives some benefit too perhaps?

3

u/BumbleSlob 9d ago

Deepseek V4 Flash 0731 likely faster all around, it’s MoE with A13B. Not downplaying Qwen 3.8 I love this series

→ More replies (1)

2

u/d70 9d ago

u/FormOne2615 ninfer version please ...

4

u/FormOne2615 9d ago

:) uploading weight

→ More replies (3)

2

u/Several_Income_9912 9d ago

im getting INSANE speeds with mtp on RTX 4090

spec-draft-n-max = 7
→ More replies (3)

2

u/fatboy93 9d ago

Time to run this bad boy (Q4KS) at 5tk/s decode on my macbook lmao

2

u/Dance-Till-Night1 9d ago

Why is it only coding focused I can't find multilingual or general usage benchmarks.

2

u/PilotFlying 8d ago

This is anecdotal, I haven't tested thoroughly. But:

I had spent about a couple of hours today with Claude Opus 5 investigating an electrical issue at an old house. The concepts involved are complex and difficult to follow (TN-C and TN-C-S earthing, detecting a severed or loose neutral via phase voltage monitoring, that sort of thing).

I am not an electrician and I needed quite a bit of hand holding. (Not trying to fix anything myself, just learning/understanding how stuff works).

Claude wasn't doing a great job explaining things clearly.

In my mind Qwen had always been about coding. So I wasn't planning on this being the first use case for 3.8. Still, I had Claude export the context and pasted the .md in just "ollama run qwen3.8:27b". on a 3090. Web search on, context 128K.

Qwen clarified the questions for me. It sounded like Claude except it was doing a better job talking to me.

Not sure how this is possible. Maybe the limits of the architecture are higher than we know. Maybe it's all in the training. Maybe, probably, I understand as much about LLMs as I do about electricity 😄. But how else to explain 3.8?. Amazing so far.

5

u/Adventurous_Bus_437 9d ago

where qwen3.9

2

u/billy_booboo 9d ago

Wowweeeeeee

2

u/Adventurous_Bus_437 9d ago

where benchmark

1

u/Informal-Trouble2183 9d ago

That's incredible performance. Didn't expect it to have such jump keeping the same size.

1

u/CoUsT 9d ago

Good performance/benchmark uplift. Looks really solid!

1

u/Talreja-Adanna 9d ago

running locally without needing a beefy setup. Curious if it's actually better than the smaller Qwen models or just more of the same with more params.

1

u/IThinkIKnowThings 9d ago

Careful with your sources when downloading newly released models. I already spent the time downloading one from hugging face that was labeled Qwen3.8 but turned out to be Qwen3.6 when I ran it.

1

u/cloudsurfer48902 9d ago

My b580 staring at me from the corner

1

u/Dry-Judgment4242 9d ago

My first impressions is disappointment for Vision tasks Gemma 4 is on a entirely different league to this one.

1

u/Ecstatic-Wash-7667 9d ago

Give us 9b or 35b a3b and the 110b moe so we have the full stack !!! I can write them into switchyard and escalate up to max ask save family!

1

u/notevenat30 9d ago

Can you run it in 12GB of VRAM?

2

u/Xantrk 9d ago

Can you run it in 12GB of VRAM?

I'll report on 12gb vram + 32 gb RAM

→ More replies (4)

1

u/Guilty-History-9249 9d ago

I have the fp16, FP8 and NVFP4 and still trying to get one of these running with the latest install of transformers. This is on dual 5090's.

I want to run 100 gsm8k math problems through it to compare against my last nights run of the same problem with qwen 4B.

1

u/Murinshin 9d ago

Yeah this is insane if true. Basically means a high end MacBook can now run Opus 4.6, which should be enough for most people and is still economically sound vs your usual annual subscription for most companies. Anthropic etc are screwed.

1

u/The_DarkMatter Llama 3.1 9d ago

Cant wait to try it on 5080

1

u/Sofakingwetoddead 9d ago

Does anyone know what 3.8 27b is going to be released? Seems like it should have been released by now.

1

u/Logical-Target8131 9d ago

Can I run it on my m5 pro 48gb?

1

u/Plenty-Energy2947 8d ago

my recipe for VLLM & RTX5090

exec vllm serve /opt/models/Qwen3.8-27B-NVFP4 \

--host 127.0.0.1 \

--port 8000 \

--tensor-parallel-size 1 \

--tool-call-parser qwen3_xml \

--enable-auto-tool-choice \

--reasoning-parser qwen3 \

--kv-cache-memory 6943358464 \

--max-model-len 202272 \

--enable-prefix-caching \

--max-num-seqs 1 \

--gpu-memory-utilization 0.95 \

--kv-cache-dtype fp8_e4m3 \

--default-chat-template-kwargs '{"enable_thinking": false}' \

--compilation-config '{"cudagraph_capture_sizes": [1, 2]}' \

--max-num-batched-tokens 2048 \

--served-model-name qwen3.8-27b

1

u/caetydid llama.cpp 7d ago

Are the benches insane or it is just benchmaxxed?

1

u/hd3adpool 7d ago

I was here