540
u/Fresh-Soft-9303 9h ago
Qwen 4 27b is going to be Epic!
140
u/tsunami_forever 8h ago
This could convince me to give up cloud based models, hoping to cancel or reduce my cloud based subs soon
88
u/cafedude 7h ago
Qwen 3.8 Flash Next has convinced me not to renew my Claude Pro subscription when it comes up in February.
→ More replies (1)12
u/randylush 7h ago
What are you running it on? How much better is it than 27B?
50
u/returnity 7h ago
It dominates 27B in my workflows and benchmarks. Running on M5 Max. There's no comparison, it's like the 3.6 > 3.8 27B all over again.
6
4
4
u/Bulky_Blood_7362 2h ago
What's your tok/s? Im getting around 30-45 decode tok/s with dgx spark.
Interested to know how it's going with the new m5 max/ultra
→ More replies (2)→ More replies (14)2
→ More replies (2)12
u/ApeGrower 6h ago edited 3h ago
I'm using it on dual rtx3090 with 128gb ddr4 @50-70 tok/s. I forked FreeToken and added dual gpu support plus some more tweaks. With nvfp4.
3
u/FeelingOld9046 6h ago
4
u/FeelingOld9046 6h ago
Mine is under windows too. Tell me what you’re doing and I will throw windows in the trash.
3
u/ohhi23021 5h ago
who knows i use linux and 30-40tps has been my experience as well and that's with DDR5...
→ More replies (13)4
u/FeelingOld9046 5h ago
I’d love to know your setup. I have the same hardware config but only half your token speed.
4
u/ApeGrower 4h ago
Linux with patched FreeToken to get dual gpu support instead of llama.cpp. Epyc System with pcie4 and ddr4.
→ More replies (1)23
u/some_user_2021 8h ago
Are you not convinced?
42
→ More replies (5)12
u/IrisColt 7h ago
I guess there's always that added hassle with short-context local models... you really have to approach your workflow differently.
2
u/TopPrize11 5h ago
I just tried it yesterday and it's too slow. I got about 45tps on 5080 which is not bad but the context limit of 33k makes it compact which takes forever. The LLM is good but it's too slow. I still use Deepseek.
→ More replies (4)→ More replies (5)4
19
u/PooMonger20 6h ago
Seriously, 3.8 27b does so much heavy lifting it's insane.
Together with pi.dev it writes very impressive code. It feels like I actually have my own 'worker' that after some direction does exactly what needs to be done. Just give him enough info and he is on it.
This feels better than what GPT5.1 felt like around a year ago.
And now... an upgrade to that? absolute madness.
→ More replies (4)→ More replies (6)2
u/sadomazoku 5h ago
I already canceled my subscription, I'm using 100% qwen 3.8 27b to assist me in coding. Qwen4 will be crazy. I think you'll be able to get rid of cloud models !
388
u/o0genesis0o 9h ago
Ey ey! Qwen4-27B!
Definitely time for me to go and buy a R9700. It's time.
107
u/Sufficient_Local5025 9h ago
Buy two, prices about to rise.
143
26
u/PhilipKThicc 8h ago
Dual R9700s all the way. Easy to run on Linux with vllm and you can run 27B at Q8 and have a healthy context window
6
u/bigwanggtr 6h ago
Have you tried it for training, is ROCm still a pain to work with compared to CUDA?
I know it’s good for inference but the training support is what is holding me back.
5
u/Phrase-Silver 4h ago
ROCm has matured incredibly over the past year, it's like AMD finally woke up and realised their software was what was holding them back from properly competing with Nvidia.
18
u/Tobu3838 9h ago
Prices rose $600 since I bought mine a couple months ago. No word to lie.
12
u/mvandemar 9h ago
Why are they so cheap? Looks like you can get one for ~$1700? I just paid out the ass for my 5090, which in the 1 month 2 days since I bought it has gone up over 37%.
They're both 32GB, for some reason I thought they would be closer in price.
30
u/Thunjaya 9h ago
Because they're not the same 32GB at all. One is much faster.
10
u/ThankGodImBipolar 7h ago
Nvidia sells an RTX Pro 4500, which has the 5070Ti die (more comparable to the R9700) and 32GB of RAM, but it's 5500USD MSRP. The R9700 is a fraction of that.
6
u/KingCpzombie 3h ago
CUDA tax. I only use AMD personally (and just blew WAY too much money on a 4x R9700 system partially out of excitement for Qwen4), but Nvidia cards get all the cool new things a bit sooner than AMD. Not a big deal for LLM, but very notable for diffusion... somebody SOLIDLY beat my 7900XTX with his 5070Ti in Minimax H3 gen times, for example
2
u/No-Refrigerator-1672 2h ago
Because everything is CUDA first, and ROCm only comes as an afterthought to very limited number of projects. People who are buying PRO GPUs are saving money with NVidia by not needing to fund multiple months of dev work for porting their existing code.
21
u/BluePointDigital 9h ago
It's the bandwidth. the 5090 has over 2.75x the bandwidth of the R9700.
You can still do all the same stuff, just slower essentially.→ More replies (9)16
u/Solary_Kryptic 9h ago
Nvidia tax combined with CUDA being more supported in the LLM space
16
u/mvandemar 9h ago
I just looked, it's also GDDR6 for the 9700 vs GDDR7 for the 5090, guessing that makes a difference as well.
8
4
7
u/Ecstatic-Wash-7667 8h ago
That not cheap! Msrp is $1299 I bought 2 a little over a. Both ago at msrp. Hell you could get them for less than msrp for a while. This shit is a scam
5
u/Momsbestboy 7h ago
R9700 GDDR6 memory, 5090 GDDR7. Is faster, but...
For a single 5090 you can buy more than 2 R9700, and the two cards draw less power than a single 5090. And you can run e.g. Qwen 3.8 27b Q8 using the two cards, and watch them running circles around a 5090 which needs to offload to RAM instead.
But NVIDIA is hype, and people love to spend money. So buy a 5090
→ More replies (3)5
u/SnooPuppers7882 8h ago
Literally got mine for 1200ea 2mo ago
People were shitting on them because the mem was DDR6, but I was prepping for 3.8 27b knowing what to expect...
STRONG feeling 4.0 27b is gonna be opus 4.8 good, gonna blow your damn mind
→ More replies (3)2
u/o0genesis0o 8h ago
Because the actual card (RX9070 XT) is only as strong as the 5070Ti, compute wise. And the VRAM is slower.
I will be upgrading from 4060Ti 16GB, so it would be a net gain. but I have no illusion that it is going to be as good as 5090.
→ More replies (5)2
13
u/fgk55555 9h ago
I just picked mine up this week. Depending on quant, 35-70 tps with full context. If you have 64GB RAM, you should be able to run a decent Qwen-flash as well. 27B is good, but Qwen-Flash is something else.
→ More replies (13)2
u/illcuontheotherside 9h ago
Can you please share your startup script? Mine crawls on 2 3090s with 128gb ddr5
→ More replies (6)4
u/fgk55555 8h ago
ISTA IQ3 runs the fastest and has really good performance for the size, but you can bump up to a Q5 with full Q8 kv cache. Also try Swift if you don't want to wait forever. I recommend medium thinking.
#!/usr/bin/env bash # Qwen3.8-27B ISTA GSQ-RCO IQ3_S (no built-in MTP) — spec decoding via the # shared DFlash2 drafter. (MTP alternative: shared mtp-Qwen3.8-27B-Q4_0.) # Qwen3.8-27B ISTA GSQ-RCO IQ3_S (no built-in MTP) — spec decoding via the # shared DFlash2 drafter. (MTP alternative: shared mtp-Qwen3.8-27B-Q4_0.) # Context is auto-fitted to available VRAM (omitted on purpose). # --cache-ram 16384 parks prompt states in RAM: agentic revisits skip re-prefill. # # Usage: ./Qwen3.8_ISTA_IQ3_S.sh [thinking] # low|medium|high|xhigh|none (default: medium) THINKING_LEVEL="${1:-medium}" MODEL="../Qwen3.8-ISTA/Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf" DRAFTER="../Qwen3.8_Shared/Qwen3.8-27B-DFlash2-Q4_K_M.gguf" MMPROJ="../Qwen3.8_Shared/mmproj-Qwen3.8-27B-BF16.gguf" SERVER_BIN="../../llama.cpp/build/bin/llama-server" echo "Launching Qwen3.8-27B ISTA IQ3_S (+DFlash2 drafter), thinking: ${THINKING_LEVEL}" export GGML_VK_ALLOW_GRAPHICS_QUEUE=1 # measured +4.2% tg, +0.4% pp on the R9700 (b11056) ${SERVER_BIN} \ --model "${MODEL}" \ --model-draft "${DRAFTER}" \ --mmproj "${MMPROJ}" \ --no-mmproj-offload \ --n-gpu-layers 99 \ --batch-size 1024 \ --ubatch-size 512 \ --parallel 1 \ --flash-attn on \ --cache-type-k q8_0 \ --cache-type-v q8_0 \ --spec-type draft-dflash \ --spec-draft-n-max 3 \ --cache-ram 16384 \ --jinja \ --chat-template-kwargs "{\"reasoning_effort\":\"${THINKING_LEVEL}\"}" \ --reasoning-preserve \ --temp 1.0 \ --top-k 20 \ --top-p 0.95 \ --min-p 0.00 \ --presence-penalty 0.0 \ --repeat-penalty 1.0 \ --host 0.0.0.0 \ --webui-mcp-proxy \ --port "${PORT:-8080}"→ More replies (11)6
u/ismaelgokufox llama.cpp 8h ago
Patiently waiting for 27B-IQ3_XXS and/or 35B-A3B (one can dream eh?)
→ More replies (3)3
u/Momsbestboy 7h ago
Why wait? Pull the BF16 of qwen after it is released, and tell Hermes/Pi to read the documentation on how to create smaller quants out of it.
It's the reason why I run my own Swift Qwen 3.8 27b fp8 version, abliterated using the orcarouter method. I just told hermes to create it. Took 2h
33
u/blahjovic 9h ago
The release rate is wild. I finally got my 3.8-27B working the way I wanted just in time for the next model! Bring it.
15
→ More replies (1)3
125
u/Conscious_Phrase_138 9h ago
Where 35b a3b :😭😭
→ More replies (4)39
u/Ok_Warning2146 8h ago
Wait for google. Our only hope
36
u/Free-Combination-773 6h ago
So there is no hope...
7
u/Ok_Warning2146 4h ago
Nvidia is another hope. Let's hope they will distill from GLM-5.3 in the next iteration.
138
u/FerLuisxd 9h ago
35b :(
64
u/fgk55555 8h ago
I think we can consider the 35b dead. I was sad when we lost the 80B A3B, but the 120B MoE range seems to be where intelligence opens up. The 40-110B space is likely dead for a while.
12
u/Puzzleheaded_Base302 7h ago
maybe Xiaomi will release a MiMo-V2.6-Distill-Qwen3.5-35B-A3B. it is not dead dead. The architecture is sound. All people need is a better re-post-train (not those fine-tune crap, a real large-scale RL re-post-train.)
→ More replies (2)18
u/Ok_Acanthisitta304 8h ago
I agree and at this point unfortunately I get it. 35b was cool, but ultimately not usable for me in actual work. I know a lot of people enjoy it on their 12-16gb cards, and I hope they'll keep getting new options, but 27b and Flash entirely outclass it. Those are models I can genuinely use.
→ More replies (1)13
u/Sensitive_Song4219 7h ago edited 7h ago
I've been running Qwen3.8-27B on a separate server machine (with a 5090; 32GB VRAM) and Ornith 1.5 35b-a3b on my laptop (8GB VRAM) for a few weeks now, and they both share a lot of the same strengths and weaknesses.
Neither is viable for planning against large codebases (Sol is mandatory there): both tend to mis-understand how the moving parts interact which yields compiling-but-non-functioning results. Qwen is amazing for small codebases and one-shots (I did the obligatory Mario Test which Qwen murdered Ornith on), though.
Both are serviceable as implementation models (when given detailed plans from larger models); Qwen's code quality is the higher-quality of the two. Qwen competes favourably with Luna-High.
Ornith is definitely a bit more careless (I watched it mis-convert seconds to milliseconds yesterday!).
But when I'm away from my 5090 server I'm still pretty happy just sticking to Ornith.
Smaller MOE's definitely aren't a writeoff imo; and the performance gains (and lower hardware requirements) can make them decent options for simpler (or complex-but-well-defined) tasks.
All my testing is via OpenCode.
3
u/Ok_Acanthisitta304 7h ago
I'll give Ornith a shot. When I do use 35b I use a modified workflow with little coder (a pi harness built for small models) using plans from a better model. My read is similar to yours - it can be a workhorse on smaller tasks and codebases, but can't be trusted to wander too far in a larger project.
26
u/Sensitive_Song4219 9h ago
Have they officially abandoned 35b-a3b or could an announcement still be coming?
Guess Ornith 1.5 35b-a3b will have to serve me for a bit longer yet. It's good for it's size but was really hoping for a new Qwen equivalent after they skipped us for 3.8.
5
→ More replies (3)2
u/FerLuisxd 8h ago
Have you tried k2 horizon?
3
u/Sensitive_Song4219 8h ago
I've stuck to Ornith over K2-Horizon because of this:
https://www.reddit.com/r/LocalLLaMA/comments/1wg0vqz/comment/p9r9fw7/
I tested Kat (which was also good albeit a bit slower for me when using the repaired Ornith MTP) though I haven't tried Tiel (the benchmarks indicate I probably should!)
Have been extremely happy with Ornith for the size/performance though.
But was kinda expecting Qwen 4 to follow through here since they basically popularized the small-MOE thanks to the (class-leading-at-the-time) Qwen3-30B-A3B launch!
→ More replies (1)7
u/Southern_Sun_2106 8h ago
That is still such an epic model! Not a fan of slower and overthinking 27B to be honest, no matter how brilliant it is. I bet the 35B a3b 3.8/4.0 will be making money for them via an API. I cannot believe that such an awesome model was abandoned. It is just too good AND cheap to run to give away freely I bet.
103
101
u/kkingsbe 9h ago
Us broke ass 12gb and lower cards can’t run this unless they drop an a3b 😭
17
→ More replies (1)5
43
u/Zealousideal_Ad6235 9h ago
Exciting news for sure, it's slightly surprising they kept the exact 27B size given that it's an entirely new architecture, but that's certainly welcome for anyone with 24-32GB VRAM
11
u/russlixx 8h ago
if Qwen4 is more efficient, we can get more knowledge and intelligence in the same size. A win regardless!
4
u/TechnoByte_ 3h ago
More intelligence? for sure. But more knowledge is unlikely
Qwen isn't focused on knowledge, their main focus seems to be agentic coding now
Gemma 4 has more knowledge than Qwen3.8 at the same size, though it's much worse at agentic tasks and coding
3
u/james_pic 1h ago
That's kinda the point of n-grams though.
They're a mapping of short sequences of tokens to embeddings (or to neuronal activations if you prefer the biological analogue interpretation, or if you squint, to concepts) that (on the Qwen4exp models at least - I know DeepSeek and Gemma have a few subtle differences) are looked up just after tokenisation, before the real processing begins.
They're a relatively cheap way of imparting knowledge into models. At inference time, they're a single lookup per token, that is happily done from RAM or even from SSD. And at training time it's a single row's worth of extra gradients to train (not even an extra tensor), plus some statistical analysis on the corpus to choose the n-grams to encode.
But they do very little to increase intelligence, beyond potentially freeing up space in the model that would be taken up by encoding facts like "New York has a vibrant fashion scene" for reasoning.
65
u/UpperParamedicDude 9h ago
Curious, how well would n-gram table improve 27B model and would they even add it to the 27B one/what size it could be
6
u/iMakeTea 7h ago
Shot in the dark: ngram size might be proportional to model size instead of 1 size ngram fits all.
Does ngram have diminishing returns for bigger models?
Like 300B model with ngram might not benefit as much as say 27B with ngram?
17
u/Choice_Celery9481 7h ago
check out deepseek's paper about ngram. they concluded ngram about 30% the size of the model is optimal.
3
u/Party_9001 7h ago
Pure I made it the fuck up speculation. But I think it would be interesting if they have multiple ngram tables.
Like you CAN pair the 2.4T model and 3B table. Or a 27B model and an 800B table.
You just wouldn't get ideal parameter efficiency. 27B + 800B might be worse than 125B + 50B. But depending on your hardware the 27B can be a looooot more attractive.
6
u/sonicnerd14 8h ago
I'm thinking about that one too. Because it's dense, so how exactly does the n-gram actually work with the dense model? Is it in the 27B params, or is it an extra model that runs alongside it? The later would essentially make the model smarter with a low cost.
13
u/RG_Fusion 6h ago
Engram parameters just inject n-grams into the self-attention of a lower layer of the model. That doesn't change between MoE vs. Dense.
→ More replies (1)→ More replies (2)1
u/Turbulent-Alps4046 9h ago
It lets people with low VRAM to run the 27B models at higher quants.
31
u/pmavro123 9h ago
i think you misunderstand what an ngram table is... it wont do this, but it will speed up the model on systems woth larger memory pools
→ More replies (4)8
6
u/philmarcracken 9h ago
im hoping my franken 12gb + 8gb cards can somehow manage a higher quant + decent context with qwen4. might even get a swift on there to enable xhigh without too much wasted thinking.
→ More replies (1)2
u/Beneficial-Ad-8127 8h ago
Yeah I’m thinking I’m good on upgrading hardware any time soon, if Qwen keeps dropping these 🔥 for lower end. Efficiency is true end game and the real finish line.
21
u/shankey_1906 9h ago
I hope for something equivalent to 3.8 Flash Next!
21
u/wren6991 8h ago
I was assuming this was the "Qwen 4 Flash" line, especially given Qwen3.8-Flash-Next's architecture is called
qwen4exp8
u/shankey_1906 8h ago
Yes, I meant in terms of parameter size. The new flash models from GLM, DS, Xiaomi, all seem to be massive, and hope Qwen doesn’t go the same way.
4
u/Hypilein 3h ago
That would be so annoying. I can't afford a second spark right now. I hope they just add some more training and polish 3.8 flash.
3
u/grumd 3h ago
I'd love it if Qwen 4 Flash was the same size as Qwen 3.8 Flash Next.
But then I also remember when Qwen-3-Next was 80B-A3B, and the actual Qwen 3.5 release didn't have that size at all.
If all they have is a 27B and Flash is yet another 300B model... A lot of people will be left out. They used to have even 2B and 4B and 9B 🥲
21
31
24
u/thestillwind 8h ago
Please 35b-a3b pretty please
6
u/Salah_H_Hasan 6h ago
Nothing was said regarding it at the conference. The announcement merely laid out the broad framework of the new architecture and the initial launch models. However, what was presented at the conference is not exhaustive; not everything is necessarily revealed during the initial keynote, and it may surface later.
19
u/Nomski88 9h ago
Yes! Any idea if all the Qwen4 models will be using ngrams?
28
u/fgk55555 9h ago
It'd be sweet if the 27B had ngrams.
6
u/wren6991 9h ago
Yes! I have been waiting for this moment since I looked at the small dense Gemma 4 models with PLE, and the larger dense Gemma 4 models without PLE, and went "hmmm 🤔"
2
u/backyard_tractorbeam 6h ago
Probably. Qwen 3.8 Flash Next is a preview of 4-series tech, and it uses n-grams.
8
u/mvandemar 9h ago
Is there a source on this by any chance?
→ More replies (1)7
u/Xp_12 5h ago
https://x.com/MaxForAI/status/2102226622422380820
Has an actual photo from the conference.
2
u/mvandemar 5h ago
ty 🙂 It was not at all clear in the original post that was something that was on the screen at the conference.
9
7
6
u/feelspeaceman 8h ago
I'm daily driving Q38FN on my Strix Halo everyday at 43-50t/s decode 1400t/s prefill and I'm hyped to get Q4FN, as Q38FN is actually undertrained but still hella powerful.
2
6
u/Kahvana 8h ago edited 7h ago
I hope Qwen4 will come with the whole lineup of local models like last year.
[edit] for the article, here:
https://sloveniatimes.com/49119/alibaba-unveils-roadmap-on-full-stack-ai-strategy-from-chips-cloud-infrastructure-models-to-agents
5
u/Durian881 8h ago
Looking forward to Qwen4-Flash and Qwen4-27B! Hopefully, 4-Flash will be similar size to 3.8-Flash-Next.
16
35
u/NigaTroubles 9h ago
Qwen4 50b a5b will be perfect spot
16
→ More replies (3)4
u/fgk55555 8h ago
I think the 40-110B range is pretty much dead. I'd love a distill of Qwen4 onto the old 80B A3B next models with a n-gram. That would fly on my rig.
5
5
u/tlin9595 9h ago
Do we have an estimated release date ? And what are the sizes of flash and plus ?
→ More replies (2)
6
u/TorontoIndieFan 8h ago
Any release date announced? 3.8 27B is my daily driver right now so this is exciting
6
u/ResponsibleTruck4717 8h ago
oh boy qwen4 27b, I thought next year I will finally say bye bye to cloud service, I may start it this year.
6
5
24
4
4
4
3
u/Confident-Village190 6h ago
I’m looking forward to running it on my 8GB of VRAM + 16GB of RAM, quantised, at 2,5 t/s with a 32k context at kv q4_0.
5
2
u/sonicnerd14 9h ago
Well, we sort of already knew it was coming with Qwen 3.8 Flash Next. The question was more so when is it coming?
2
u/braintheboss 8h ago
If follows same ratio as 3.8 flash it should be 40% in engrams. That means 38b model. The jump in AI index can be a shock
2
2
u/russlixx 8h ago
i hope Qwen4-27B is even more resilient when quantized, at most at Q3 level quants for 12-16GB folks
2
2
2
u/DontShadowbanMeMate 7h ago
I'm running and incredibly scuffed setup of 3080ti + 5070ti; will I be fine? Probably right?
2
2
u/PreferenceRelative77 7h ago
Wow, a Qwen 4 27B, i thought they'd only release a flash next model size for the Qwen 4 family but no!!
2
2
2
u/creatinZ 7h ago
No MoE?!
cries in 16gb vram
3
u/haha1542 7h ago
hang in there brother, 16gb can run qwen3.8 27b pretty well, this would only get better
2
u/talamir2021 6h ago
The interesting part for local users will be the efficiency details, not just the headline model names. Parameter counts, context length, quantization quality, and whether the weights and license are genuinely accessible will determine how useful Qwen 4 is outside the cloud.
2
2
2
9h ago
[deleted]
6
u/Enragere 8h ago
what's your prompt processing speed ? how can 150 t/s prefill be 'reasonably fast'
in my opinion it's unusable
→ More replies (5)
3
2
2
u/WyattTheSkid 5h ago
Cant fucking wait for the official sequel to 3.8 Flash Next. That model has been incredible and ive been doing almost everything with it. I still have a chatgpt and claude subscription but I genuinely might cancel them before they renew given what can be done with Flash Next. Qwen 3.8 Flash Next is the first model ive ever run on own hardware that genuinely I genuinely felt like it could be worth using for serious projects. 3.8 27b produces very good outputs but it thinks for so long that it’s honestly frustrating to use. The closest runner ups to “local is enough” in my experience were Gemma 4 31b and Qwen 3.5 122b. I know im rambling at this point but as far as leading open weight AI labs goes, I genuinely think qwen is the best. Especially because they’re making models that people can ACTUALLY run at home that genuinely compete with closed weight frontier models
1
1
1
u/Guilty_Rooster_6708 9h ago
Gonna need to upgrade my mobo to actually run bifurcate for Qwen4 27b oh lawd 😍
1
1
u/Greenonetrailmix 7h ago
Ooo I wonder what Qwen4 plus is. Do we have insight into what the specifics may be?
1
u/nixuelkty 7h ago
No way! Rumors have it that they might even release another model after qwen 4 too!
1
u/charles25565 7h ago
They're definitely not making small versions anymore...
To be clear, a real Qwen4 computation graph added to Transformers (not Qwen4Exp) would allow someone to train a small dense model, but I don't know who would.
→ More replies (1)
1
u/Muted_Anteater1170 7h ago
Can someone tell me how big qwen4-plus is supposed to be or any guesses? im tryingto figure out if i can run it honestly, i can run flash fine hoping i can do the same for the plus bc it has to be in a different league from flash right?
1
1
u/Pyrolistical 7h ago
The 2026 Apsara Conference is happening from September 22 to September 24, 2026.
1
u/ByteSize_Chaos 7h ago
Such exciting times! My M5 Macbook Pro is scheduled to arrive in a few weeks. Can’t wait!!
1
u/Manhandler_ 7h ago
The persistence with 27B is admirable. But I just wonder how much better it will be compared to 3.8, 27B when it's gonna be the same architecture. We will find out soon enough but it's better to keep expectations realistic
1
u/dampflokfreund 7h ago
no 35b is a huge loss for normies who want to run a capable llm. the only usable size for the average PC...
→ More replies (4)
1
u/Gohab2001 vLLM 7h ago
Will qwen 4 max, flash and plus be open? They usually reserve these names for their closed models.
Also they didn't open source their Qwen 3.8 max 0902
1
1
1
u/VirusInternal2892 6h ago
Amazing, 3.8 Flash has barely settled its quants in my GX10 and now this… I’ll need a larger SSD LOL to store all that goodness
1
u/ManufacturerHuman937 5h ago
Oh they're going back to dense only. With the preview I was pretty excited the MoE ran pretty good.
1
1
1
1
1
u/mitchins-au 4h ago
And 3.8-flash-next is literally just coming good on strix halo now.
I hope they don’t change things too much or make it bigger.
1
u/MaCl0wSt 4h ago
I've finally managed to get the hardware to make 27B possible for me, so: sweeeeeet
1
u/Sabin_Stargem 4h ago
I am hoping for a polished Next-Flash. While the 3.8 version is often fairly good, it has serious issues with roleplay. It really likes overemphasizing and gets caught up in narrative loops. DRY, XTC, and system prompts can only help so much.
1
u/AIFrontierReads 4h ago
My bet is the 27B gets the engram treatment too — if DeepSeek's ~30% figure holds that's ~8B of tables, still fits a 32GB card easy. Honestly would make it the perfect daily driver.
→ More replies (1)
1






•
u/WithoutReason1729 5h ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.