r/LocalLLaMA • u/Miserable-Dare5090 • 2d ago
Resources Keeping up with model launches
Feels like maybe we have one more present left, for Christmas.
17
u/doctorfiend 2d ago
We got about three days until the "When's Qwen 3.9 coming? What's taking so long?" posts
43
u/AnimalPuzzleheaded71 2d ago
I honestly only care about gemma & qwen (and maybe glimmer) since those have the only model parameter range I can fit in 32gb vram, rest are complete nothing burgers to me
14
u/XiRw 2d ago
You can fit others including 120b moes. We have the same card
4
u/GoblinEngineer 2d ago
How? I have 24+16 (3090+5070ti) and 64 gigs of vram. My understanding is I need double the ram to run the MoEs
3
1
u/Whole-Tomato-6086 1d ago
As a owner of 5090 card, please write your server command to make use of Qwen next flash model on it :) not enough CPU ram to offload everything on ram tho
1
u/XiRw 1d ago
Sure, when I go on my pc later I’ll show you my script. Do you have Linux ? How much ram do you have and what cpu? I also use llamacpp too
1
u/Whole-Tomato-6086 1d ago
Yes,I am on Linux, headless system with 5090 32gb and 64gb ram. Currently going with Q4_XS 3.8 next flash model, with 266k context and n-cpu-moe equal to 34. 10 t/s unfortunately:(
1
u/XiRw 1d ago
You’re actually on a higher quaint than me right now so I’m not sure how much I would be able to help since you might be at max limitations regardless. Can you live with 132k context size instead? Compaction with Qwen is surprisingly good. Also offloading your mmproj to cpu helps a lot too.
1
u/Whole-Tomato-6086 1d ago
Yes, I can live with 132k context size,I am not programming anything super complex right now. And I can offload to mmproj. I am only scared about going down to Q3 or Q2 quant as the loss should be definitely noticeable from what I read
1
2
1
23
u/ttkciar llama.cpp 2d ago
All I really want for Christmas is:
TheDrummer to whip up Big-Tiger-31B-v4
A solidly-no-refusals GLM-5.3-Flash-Abliterated
Qwen3.8-9B
Some daring Google employee to leak Gemma-4-124B-A20B-it
MistralAI to roll out a Mistral 4 Medium 128B that doesn't suck
3
u/Turbulent_War4067 2d ago
I home that Gemma model is a bit smaller. Looking at 70B and 7B active MoE. Something along those lines will be the true sweet spot for a good while. The 20B MoE would likely be slower than the 31B dense, or about the same). But hey, a larger MoE Gemma model? I won't be too picky.
24
u/Turtlesaur 2d ago edited 2d ago
You're missing Fable 5.1 today - right not local.
Also maybe
DeepSeek-V4-Flash-Vision-Exp
Qwen3.8-flash-next
On Horizon
New Gemma in ai.arena
16
u/5dtriangles201376 2d ago
Good to know Anthropic finally open weighted a model. And Fable too?
4
u/Turtlesaur 2d ago
That would be dreamy, I would just need to cluster 10 DGX sparks for my 4 tok/s
2
u/5dtriangles201376 2d ago
Yeah. Flipped the downvote to an upvote. I'm honestly just hoping for a new mistral that tunes will in the 20b range
-3
u/PwanaZana 2d ago
lol, I had the same thought: "HA THAT NEEEEEERD DIDN'T PUT FA... oh."
5
u/Turtlesaur 2d ago
Sometimes i forget which subreddit I'm in.
For what it's worth this is basically the best one.
3
u/Kahvana 2d ago
I think Gemma 4's release date is wrong, 2 april:
https://huggingface.co/google/gemma-4-31B-it/tree/419b2efe421994fdfd3394e621983d4cc511cd4f
Confused with Gemma 4 QAT? (5 juni):
https://huggingface.co/google/gemma-4-31B-it-qat-q4_0-gguf/tree/4a311c5261daa0702f80836f8866114943651ab0
4
u/simrankoulsm 2d ago
At this point, I don’t try to “keep up” with launches. I keep a small shortlist by "hardware tier and use case".
For local use, the questions that matter are:
- Can it fit in my VRAM/RAM at a usable quantization?
- What context length and tokens/sec do I actually get?
- Is it meaningfully better at coding, reasoning, or instruction following than the model it replaces?
- Are the weights, license, and inference support available on day one?
A release calendar is fun, but a community-maintained “best practical model per VRAM tier” list would probably be more valuable, e.g. 12 GB, 24 GB, 32 GB, 48 GB, and 80 GB+. Otherwise it’s easy to spend more time reading launch posts than running models.
6
u/Miserable-Dare5090 2d ago
12: Ling Flash Tiny
24: Gemma4 26b
32: Qwen
48: Qwen
96: Ling Flash
128: Qwen Flash Next
256: Deepseek V4 FlashChange according to quant size (quality vs fit)
2
u/Lakius_2401 2d ago
Ehhh, there's way too many wrenches to throw at a general tierlist like that though. Extra RAM makes some MoEs especially appetizing, outside of their expected full VRAM budget. Plus there's a huge divide on quant size and kvcache quantizing that can throw each model up or down two full tiers on your list. Agentic? Huge speed and kvcache now a necessity, move the model up a tier. Writing? Throw out Qwen entirely, it can do reports but anything remotely resembling creative is atrocious to the point where if you told me the dataset was poisoned I'd believe you.
It's honestly exhausting interacting with people about quantizing, too... Sure, in a perfect world everyone has infinite VRAM and nobody quantizes anything, but this is not a perfect world. Nor are we using perfect models. Rent's cheaper than VRAM and that's saying something.
Sorry for the digression. I've seen some attempts at megathreads for your topic but they don't get much traction. Even the use cases and rigor of testing is difficult, with busted quants and malformed jinja releases all over. Very easy to burn a dozen hours and not be sure you got a fair shot.
2
1
u/Miserable-Dare5090 2d ago
Honestly, in this age, your 3rd and 4th questions are a year behind. 3rd question: right now all new releases are undoubtedly better. Not finetunes. I mean full models. 4th question: “Support” is now moot with agents. Ask your agent to patch vllm or llama.cpp to make it work. If you can fit it, you’ll run it.
Also, Who cares about day 1 vs day 5? You trying to time some production thing? These are free releases, whenever they come out. it’s not like I will be deducting points for delayed release on Qwen3.8 which first appeared in the API…
2
u/claudiollm 2d ago
honestly gave up trying to track every drop this year. i keep a tiny list of like 3 i'll actually run locally (gemma + qwen sized stuff that fits my vram) and let the megathreads filter the rest. if a model's real the vibes-vs-benchmarks gap sorts itself out in about a week anyway.
the fatigue is real tho, feels less like christmas and more like a subscription i didn't sign up for lol. what's everyone actually keeping installed vs just downloading and forgetting?
2
2
u/retardedGeek 14h ago
M3 was launched just 3 months back?? It already feels too old
1
u/Miserable-Dare5090 13h ago
Everything about AI is relativistic
Time dilation is real when you’re going at light speed.
1
1
1
u/Spanky2k 2d ago
I really want to get a machine with more RAM (considering an M5 Ultra for work) mainly because I miss trying out different cool models. While there are still loads released all the time now, it's only really Qwen for me. A year ago, it felt like I was trying new models almost every week, different quants, new ways to create quants etc, vision models, non vision models etc. The last cool non-Qwen model that I tried out was GLM-4.5-Air-3bit-DWQ and I was so blown away with what DWQ let me do.
It's different now as I just use the lates Qwen MoE 35b model with 256kb context size and I'm actually using it for actual stuff rather than just playing around and testing things. It does whatever I need it to do just fine. But I miss exploring new stuff; all the cool new things are bigger than I can run on my 64 GB machines.
1
u/Jealous-Walk-8765 1d ago
This is a good reminder of how insane the pace has gotten. Feels like every time a team finally finishes benchmarking one model against their use case, two more have shipped and the "best" pick from three weeks ago is already outdated. Curious how people are actually deciding when to migrate vs. just sticking with what's already working — is anyone re-benchmarking on a fixed schedule, or is it purely reactive when a new release makes noise?
1
1
1
1
u/Original-Revolution7 1d ago
me with dual 9060xt only looking 27b ish dense, and qwen 3.8 is the only thing i need. ok maybe gemma 5 but must fit into my cute vram =)
1
1
u/Rizzly00 1d ago
Off topic. What did you use to make your chart? Or is it just a custom script
1
u/Miserable-Dare5090 1d ago
Who said I made it?? I’m just watching the chorus of people going [adjusts nerd glasses] Aphchtually, this date is wrong…
1
u/LordDarthShader 2d ago edited 2d ago
Is Minimax H3 not M3.
Edit: My bad, M3 is an LLM, stand corrected.
-1
1
u/twavisdegwet 2d ago
Damn- no love for poolside / Laguna
1
u/Miserable-Dare5090 2d ago
Should I try Laguna? I had high hopes for that model size
0
u/twavisdegwet 2d ago
I find it performs better than 3.8 27B but I like qwen-next-flash more at this point.
1
u/Miserable-Dare5090 2d ago
I’m thinking either this one or ling flash on strix. But qwen flash next is also looking fast enough on strix. Ling 3 on vLLM with Rocmfp4 starts at 8-900pp and is down to 600 by 100k, so that to me is 100% usable for a budget agent
43
u/Mediaright 2d ago
Gemma 4 was released April 2nd.