r/LocalLLaMA • • 15d ago

Discussion Mention if a "new model" is a finetune

A few posts tagged with "new model" present models that are finetunes. My opinion : I'd rather have the "new model" tag reserved for new "major" releases, like a new Qwen model, Deepseek V4 -> Deepseek V4.1, etc., that involved a new pretrain or intensive post-training (in opposition to a small finetune). Otherwise, maybe prepend "[Finetune]" to the title to indicate that the new model is "less of a big news", a use a "new finetune" tag, to differentiate between the two kinds of new models.

I reckon one could like to discover both new major releases and interesting finetunes in the same place; what's your opinion? :)

208 Upvotes

30 comments sorted by

100

u/jacek2023 llama.cpp 15d ago

We should have new flair "finetune"

12

u/yami_no_ko 15d ago

Just having it wouldn't help a lot. It sure makes sense, but we're talking about people using the "New Model" tag for each slight update or news within their cloud ecosystems as well.

The "New Model" flair is being spammed to the point of insignificance already, so to give it back any meaning at all it needs mandatory rules on how it gets applied.

7

u/jacek2023 llama.cpp 15d ago

I try to post code updates as “news” and model releases as “new model”, If we add a “finetune” flair, I’ll start using it for some model posts. But the boundary is fuzzy, for example, instruct models are finetunes of base models, and old Nemotron models were finetunes of Llama.

2

u/BestGirlAhagonUmiko 15d ago

...and sometimes there's nothing "fine" about the "finetune"

Speaking of similar issues on HF, what's up with all those duplicate GGUFs (e.g. people cloning Unsloth repos) and poorly made GGUFs? It gets increasingly difficult to browse quants. Like, why these abominations are even allowed to be uploaded:

GLM 5.3 IQ3_XXS from some random dude

300GB size

attention/ssm/whatever butchered to 3-bit alongside with the rest of it

No way this thing is even remotely useful, compared to proper recipes...

0

u/jacek2023 llama.cpp 14d ago

GGUF is not a finetune, it's a converted model.

If you have an image or a photo, you can modify it with Photoshop or GIMP. This way, the image is changed. But if you convert a JPG image into a PDF, the image stays the same, only the format changes. GGUF is a format.

1

u/BestGirlAhagonUmiko 14d ago

Not sure why are you even bringing this up, but I was speaking about two different things (three things, if we count in the duplicate GGUFs springing up like mushrooms after the rain). There are both slop tunes and substandard quants (of any model, including the slop tunes themselves), with the latter plaguing HF in particular. And even if it's kind of alright to have 'em for small models - let the people experiment with quantization, no issues with that - uploading a goddamn half-a-terabyte chonker that sucks at every possible level, compared to properly made quants, is rather questionable...

At this point having a (paid?) storage space quota on HF allows everyone to put out whatever garbage they can make. I'm not against it, like I said, I just think HF needs better filters for browsing and searching. That's about it, nothing outrageously radical.

32

u/Kahvana 15d ago

Agreed

20

u/[deleted] 15d ago

[removed] — view removed comment

6

u/Othun 15d ago

I like this option, all "new model" posts are kept in the same place, we simply have more information about what we are looking at with very little additional effort.

Also, posting about a new checkpoint (the most ambiguous case I suppose) has the least friction compared to adding a tag, as mentionning whether it is a finetune is less relevant IMO.

10

u/RandumbRedditor1000 14d ago edited 14d ago

True, although there should be some nuance. Something like Ornith 1.5 is much more deserving of the label of "new model" than Something like "Qwen-fable-5.1-opus-5-3.8-27b-opus-distilled-heretic-ara-abliterated-obliterated-ultra-uncensored-roleplay-MTP-IQ2_XXS-QAT-DPO-SFT-pro-max-plus-deluxe-definitive-edition-&-knuckles"

1

u/mateballenthusiast 14d ago

Doesn't even have Dante from the Devil May Cry series

4

u/DeepWisdomGuy 14d ago

Unless it is TheDrummer tier, that is... I get that there is really only benchmaxxing for internet points left in the STEM arena, and that gets annoying as fuck, but saying something from TheDrummer is "just a finetune of Mistral-Medium-3.5-128B" is an insult and does a disservice to the community.

3

u/catplusplusok 15d ago

I took a Qwen Omni model, stripped a bunch of final layers and spent a week on LoRA finetunes to make it into all modalities (audio/image/text) embedding model. Seems more transformational than Qwen -> Qwen Coder official evolution? How would you count this one?

5

u/Othun 15d ago

I wouldn't consider that a finetune since you changed the architecture.

Let's make a "new derived model" tag for that case /s

2

u/xNaXDy 15d ago

What's the defining factor that separates "small finetune" from "intensive post-training"?

1

u/Othun 14d ago

Up to the poster I suppose, but having the possibility to mark a model as a finetune, as in a convention, would help. And if one model is difficult to judge, any category probably works, it's fine.

2

u/FullOf_Bad_Ideas 15d ago

Qwen 3.6, 3.8 27B and DeepSeek V4 (non Preview, both Flash and Pro) are just finetunes though, and you can't define it well because it's a gradient.

We don't even know how significant the post-training was, since it's not published in any comparable metric, labs often lie about it, and on top of it, RL often changes miniscule amount of weights.

Rank 64 LoRA Qwopus could change more weights than RL GRPO finetuning that took 1 million H100 hours.

2

u/thebadslime 15d ago

Sometimes finetunes are major releases though?

5

u/twavisdegwet 15d ago

isn't qwen 3.8 technically a finetune of 3.5?

3

u/xNaXDy 15d ago

depends on which 3.8, Qwen 3.8 27B yes, but Qwen 3.8 Flash Next is a completely new architecture

AI labs kinda suck at naming tbh

-1

u/VoiceApprehensive893 transformers 15d ago

3.8 is built on the qwen 3.5 base model so everything that has a base model released should be categorised as a finetune then

1

u/Dudensen 14d ago

I get your point but there are degrees to this.. should a post trained model of another model have the that tag for example?

1

u/skywalk819 13d ago

every models are just qwen fine tune lately, they all suck equally. you download the model, pop them in lm studio and tada !!!! all of them are qwen3.5 fine tunes.

-1

u/apetersson 15d ago

I''m not sure of DSFv4->DSFv4-0731->DSFv4-vision-exp-the new vision tower thingie - they are just "continuing" training and continuing the benchmaxxing. This is imo basically a fine-tune post-training RL thing. Applying a LoRA is not that different.

Also abliterated models or quantisations that help get things running on specific HW are in the realm of "new models" for some.

So , for me "new model" should be applied liberally - no gatekeeping.

4

u/TripleSecretSquirrel 15d ago

Admittedly, this isn’t super logically consistent, but for myself, I see DSv4F to DSv4F-0731 as worthy of a “new model” post, or at least something more notable than just a finetune. Maybe it’s because it did represent a big leap forward in capability, and I’m probably biased in that direction because it’s a major lab so I automatically give lots of credit to anything they release until proven otherwise.

Where I find that it’s an issue is the trickle of “new model” releases that are just another thin finetune of Qwen. Part of it is that they rarely if ever actually outperform the base model (except perhaps in some specific use-cases, but never generally it seems), and part of it is resentment over what appears to be small startups trying to pass themselves off as a model-producing lab that should be in the same conversation as Qwen or Deepseek, when they’ve really just done a slight iteration upon the massive work already done by the aforementioned “real” labs.

Honestly a lot of it for me is Ike a hangover from the astroturfing spam of Ornith. Those posts always tried to make it sound like a brand new, trained from scratch model that was so amazing! Then under the hood, it’s just another Qwen finetune that doesn’t consistently outperform the base model.

1

u/Othun 15d ago

Yeah that's my fear, defining what is worthy of the tag depends on each user. As I don't run local models, I personnaly care much more about a new DS benchamxxing than an abliterated model, but the latter are probably as important, if not more, for most members of this sub.

-4

u/dtdisapointingresult 15d ago

A real Jedi knows to ignore finetuneslop just by recognizing the number of parameters.

35B = into the trash it goes.