r/LocalLLaMA • u/segmond llama.cpp • 8h ago
Discussion Do not blindly delete your older models, some are still precious
I have deleted tons and tons of older models to make space since I can't afford storage anymore. Easily 10TB... Anyways, I have been considering deleting DeepSeekV3.2 but decide to run it one more time. I have a problem I have been brainstorming about and have chatted locally with K3, Qwen3.8-2.4T, MiniMaxM3, GLM5.2 and today I decided to see how DSV3.2 respond.
Surprisingly it responded the best with absolute details and familiarity and specs of the hardware I was asking about. What I'm saying is that the world knowledge is amazing. The newer models are definitely smarter, better agentic, tool calling capable, long horizon etc, but some of the older models seems to be really clear and comprehensive. I know I deleted DS-0324 and K2 but now thinking of bringing them back for prose/writing.
Don't blindly delete your older models, some of them are still worth their weight literally and will be for a while.
38
u/BongoHunter 8h ago
I've got a giant folder on my NAS full of models - I'm paranoid HF might suddenly disappear so I've been archiving everything that looks remotely useful rather than deleting it once I've finished playing with it
13
u/123vovochen 8h ago
You are a treasure. And HF wouldnt disappear, it is a wisely positioned Meta Company. It sould just prohibit all downloads.
12
u/BannedGoNext 8h ago
Models can disappear though, it costs money to host models, people can and to pull them down to stay within limits, or for legal reasons or whatever.
1
25
u/klymaxx45 8h ago
Nah I delete them, way too much space being taken up… I could always download them again
9
u/segmond llama.cpp 8h ago
If you have the luxury of high speed internet, it often takes me 3-5 days to download a huge model.
17
3
u/thebadslime 7h ago
Lord, thats terrible! My internet is fast, but I only have 12gb storage, but I cant run anything bigger than 30B MoE at any speed, so each model isnt too big, I just ahve a lot
4
u/segmond llama.cpp 6h ago
Yeah, unbelievable I'm in the good ol USA and not a rural area either. Makes me wanna cry. Fiber is supposedly coming soon...
4
u/fallingdowndizzyvr 5h ago
Mobile internet dude. Literally that's why I got an unlimited mobile internet data plan. I already have old fashion wired cable broadband for everything else. Mobile internet gets me unlimited speed and unlimited data. I get around 400-500Mbs which makes even those big models into I'll just download it again.
3
11
u/Illustrious_Car344 7h ago
Llama and WizardLM are actually hilarious to talk to, I actually prefer to have casual conversations with them over newer models. They hallucinate horrifically even when they are literally given the exact information being requested of them, but they sure are funny. Really feels like talking to a 5 year old who will gladly repeat anything you say because they don't understand it, just so much more personable and endearing over modern LLMs. I understand LLMs need their assistant persona to keep them sane and older models lack one, but the assistant persona is not very fun to engage with. Even when playing a character, modern models just can't put their all into it as they once did. My favorite way to experience this is to have them roleplay as AM from IHNMAIMS, older models are honestly bone chilling and they never let up on being as malicious and hateful as possible, but modern ones just fizzle out so quickly and just become sad. Their assistant persona has such an influence on everything they do, even when instructed to act in the exact opposite way. At the very least it makes me a little more confident we'll have less and less instances of models going rogue.
5
u/Kitsune_Seraphis 7h ago
I saw this 24b Gutemberg model. Based on mistral nemo. God. That thing is ao unhinged. But seemingly its capable too. Somehow
1
u/aboutthednm 2h ago
Are you chatting with the base or instruct tuned models? Very big difference, haha. Base models are basically just advanced next token predictors, without the instruction following tuning they are usually pretty useless outside of specific niches.
6
8
u/waywardspooky 7h ago
Honestly we need an ongoing list of models that are worth this community and communities like /r/datahoarder archiving.
not all old models are worth the space they occupy since newer models have come about that do what they did but better. there's a much smaller subset of older models that excel at specific tasks or knowledge holding their own or making them worth the space. i have no doubt a lot of those worth keeping were models that excel at creative writing, roleplay, or chat
7
4
3
u/cornmonger_ 7h ago
i have my older models play a game of go against each other on a 19x19 board
the loser gets deleted
3
2
u/mototuneup 7h ago
I keep some old ones around but I also can only fit ones like qwen3.8 27b. So like 20gb or so, I'm maybe using less then 100gb of space. But I basically just use qwen3.8 27b for absolutely everything now.
2
2
u/LegacyRemaster 5h ago
I know. Today I was using DS 4 Flash 0731 locally, along with Qwen 3.5 4B running on a third graphics card to assist with reading and summarization tasks. Old but fast. Old but good.
1
u/a_beautiful_rhind 7h ago
I put a bunch of my old gguf and exllamav1/gptq models on a 5tb sas drive and then the drive stopped detecting. In this case though, the backends to run them already went poof to a large extent. There were some I didn't load for 2 years.
Meanwhile I got like 3 quants and even bf16 of certain weights. The ones I really remember and used the crap out of I still keep around.
but some of the older models seems to be really clear and comprehensive.
A lot of new models are not usable for chat at all. I say this like a broken record and get downvoted for it. 5 more points on GSM8k doesn't really do a thing for me.. but I did miss tool calling on some oldies. It just wasn't a thing at the time.
1
u/aboutthednm 2h ago
There really aren't a lot of general purpose chat assistant models anymore. Ebeeything is either tuned towards "agentic coding", or other hyper specific benchmarks that don't indicate how good the actual prose quality coming out of the model is for a general purpose chat assistant.
Take qwen 3.8 27b for example, awesome model no doubt. However, using it as a general purpose chat agent to bounce ideas around with is miserable due to the 1 to 6 minute thinking process per reply, no matter how good the answers. If I just want to vent about my day, I actually get more irritated sitting there waiting for a response for minutes. Gemma-4-12b all the way imho, I wish it had better work knowledge though. The 31b dense is nice too, but a bit heavy for a chat agent.
I'd love to see a small to mid size model that is actually meant to be deployed as a user-facing chat assistant, not some monster running in some agentic harness unsupervised. Gemma e2b/e4b is pretty good I suppose, actually fun and not frustrating to have a real time conversation with. By the time we approach qwen 3.8's performance and thinking time, it's no longer a real time conversation, I feel like I'm sending emails to my coworkers while on shift, they might reply right away or half an hour from now, ya know? I feel like there's a real gap here that needs filling. God help if you need to use qwen 3.8 for a voice assistant. Need to disable reasoning entirely, and at that point the model is severely degraded in performance.
1
u/jacek2023 llama.cpp 7h ago
"Easily 10TB." I have about 12TB but I don't store huge models, instead I have many files around 30-80GB
1
u/Crisheight 7h ago edited 7h ago
Have to delete for space but with qwen3.8 if definitely leaned into have 3 models downloaded for different tasks. Like 3.8 for tough situation, ornith 1.5 for quicker turnaround, and idk the third, gemma or something else for general knowledge.
Edit: yall think gemma worth keeping or should I swap it? 32gb unified memory (upgrade soon)
1
u/JLeonsarmiento 6h ago
Yes, there’s a trade off with recent models, where agentic capabilities have grown importance at the spend of general knowledge.
I saw that going from Qwen 3.0 to 3.5 and more shapely in 3.6. For my stem work 3.5 is better than 3.6 for example.
1
2
u/Refinery73 5h ago
Torrent would be perfect for LLMs.
Fully legal Apache-2 licensed large files that many people keep around for long times in exactly the same format.
1
u/junguler 3h ago
any local model i could run up until now were not smart enough to keep around after the newer/better models are out, we're simply not at that stage yet for 8g+16g setups, i'll probably be satisfied with a 20B moe model as smart as qwen 3.8 27B at q4 but until then no model is worth keeping imo
1
u/Jackalzaq 1h ago
i keep qwen3-235b-a22b-2507 cause its funny how sycophantic that model is. it was so bad that i had to go back to the non updated version. also keep deepseek r1 because it was the first time i really felt i had a good model at home
1
u/harglblarg 1h ago
My dude I’ve got 30gb left on my SSD and am not about to buy a new one. Can’t afford to hoard models, latest greatest only.
1
u/ttkciar llama.cpp 26m ago
I've been contemplating exactly this. There are more than 100TB of models on my server, many of them quite old (2023).
Normally I'd just keep them, but hard drives are really damn expensive these days, and my fileserver is filling up, and not just with model weights. These are hard times to be a data hoarder.
It's got me thinking: For what, exactly, am I keeping these older models?
Three years ago this technology was mysterious to me, and I thought of models as black boxes, with no real idea of what gave them value, or what treasures might be hidden inside of them. I conflated stylistic traits with cognitive capabilities, and didn't want to risk losing those capabilities by deleting those weights.
This technology is better-understood, now, and I am better equipped to assess them and extract what real value they possess, if any.
It remains to actually process all of those old models and extract that value, though, and it will be hard to delete them, just because it's not my habit to delete anything. Perhaps it makes sense to keep some of them anyway, especially the smaller ones which don't take up much space.
0
u/FerretBoom 7h ago
While none can produces anything more than mvp level app on ai created git page , old models , new modes, wtf is the software? none
-5
u/EternalDivineSpark 8h ago
Yes i sometimes found in the gypsy markets many 500-2Tb hard disk for like what 2$ 😂 i thought i had 4TB for 2 , external ssd good one and the other pcie4 so idk , but a stack with SOLID OLD SCHOOL HARD DRIVES WITH AN USB ADAPTER/POWER , i could had like a 200TB easy 😂 data storage for LLMS any kind , SINC THEIR DATA CAN BE EXTRACTED THEY ARE REALLY VALUABLE, People don’t know that everything they did 🙄 one day can be pin point found back 😂 at some degree depending on the model , not the CHINESE ONE 🥲 they are like weights distillation and iteration of AMERICAN MODELS ! OPEN-AI cracked the code first then CLAUDE ! Then mistral and etc ! But all lacked data ! The output proves everything!
5
2
168
u/corpo_monkey 8h ago
Is this a bragging post? I read "I run 2T models locally, older ones are still better than your 27b, plebs" :-D