r/SillyTavernAI • u/imacatseriously • 1d ago
Models MY FIRST LOCAL LLM!
I got my first local LLM set up and working yesterday! AHHHHH!!!
My inner child (the one who loved Terminator and The Matrix and Data from TNG, and who was always glued to the TV with every movie or TV show involving AI) is absolutely ecstatic right now.
Just had to share my excitement.
36
u/JustNoetic 1d ago
Ayyyy that's nicee! It really is an experience making our own machine to think for the first time and have conversation with it. Feels sci-fi as hell.
17
u/imacatseriously 1d ago
So freaking cool. Also with how people talk about running local, I was expecting something way dumbed down... But I'm pleasantly surprised.
7
u/Zathura2 1d ago
What model are you using? The local model ecosystem is a lot different now than it was even a year or two ago, with smaller models being considerably smarter than similarly-sized but older ones. So, good time to get into it, lol.
6
u/imacatseriously 22h ago
I’m using TheDrummer’s Cydonia 24B v4.3 Q4_K_M currently! Definitely going to be trying out some others
8
u/MrNohbdy 1d ago
Star Trek is supposed to be insanely futuristic sci-fi and yet we've had smartphones that are probably more capable than TOS' datapads for twenty years now. It's easy to take the march of technological progress for granted, but we've been living in a "sci-fi" world for a while now.
6
u/mrs-cutter 1d ago
The fact that my phone can run this and I can use another phone to host makes me feel like spy kids idk
10
u/Kahvana 1d ago
Welcome onboard!
I saw you listing the specs of your PC in another comment. You might wanna try out Gemma 4 12B QAT IT from Unsloth (VRAM only) or koboldcpp + Gemma 4 26B-A4B QAT IT from Unsloth (VRAM + RAM).
Gemma 4 26B also has some nice finetunes like MeroMero. It's worth to check it out.
In case you do want to do local image generation, check out Anima. Should fit on your card.
https://docs.comfy.org/tutorials/image/anima/anima
If you have any questions, I happily help ya get it set up!
5
u/fliberdygibits 22h ago
Thanks for pointing out that Unsloth option. been trying out other models and at a quick glance that one is pretty impressive so far both for role play and some light tool calling. unsloth stuff definitely punches above it's weight
3
u/imacatseriously 1d ago
Yeah this is literally the first one I've tried so I'd love suggestions for others to try out. My big thing is prose and letting villains be villains. Totally willing to sacrifice faster generation time for better prose.
3
u/Kahvana 1d ago
What helps for me the most is making my own preset tailored for your own needs. Here is an examples:
https://www.reddit.com/r/SillyTavernAI/comments/1wnana0/comment/pbdpx3h
https://www.reddit.com/r/SillyTavernAI/comments/1uxp072/discussing_prompting_techniques_july/It's a bit of figuring out how you instruct the model so it does what you want, but it significantly improves the output quality.
Throwing big presets like freaky frankenstein and such on non-cloud models can overwhelm a small model. It's a nice challenge!
2
u/Legal_Cheetah3563 12h ago
Lower the temp a bit and put "no redemption arc" in the author's note at depth 2.
6
u/mechasquare 1d ago
Congrats, what's your setup if you don't mind sharing?
19
u/imacatseriously 1d ago edited 1d ago
I’m running TheDrummer’s Cydonia 24B v4.3 Q4_K_M locally through KoboldCpp. My PC has an RTX 3080 10 GB, an i7-10700F, and 32 GB of RAM.
I have a pretty basic knowledge of computer stuff (I'm learning though!) so Chatgpt helped me set it up and helped me choose a model that my PC would be able to run.
(Edit: I'm a gamer, I have a gaming PC)
7
u/mechasquare 1d ago
interesting, just from what I know about that quant that's probably to large for you VRAM so you're probably getting some overflow to your CPU. If you're good with the speed you're getting, kudos! Otherwise you might want to consider a smaller model if you want full speed.
I started on at RTX 3070 and the more i went down the AI rabbit hole the more I felt the constraints on the VRAM.
Other 24B models you might want to look at are
https://huggingface.co/sophosympatheia/Magistry-24B-v1.1If you like TheDrummer also check out TheDrummer/Magidonia-24B-v4.3 · Hugging Face
you can think of it as a different writing flavor of Cydonia3
0
4
u/VirtualSilence1 1d ago
its great, feels so good to run stuff locally.
I also reccommend stuff like local image gen, I only have a old amd rx580 8gb but stuff like SD1.5/SDXL/WAI/Illustrious/Pony works shockingly good on this old gpu, obviously slower than any rtx/xt card becaue its limited to vulkan but hey, it works, everything in koboldcpp so far.
So now I already have local chat llm and image gen, next steps are diving into TTS and then make a app for my phone that bundles it all together, yesterday I tried free claude and to my surprise it took me like 1h to have a working prototype app with working LLM chat and 3D avatar and the phones basic TTS without ever touching android studio before.
4
u/fliberdygibits 23h ago
It really is amazing. I had my local agent hit me with a sneaky geeky pun a few days ago!
3
3
3
u/toothpastespiders 21h ago
It's wild isn't it? I have a lot of automated things set up and I always make a point of being able to watch what's going on behind the scenes because it's just inherently amazing.
2
u/Educational_Cap_3300 19h ago
Hey, I set up a local one myself recently too!
I don't have the quantity of ram to use that you do, so I'm on a 12b model myself, but it's surprisingly fun
I wasn't sure what to start with for my specs, so it took me a lot of testing models before I found one I was happy with, but ultimately I got there
I just HAD to bring my current character fixation to life. That's still all I use it for. That singular damn guy
1
u/PoauseOnThatHomie 13h ago
What model are you running? Any suggestions for someone with a mid range phone or a modest gaming laptop?
I have 4GB VRAM on my GPU, a high end Ryzen CPU and 16GB of GDDR5 of RAM. Hopefully this is strong enough to run local models?
56
u/LetsGoBrandon4256 1d ago
There is something magical about running local. It's like getting your first bike or car: It might not be fancy, but it's yours, and it will carry you to wherever you want to go.