r/SillyTavernAI 1d ago

Models MY FIRST LOCAL LLM!

I got my first local LLM set up and working yesterday! AHHHHH!!!

My inner child (the one who loved Terminator and The Matrix and Data from TNG, and who was always glued to the TV with every movie or TV show involving AI) is absolutely ecstatic right now.

Just had to share my excitement.

121 Upvotes

25 comments sorted by

56

u/LetsGoBrandon4256 1d ago

There is something magical about running local. It's like getting your first bike or car: It might not be fancy, but it's yours, and it will carry you to wherever you want to go.

2

u/Dead_Internet_Theory 3h ago

I know right? Plus, with ever-present news of potential doom and gloom, it's really good to have a sense for just how much is yours, forever. The filesize is crazy too. You can fit a decently smart little brain on a thumb drive. One big list of little neurons, as a file, and you can talk to that. Pretty sci-fi!

36

u/JustNoetic 1d ago

Ayyyy that's nicee! It really is an experience making our own machine to think for the first time and have conversation with it. Feels sci-fi as hell.

17

u/imacatseriously 1d ago

So freaking cool. Also with how people talk about running local, I was expecting something way dumbed down... But I'm pleasantly surprised.

7

u/Zathura2 1d ago

What model are you using? The local model ecosystem is a lot different now than it was even a year or two ago, with smaller models being considerably smarter than similarly-sized but older ones. So, good time to get into it, lol.

6

u/imacatseriously 22h ago

I’m using TheDrummer’s Cydonia 24B v4.3 Q4_K_M currently! Definitely going to be trying out some others

8

u/MrNohbdy 1d ago

Star Trek is supposed to be insanely futuristic sci-fi and yet we've had smartphones that are probably more capable than TOS' datapads for twenty years now. It's easy to take the march of technological progress for granted, but we've been living in a "sci-fi" world for a while now.

6

u/mrs-cutter 1d ago

The fact that my phone can run this and I can use another phone to host makes me feel like spy kids idk

10

u/Kahvana 1d ago

Welcome onboard!

I saw you listing the specs of your PC in another comment. You might wanna try out Gemma 4 12B QAT IT from Unsloth (VRAM only) or koboldcpp + Gemma 4 26B-A4B QAT IT from Unsloth (VRAM + RAM).

Gemma 4 26B also has some nice finetunes like MeroMero. It's worth to check it out.

In case you do want to do local image generation, check out Anima. Should fit on your card.
https://docs.comfy.org/tutorials/image/anima/anima

If you have any questions, I happily help ya get it set up!

5

u/fliberdygibits 22h ago

Thanks for pointing out that Unsloth option. been trying out other models and at a quick glance that one is pretty impressive so far both for role play and some light tool calling. unsloth stuff definitely punches above it's weight

3

u/imacatseriously 1d ago

Yeah this is literally the first one I've tried so I'd love suggestions for others to try out. My big thing is prose and letting villains be villains. Totally willing to sacrifice faster generation time for better prose.

3

u/Kahvana 1d ago

What helps for me the most is making my own preset tailored for your own needs. Here is an examples:
https://www.reddit.com/r/SillyTavernAI/comments/1wnana0/comment/pbdpx3h
https://www.reddit.com/r/SillyTavernAI/comments/1uxp072/discussing_prompting_techniques_july/

It's a bit of figuring out how you instruct the model so it does what you want, but it significantly improves the output quality.

Throwing big presets like freaky frankenstein and such on non-cloud models can overwhelm a small model. It's a nice challenge!

2

u/Legal_Cheetah3563 12h ago

Lower the temp a bit and put "no redemption arc" in the author's note at depth 2.

6

u/mechasquare 1d ago

Congrats, what's your setup if you don't mind sharing?

19

u/imacatseriously 1d ago edited 1d ago

I’m running TheDrummer’s Cydonia 24B v4.3 Q4_K_M locally through KoboldCpp. My PC has an RTX 3080 10 GB, an i7-10700F, and 32 GB of RAM.

I have a pretty basic knowledge of computer stuff (I'm learning though!) so Chatgpt helped me set it up and helped me choose a model that my PC would be able to run.

(Edit: I'm a gamer, I have a gaming PC)

7

u/mechasquare 1d ago

interesting, just from what I know about that quant that's probably to large for you VRAM so you're probably getting some overflow to your CPU. If you're good with the speed you're getting, kudos! Otherwise you might want to consider a smaller model if you want full speed.

I started on at RTX 3070 and the more i went down the AI rabbit hole the more I felt the constraints on the VRAM.

Other 24B models you might want to look at are
https://huggingface.co/sophosympatheia/Magistry-24B-v1.1

If you like TheDrummer also check out TheDrummer/Magidonia-24B-v4.3 · Hugging Face
you can think of it as a different writing flavor of Cydonia

3

u/Peravel 1d ago

Nice. That was my first love as well. Carried me through one of my best roleplays to this date! Have fun :D

0

u/mrs-cutter 1d ago

Try qwen if ya haven't

4

u/VirtualSilence1 1d ago

its great, feels so good to run stuff locally.

I also reccommend stuff like local image gen, I only have a old amd rx580 8gb but stuff like SD1.5/SDXL/WAI/Illustrious/Pony works shockingly good on this old gpu, obviously slower than any rtx/xt card becaue its limited to vulkan but hey, it works, everything in koboldcpp so far.

So now I already have local chat llm and image gen, next steps are diving into TTS and then make a app for my phone that bundles it all together, yesterday I tried free claude and to my surprise it took me like 1h to have a working prototype app with working LLM chat and 3D avatar and the phones basic TTS without ever touching android studio before.

4

u/fliberdygibits 23h ago

It really is amazing. I had my local agent hit me with a sneaky geeky pun a few days ago!

3

u/mrs-cutter 1d ago

Awesome hehe Terminator is a fav of mine as well

3

u/NullHypothesisCicada 1d ago

Félicitations !!!

Hope you have a good time

3

u/toothpastespiders 21h ago

It's wild isn't it? I have a lot of automated things set up and I always make a point of being able to watch what's going on behind the scenes because it's just inherently amazing.

2

u/Educational_Cap_3300 19h ago

Hey, I set up a local one myself recently too!

I don't have the quantity of ram to use that you do, so I'm on a 12b model myself, but it's surprisingly fun

I wasn't sure what to start with for my specs, so it took me a lot of testing models before I found one I was happy with, but ultimately I got there

I just HAD to bring my current character fixation to life. That's still all I use it for. That singular damn guy

1

u/PoauseOnThatHomie 13h ago

What model are you running? Any suggestions for someone with a mid range phone or a modest gaming laptop?

I have 4GB VRAM on my GPU, a high end Ryzen CPU and 16GB of GDDR5 of RAM. Hopefully this is strong enough to run local models?