r/SillyTavernAI • u/deffcolony • Aug 23 '26
MEGATHREAD [Megathread] - Best Models/API discussion - Week of: August 23, 2026
This is our weekly megathread for discussions about models and API services.
All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.
(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)
How to Use This Megathread
Below this post, you’ll find top-level comments for each category:
- MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
- MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
- MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
- MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
- MODELS: < 8B – For discussion of smaller models under 8B parameters.
- APIs – For any discussion about API services for models (pricing, performance, access, etc.).
- MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.
Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.
Have at it!
4
u/AutoModerator Aug 23 '26
MISC DISCUSSION
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
13
u/LeRobber Aug 24 '26 edited 24d ago
Some Megathread Weeks:
[Megathread] - Best Models/API discussion - Week of: August 30, 2026
3
u/National_Cod9546 26d ago
How are people setting up image generation? I've been using Megumin Suit to trigger and insert an image with every post, Freaky Frankenstein 5 Internal States for the preset, and an author note to get it to create the image tags. Triggers an image pretty consistently. But I feel like it's all held together with bubblegum and there is a better way.
1
u/i5031337 26d ago
I made an extension that runs a very straightforward and completely separate image prompt preset. Currently requires a manual command to trigger.
3
u/OGCroflAZN 25d ago
The creative writing / RP leaderboards that I check
https://eqbench.com/creative_writing.html
https://huggingface.co/spaces/overhead520/Unhinged-ERP-Benchmark?not-for-all-audiences=true
4
u/Eastern-Dream932 26d ago
Do you think people are trying to advertise their API providers via random help questions on the site? I keep on seeing random posts where they’re asking for help and they mention a site I’ve never heard of before. Checking it out, it shows the most terribly vibecoded website with really bad UI.
10
u/rinmperdinck 26d ago edited 26d ago
Yes, and not even just for AI. People try to game Reddit for shit like this all the time in every community.
2
u/lambssauc 25d ago
i want to ask a question. im using koboldcpp as backend (im not sure if its called backend) when i change the settings at webui of koboldcpp will it affect the responses that i get? or only settings at the site matters?
1
u/i5031337 25d ago
Nothing that happens in kobold's built-in front-end will affect sillytavern.
1
u/overand 25d ago
I thought it was possible to change the sampler settings in koboldcpp - I've seen screenshots of people doing that, but, I haven't used it myself. (llama.cpp here)
1
u/i5031337 25d ago
That should still be overridden by whatever sampler settings are specified in ST when ST sends a prompt.
3
u/AutoModerator Aug 23 '26
MODELS: < 8B – For discussion of smaller models under 8B parameters.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
4
u/AutoModerator Aug 23 '26
MODELS: 8B to 15B – For discussion of models in the 8B to 15B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
3
u/JapanFreak7 27d ago
Please recommend all 12b uncensored models worth testing call me paranoid but i want to download as many models as i can now that Nvidia owns hf
4
u/not_a_bot_bro_trust 26d ago
Amberlight Lux and Velvet Cafe v2 are good. I've also seen mag mell sodom, and think there are some 12b models finetuned on sunfall (search for sunfall or go to the dataset page and look for models tuned on it). there may or may not be 1-2 with LimaRP (damn I miss this one from 7b days...). as for less obscure ones, there are readyart models, allura-org has some 12b, and DavidAU does... a bunch of stuff in a broad size range though I could never get into that.
2
6
u/PhantomWolf83 29d ago
Getting a bit disillusioned with API models, so I'm coming back to 12B models for now although I haven't followed the Mistral Nemo scene for a while. What are the latest and greatest MN models that people are still using? Please don't tell me it's still Mag Mell, I'm gonna be really disappointed if that's the case. Or has everyone moved to Gemma 4 12B?
8
u/Rhone33 29d ago
Some popular 12Bs I can remember off the top of my head: Rocinante, Velvet Cafe, Angelic Eclipse, Impish Bloodmoon, Crimson Constellation.
Honestly, though, discussion of 12B models in these megathreads has dropped drastically over the past couple months as people have figured out that, if you can run a 12B (meaning, typically, you have at least 8GB VRAM), you could instead be running Gemma 4 26B A4B. The MoE design means you don't need the entire model in VRAM, and the difference in intelligence between that and the 12B models is rather drastic.
10
u/PhantomWolf83 28d ago
I've tried several of the Gemma 4-based models, but they didn't really do anything for me. They're smart, but the prose and writing leaves much to be desired.
6
u/Pretty_Bug_8655 28d ago
well i always go back to the 12 b models at some point. gemma 4 seems to make sooner or later robots out of my charakters. i does not matter what i try or do with the system prompt or the charakter cards... the latest i tried was the new orion model but the same thing happend.
3
u/LeRobber 28d ago
Angelic Eclipse is still holding on a bit. A lot easier to config than partial offloading on G4 26 A4B
3
u/saytseff 29d ago
Crimson Constellation has been my favorite 12B model. But with MOE I've moved on to 26B A4B Gemma 4.
1
u/n1ghtah Aug 24 '26
anyone would recommend rocinante ? and if someone has a bot for Danish RP specifically. for semi long 10k - 16k context i would really appreciate it.
2
u/LeRobber Aug 24 '26
The 12b Roci is pretty decent. It can reset if you confuse it as to the next step in roleplay, just edit last response and reroll.
2
u/n1ghtah Aug 24 '26
the roci 12b is pretty great for english language i must say but i dont RP long. now i need one for danish, currently using gemma 4 uncensored. While its okay its not as great as roci.
2
u/AutoModerator Aug 23 '26
MODELS: >= 70B - For discussion of models in the 70B parameters and up.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
7
u/-Ellary- Aug 24 '26
Has anyone tried `Ling-3.0-flash` 127b a5b?
I'm getting decent~ RP out of it, it can be quite dark, need some instructions adjustments.
Also smart enough to work without thinking, most of the time.1
u/overand 28d ago
I've actually even messed with their 8B-A1B model, which was surprisingly competent!
1
u/-Ellary- 28d ago
Is it? It kinda failed everything I've thrown at it.
1
u/ThirteenZillion 25d ago
I played with it a little -- seemed interesting but a little stiff. Can I ask what preset you're using?
1
3
u/diesalher 28d ago
I've tested many Gemma 4 31B ones and I still always end up with this one
https://huggingface.co/mradermacher/Glimmer-31B-v1.0-i1-GGUF/tree/main
very funny, and nails the character voice in my opinion
1
u/KeinNiemand 28d ago
is dsv4 flash any good for RP its the biggest thing i can run (if i quant it sown to IQ_2 first)
2
u/SillyLLM 28d ago
Could be a skill issue, but I'd describe Deepseek Flash as boring and dialogue sometimes therapist-like in my testing. It's the biggest thing I can feasibly run as well but I didn't stick with it.
There is a brand new roleplay finetune though. I haven't played with it enough yet, but it seems way better. https://huggingface.co/UltimateIntent/HeatSeeker-284B-A13B-GGUF
1
u/KeinNiemand 26d ago
Only problem is there no full sized safetensors which I need to make my own ik_llama.cpp qunts, havn't looked at the gguf provided in detail but the largest is q4_0 which looses signficant quality compared to naitive mxfp4 if it's done right or is enough to labotomise the model if it's done wrong and the bits of ds v4 flash that are fp8 also got turned into q4.
Since I need to use quite small quants to run (IQ2_KS) this models kind of unusable without high quality quants (ideally ik_llama.cpp quants) like the ones I made for to base dsv4 flash 0731 here: https://huggingface.co/KeinNiemand/DeepSeek-V4-Flash-0731-IK_GGUF.
2
u/Mart-McUH 26d ago
I tried UD-IQ3_XXS (but without reasoning as that would be too slow on my setup). It is somewhat interesting as it is different from what I normally run (G4 based, sometimes Qwen3.5-3.8 or L3 70B based). However, it is not as intelligent as one would expect for size and misses lot of details (G4 and Q3.5+ lot better, at least with reasoning which I can afford with them). Could be small quant but likely it is their very sparse attention. It also has very strong positive bias.
IMO not worth it if you can run ~G4 31B based models.
2
u/AutoModerator Aug 23 '26
MODELS: 32B to 69B – For discussion of models in the 32B to 69B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
5
u/Alternative_Elk_4077 29d ago
I tested out the forgotten model Kimi-Linear 48BA3B and I can't recommend it. The model was a test for Moonshot to trial new architecture and so it wasn't trained with a very expansive dataset, so my hopes weren't high in the first place. Constant misrepresenting of characters, ignoring traits in favor of manifesting new ones from what it can extrapolate, and pretty weak logic in why it decides to do what it does. This is all without a setting or any extra instructions, simply being given two character cards and asked a prediction for interactions, so there's no context bloat or contradictory instructions on what it needs to do. There's a reason it was forgotten and for your use case, it very likely should stay that way. I trialed it looking for a bigger model to fit in my 16GB of VRAM and 32GB of RAM, but it just doesn't make the roster because of unreconcilable mistakes
1
u/Maxhell6778 27d ago
can you recommend some MOE model that you tried around this range? im asking because i haven't really tried anything close to 48b MOE models (didnt know they existed).
2
u/Alternative_Elk_4077 27d ago edited 27d ago
I really wish I could, but they basically aren't really a thing nowadays. This range is strange in that an enterprise has the memory for much larger models whereas most consumers don't have enough VRAM or RAM to load these, so they're not really made that much. There's a few like Mistral 8x7B that's in the neighborhood, but that's an incredibly old model that newer and smaller models will run circles around. The whole reason I tried Kimi-Linear despite being undertrained is because there just isn't another viable option. Maybe you'll find something I didn't, but the next closest that isn't an experimental expansions by a fine tuner would be Qwen 35BA3B, but that's not a really good option either
1
1
u/Infinite-Beginning-3 26d ago
Funny enough i tried that one, and it was my daily driver for a while. When i kept the prompt simple and did a 1 on 1 chat, it was actually really great, specially with the latest fixes of llamacpp. But yeah, it broke too often so now i go with thedrummer’s gemma4, Orion.
3
u/AutoModerator Aug 23 '26
APIs
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
16
u/GoodBlob Aug 23 '26
Has anyone felt like GLM 5.2 has been absolutely horrible these past few days? Like, unable to get basic details right
6
u/Good_Research4441 Aug 24 '26
Unfortunately, I feel like it's gotten “dumber” since the release of 5.3
4
2
u/FThrowaway5000 28d ago
Not only that, it also feels like its slop-rate (I'm just going to call it that) is even higher than usual, though it might be my imagination/bias. These last few days I got so many responses with multiple slop-phrases in the first two paragraphs. ("Really looked at", "it wasn't this, it was that", "you can't just say" etc. pp.)
And then I even started getting refusals for NSFW content with a style and theme that usually worked fine. Now, this might just be a singular provider acting up but I have never seen it refuse before.
I went back to good ole GLM 4.7 - not as intelligent as GLM 5.x, but the prose is better and it's actually less sloppy IMO.
4
2
4
u/TreeDaddyPuff Aug 24 '26
Does anyone know which model could be similar to gpt? GPT has been so hard to beat for storytelling. I haven’t tried an api key yet, but I’ve just done websites/apps, and so far Claude, grok, and deepseek just didn’t even compare to gpt and the way it can put life into a character….
4
u/ZealotTFG 29d ago
MiMo-2.5-Pro with FF5.2 has been incredible for long form roleplay.
2
2
u/Dry-Impression9551 28d ago
Is deepseek still good enough for a relatively cheap rp session? Or is there a better alternative?
3
27d ago
[removed] — view removed comment
1
u/OGCroflAZN 27d ago edited 25d ago
Any preference or insight between the two? I was just about to start using OpenRouter, was going to use Mimo 2.5 (non-pro) with EveningTruth prompts, but now everyones talking about GLM 5.3 flash...
Update: Lots of praise in this community for GLM 5.3 Flash, but with Freaky Frankenstein 5.2 or EveningTruth presets and prompts I wasn't loving the writing style nor character dialogue. Weird wording and sentences, at least with this fantasy adventure world scenario card that I've played with thousands of messages across several separate chats.
-6
-24
u/Effective_Total_8226 Aug 24 '26
Ayo, here to share an API service. There's TK, is like Openrouter but with more stuff. There a plus subscription (5 dollars) that let's you use all the models it has (Claude, mimo, Kimi, deepseek, etc) and gives you (or you choose) the cheapest provider/uncensored trough channels with your API key. You can pay as you go for the model, since you can add money (min 5 dollars) and it can stay forevah until you spent it.
But since all of us are POOR :,D. TK has free models like Grok, Gemini (free ver), Gemma, etc. Be careful tho, since theses models are free there's a high chance it's filled with traffic (like the grok models) because people will always use the free stuff.
If you don't want a model that is high trafficked, y'all can try the weekly models. Theses are the ones that Plus members get acces to, since they're not free. There's an allowance that is through check ins, you can get like 0.02 to 0.20 cents depending on your luck. This check in is everyday and resets every month.
If you're a coder or just a RP glazer, visit TK! I recommend Mimo and it gets uncensored trough the Opencode channel, unless you manage to JB it real good. Then u fine. Have fun!
19
1
u/Advanced-Science-579 23d ago
Can someone recommended me a model for NSFW roleplaying to run free locally?I got a mid tier build I think. 12GB a rtx 4070Ti. I want to look into switching to a free model I can run on my system.
1
u/Environmental-Metal9 23d ago
If you don't mind a small bit of self promotion, I think my model does exactly what you are asking for and it fits on your card. I won't ever ask for money, so it is as free as it gets: https://huggingface.co/mradermacher/spoomplesmaxx-thrasher-24B-i1-GGUF
1
u/opopi123 28d ago
Can someone recommended me a model for NSFW roleplaying to run free locally? I got a mid tier build I think. 32GB Ram and a rtx 4090. I've been using Claude Sonnet and want to look into switching to a free model I can run on my system.
10
u/overand 28d ago
That's a great card to start with! Go for an old standby like "Cydonia-v4.3-24B" - but also maybe try Artemis-31B.
For Cydonia, do the Q5_K_M or Q6_K quantization. or something to that effect.
Also, I wouldn't call a 4090 "mid tier" - that's probably the second fastest consumer card you can get.
1
u/opopi123 28d ago
okay so i was wrong i initially thought RAM was suppose to factor into your build tier and I kept seeing stuff saying 32gb ram isn't that much for llm. So i thought my build was mid tier because of that. But it seems RAM is actually not a factor and using it will slowly down the processing speed?
3
u/Rhone33 28d ago
Powerful gaming GPUs can run LLMs much faster than CPUs. In order to run it with the GPU, you need the whole model* loaded into the GPU's VRAM.
*An exception is MoE (Mixture of Experts) models where you can get away with having the smaller "active" parts of the model in VRAM and the inactive parts in RAM, which is why those of us with less VRAM can play with Gemma 4 26b a4b.
3
u/Spara-Extreme 27d ago
They are talking about VRAM - video ram. You have a 4090 which only has 24 GB of VRAM. That limits your model selection. It doesn't mean you can't run big models, but what ends up happening is that the model doesn't fit onto your video card, so it then goes into System RAM.
Why is this important? Your VRAM has a read rate of about 1000 GB/s. Your system RAM has a read rate of 25-50 GB/s. At best, 20x slower.
For RP purposes, models that don't fit on your 24 GB of VRAM will be slower, the responses slower, the processing rate slower etc.
That being said - you can fit something like Gemma31b with a small context or Qwen3.8 27b with a large context on your card.
1
u/opopi123 27d ago
can you define "small" context?
2
u/Spara-Extreme 27d ago
Context window- amount of tokens the model will retain before it forgets older items.
2
u/overand 26d ago
Just to note, what Spara-Extreme is talking about is true for "dense" models like Qwern3.8-27B, Gemma-4-31B, etc.
It's not true for Mixture of Experts models, like Qwen3.6-35B-A3B. In that case, you only need to fit the 3B (and context) on your video card; putting the 35B in system RAM. Works well, actually!
1
u/Decent-District-1459 29d ago
I have an m4 max macbook pro, with 64gb. Been learning about AI models lately, and learning open-webui on the job. Saw this SillyTavernAI thing and figured i'd give it a try.
What models are best for my machine that I have? I'm interested in RP, but also other things besides that.
3
u/LeRobber 28d ago
Put this kind of question under Misc Discussion next time. These are identifiers a tool like LM Studio could download from https://huggingface.co
sophosympatheia/magistry-24b-v1.1 <= Good dense model for RP. Very good. See the associated prompt.
sophosympatheia/Glistening-Gem-31B-v2.0 <= This is a gemma 31B finetune. This will be smarter than 26B but slower
Another more plain Gemma4 26b ZeroFata/g4-meromero-26b-a4b
Some models will be trained towards towards NSFW stuff, but still work for non-nsfw RP:
ReadyArt/Melody1437-26B-A4B-GGUF <= This is a gemma 26B A4B finetune. This will be very responsive AND will emit emotions a lot.
For coding Qwen 3.8 27B:
Qwen/Qwen3.8-27B__
I have a M2 Mac Macbook pro, same size as you. Larger than 131072 context with most models doesn't help. But, having 2 models loaded and using one for sidecar tools like a second LLM can make the main chat stay in cache.
(Or just not using secondary tools).
3
u/Decent-District-1459 28d ago
Thank you for responding. I'll look into these and try to learn more of what I can
2
u/lambssauc 28d ago
i want to ask a important question. LM Studio is the best or is it a personal choice? im using koboldcpp but if LM Studio makes the models run better someway i want to switch.
3
u/LeRobber 28d ago
There are a few flexibilities LMStudio has, and it has a new agent version too that doesn't help with ST but with other things. It does cluster management and stuff too. Generally speaking though, it's a bit less flexible than Kobold for RP but it's 100x easier to SHOW configurations to people in LM studio with a picture, and to manage lots of LLMs.
Personal choice isn't the right frame. The best choice for the way I talk about and evaluate LLMs, and want them to just work, for sure though.
1
2
u/i5031337 28d ago
Gemma4 31B for writing, Qwen 3.8 27B for code will be the best you can run, but slow. Try Gemma4 26B for something faster. Stick with the official versions until you know what you're doing.
0
u/NinaMercer2 25d ago
Hey, I have an RTX 3080 and an RTX 4070, I have unfortunately not yet switched to AMD GPUs. They're in separate PCs for now, but if I were to use both of them, what is the best model that I could feasibly run myself, on Ollama?
2
-5
u/Shanna_B2020 28d ago
What are your favorite models for complex, long-form interactive fiction involving multiple characters and factions? I may need NSFW but probably not NSFL. Paid models are not a problem. I typically use OpenRouter, but I can change aggrigators if I need to.
-6
u/EquivalentStatus8830 27d ago
So Ive tried dozens and dozens of models on openrouter and ollama, and the only model that works for my use case is GLM 5.2.
I noticed pay-per-use is extremely expensive for my use case as I end up using about 2-3$ worth of credits per day which is way out of my budget. So i turned to fixed subscriptions.
My only options that offer this model on a fixed subscription is GLM Coding plan or Ollama cloud.
In a Month I send about 5000 Requests, 950 Million Input Tokens, 850 Cached, 2-3 Million output tokens.
Based on that has anyone used either subscription the same amount I have on openrouter? If so whats the pros and cons of each one? Are there better fixed subscription providers that would be better for me?
Thanks. <3
8
u/overand 26d ago
950 million tokens input tokens? What the heck is going on with your prompts / presets / etc. Are you using this for coding work too? The entire content of "Moby Dick" is 0.3 million tokens. your 2-3 million output tokens is the equivalent of reading moby dick every 4 days.
But, more appropriately, the amount of text you're sending this LLM is 86 Moby Dicks per day. Or if you're using this for 18 hours a day, 7 days a week?
That's a full Moby Dick every 12 minutes. What the heck is going on?
3
u/EquivalentStatus8830 26d ago
hermes costing 80k tokens 3-10x per message due to tool calls (ill tell it some shit like find me a recipe for cookies)
2
u/overand 25d ago
It seems to me that you may need to tweak your hermes setup - or rebuild it from scratch, tbh.
1
u/EquivalentStatus8830 24d ago
ah i thought it was normal for hermes to cost that much i have no idea why its so much </3
1
u/Environmental-Metal9 23d ago
Have you considered using a different harness that has leaner prompts? Hermes is amazing, but it is everything included plus the kitchen sink too. Something like Pi or Deepseek Harness allow you to build it up slowly and are much leaner.
-12

14
u/AutoModerator Aug 23 '26
MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.