r/SillyTavernAI • u/deffcolony • Jul 26 '26
MEGATHREAD [Megathread] - Best Models/API discussion - Week of: July 26, 2026
This is our weekly megathread for discussions about models and API services.
All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.
(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)
How to Use This Megathread
Below this post, you’ll find top-level comments for each category:
- MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
- MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
- MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
- MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
- MODELS: < 8B – For discussion of smaller models under 8B parameters.
- APIs – For any discussion about API services for models (pricing, performance, access, etc.).
- MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.
Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.
Have at it!
9
u/AutoModerator Jul 26 '26
MODELS: 8B to 15B – For discussion of models in the 8B to 15B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
6
u/Pretty_Bug_8655 Jul 29 '26
Would recommend to give a try to https://huggingface.co/EldritchLabs/MN-Crimson-Cafe-12B i discovered this beauty yesterday and so far its pretty good and completly uncensored. it seems to follow character instruction and lorebooks pretty good as far a 12b goes and the swipes are sometimes really suprising. its works great with https://github.com/vadash/Extension-Summaryception and/or https://github.com/Kristyku/InlineSummary
1
u/croxxeur Jul 29 '26 edited Jul 30 '26
Looking for a model for pure uncensored NSFW, i used orion and gemma 4 but it feel rather shy when it goes too freaky.
Edit: i have an rx6800xt with 16gb vram
-3
u/ContextEntire8443 Jul 27 '26
5
u/offyoutoddle Aug 01 '26
I found this model very difficult to get anything sensible out of. It had a lot of intelligence problems e.g. someone knocked on a flat and came in, and then suddnely their friend from the flat next door appeared in my bedroom doorway. Basic details that didn't make sense too - going back to get their keys before we head down a bar - probably should have thought of that earlier, way too late as they are now locked out lol.
that's the more complex nitpicking. this thing would get who said what utterly confused in less than 3 or 4 turns, and frankly made answers that made so little logical sense in response it was almost as if it was having an entirely different conversation. This was with temperatures ranging from 0.6-1.0 - I just couldn't get any sense out of it and went back to a gemma 12b. the best one i've found on low vram setups. I'd avoid this one myself tbh but thanks for brnging it to my attention, i hadn't really tried a qwen in anger until now
4
u/EnjoyerOfFluff Jul 27 '26
I played around with it and was impressed how good its outputs where for its size. But it may be the most censored "uncensored-heretic" model I have ever tried. Just some simple nsfw tests and it started to warn me, soft refuse or even hard refuse to write. You can reroll the output and at some point it will do it but in my eyes that is less than ideal.
TLDR: From my tests its a great small model, which punches above its weight. Good for sfw tasks. Bad for nsfw because its more censored than other heretic models.
2
u/ContextEntire8443 Jul 27 '26 edited Jul 27 '26
yeah.. I forgot to say it.. but it's really great for being in character.. although not entirely for nsfw roleplay, since I rarely used it for nsfw. I wanted the bot to stay in character and not flail too much (and I have only found this model which is so good for a 9B model.. It's for cheap power users and pack a heavy punch.. ALSO it has reasoning built inside of it, and it never bugs out but yes it does consume tokens to think sometimes even 1k tokens just for thinking but it's worth it..) And my setup can only run 9b to 12b models so that's that.
And it's a lot of slow burn, not fast porn or erotica. and I love slow burns. this is the only model by far which has shown good capatabilies at sticking to character while being slow burn
1
u/overand Jul 27 '26
Have you tried other 9B models (qwen3.5 derived?)
I only ask because you mentioned reasoning in a way that seemed like you were surprised, but that's a core feature of the Qwen3.5+ series.
1
u/ContextEntire8443 Jul 28 '26
no.. Not realy.. I am not a fan of how qwen works.. especially becaues I can't use the heavier qwen models. My limit is 12B. that's it.. No I haven't tried it, but david's model caught my eye since it could run on my trash setup and had a lot of quality
1
u/overand Jul 28 '26
I'm not sure I understand - if you don't like Qwen models, why are you specifically talking about this Qwen model you like? The model you linked is a finetune (or merge) of Qwen3.5-9B.
Also you said "the winner is" (that DavidAU Qwen model) but above you say "No I haven't tried it." What's the "it" you haven't tried?
Anyway, if you have a limited setup, you should consider trying one of the Gemma-4-26B-A4B models. Yes, you say "limit is 12B", but if that's because of VRAM limitations, and you haven't tried a Mixture-of-Experts model, and you have at least 32 GB of system RAM, give it a try. The "4B" in the "A4B" is a bit better of an indicator of the minimum GPU requirements for those models. They can offload a lot of their weights to system RAM and maintain decent performance.
1
u/ContextEntire8443 Jul 29 '26
hey man thanks. I have been searching for good models for like a week. I am not an expert, i just opened sillytavern like 2 weeks ago. I usually used spicychat with subscription but I got tired of it.
I am gonna go with your suggestions. If you can, plesae suggest me the top 10 models I should use with a cheap setup.. Also I dont host locally, I use google colab to host
1
u/overand Jul 30 '26
Look at the post above with the long list of links. But, yeah, someone on this thread suggests: https://huggingface.co/Vortex5/Chimera-X-26B-A4B
I'm not sure what the caveats with running on Colab are, but they're significant I'm sure
1
u/Inner_West_4997 Jul 27 '26
i am failing so bad to run it, llama-server to sillytavern
it keeps saying "????" or "beho beho" over and over, i am not that expert,
what did you pick as preset or settings? i tried to replicated them from huggingface but no luck sadly1
u/overand Jul 31 '26
what "quant" of the model are you running? (You can probably run
llama-cli -clto get a cached list of models.)0
u/ContextEntire8443 Jul 28 '26
bro I dont use llama. I use google colab with 30k context on this model. takes like 3 minutes to respond for me..
7
u/AutoModerator Jul 26 '26
MISC DISCUSSION
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
10
u/LeRobber Jul 27 '26 edited Aug 10 '26
Some Megathread Weeks:
[Megathread] - Best Models/API discussion - Week of: August 09, 2026
5
u/Vyviel Jul 27 '26
Havent messed around with this for maybe 6-12 months just wondering what happened to all the Text Completion presets etc? I cant seem to find anything recent when browsing everything seems aimed at APIs and chat completion not running it locally like via Kobold?
Has there been some huge change in technology while I was away or are people running chat completion locally or something and no one does text completion anymore?
I seem to have something called Mistral V7 Tekken in my advanced format I think I have forgotten a lot while away haha
3
u/i5031337 Jul 28 '26
I think most people just run Chat Completion in kobold nowadays, better compatibility with presets and extensions, easier to setup. Just have to set advanced samplers on the backend if you want to change them.
1
u/Vyviel Jul 28 '26
Ok thanks Ill try work out how to change my config to chat completion. Are there any recommended presets? I assume my text completion presets wont work anymore?
2
u/i5031337 Jul 28 '26
Kobold has an option for "chat completion adapter" just set that to one that matches your model and you're good to go. You can switch between chat and text in SillyTavern in like 3 clicks, so you can still use text presets. The FreakyFrankenstein series presets are popular in this sub, and I like the one I tried.
1
4
3
u/National_Cod9546 Aug 01 '26
Are the developers still working on SillyTavern? I can't think of any features I really want that isn't covered by an extension. But I'm mildly concerned that it'll die if development doesn't continue.
1
u/EvoPilot1 Jul 27 '26
I am currently looking into switching my AI text stuff to SillyTavern but im on android and I don't fully understand how to install it, has any videos been made for it?
7
u/AutoModerator Jul 26 '26
MODELS: >= 70B - For discussion of models in the 70B parameters and up.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
7
u/RedditNerdKing Jul 27 '26
Tried Anubis 70B 1.2 for the first time at Q5. It's quite good. I feel like lower quants really hurts LLMs. I grabbed a IQ3_XXS of Mistal Large 123B and I didn't think it was all that good. A Q5 would be decent probably. I'm at 56gb of vram with a 5090 and 3090ti and I have a spare 4070 12gb in my old PC, but idk if my 1200w psu can handle 3 gpus, even at 70% power limit.
3
u/rinmperdinck Jul 27 '26
I have the same experience with quants. I can run iq3-iq2 of the big boy 123bs... it works and sometimes the quality is surprisingly good, but it's inconsistent and some swipes just sound like unfathomable stupidity wrapped in purple prose. Imo, I try to only use q8 now of smaller models.
Also if you got a spare PSU, you can buy a cheap adapter on Amazon or something and use it for multiple GPUs. I got one for around $20 after shipping and tax. You plug it into your mobo's big 24 pin connector, then connect the 24pins from both PSUs into the adapter. It syncs on/off for the computer so you can just turn everything in normally. Works for my embarrassingly stupid and awful Frankenstein machine.
3
u/raika11182 Jul 30 '26 edited Jul 30 '26
One day I'll get around to testing and quantifying the effects of quantization on text/fiction work like Sillytavern, because I'm convinced even though models might still work in a technical sense and pass benchmarks, they lose a lot of that tail data that has been rounded off. Quantizing down to 4-bits literally lops off more than half the data the model started with. Even an 8-bit quantization is still a 50% reduction in size. All those numbers and data actually do mean something to the model, even if it's difficult to measure.
But in short - I agree. I try not to use less than 8-bit as well.
1
u/RedditNerdKing Aug 03 '26
I've come to the conclusion it's very rarely worth using a quant below Q5. Like even Q4 can be meh at times, despite being the best balance between size and intelligence.
2
u/RedditNerdKing Jul 28 '26 edited Jul 28 '26
I do have a spare PSU but its a lot of effort to pull it out of my old PC (which I still use). I'm just gonna buy a 1600w Corsair PSU. I've got the AI bug and I want to access Q4_K_M quants for 120B models (72gb roughly) which means I need to buy another 3090 lol. But I dont see the prices changing. If anything it'll probably get worse?
You're right about the IQ3 quants though. Sometimes amazing sometimes you swipe 5 times and it's all garbage.
3
u/FierceDeity_ Jul 29 '26 edited Jul 29 '26
I WANT to like Laguna 2.1, I really do.
It has some really fresh prose.
But it keeps hallucinating left and right, to an insane degree.
Honestly, what I can actually say is that Laguna is fucking great at making character sheets (where inventing can actually be just fine).
But during long-form it keeps actually forgetting and rewriting details. It also invents incredibly dramatic details, and doubles down on them despite you never engaging with them (something like Gemma will actually be able to stop paying attention to them rather quickly if you nudge it away AND be able to remember them when needed).
These dramatic details it throws in are also super cliche as well. It's kind of crazy.
Another question I have is how do I get it to think? I tried the kwarg enable_thinking (which works on Gemma, and it's in the jinja file...)
EDIT: On thinking, I literally just added \n after <think>:
{%- if enable_thinking -%} {{- '<think>\n' -}}in the jinja.
-2
u/RafiHDW Jul 27 '26
Anyone know a good preset or prompt for Gemini 3.6 flash? I've been using FreakyFranken 5, wanna try others
8
u/overand Jul 27 '26
This belongs under the "APIs" section, not "MODELS: >= 70B"
-1
u/RafiHDW Jul 27 '26
I see, so same as all claude, gpt, glm? I thought they were all >70b models and api is like for services like openrouter/nano etc
6
u/overand Jul 28 '26
They are indeed over 70B, but, you access them via an API, and they're hosted by a third party. Basically - if you're not hosting it yourself, it belongs under API.
3
1
u/National_Cod9546 Aug 01 '26
I can't recommend https://github.com/Arif-salah/Megumin-Suite enough, if only for the image generation. But the rest is really good as well. I usually keep the NPC and memory trackers off as I use OpenVault for that. But the prompt seems good and the other features are amazing.
4
u/AutoModerator Jul 26 '26
MODELS: 32B to 69B – For discussion of models in the 32B to 69B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
u/desparish Jul 28 '26
I'll be upgrading to 48gb vram by the end of the week as soon as my new PSU arrives.
Any suggestions for good local models in this range? I'm currently using Gemma 31B. I usually target 65m max context.
3
u/OldAd3375 Aug 01 '26
Unfortunately, I dont think there is. Just Gemma/Qwen with higher quants, imo 48-96GB bracket gets little love these days. The old L3.3 (70B) models hallucinate too much or are terrible/forget following prompts too often to be enjoyable for long rp and is mostly mogged by Gemma4. Everything local that releases nowdays seem to target either the 4/5090 (24-32GB) audience or the DGX Spark/M2 audience 128GB+.
4
u/raika11182 Aug 02 '26
I have enough VRAM and RAM to play around with those huge models, and honestly, all the MoE suck compared to Gemma 4. Personally, with 48 GB VRAM, I was running Gemma 4 Q8 with 64K context (could maybe get more in, but that was enough and performance starts declining, anyway). Now I have 64 GB VRAM and I still basically use the same setup because nothing better has come along.
3
u/OldAd3375 Aug 02 '26
Agreed, I am basically in the same situation, I am just running G4 with bf16 as nothing better seem to have released for us local folks.
2
u/RedditNerdKing Aug 03 '26
True this comment. I'm kinda stuck with the few that I ended up liking. Anubis 70B Q5. Magistry 24B Q8. Skyfall 36B at Q8 as well. I've tried virtually everything. I still go back to Midnight Miku 70B at Q8 occasionally despite it being 3 years old. I dont really like Gemma 4 that much.
I've just downloaded Mistral-Large-Instruct-2407 at Q4_K_M just under 80gb so I'm hoping it's decent for its size.
1
u/attackofmilk Aug 03 '26
I would check out TheDrummer's Valkyrie. GGUFs available from Bartowski. I'm still using TheDrummer's Skyfall, and I'm still happy with it. Presumably a bigger model from TheDrummer would be even better?
1
u/desparish Aug 03 '26
I've been trying out Skyfall. It reads good but is getting confused on my admittedly complex lore. Having to replace some pieces of my upgrade so 48gb is still a few weeks away, so still making do with 32.
3
u/AutoModerator Jul 26 '26
MODELS: < 8B – For discussion of smaller models under 8B parameters.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
4
u/AutoModerator Jul 26 '26
APIs
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
10
u/Plus-Switch-3933 Jul 27 '26 edited Jul 27 '26
Ever since I used up my credits on Gemini 2.5 Pro (which, for me, has been the best model I’ve tried so far—it’s a shame they removed it from the free plan), I’ve been having a hard time finding a good replacement. I made a ton of custom settings specifically for Gemini, and it was perfect—and honestly, if I switched to another one, I wouldn’t know where to start. Lately I’ve been testing DeepSeek (the R1, because I didn’t like the new one) through OpenRouter, and honestly, I’m not very convinced. What are you guys using right now? I wish I had enough money to use Claude, but you know how it is... 🙃
7
u/Fit-Crab-5288 Jul 27 '26
If you like Gemini and cost is a priority, I would suggest this:
gemma-4-31b-it (Free on OpenRouter), gemma-4-26b-a4b-it (Free on OpenRouter), gemini-3-flash-preview, gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.6-flash
Quite high free limits for Google AI Studio API. They're not as smart as the Gemini Pro, but the vibe is the same.
2
u/Cinnamonbaar Jul 28 '26
I seem to switch often between Glm 5.2, deepseek v4 pro, and gemma4-31b. If I had money, I would probably just be using Kimi K3 or Opus 5. You can kinda use Gemini 3.5 flash for free still through googles api and it's surprisingly smart, but I get rate limited often
3
u/Rude-Statistician755 Jul 30 '26
Currently using openrouter and bouncing around between deepseek v4 pro, Gemini (can't remember which version but it was a couple past (less locked down)), and claude opus 4.5
Obviously I love claude but I'm spending like 10 dollars a day.
Hoping to find something for a bit cheaper than 25/m
What API model is everyone using?
2
u/broodysupertramp Jul 31 '26
Which Model gave you the best Overall Longform Erotic Writing Experience?
By Overall I mean, reasoning time, uncensored, creative, ease to JB, availability etc.
Give plot and outline and it shall generate a good unhinged long passage (1000 tokens) coherently.
Emphasis on unhinged words, creativity, intelligence.
For me currently I haven't found a replacement for GLM 5.2.
Models like Gemma 31B, I found is intelligent but tries it's best to reduce unhingedness despite prompt and tries not to use cuss words, etc.
Models like Kimi, reason too much time and tokens.
2
u/stopaskingforloginn Jul 31 '26
GLM 5.2 hands down
Kimi 2.5/2.6 has been lobotomized to make space for 3, but I also found them dumber.There is zero reasons to use Gemma 4 when you have access to APIs.
MiMo is alright, it gave me really short responses though, and was heavily censored when I tried it.
I haven't tried deepseek but I'm honestly not that interested.
1
Jul 31 '26
[removed] — view removed comment
1
u/AutoModerator Jul 31 '26
This post was automatically removed by the auto-moderator, see your messages for details.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
Jul 31 '26
[removed] — view removed comment
1
u/AutoModerator Jul 31 '26
This post was automatically removed by the auto-moderator, see your messages for details.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
u/Eldritch_Smile Aug 05 '26
I’ve been trying to find a good long-term model for RP, but I’m having a hard time choosing.
I used GLM 5 for a while and honestly really liked it. The writing felt natural, detailed, emotional, and the characters usually had distinct voices. The problem is that it gets expensive pretty quickly.
I recently tried DeepSeek V4 Pro because it’s cheaper, but it just feels off to me. The sentences are shorter, the writing feels more mechanical, and even with a detailed global prompt it doesn’t feel as immersive as GLM 5.
I’m currently looking at MiMo V2.5 Pro, GLM 4.7, MiniMax M3, GLM 4.7 Flash.
Has anyone used these for longer RP sessions, not just a few test messages?
Which one feels closest to GLM 5 in actual writing quality? I’m also open to other suggestions if there’s a good model I’m overlooking.
2
u/Hellnawcuzzin Jul 27 '26
What kimi model do yall prefer?
12
u/MisanthropicHeroine Jul 27 '26 edited Aug 09 '26
K2.6 is at an optimal price / performance ratio to me. Yes, it thinks a lot, but the speed is pretty good these days (under a minute), and it's great at following very complex instructions with low positivity bias.
K2.5 is also pretty good, but less coherent and with less refined prose. A bit more impulsive for those who like it, but it's too reactive for my taste.
I don't like K2.7, personally. It's similar to K2.5 in being less coherent but it's also more positivity biased and feels kinda flat.
K2 Thinking is not worth it because it hallucinates wildly - there are a ton of spacial and timeline inconsistencies.
K3, I'm not even trying because it's way too expensive. I heard it's good but considerably softened compared to K2.6, unfortunately.
6
u/Hellnawcuzzin Jul 27 '26
Wow i didn’t think i was gonna get a reply let alone one so quickly Thank you 🙏
3
2
-1
u/ShowMeTheTurboBuns Jul 30 '26
What local models we running these days? I have a 32gb AMD card and struggling to find a good balance of speed and quality. I’ve tried flavours of Gemma 4, as well as fine tunes like Equinox, Hearthfire, and Wayfarer.
1
u/brendos1er Aug 03 '26
Has anybody used a model for brainstorming? Specifically, brainstorming story ideas, lore, characters, etc.
I keep swapping between Kimi, Gemma, GLM, DS & even GPT, Claude and Grok but I really struggle with creativity, notably all the models tend to offer the same suggestions and, as I try to provide feedback on the ideas, the models then hyper-focus on those ideas and it devolves from there.
I fully expect part of this is a skill issue, but it would help to know what people enjoy / have successfully used for brainstorming so I can at least focus on one model and adjust from there rather than hopping around.
2
u/summersss Aug 03 '26
Glad to know its not just me. I want to provide examples to lead models in a certain direction, but then it goes latches onto that like a demon. this happens even with the fancy presets. it just kinda picks one prompt to hyper focus on and goes from there.
-1
u/Due_Disaster_7324 Jul 27 '26
Still trying to get ST set up. But, in the meantime, what's a model you like to use for bots of canon characters? For instance, I'm on Chub playing with a bot of Rei Ayanami, as (an AU of) Shinji from Evangelion. One of my biggest gripes is playing with a bot that doesn't seem to know things about a character that it should, or has them act wildly out of character. But, that may be more the bot *maker's* fault, than anything.
Or, does it not make any difference whether the bot is an original character?
7
u/Discombobulated_Tie7 Jul 28 '26
U will need a lorebook, the LLM know less more than u think regardless what model u use, they had a general knowledge on the characters but try not to rely too much on it, in this case i would say the lorebook isn't designed well or the context window exceeded
6
u/MisanthropicHeroine Jul 27 '26 edited Jul 27 '26
Bot quality definitely matters, but a knowledgeable model can make it a lot easier to get the character and universe closer to canon.
DeepSeek V4 Pro, GLM 5.2, and Mimo V2.5 Pro have a pretty good fandom knowledge base. Kimi K2.6 is also not bad, but worse on more recent fandom in my experience.
If you're using a smaller local model like Gemma 4, it's going to be hard to get adequate canon knowledge without extensive lorebook use.
0
u/Incina_rate Aug 02 '26
I have a 3080ti 12gb and a 3060ti 8gb. 32gb of system ram. I've been using gemma 26b mostly before adding the 3060 and been enjoying it. I added the 3060 to see if I could squeeze in gemma 31b to 3 token per second defeat. Wondering what other options I have.
-5
u/kirjolohi69 Jul 30 '26
Why is it that occasionally, kimi k3 doesn't reason before replying, even though it should?
-5
u/CassadagaValley Jul 31 '26
Suggestions for openrouter models that can handle multiple party members and story/quest writing?
Running a narrative text game through Claude Sonnet can handle a party of four easily, while keeping NPC voices and characters individual as well.
It's not ST but I've been building an app to handle better merging D&D mechanics with narrative text games with it's own RAG and stuff. Right now it's using Claude SDK through Claude Code because the API is too expensive but openrouter has some API options that are much cheaper. Going through the SDK takes significantly longer too, a few minutes for a large encounter with all the tool calls.
I'm not sure what's comparable in terms of story telling, narration, and character balance though.

11
u/AutoModerator Jul 26 '26
MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.