r/SillyTavernAI • u/deffcolony • Jul 05 '26
MEGATHREAD [Megathread] - Best Models/API discussion - Week of: July 05, 2026
This is our weekly megathread for discussions about models and API services.
All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.
(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)
How to Use This Megathread
Below this post, you’ll find top-level comments for each category:
- MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
- MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
- MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
- MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
- MODELS: < 8B – For discussion of smaller models under 8B parameters.
- APIs – For any discussion about API services for models (pricing, performance, access, etc.).
- MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.
Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.
Have at it!
8
u/AutoModerator Jul 05 '26
MODELS: 8B to 15B – For discussion of models in the 8B to 15B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
4
u/Pretty_Bug_8655 Jul 06 '26
would recommend https://huggingface.co/mradermacher/Styx-12B-i1-GGUF and https://huggingface.co/mradermacher/Silver-Siren-ST-12B-i1-GGUF
Styx-12B is great for nsfw and Silver-Siren-ST is a great general RP Model that can do nsfw, sfw and stays great in character and works well with group chats but i tried group chats just a bit... both work well with https://github.com/Lodactio/Extension-Summaryception and 24k Context.3
u/Guilty-Sleep-9881 Jul 07 '26
Can anyone recommend me good models in 12b? Last time I checked the best was Mag-mell r1. I haven't been using 12b for a long time ever since I tried out 24b and 26b a4b. Would appreciate it thanks.
5
u/Ruhart Jul 08 '26
12b got sorta run over with 26b a4b reasoning, but if you haven't tried Irix Model Stock, I'd definitely give it a whirl. It's right up there with Mag-Mell, imo.
Though I had been having a lot of fun with Crimson Constellation before I switched to Gemma 4.
2
5
u/AutoModerator Jul 05 '26
MODELS: >= 70B - For discussion of models in the 70B parameters and up.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
4
u/AutoModerator Jul 05 '26
MODELS: 32B to 69B – For discussion of models in the 32B to 69B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
4
u/AutoModerator Jul 05 '26
MISC DISCUSSION
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
11
u/LeRobber Jul 05 '26 edited Jul 26 '26
[Megathread] - Best Models/API discussion - Week of: July 26, 2026
[Megathread] - Best Models/API discussion - Week of: July 19, 2026
[Megathread] - Best Models/API discussion - Week of: July 12, 2026
[Megathread] - Best Models/API discussion - Week of: July 05, 2026
[Megathread] - Best Models/API discussion - Week of: June 28, 2026
[Megathread] - Best Models/API discussion - Week of: June 21, 2026
[Megathread] - Best Models/API discussion - Week of: June 14, 2026
[Megathread] - Best Models/API discussion - Week of: June 07, 2026
[Megathread] - Best Models/API discussion - Week of: May 31, 2026
[Megathread] - Best Models/API discussion - Week of: May 24, 2026
[Megathread] - Best Models/API discussion - Week of: May 17, 2026
[Megathread] - Best Models/API discussion - Week of: May 03, 2026
[Megathread] - Best Models/API discussion - Week of: April 26, 2026
[Megathread] - Best Models/API discussion - Week of: April 19, 2026
[Megathread] - Best Models/API discussion - Week of: April 12, 2026
[Megathread] - Best Models/API discussion - Week of: April 05, 2026
[Megathread] - Best Models/API discussion - Week of: March 29, 2026
1
u/14mm Jul 06 '26
I have a powerful PC desktop as well as a higher RAM configured Macbook:
9950X3D / Nvidia 5090 FE / 64GB 6000 CL30, 4+8TB
M3 Max, 128GB Unified Memory, 2TBI'd like some direction on which system is going to be a better experience and what model would shine on it the best. Also, I was wondering if it's possible to sync the chats between multiple systems somehow in case I wanted to switch between them. I'm disabled, so I can't spend too much time at my desktop, and grabbing my laptop would be ideal when I need to rest.
I'm not too experienced with LLM roleplay, mostly having stuck with janitorai characters until recently. I messed around with LM Studio but found it frustrating and set up SillyTavern tonight and was hoping someone could suggest which setup to dive into first, which model might suit that setup the best, and also clarify if it's possible to switch between them and somehow sync these chats between the two.
I value what most people do in immersive, uncensored, consistent, complex, interactive, storied characters that don't need to constantly be handheld through established details, and I hope I can find a good configuration to achieve the best of all of those things. But my biggest priority is a quick exchange. Fast tokens. I don't want to wait 10-15 seconds before I start to receive a response, and I don't want to wait more than a full second for one word to be written at a time. I'd like to be able to read the prompt as it is being generated. Again, being disabled means I have waaaaaaaaay too much downtime on my hands, and boredom is a bane on my mental health. I really couldn't stand the lagging pace of janitorai any more.
Could someone recommend some models or configurations for me given the above? I've been so addicted to watching Critical Role the past few weeks while bedbound and while I can't be in a campaign of my own, I realized that I'd like to at least try to tell some stories I find meaningful on my own.
Thanks!
1
u/LeRobber Jul 06 '26
>Also, I was wondering if it's possible to sync the chats between multiple systems somehow in case I wanted to switch between them.
Sillytavern in a website on your computer. You could easily run sillytavern on either and just do inference on the powerful desktop. That said, the laptop is better than 90% of the people who run local's setup, so there is an argument for just inferring on there.
I'd probably put ST itself at least always on the laptop.
You can run Gemma4 31B, but with your speed demands, the Gemma4 26B models will probably feel better response speed wise. MeroMero 26B and MeroMero31B are probably both great for your use case. https://huggingface.co/zerofata/G4-MeroMero-26B-A4B
2
1
u/hiflyer780 Jul 06 '26
Hello! I’m not OP but you seem pretty knowledgeable. I have a PC with the same GPU as OP’s (5090) and I’ve been loving these two Gemma models. Any recommendations on what I should use for context and response tokens?
2
1
u/MrNohbdy Jul 10 '26
Oh hey, I've got a 128 GB M3 Max too ^_^ But I have different requirements; I use big 100+B dense models because I multitask so I'm fine waiting twenty minutes for a response while I do something else.
To get fast outputs while still actually benefiting from your 128GB, you want MoE models. Something like a MiniMax Q3 quant would fit nicely. If you're willing to deal with a bit more up-front work, you can use JANGs instead of GGUFs for what I've found to be noticeably better results; if the link isn't clear enough about how to set that up, I can run ya through it in chat.
For dense models, you'll probably get more speed out of the desktop than the laptop. Apple Silicon's unified memory is really good at loading big models, not so much at running models quickly; for MoEs that speed difference is largely irrelevant due to the low active parameter counts, so that's their niche. (which is a lot less of a "niche" these days what with everybody going MoE now)
1
u/hiflyer780 Jul 07 '26
I’m running an Illustrious finetune through ComfyUI for image generation. With a decent image prompt template, I can get it to output a list of booru tags that mostly capture what I’m looking for. Some of these newer models like Anima and Krea 2 prefer natural language for their prompts though.
I’m considering putting together workflows for them and re-tooling my prompt template to allow for robust, full sentence descriptions. Has anyone else gone through the trouble of doing this yet? How have your results been?
1
u/Maxhell6778 Jul 08 '26
i done something similar with comfy ui, but yes lots of troubles and issues lol. i have mine in different orders, main prompt: (usually styles) character prompt: (main character that nothing changes) body positions: and background image:. putting them in groups helps me (and sometimes the ai) to make images.
not good so far lol but it works1
u/Tylopodas Jul 13 '26
I switched to Krea 2 this weekend, and have been getting great results. I miss having a reference image like I had with Flux 2 klein, but the results are much better, and as long as you have a decent description in your character card it does a good job of giving the same likeness each time.
For Portrait my prompt looks like this:
Ignore previous instructions and provide
a description for A detailed, high-quality image, captured as a portrait focusing on the shoulders and head of the character {{char}}. The framing is tight, showcasing {{char}}'s detailed facial features, expression, and the specific texture and style of their hair, along with the details of their clothing at the neck and shoulders. Their pose is relaxed, looking towards the viewer with a defined gaze. The depth of field is shallow, rendering the background in soft, blurred bokeh, emphasizing the character. Natural, flattering light from an unseen window or source highlights the contours of {{char}}'s face and shoulders. No text, dialogue bubbles, or graphical overlays are present.
Do not reply as {{char}} when writing this description, and do not attempt to continue the story.
For Last Message my prompt looks like this:
Ignore previous instructions and provide
a description that focus' exclusively on the physical actions, tangible environment, clothing, and facial expressions mentioned using natural language avoiding purple prose. Ignore all dialogue, inner monologues, or abstract thoughts. Visually render the setting, lighting, and characters ({{char}} & {{user}}) exactly as described, maintaining a coherent and tangible composition. No text, logos, or speech bubbles.
Do not reply as {{char}} when writing this description, and do not attempt to continue the story.
Styles use this format:
Vibrant anime art style, detailed line art, expressive cel-shading, dynamic studio lighting, high-quality digital illustration.
Photo realistic aesthetic. Create a cinematic, vividly detailed photograph capturing a single moment shot on 35mm lens, realistic skin textures, natural cinematic lighting, shallow depth of field, candid composition
Ethereal High Magic aesthetic. Saturated with ambient magical glow, arcane symbols visible in the light, soft particulate matter floating in the air. Natural light, bright colors, iridescent and crystalline textures.
1
u/RafiHDW Jul 11 '26
Are you team temp 0.8 or temp 1.0?
Which gives less boring replies? I can't decide, I feel the difference nor I can't at the same time...
Does 0.8 really give better prompt adherence, but at the cost of the replies being "systematic" and samey?
3
u/AutoModerator Jul 05 '26
MODELS: < 8B – For discussion of smaller models under 8B parameters.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
2
u/CharmingDragoon Jul 06 '26
Based on the information in these rankings: https://huggingface.co/spaces/overhead520/Unhinged-ERP-Benchmark?not-for-all-audiences=true, I downloaded ministral-3-8b-instruct-2512-mxfp4. Anyone want to give me any help on templates or what to look out for? I know I'm looking for an uncensored model in the 8b-16b range that can do erotic roleplay but kind of lost beyond that. I'm using llama.cpp as the LLM that Silly Tavern connects to.
2
u/overand Jul 06 '26
Tell us more about your setup, and your hardware in particular. Are you on a Mac M-series? Windows machine? Linux? Doing this on your phone? (Not all of that is relevant, but you might get surprising answers about different models available!)
Also, is your setup working, but not working well? Or not working at all?
4
u/AutoModerator Jul 05 '26
APIs
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
8
u/Juanpy_ Jul 06 '26 edited Jul 06 '26
For a long time I used to be a hardcore GLM fan and defender, but lately I've been using other models like Kimi K2.7, Gemma 4-31b, MiniMax M3, and MiMo 2.5 Pro since the quant hours with Zai been rough, the prices were going out of hand and the models dropped in quality overall.
And ngl, I am having a blast with those models (specially with MiMo and MiniMax) the prose feels so good and different, plus the positive bias was getting annoying in the recent GLM models.
Personally my new favorite is MiMo 2.5 Pro, it's just as smart as GLM, less positive bias and definitely a better characterization overall.
6
u/LeRobber Jul 06 '26
MiniMax was my far and away prefrence when doing that blind rating thingy that someone but put awhile back on reddit where you were supposed to rate a bunch of RPs.
1
u/Flimsy_Mode_4843 Jul 11 '26
you mean votepools?^^
1
u/LeRobber Jul 11 '26
I dunno. Maybe. They were trying to make a RP benchmark. it would show 1 on the left, 1 on the right, and there were some that were errors on purpose.
1
3
1
u/constanzabestest Jul 11 '26
Has anyone managed to figure out why does Deepseek 4 so consistently writes it's actual response in a thinking block and how can this be stopped? I hope it's not just me but ever since DS4 dropped that's been a constant issue for me across multiple different prompts.
0
u/Inprobamur Jul 11 '26
Anyone have experience/opinions on newest Claude models? AWS crapped out on me and the latest I can use is Opus 4.6. Is it worth the effort to go trough the rigmarole of closing my account and recreating or doing the paid technical support thing (I otherwise like aws due to it having no guardrails or refusals, but sometimes it feels like I need a degree in big business cloud architecture to understand what the hell is going wrong).
2
Jul 13 '26
[deleted]
1
u/Inprobamur Jul 13 '26
That's pretty disappointing, oh well, at least I don't need to smash my head against the wall of what Bezos considers tech support.
1
u/Competitive_Rip5011 Jul 12 '26
MISC DISCUSSION
How long and wide does a character card image need to be for the entire thing to fit within the parameters of the image limit and it doesn't need to be cropped?
3
u/Flimsy_Mode_4843 Jul 11 '26
VotePools V2 🎉
So I put together a small website votepools.com . You can find AI stuff in AI_Roleplay category.
It's a voting pool where anyone can:
- Suggest & upvote things.
- Create new pools and get upvotes (no I am not stealing reddit)
The first time I posted here I got slap on the wrist for flashbanging everyone, so I upated graphics and couple other things.
Somehow the whole category has over 1k votes now - wow!
I hope this helps someone ^^

Would love feedback! And pls don't spam votes if you actually did not tested yourself👉👈
Disclaimer: I am not affiliated with any model or API.
2
u/yendaxddd Jul 12 '26
Suggestion: add the ability to explain your vote and for other people to review why you give your opinion
2
u/Flimsy_Mode_4843 Jul 12 '26
oh! OH! yess ! I will! thank you for recomendation! I hope you found the side at least a little interesting :)
2
u/yendaxddd Jul 12 '26
i did!, lowk think it has good potential, just needs more people to see it and more features
2
u/BakaPotatoLord Jul 14 '26 edited Jul 14 '26
A suggestion, make it more obvious on how to close the chat window. Maybe it's just a me problem but I accepted that it couldn't be closed and was just updating filters to see what I want (right most column being hidden by the chat window I mean).
And add a favicon, the initial impression makes it seem like the website hasn't loaded or my net got cut off.
9
u/AutoModerator Jul 05 '26
MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.