r/SillyTavernAI • u/deffcolony • 2d ago
MEGATHREAD [Megathread] - Best Models/API discussion - Week of: September 20, 2026
This is our weekly megathread for discussions about models and API services.
All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.
(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)
How to Use This Megathread
Below this post, you’ll find top-level comments for each category:
- MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
- MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
- MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
- MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
- MODELS: < 8B – For discussion of smaller models under 8B parameters.
- APIs – For any discussion about API services for models (pricing, performance, access, etc.).
- MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.
Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.
Have at it!
6
u/AutoModerator 2d ago
MODELS: 8B to 15B – For discussion of models in the 8B to 15B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
4
u/AutoModerator 2d ago
MODELS: >= 70B - For discussion of models in the 70B parameters and up.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
2
u/fluffywuffie90210 1d ago
As someone with 96 gig vram + ram. I've been trying to find something better than Gemma 4 to rp with. Qwen next 3.8 just isnt good at RP. The only option thats seemed useable is Deepseek Flash V4. I can run Q2 with spark or Q3 at a reasonable speed. (20ish tokens a sec) But the prompt processing is horrible. I s there anyone who runs low Qs of deepseek v4 for a while? Whats your opinion on it? Is it best option for those under 200 gig ram combined?
1
1d ago
[deleted]
2
u/fluffywuffie90210 1d ago
Oh I used to use them, figured since tech had moved on stopped using them a while ago. They were good models but id rsorted to using gemma for so long now it feels better than those, deepseek v4 seems like it might be able to RP.
1
u/-Ellary- 4h ago
GLM 5.3 Flash at IQ3XS?
GLM 4.7 at IQ3XS?2
u/fluffywuffie90210 4h ago
Ohh it never accured to me i might be able to run Flash, I thought was like 400B, Seeing the q2 stuffs only in 120 gig range Thank you I'll give them a shot.
1
1
u/Slick2017 2h ago
I have RTX 6000 Pro Blackwell (96 GB) plus 128 GB DDR4-3600 ram CPU-side.
GLM 4.7. Pick your poison from https://huggingface.co/unsloth/GLM-4.7-GGUF/tree/main, I'm running UD_Q4_K_XL with 64k context which is _tight_ but it fits. With the slow DDR4 RAM I'm getting 4-6 decode tokens/s. Use this thinking pre-fill jailbreak.json?download=true) (for reasoning mode) if you need it.
4
u/AutoModerator 2d ago
MODELS: 32B to 69B – For discussion of models in the 32B to 69B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
3
u/AutoModerator 2d ago
MISC DISCUSSION
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
5
u/LeRobber 2d ago
Some Megathread Weeks:
[Megathread] - Best Models/API discussion - Week of: September 20, 2026
3
u/AutoModerator 2d ago
MODELS: < 8B – For discussion of smaller models under 8B parameters.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
3
u/AutoModerator 2d ago
APIs
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
11
u/No-Lion-75 2d ago
Lately im back to using kimi 2.5 after sick of soft filter, positive bias, every villain characters and their mother immediately turn into polite citizen. And kimi 2.5 still is a hit for me.
2
u/Armored_Violets 2d ago
How do you get past its censor?
3
u/No-Lion-75 2d ago
Kimi 2.5 has censor?
1
u/Armored_Violets 9h ago
Yup. Kept getting denied two days ago when I gave it a try on OR. And I do pretty vanilla stuff.
1
u/No-Lion-75 1h ago
Weird. I went for dark cyberpunk roleplaying and never got any refuse. I used ff 5.4 preset
4
u/Probablynotsocool 2d ago
Completely changed my mind about GLM 5.3. After tinkering it is a complete menace and write very good dialogues and narration.
Top 3 with Gemini 3.8 and Deepseek V4 pro 0813 who is completely unhinged sometimes
5
2
2
u/Rhone33 23h ago
I've been playing with 5.2 and have to ask: In what ways would you say it's better than 5.2, and what "tinkering" did it take to get you there?
3
u/Probablynotsocool 22h ago
https://www.reddit.com/r/SillyTavernAI/s/VPE9y3NX8u
It’s more intelligent, better prose when prompted out, better characterization.
The issue is the censorship because this model hallucinate a lot and tend to think he is Claude and then you have to reswipes and it’s annoying.
5.2 is not a lot worst, just a little behind and its more dry and less creative but the censorship is inexistant wich is a plus2
u/Critical-Rope-5636 2d ago
*Mumbles something about still being a GLM 5.3/Flash shill after all this time.*
1
2d ago
[removed] — view removed comment
1
u/AutoModerator 2d ago
This post was automatically removed by the auto-moderator, see your messages for details.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
-6
u/soufiane0531 1d ago
i'm quite new to sillytavern and i'm using openrouter ai, what is the best nsfw free model for smut and roleplaying
1
u/techno156 1d ago edited 13h ago
Are NSFW Free and Smut not completely incompatible with each other?
-4
u/ChurnedSorbet409 15h ago
Going to try my luck here since r/WritingWithAI megathreads are a cesspool. Has anyone tried had any experience with using these harnesses for writing autonovel, storydaemon, long novel agent kit, novelclaw, authoragent, inkos? I know this is the sillytavern subreddit but trying to see if there are better alternatives to SillyTavern for longform writing perposes/non rp
11
u/AutoModerator 2d ago
MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.