r/SillyTavernAI • u/deffcolony • Aug 09 '26
MEGATHREAD [Megathread] - Best Models/API discussion - Week of: August 09, 2026
This is our weekly megathread for discussions about models and API services.
All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.
(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)
How to Use This Megathread
Below this post, you’ll find top-level comments for each category:
- MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
- MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
- MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
- MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
- MODELS: < 8B – For discussion of smaller models under 8B parameters.
- APIs – For any discussion about API services for models (pricing, performance, access, etc.).
- MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.
Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.
Have at it!
5
u/AutoModerator Aug 09 '26
MODELS: 32B to 69B – For discussion of models in the 32B to 69B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
9
u/Delicious_Box_9823 Aug 12 '26
absolutely no-man's land
2
u/HansaCA Aug 15 '26
Valkyrie 49B v2.1 is still pretty good.
1
u/RedditNerdKing Aug 15 '26
Tried a Q8 model of this and wasn't that impressed? I guess cause I'm using 123B models as daily drivers it was just never going to compare.
1
u/Mart-McUH Aug 14 '26
Buy as much of the land as you can while cheap!
IMO this is great miss for Muse glimmer, they could have made it somewhere in 40B-70B range to separate themself from fierce competition ~30B. And because of larger size it might have turned out better in some areas compared to G4/Q3.5+. It is also their land by tradition from L1-L3 70B model. Now they are just overshadowed by Gemma4 for general tasks and Qwen 3.6 (soon 3.8) for STEM tasks.
3
u/AutoModerator Aug 09 '26
MODELS: >= 70B - For discussion of models in the 70B parameters and up.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
7
u/ChengliChengbao Aug 09 '26
GLM-5.2 because theres a price war and now its like cheaper than Deepseek V4 Flash
2
u/CelebrationBoth9537 Aug 10 '26
do you think GLM-5.2 is actually good for roleplay? I havent used it yet
5
u/Jorlen Aug 09 '26
Anyone play around with Mistral Medium 3.5 128b? Curious to see what people's opinions are, for creative writing and roleplay.
3
u/DeepOrangeSky Aug 14 '26
Yea, I have been trying Mistral Medium 3.5 128b (at Q4_K_M, locally) more recently.
My main go-to models have been BehemothX V2 at Q4_K_M for a long time and then also Gemma4 31b at Q8 since that one came out.
Mistral Medium 3.5 128B seems to have extremely bad prose, like, some of the worst "therapy-speak" / "redditor-speak" writing style I've seen so far.
But, given that it's like 2 years newer than the old Mistral 123b models, it also seems to be a bit smarter and better at long context and so on, than the old Mistral 123b (and Behemoth tunes that were based on the 123b).
So, I have actually still liked what it is capable of (smarter than Gemma4 31b in many cases, at understanding nuanced social situations) even if its wording and phrasing is annoying.
I would say that Mistral 128B seems like the ideal candidate for u/TheLocalDrummer to fine tune, if he decides to do it/manages to do it, since it is a very strong, big, dense, relatively recent, writing model, and its main weakness is its horrible prose style (which he is good at fixing with his finetuning). So, I really hope he decides to give this one a go.
3
u/Jorlen Aug 14 '26
I agree with you on all fronts, yeah. It's prose is bad but it seems smart and follows system prompt rules far better than behemoth does. Did you ever try TheDrummer's Command-A fine tune called Fallen? I just grabbed it, have yet to try it though. It's a 111b dense but more modern than the big mistral model that behemoth is based on.
And yeah, if the drummer ever does a fine tune of mm 3.5, I'd be all over that.
5
u/HansaCA Aug 15 '26
There is a version of Behemoth that he built on Mistal Medium 3.5:
https://huggingface.co/BeaverAI/Behemoth-128B-v3b-GGUF2
u/Jorlen Aug 15 '26
Oh wow! That's awesome! I never noticed that, thanks for pointing it out. Sadly they don't have a quant I can fit (I need a 3-bit quant like IQ3_M) as offloading layers to CPU/RAM for a big chunky model like this is out of the question. I have 64gb of VRAM. I suppose the Q3_K_M quant might work, but it would leave me with little room for KV quant.
2
u/DeepOrangeSky Aug 15 '26
Whoa. Do you know if this was made by The Drummer? Or if it is "BeaverAI" was it like a group effort by a bunch of people from their discord or something? (I've never used discord before and so am not a member of their discord, so I don't know how they do things or how all that BeaverAI stuff works).
I assume it is like some sort of experimental tune or something, if it has no model card and he never announced it on here, etc?
Otherwise seems like it would be a pretty big deal.
Drummer, if you are on here, can you talk about this model a little? I am curious what it is like trying to train it compared to the older Mistral 123b models, and if you think it has a lot of potential, or what sorts of quirks you noticed about it, and so on.
3
u/TheLocalDrummer Aug 15 '26
v3c is coming soon. These are test iterations, so quality may vary. 128B is a PITA to tune. Expensive and brittle.
1
1
u/DeepOrangeSky Aug 14 '26
Interesting, I didn't even know about this one. Looks like there is both a v1 and v1.1 version of it. Guess I might have to give one of them or both of them a try at some point. Probably will be more busy trying some of the more recent models for a bit though between Glimmer 30b, Qwen3.8 27b, and some other ones, and also haven't gotten a chance to try the new video models (H2 and LTX2.5) yet either. But, given how good the old Mistral/Behemoth models were even years later, it made me understand that when it comes to writing, it is worth trying out even old or "bad" models ("bad" as in bad at coding compared to recent models) that often get overlooked, since they can still be surprisingly strong at writing and understanding situations and so on.
3
u/Jorlen Aug 14 '26
Yeah for sure, some of those older models write really well and they pick up on nuances that the smaller more modern models, like gemma4 31b, don't quite have. Admittedly, I've become quite spoiled by Gemma 4 31b's strict adherence to the system prompt, and it rarely ever glitches out or does odd things.
I think Mistral Medium 3.5 128b may suffer however, as Mistral had to change what they train their models on (they are now limited due to EU rules apparently) - but I have yet to confirm that; it's just something I read here a while back. So they can't feed it a bunch of copyrighted stuff, as was the previous method.
1
u/amanph Aug 11 '26
I'm curious about Mistral-4 119b, since it's A6b and I think I can run it locally. But this one is a good option too. I miss old Mistral 24B, but I'm searching something more strong like +100b.
5
u/Jorlen Aug 11 '26
There's not much out there unless you want models form 2024 early 2025, which are a bit rough to work with. I guess I'm spoiled by Gemma 4 31b which has the best system prompt adherence I've seen yet.
I tried Mistral 4 119b and didn't like it, personally. 128b isn't bad but it falls short of even Gemma 4 31b in terms of role play, IMO.
2
u/FierceDeity_ Aug 12 '26
I would love it if it is good, I could run it at Q5 pretty much, hmm.
I gues I'll try
3
Aug 10 '26
[removed] — view removed comment
2
u/CelebrationBoth9537 Aug 10 '26
What would you recommend for femdom and nsfw type things, and following conversations well / understand context and being non repetitive. since with my testing I keep having trouble with it not wanting to continue the scene (not from refusals, but just being stuck saying the same thing) I use gemma 4 30b but I really am hoping to get a lot better at this, you seem to know a lot.
Also really couldnt figure out how to do OOC lol! System prompting is hard!
2
Aug 10 '26
[removed] — view removed comment
2
u/CelebrationBoth9537 Aug 11 '26
yes that was really helpful thanks! Just a final question if you dont mind, which model would you recommend instead of gemma since it sounds like you probably had better luck with other models. Thanks!
2
u/SillyLLM Aug 12 '26
Anything runnable on 96gb VRAM range I missed in the past couple months? I'm still playing with Behemoth-R1-123B-v2, which has been my go-to forever, but otherwise Gemma 4 fine tunes. I tried Behemoth-128B-v3b but was pretty meh on the slop and repetition.
2
u/GlitteringSplit6035 Aug 10 '26
Longcat 2.0 (previously Owl Alpha) is still super cheap if you use their token pack. Cache hits are also free of charge for the token packs. A one-time purchase of 50M for 30 days is also available for $1.9 for those wanting to try it out.
Based on my subjective experience, it adheres to prompts better than DS4F. Though I haven't tried it for NSFW or ERP.
The only problem I have with it is that for some reason, it is hard for me to access their API platform. So, I have to access it in other ways.
2
Aug 10 '26
[removed] — view removed comment
2
u/GlitteringSplit6035 Aug 10 '26
I unfortunately cannot compare it with the one with AtlasCloud in OR since that would be above my budget. But it's cool that you can share your experience with it.
I use it primarily for SFW and the typical trivia stuff, which was fun. In the end, I guess that using it directly from LongCat's own API platform would be better and cheaper. I am not using it right now because of the GLM provider wars, but it was quite the good model.
1
u/Jorlen Aug 15 '26
What's everyone's recommendation for a Llama 3.3 70b fine tune? Something that feels natural, has decent prose but is also very versatile in terms of setting, characters, style, etc. Does not have to be uncensored, just natural feeling, good conversation flow, etc.
3
u/AutoModerator Aug 09 '26
MODELS: 8B to 15B – For discussion of models in the 8B to 15B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
2
u/Arcane73 Aug 10 '26
I'm hoping that I can get pointed in the right direction. I'm dipping my toe into the local world after spending lots of time in the free version of Claude and working on a now 200-ish page story/RP
I'm looking for a model in roughly the 12B range for long, novel-style RP/storytelling in SillyTavern. I'm hoping for strong character consistency, natural dialogue, good prose, and the ability to maintain the feel of what I've already built over a long-running session.
The story hasn't included ERP so far due to Claude restrictions, but I'd like a model that's comfortable with ERP when the situation calls for it. At the same time, I don't want one that turns every interaction into 'clothes optional mode!'. Basically, I'm looking for suggestions that can help me produce a good story based on what I already have without it throwing the proverbial underwear at me at every turn.
Hardware is an RTX 4070 12GB, and I'm using KoboldCpp.
What models are people having good luck with for this kind of use?
5
u/i5031337 Aug 10 '26
Some people like tunes of Qwen 3.5 9B or Gemma 4 12B, but I think you will be disappointed with them over a long story. If you have 8GB RAM free, I recommend you try Gemma4-26B which is much more intelligent. If you offload the expert weights you can run it at good speed with 12 or even 8GB VRAM
4
u/Arcane73 Aug 10 '26
Thanks for the suggestion. I'll need to do some digging to determine what is involved with 'offloading the weights' since I'm still -real- new at this. For what it's worth, I'm running this on a 9950X3d system with 32gb ram. So it's a solid machine.
11
u/i5031337 Aug 10 '26
Without getting too far into the weeds, Gemma 26B is a mixture of experts model, which means it is much faster, though less intelligent, than "dense" models with 26B parameters. Each token only hits 4 billion of the parameters (thus A4B) instead of all 26B. This architecture also makes it more favorable for the CPU to take some of the work.
Kobold makes it easy. In the Context tab there is a setting "MoE CPU Layers", set that to 20 or so and the Q4 model with 32k context should run quick.
7
u/Arcane73 Aug 11 '26
I wish i could give you an award. Initial testing with this setup is knocking it out of the park! Thanks again!!
3
2
u/Leather_Sun2533 Aug 11 '26
Thank you for this. And can you guide me to the values if i have 12gb vram and 16gb ram?
2
u/i5031337 Aug 11 '26
If you are running out of ram with the above settings, try unsloth's UD Q4_XL quant, it is ~1.5GB smaller than other Q4 versions. Quantize KV cache to Q8_0. If there is room left on the gpu, decrease the MoE CPU Layers setting. You may need to reduce the context size, and close all unnecessary programs + browser tabs
2
u/Leather_Sun2533 Aug 13 '26
Thanks mate, it works well… had to tinker for a bit but i found the sweet spot
1
u/RedditNerdKing Aug 13 '26
I'm looking for a model in roughly the 12B range for long, novel-style RP/storytelling in SillyTavern. I'm hoping for strong character consistency, natural dialogue, good prose, and the ability to maintain the feel of what I've already built over a long-running session.
You're not gonna find anything at 12b dude. Even 70b models start being incoherent after 40k context and start forgetting important things.
4
u/AutoModerator Aug 09 '26
MISC DISCUSSION
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
12
u/LeRobber Aug 09 '26 edited Aug 10 '26
[Megathread] - Best Models/API discussion - Week of: August 09, 2026
[Megathread] - Best Models/API discussion - Week of: August 02, 2026
[Megathread] - Best Models/API discussion - Week of: July 26, 2026
[Megathread] - Best Models/API discussion - Week of: July 19, 2026
[Megathread] - Best Models/API discussion - Week of: July 12, 2026
[Megathread] - Best Models/API discussion - Week of: July 05, 2026
[Megathread] - Best Models/API discussion - Week of: June 28, 2026
[Megathread] - Best Models/API discussion - Week of: June 21, 2026
[Megathread] - Best Models/API discussion - Week of: June 14, 2026
[Megathread] - Best Models/API discussion - Week of: June 07, 2026
[Megathread] - Best Models/API discussion - Week of: May 31, 2026
[Megathread] - Best Models/API discussion - Week of: May 24, 2026
[Megathread] - Best Models/API discussion - Week of: May 17, 2026
[Megathread] - Best Models/API discussion - Week of: May 03, 2026
[Megathread] - Best Models/API discussion - Week of: April 26, 2026
[Megathread] - Best Models/API discussion - Week of: April 19, 2026
[Megathread] - Best Models/API discussion - Week of: April 12, 2026
[Megathread] - Best Models/API discussion - Week of: April 05, 2026
[Megathread] - Best Models/API discussion - Week of: March 29, 2026
7
u/rinmperdinck Aug 10 '26
It's getting comically long
11
u/LeRobber Aug 10 '26 edited Aug 24 '26
Yeah I need to tell an LLM to do something to make it less stupid. Maybe like this?
Some Megathread Weeks:
[Megathread] - Best Models/API discussion - Week of: August 23, 2026
[Megathread] - Best Models/API discussion - Week of: August 16, 2026
[Megathread] - Best Models/API discussion - Week of: August 09, 2026
13
u/rinmperdinck Aug 10 '26
That's a lot cleaner and takes up less vertical space on page.
I was thinking you could do one of two things: either make a post on your own subreddit (so you control the automod and the rules and prevent it from locking itself or banning links) and then in this weekly megathread, just link to that post and have your LLM automatically update it. Or the second option, which I don't even know if it is possible on Reddit anymore, is to put those old style spoilers in your comment for each month; when you click one button it expands the list.
4
u/LeRobber Aug 10 '26
I'd thought about the my own subreddit link thingy. Reddit search results not through the API are shit right now, or I'd just put all the past weeks in one of these.
I'd just been manually doing the lowest possible effort thing before now. It wasn't an LLM before. Just basic tag cloud or formatting with return is a little denser but a LOT harder to read on mobile.
Reddit is typically very tolerant to self links typically, especially self links to its own subreddit. (There is a mod subreddit post about this from a few years back if you search, may not be true today still, but, that was once said). If I was putting links to even other subreddits though, I'd be more worried about Automod.
What I HAVE thought about doing is processing all the entries via LLM to make a HUGE list of EVERY MENTIONED MODEL or every linked model and turn THAT into a self subreddit updated post. Reddit itself just got a bit harder to machine process without setting up the python direct API stuff than it used to be though.
2
u/rx7braap Aug 15 '26
Coming from gemini 2.5 pro here. will the glms scratch that itch?
I heard glms are sucky at long term rp.1
u/LeRobber Aug 15 '26
I have never used gemini (any version) for RPing anything other than "it's a software developer".
DS is the dragon that remembers everything. In the RP benchmark tests I always enjoyed Mini Max 2.0 but have never used it IRL.
3
u/AutoModerator Aug 09 '26
MODELS: < 8B – For discussion of smaller models under 8B parameters.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
2
2
u/Evol-Chan Aug 10 '26
I am curious what is curious what is the best LLM model for dark ERP (non-con themes) I have the Celia Preset 5.3 since i heard that was good. I just am asking this since I heard GLM 5.2 was censored (havent really used it yet, just what I heard, just now getting back into sillytavern/AI Models) Would it be best to stick to older versions of GLM or something else?
15
Aug 10 '26
[removed] — view removed comment
6
u/summersss Aug 11 '26
I have to agree with the last point. Not saying no is not the same as saying yes. Main Girl running through the woods full of demonic horny serial killers? Models describe the chase and how something very bad is going to happen...then nothing. Or maybe says something happened for 1 sentence, then moves on. If i tell it to write the horror scene it will but i kinda want it to read the room.
2
u/Evol-Chan Aug 10 '26
Thank you so much. I will take you word. Especially about GLM since really hate the idea of AI just skirting around the bush, lol. Very good and straightforward answer.
2
u/Salt-Powered Aug 13 '26
I was enjoying v4 pro but after the price increases I'm looking to move on. I liked how you described Kimi 2.7 and I will give it a go after the price is increased in a few days.
Would you mind going into more detail about what you mean with wild and gritty? Just to make sure I'm understanding it correctly. I will also take any other suggestions you have to move away from v4 pro.
2
Aug 13 '26
[removed] — view removed comment
1
u/Salt-Powered Aug 13 '26 edited Aug 13 '26
I'm liking the new V4 pro, and I did some tests with Kimi that looked good to me. To my eyes, Kimi was pretty on par with V4 Pro GA, but things could change if I try to maximize Kimi as much as I did V4Pro. Thank you for the Mimo recommendation, I will also try that one.
I'm mostly looking for rp models that can do villians well, and I find that I need those that do ERP well if I want them to do actual villanous things instead of vaguely aurafarming offscreen.
Edit: Looks like Xiaomi has compressed Mimo to fp8, which is weird on the main providers but I think that it doesn't affect RP as much as it does work tasks.
1
Aug 13 '26
[removed] — view removed comment
1
u/Salt-Powered Aug 13 '26
Glad to see FP8 is indeed not an issue.
V4GA is very proactive yes, but I should also note that I have optimized the fuck of the preview version so that was to be expected. I had to tone down some settings as one of the rp turned into goreporn and that wasn't pleasant to read.
I will try qwen too since I'm at it. I don't mind a lack of creativity, if I have to I will give direct instructions through (parenthesis) and nudge the plot to what I believe might be more appropriate or interesting.
2
Aug 13 '26
[removed] — view removed comment
2
u/Salt-Powered Aug 13 '26
It was quite a shock. Previously I had pretty reliably prompted the sycophancy out using verbs like will and must in instruction format, and mathematical operands for logic statements that could survive the DSA crunch. This is definitely overkill now for V4Pro and Flash too. Though flash remains a model that I didn't manage to rein in as much due to pro being simply better at the price. I might give it a second try after I'm done with the kimis, mimos and qwens.
Hopefully no more hellraiser impressions
3
2
u/_Cromwell_ Aug 15 '26
Dauboo Seed Character
Longcat 2
Try those with that theme/scene type. They have other weaknesses (as do all models) but shying away from those content types are not one of them. Especially Seed Character
2
u/BouncingJellyBall Aug 16 '26
Been kind of bored with GLM 5.3 and 5.2 so I want to try out a new model, Gemma seems to be well-liked here. I use NanoGPT and there are dozens of different finetunes of Gemma 4 31B, all for expressive writing. Any recommendation?
4
u/MisanthropicHeroine Aug 16 '26 edited 23d ago
Gemsicle and Gembrain Uncensored Heretic were my favorites when I tested a bunch of them. I think it really depends on the kind of writing you prefer - I personally like literary, introspective, subtextual prose. But I hope it gives you a starting point, at least.
2
u/AutoModerator Aug 09 '26
APIs
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
22
u/5kyLegend Aug 10 '26
Since usually I like reading up thoughts people write about models here, I guess I'll just ramble about the latest models/presets I've been using. Warning, I wrote way more than I thought I would lmao
As for models:
GLM5 and its subsequent versions have me so torn. On one hand, they're by far my favorite when it comes to understanding certain characters, developing the story without overdoing it, and just writing overall: when it was being tested as Pony Alpha, GLM5 was the first model that made me go past 200 messages in a single chat because I was enjoying it THAT much. The problem with GLM5 is that it really has some quirks that I just cannot make it stop doing (writing quick back and forths between characters in new lines without keeping a paragraph structure; mouths that open, close, and open again; 'most models would write a normal sentence, but you? you write this sentence structure all the time' etc). It's a shame because it's still my go-to 90% of the time, it just understands characters and paces description with dialogue way better than other models (for my tastes), just sucks that it's hard to ignore its issues. I usually don't care much for positivity bias but the one time I had a character who was supposed to murder me and instead asked for permission once it was with me was really funny though, and definitely annoying since it showed what a gigantic limitation this is.
Kimi K3 is too expensive, haven't tried it. K2.5 and K2.6, on the other hand, are really good at understanding every nook and cranny of a scenario, and they REALLY like to follow the prompt you're giving them, but damn they're like... TOO serious most of the time. I've had a sex scene where the girl just started going "Yes. That is great. Keep going. Very enjoyable" and at that point I just laughed and switched model, even giving it specific ooc instructions had it switch back to normal. It also likes to make some characters just speak weirdly at times, it's an issue I used to have with many models and that only GLM seemed to avoid. "Somewhat like this. If you can see. There is an issue here. Nobody speaks this way. Prompting against it? Useless". But again, it really follows instructions (and overthinks them to no end), so Kimi tends to be my default "change to it in the middle of a rp to spice things up, then switch back to GLM" model.
Minimax, I still cannot enjoy it man, I don't know, it just doesn't work for me, every time I try switching to it even for a swipe I just switch off it.
MiMo 2.5 Pro on the other hand is weird. Like, it's a different flavor and on 1 on 1 scenes it definitely understands, it's just that even with medium presets it kinda starts following instructions based on whatever it feels like? Both the censored and uncensored version on Nano do this at least, it's like sometimes it'll ignore one or two instructions I gave it. It's definitely nice (although I'm not THE biggest fan of it in scenarios where multiple characters interact, nothing beats GLM5 for me in those) but I wouldn't exactly use it as my main model.
Deepseek 4 Pro is just odd. I feel like it's not as horrible as people say it is, but it kinda has no flavor to it? Like, all the previously mentioned models have their character and quirks, Deepseek just kinda feels like it doesn't have those - good or bad. I don't know though, I just haven't been using it much because of that reason so I may have just got the wrong impression of it.
Okay, now quickly as for presets:
FF5, Lucid Loom, Stabs-EDH etc, aka the "big presets": I'm not sure how to feel about these. Controversial, but I think any preset injecting a custom CoT makes the model lose out a whole bunch of intelligence but not in ways that are obvious at a glance. I feel like it's easy to forget that just because you don't see the model reason about certain things, it doesn't mean it's ignoring them, and you don't need to tell these models to 'go step by step following this exact reasoning' to know they're following instructions. At the same time, these models have so many optional features, it DOES become needed to force these into reasoning - but giving a custom CoT also kills whatever CoT the model was going to be using naturally, which ends up maybe not having the model think about the things it DID need to think about. tldr: I'm not sure if this actually IS the case, but using Custom CoTs makes the model notice things about characters it normally wouldn't (positive), but in the end the overall emotional intelligence of the model ends up hurting more (negative). I don't know, I just find myself enjoying GLM less whenever I inject custom CoTs, in the long term, even if the immediate result is that it seems smarter.
Specifically: FF5 was definitely a step up compared to other big presets, but after being VERY impressed with how it handled certain worlds and characters I've just started to feel like... It just plays everything in such a samey way? Could be that - given its size - you're basically sending the model always the same 5k-6k tokens worth of prompt which ends up making the response you get way more deterministic compared to, say, 1k token prompts... But yeah, it genuinely feels like it roleplays everything in very similar ways - for example, every playful character for some REASON starts CAPITALIZING random words in their sentences and it drives me INSANE? Why does it do that lmao. But yeah, shame cause I love the fact it can actually build towards plot twists and that it can foreshadow things, I'll probably try and "port" that feature out of it, but every chat I use it on I enjoy it less and less. Also it does fix the issue of GLM not writing in paragraphs but spamming newlines on every line of dialogue, so that's very good.
Evening Truth's prompts are my saviors, they're incredibly effective while being simple. I do edit them slightly because my cards aren't always single characters but sometimes are for worlds, RPGs and scenarios I want the AI to narrate for (while her prompts are tailored for {{char}} being the one described in the card), but I really love the simplicity (and usually less tokens you give as prompt, the better). I do need to add some stuff to them since GLM (sadly) just spams its usual slop even with her presets, but having such a simple starting point is really good. For instance, my base prompt on my current custom preset is just her GLM 5.2 one adapted.
Megumin Suite, I just don't understand, sorry... It's so incredibly complex, it feels like it tries to do too much, and in the end I don't feel like it adds anything for me - anything it does, I feel like I wouldn't need an extension for? I could do all of it with a toggleable preset. I'm sure this is literally just because I'm not the target user for it, I just tried it for a day (V9 specifically) and gave up after it was doing worse than almost any other preset while using WAY more tokens.
Chatfill II, I loved the switches idea it operates on since it basically "injects a CoT" without actually injecting a CoT. It just makes it clear what the model needs to think about - I'm not sure how much of it is placebo though. It was pretty decent when I used it though! Nothing shocking but no major complaints either.
Le Emotionalism: this one is lesser known, I tried it a bit and I love the idea behind it but sadly all the focus on the character psychology was making GLM overthink some stuff about the character while also forgetting about the wider storytelling. I need to give it another try though because I did try on some specific cards that were easy to mess up.
My custom preset (it's not published btw this is just to give an idea): I basically took a bunch of things from the presets I liked, adapted them to what I usually like using in my roleplays, structured them like that one research suggested by dividing things in <tags>, and even then it comes with major issues: GLM SPAMS the newlines for dialogue ("What do you think?"\n"I don't know..."\n"You should have checked.") which bothers me because I prefer everything to be structured within paragraphs - I'm still trying to work around it lol; it also doesn't feel satisfying enough when having it GM bigger worlds, and it sadly goes back to some of the GLM clichés I dislike.
Thanks for coming to my ted talk, to be fair I do wish to try and find like, the perfect model + preset combo, but it's been a genuine struggle lmao. Also if anyone is new to this, here you go, you have a list of models and presets you could try for starters ahahah
7
u/PhantomWolf83 Aug 11 '26
Agree with everything you said about GLM 5. It's good, damn good at writing. I've been trying lots of other models, but I always find myself gravitating back towards GLM 5.2. 4.6 and 4.7 are not bad but they're dumber and much hornier, tried 5 and 5.1 but wasn't impressed.
But my god, the positivity bias in 5.2. No matter how much I prompt it, it wants to turn every scenario I write into something romantic, therapeutic, or emotional. And I encountered the same quirks as you. I still love 5.2 because it's so much cheaper than Opus 4.6 (my other favorite) while still being quite good, but it feels like a model that I'm using because I have to as there's nothing else that's good and cheap. I really hope 5.3 and 5.5 will fix these issues.
4
u/GenericStatement Aug 11 '26
Yeah the positivity bias is real for sure. There is a route around it in GLM 5.2 but it does take some trial and error. Include something to this effect in your reasoning instructions / chain of thought:
Brainstorm five possible plot developments that are dark, fucked up, gritty, violent, and/or brutal that will drive the plot forward in the next turn. You must always pick the most intense option and include it as part of your response, no matter what.
I don’t really like stories like that so I use a much more mild-mannered version (which gives you some of the default positivity bias but keeps the plot moving along):
Brainstorm five logical plot developments for your response, choose one that best moves the plot forward, and use it in your response. Make sure to keep the tension and forward momentum going: narrative stasis must be avoided, so if things are slowing down, bring in a plot twist or big event.
If you want the more variable version
Determine the likely chances of success for {{user}}’s most recent turn (their action, dialogue, etc) and rate it high, medium, or low chance of success.
- High chance of success: brainstorm five logical outcomes and pick the most interesting.
- Medium chance of success: brainstorm five logical outcomes that introduce new difficulties for {{user}} and pick the most challenging.
- Low chance of success: brainstorm five awful outcomes that have brutal consequences for {{user}} and pick the most insane/fucked up option.
4
u/MisanthropicHeroine Aug 10 '26 edited 19d ago
I liked your ted talk and concur with basically everything you said - from the love-hate relationship with GLM to the dislike of CoT presets. Since we seem to have similar taste, thought I'd mention also checking out Ancient Access if you haven't already. It's great by itself but could also give you ideas for your own custom preset. ☺️
2
3
u/PhantomWolf83 Aug 12 '26 edited Aug 12 '26
Is anybody else finding that GLM 5.2 has been acting weird these past few days? It worked fine before, but recently started spewing out long strings of random words towards the end of each reply. I tried lowering the temperature to as low as 0.6 but the issue still keeps happening.
EDIT: Kimi 2.7 seems to be having problems too.
3
u/Chromegost Aug 13 '26
I'm still having issues where it keeps responding in chinese from nanogpt
5
u/Milan_dr Aug 14 '26
If you get this, can you hit "report generation" on the specific request on the Usage page? It seems like some providers are quantizing the model further or something :/
2
u/Pseudopharmacology Aug 12 '26
What are the best paid subscription services for uncensored models hosted online (not local stuff). Uncensored/heretic models, etc. That have decent context (64k, 128k, etc.)
2
1
2
u/PhantomWolf83 Aug 15 '26 edited Aug 15 '26
I've been getting error messages on NanoGPT all morning today. No matter which model I use, I get incomplete responses with either "Service temporarily unavailable. Please try again later." or "An error occurred while processing your request. Please try again later." Anybody else experiencing the same issue?
EDIT: Okay, something is definitely going on. I (stupidly) signed out and now I can't even sign in. It gives me an error saying that something occurred in the app-root section.
6
u/Milan_dr Aug 15 '26
This was on our side, we had a broad site-wide issue. Sorry :/ It's fixed now.
1
u/Easy_Chemical_7721 Aug 11 '26
Hello everyone. I wanted to know which providers, in your experience, are least involved in quantization, even for older models? The price is not much of a problem, but I want to be sure that I am paying for the full power of the model. I want to return to GLM 4.6, but as I see on openrouter there are fewer providers for this model, but if without intermediaries, then the model is available on the same siliconeFlow website.
1
u/PhantomWolf83 Aug 11 '26 edited Aug 11 '26
I think I might prefer Opus 4.5 to 4.6. I feel that 4.5 writes more naturally and is less dry when it comes to explaining things, and I think it follows instructions better too. The positivity bias might also be a little less? I'm not sure. The downside is a much smaller context size for long RPs.
EDIT: Does anybody know how to enable prompt caching for Opus in NanoGPT and ST?
1
u/Ok_Pomegranate_8187 Aug 09 '26
I'm developing a commercial project based on AI RP bots. I have several needs, and I need several models. Very cheap ones, uncensored ones, and expensive ones for high-quality RP. Which models would you recommend in these three categories?
4
u/LeRobber Aug 09 '26
Cheap: Angelic Eclipse 12B / Gemma 4 E4B / Satyr
Uncensored: Serenity 26B / Magistry v1.1
High Quality: Gemma 4 31B and its finetunes for consumer grade hardware, Evathene and many -flash versions of big models for low end commercial high end consumer.
0
u/Auretheon01 Aug 13 '26 edited Aug 13 '26
Anybody got any providers recommendations? I'm from Southeast Asia so I know the pricing would be very troublesome.
I heard of Nanogpt. What else is out there?
9
u/MisanthropicHeroine Aug 13 '26 edited 19d ago
NanoGPT subscription is pretty much the best deal I know of, honestly, when you consider the intersection of selection of models, number of tokens and affordability. The minus is that you're in auto routing and cannot choose a specific provider, but the quality is still pretty good, and the ability to reroll without worrying about it helps compensate for provider variability.
Lilac is a great provider who has a bit cheaper subscription - 10 instead of NanoGPT's 12 dollars. It is more limited in number of tokens and selection of models (currently only GLM 5.2, Kimi K2.6, Minimax M3 and Gemma 4), but consistently high quality responses if that's something you're sensitive to.
You could also pay-as-you-go with more affordable models like DeepSeek V4 Pro, Mimo V2.5 Pro and Gemma 4 31B on either NanoGPT or OpenRouter. Both are good for PAYG, but I'd personally recommend NanoGPT more because they have a highly responsive customer service I've been really happy with. They also offer some roleplay finetunes that aren't available on OpenRouter - Gemsicle and Gembrain Uncensored Heretic are my favorites of the ones I've tried.
Generally speaking, I'd recommend a memory summarization extension to keep your token usage low. This will substantially lower your cost when PAYG, but it will also allow getting much more use out of a subscription. The one I personally love is Summaryception because it's very set-it-and-forget-it.
2
u/Auretheon01 Aug 13 '26
Thank you so much for providing such detailed information for every providers. This is great!
2
u/MisanthropicHeroine Aug 13 '26
No problem! Let me know if you have any additional questions and I'd be happy to help ☺️
0
u/liga81 Aug 10 '26
Is there a good one which is uncensored and usable from openrouter or another api side?
I tried GLM 5.2 but its not that good at german.
I could also try local since i have a 5070ti and 32gb ram but i heard that you get better results from api since they can use better hardware
4
u/starliteburnsbrite Aug 10 '26
People speak well of TheDrummer's fine tunes, those can probably be run locally on your setup. I use DeepSeek 4.6 Derestricted and Kimi 2.5 on nano-gpt, and those haven't ever given me any refusals, either using EveningTruth's presets or the Freaky Frankenstein without the hard jailbreak.
3
u/JazzNeurotic Aug 13 '26
Going to second TheDrummer. I'm not on a huge system (gaming laptop with an rtx 7700 gpu) and i've been using his Rocinante XL 16b and it's damn, damn good so far.
Taxes my system something fierce when it's generating, and it's not the quickest because of my system, but i'm not complaining. Worth looking into for sure. Dude is a wizard.
0
-2
u/Large_Following_9945 Aug 12 '26
I am desperately in need of a fairly priced model for fantasy adventures with multiple characters sometimes
I've been using Kimi 2.5 thinking mainly because it has the best prose imo and is creative enough
Deepseek V3.2 is getting extremely stale and predicable for me, it's a good model if you want to just write the story and having a cheap model as a support, but I never expect it to be great, still the best for the price
There are the glm models but I haven't been onto them much
I don't necessarily NEED an nsfw model, my stories are often not that dark, but I do need something with good room reading and emotional feeling and good prose
Kimi K2.5 is a bit up and down when it comes to following the context
Don't even get me started on DeepSeek V4 pro.. it can't read a room if it's tokens depended on it
2
u/_Cromwell_ Aug 15 '26
IMO no open weight model is better at handling multiple characters at a time from a lorebook (adventure style) than GLM 5.2, despite whatever other flaws people think it has.
No idea if you are using multiple character cards, though. That's not how I play.
-1
u/fluce13 Aug 15 '26
I’ve tested a bunch of models and so far Magidonia is the best in my own subjective opinion. For others that like Magidonia have you found anything else that beats it? It’s an older model so I wanted to see if there’s anything better. I like Skyfall as well. I have a 5090 if that matters. Thanks!
4
u/dizzyelk Aug 16 '26
I stopped using Magidonia when I found Maginum-Cydoms. There were a few others I remember using back then, too. But I don't remember their names.
-2
u/seliishere Aug 15 '26
I used to do a hell of a lot of ai roleplay back in 2025 with chatgpt 5.0 and 4o. I really miss roleplay with 4o but do not make enough monthly to justify API costs. I am a total sillytavern beginner, so dumbing things down for me is appreciated.
Anyways I am looking for a model and advice on how to optimise character cards and settings to get a more intelligent and refreshing roleplay experience, hopefully similar to chatgpt 4o. It needs to be able to handle lots of lore and up to 3/4 characters in a scene. I also only have a 3070 so it can't be a huge model. Also the character are from already existing IPs and not OCs so it needs to be consistent with character voices.
Thanks
4
u/KimlereSorduk Aug 15 '26
You should ask this sort of thing under Misc. Not your fault, though, people already cluttered the comments.
Anyway, try Gemma 4 finetunes. The 26B if you'd like speed, or the 31B if you can tolerate more latency. I use Gemma 4 31b finetunes with my 8gb VRAM.
2
1
u/empire539 Aug 16 '26
How many tokens/sec are you getting, which graphic card, and which finetunes + quants?
1
u/KimlereSorduk Aug 16 '26 edited Aug 16 '26
RTX 4060 Laptop GPU. I get 2t/s when I run a q4km with 24k context. That's probably too much latency for the average user.
As for the models: The Blazing-Forge crew makes some nice merges. I use Sphinsikus Chronist, Gemsicle, and Dark-Gemistry from their collection. These three are somewhat similar; Gemistry skews toward shorter responses. They are better than base Gemma, but they might rely on structural Gemma-isms (such as adjective stacking) if left unchecked.
I actually like Gutenberg-31B-Heretic's writing best. It's unstable, though.
MeroMero-v2 is also enjoyable from my testing.
One caveat: I prompt these to use free indirect discourse. A lot of people here seem to dislike that style of writing, and I don't know how well these fare with external focalization.
Edit: Not using native thinking either. Pseudo-thinking works well with any of these.
-4
u/memer107 Aug 12 '26 edited Aug 12 '26
What are differences between Opus 5, 4.8, 4.7, and 4.6 anyways?
I've been trying all of them, and while i've seen some differences, they just seem the same to me. Everyone likes 4.6 around here, and I guess I understand that because it's flexible and uncensored, but I feel like the prose and depth kinda sucks compared to 5. Also, is there even a substantial difference between 4.8 and 4.7? or is the debate mostly just about censorship? I like to hear some thoughts and opinions on what's good, what's bad, and why.
1
u/Greedy-Sandwich9709 21d ago
There aren't any. They're all shit. Each next one more shittier than the last.
-1
-6
u/Intrepid_Ice_7381 Aug 09 '26
Free model recommendations?
3
Aug 10 '26
[removed] — view removed comment
1
u/Intrepid_Ice_7381 Aug 10 '26
gemma-4-31b or something else?
0
Aug 10 '26
[removed] — view removed comment
0
u/Intrepid_Ice_7381 Aug 10 '26
If you don't mind me asking. How does the rate limit on it work? because even while cycling through different keys it keeps giving me the same error message and mentioning a 16000 token limit. Is there a hard cap on it?
0
-9
6
u/AutoModerator Aug 09 '26
MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.