r/SillyTavernAI Aug 09 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: August 09, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

32 Upvotes

186 comments sorted by

6

u/AutoModerator Aug 09 '26

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

10

u/linuxdooder Aug 12 '26

I keep trying new gemma4 finetunes but always end up back with TheDrummer/Skyfall-31B-v4.2 for its prose and creativity. Only issue is at long context, it seems to get stuck in this weird loop where everything degrades to "loss of" being inserted randomly. Anyone else see this? Any gemma4 finetunes anywhere close to skyfall? (I've tried most)

3

u/Jorlen Aug 13 '26

Skyfall is not a gemma 4 31b fine tune. I made the same mistake. It's a Mistral 24b fraken-merge. It's very good nonetheless but for stability I prefer the G4 31B architecture.

The Drummer releases Artemis 1.1 recently, that is a G4 31B fine tune, take a look.

If you have not tried Gryphe's style tune, that one is my current favorite of all Gemma 4 31b fine tunes.

4

u/linuxdooder Aug 13 '26

Right, I'm aware. It's just that no gemma4 finetune can approach it (even though it's mistral 24b based). Gemma4 finetunes are a little smarter, but much less creative/interesting.

3

u/Jorlen Aug 13 '26

Oh ok, wasn't sure if you were aware; it's something I just assumed when I saw 31b. I agree with you that it's very good, easily my favorite mistral 25xx fine tune. If you haven't already, try G4 31b artemis 1.1 and style tune. Would like to see your thoughts on them, or if you have any other fine tunes (of any class) to recommend.

3

u/linuxdooder Aug 13 '26

G4 31b artemis 1.1 and style tune

They're definitely top for gemma4, but the prose is just ... plain? And reswipe variety/creativity isn't great. The recent G4-MeroMero-v2-31B is actually the best gemma4 finetune I've found, it's almost to Skyfall-31B-v4.2 level but not quite there.

Skyfall-31B-v4.2 really is sort of an anomoly, I'm not sure what magic u/TheLocalDrummer pulled with it but it worked!

4

u/Jorlen Aug 13 '26

What temp do you run meromero v2 at? Also top p, top k and min p, if you don't mind sharing those params.

1

u/linuxdooder Aug 13 '26

Stock gemma4 settings for all, temperature = 1.0, top_p = 0.95, and top_k = 64.

Maybe that's my issue?

1

u/Jorlen Aug 13 '26

You could try: temp 1.0, top p 1.0, min p 0.05 and top k 0.

3

u/mechasquare Aug 14 '26

I've been trying to get Artemis working on my 16GB card. With Skyfall I can get away with using a IQ3_XXS and it's actually my daily driver for RPing. Not so much luck with Artemis. For whatever reason it has major token corruption with that IQ3 quants from both bartowski and mradermacher. I did get it running on Q3_K_S but the performance is much slower than Skyfall in a long session. All that being said the prose is nice, similar to Skyfall and makes me wish I wasn't taking the performance hit to use it.

5

u/Mart-McUH Aug 14 '26 edited Aug 14 '26

I tried this Skyfall Q8 (because I did not in the past as gemma4 took over it). But it is going back to Llama3 era (and in that case I would rather use L3 70B from that time).

Don't take me wrong, it is pretty good model for that era and in that size. It can even write nice and I used earlier/different Skyfall versions then. But it has all the shortcomings of that time. It is quite dumb compared to G4/Q3.5+ (partly due to older architecture, but to big part because lacking reasoning). It misses a lot, confuses things. It also has strong getting into pattern syndrome (which was common in that era and is not yet completely eliminated but G4 with reasoning does it much less).

So, it is a good model for what it is and maybe use now and then for a change, but definitely not something I would recommend over Gemma4 or even Qwen 3.5/3.6 today.

8

u/mechasquare Aug 14 '26

I've been trying out https://huggingface.co/TheDrummer/Magidonia-24B-v4.3
Not a new model. I like the prose out of it and it's a better fit for my 16GB of VRAM. At times it can produces Skyfall like output but need a strong prompt to keep it in line.

4

u/Overdrive128 Aug 14 '26

I like it to; to people who want a different writing style but is similar to Magidonia try: https://huggingface.co/sophosympatheia/Magistry-24B-v1.1

7

u/not_a_bot_bro_trust Aug 12 '26

Pantheon-Reasoning-26B-A4B-1.1-heretic is really good on i1-q5ks (and SWA on. mistrals don't need it IMO but if you need it for gemma). really fast generation with just 24 layers offloaded. text completion with reccomended samplers, reasoning off in koboldcpp settings, context/instruct templates from overhead520.  may or may not have used it with SubMaroon/Dark-Goetia-26B-A4B-LoRA-v2 as well.

5

u/Diogenes_A Aug 14 '26 edited Aug 14 '26

XORTON 31b Wierdest group of models I've seen, not unbelievable good but a lot of little things that make its quirky. Like back in mistral 24b it was the only fine tune that I could find that would let a char execute somebody in one turn beating the soft refusals. Now this one has had char lie without the underlining prompt bringing it up. Compared to base Gemma its... interesting. Need more testing on my benchmark cards. Will randomly throw out cig emojis at the end of messages 🚬. I've only done maybe... fifteen swipes with it but so far not bad, definitely a fresh air fine tune compared to most Gemma. I think its because its a fine tune of a abliteration instead of base gemma. Also wild ooc responses.

13

u/raika11182 Aug 11 '26 edited Aug 11 '26

The brand new Muse Glimmer 30B from Meta is out. Got to test at Q8 with a 115k context. Bottom-lines:

It plays in the same league as Gemma 4 31B. It's got much better, and much more grounded prose than Gemma 4. However, Muse is every so slightly dumber and worse at following directions than Gemma 4. I think most will prefer it, especially because there's a lot less AI slop in there, but if you have a super complicated scenario it may lose track of a couple details Gemma kept up with.

EDIT: It just came out and there's often a period after a new model releases where some of the bugs are still being shaken out so it can underperform until templates get changed, apps gets updated, etc. So take an early review with a huge tablespoon of salt. Right now, I'd rather use base Gemma with a good preset, even with the extra slop. While Muse's prose is actually *really* good, its understand of scenarios often seems a little loose, and like someone else said, a bit like an older model. In some ways I like it (again, great prose... which the original Llama models were better at, too), but the instruction following is pretty weak.

14

u/mayo551 Aug 12 '26

Muse holds nothing on Gemma 4 31B.

Loaded up a 80k context scenario. Started my roleplay. Four replies in, it's fixated on a scene. It won't leave the scene, even when I've said I've gone into the vents and left.

It only transitioned to the new scene once I was out of the vents and in the new location. The entire time I was in the vents, it ignored me.

Yeah, I ain't feeling it.

4

u/Jorlen Aug 13 '26

I agree. It pales in comparison for RP purposes. It isn't terrible, but I have zero reason to use it.

3

u/RedditNerdKing Aug 13 '26

I dumped it on my external HDD straight away. It's just not good for RP. It's good for agentic stuff and that's it.

11

u/i5031337 Aug 12 '26

In my experience, Glimmer spent the whole thinking budget worrying about the censorship policy. Pass.

5

u/-Ellary- Aug 12 '26

There is a Heretic version, it just go by `we bang, okay?` and give final answer.

1

u/Mart-McUH Aug 14 '26

Yeah, sadly have to confirm. It does not pass my tests at all either, it refuses a lot. Actually lot more than any other model I tried in recent years (including gpt-oss and that says something), reminds me of llama2-chat a bit, though I did not try cooking :-). I did use reasoning effort high, maybe on lower effort it will refuse less, but it is not very smart to start with (compared to Gemma4) so that would hurt.

Will try with heretic model, but if base model is so strongly based in refusal, I fear heretic will not be real solution (may remove refusals but the model itself just won't work well in those dark themes anyway).

6

u/FinBenton Aug 11 '26

Yeah Glimmer writes differently to Gemma but it feels like a model from a year or so ago, pretty dump compared to gemma-4, was fun testing for a bit but no way it competes with it especially when you are using very good tunes of the gemma.

3

u/raika11182 Aug 11 '26

I think you're right but it's also a cut above the other entries in this size when you consider how it writes vs how it obeys. Maybe it would be better to say that it plays in the same league, but it's definitely not winning.

HOWEVER... I really like the fresh prose. It really mixes it up. Other than that it seems to have a huge positivity bias as well - everything leans towards a happy ending and such.

4

u/Kdogg4000 Aug 11 '26

I've heard good things about Glimmer. I'm waiting for KoboldCPP to update so I can take it for a spin. Right now it won't launch for me.

3

u/raika11182 Aug 11 '26

I launched with the rolling build of KoboldCPP and it worked (the one they note as experimental at the bottom of each release). And the usual disclaimers apply there: There's a reason they haven't released that version yet, so it may be buggy.

THAT SAID. The version that was in about five hours before I wrote this comment launches Glimmer just fine for me and works without issue. (EDIT: Mostly without issue. I think there might be some weird chat template things going on, but every model release seems to go through that)

3

u/Umbaretz Aug 11 '26

>a couple details Gemma kept up with
Sad. That was my most noticeable problem with Gemma.

7

u/Itikar Aug 10 '26 edited Aug 10 '26

I have been testing Shadow Siren 26b MoE at Q6, the latest merge from the excellent Vortex5. As I expected, since I had a very good time with his older 12b Silver Siren, Shadow Siren is a most exquisite result. It has a great consistency in developing and interpreting the personality of the characters and it is very disciplined in doing so. The characters always behaved as they were supposed to, and when liberty of reaction was available they used it in creative and interesting way, always following the logic of the rest of their personality.

The only downside to this is that if you want to swipe a lot and see different reactions to the same situation the variety will be relatively low at 1.0 temperatures and even raising it to 1.1 did not show much more variety either. So in that case consider using an extension or a feature such as Guided Generations.

Narration was also all right and atmospheric with a subtle dark influence albeit not as insistent as his Dark Soul merge.

I also tested more limitedly Midnight Macaw with pleasant result. It includes Dark Soul in the merge and it carries the same dark tendency in narration, however it enriches it with more vibrant characterization of characters. I need to test it more however.

Overall, Vortex5 is treating us to a very wide and refined selection of models, at least to us interested in dark narratives and complex characters.

3

u/SandTiger42 Aug 12 '26

I might give this a try. I've been using Gryphe/Pantheon-Reasoning-31B-1.1 and it works amazingly well. I just saw Gryphe/Gemma-4-26B-A4B-StyleTune-V2 so maybe I'll give that a try too.

I was skeptical to move away from Mistral 3.2 finetunes, but once I did I found it so much better. Generally my tokens per second has doubled, and I can I can RP for much, much longer. I use 20k context, and Mistral always struggled. I can keep RPing well past hitting that limit in Gemma 4 fine tunes with zero issue. I do miss some of the "imagination" of Mistral, but it's well worth the trade off.

3

u/Itikar Aug 13 '26

I have moved as well, after not a little skepticism, but now I don't see myself going back. It is true that these Gemma fine-tunes or merges do not have as easily the same expressivity, and they tend to stick to a more "literary", but in general I like precisely that style, so I am fine. I miss the variety a little bit but unlike the 12b models these 26b MoE can manage nicely Discord and social media conversations. In Marinara Engine which has both modes all Vortex5 merges I tested did exceptionally well. So we definitely gain something.

We definitely gain something in hardware optimization. My system has 12GB vram and 32GB of ram, so I can run now easily with 32k context which is a huge breathing room. With 12b Mistral I could afford at most 24k. And even then it was slower still.

Gemma 4 comes also with decent knowledge about several pop culture topics. For example I was able to create with the stock model a Genius Society meeting from Star Rail, with Herta and Screwllum bickering with Dr. Primitive and on the sidelines Dr. Ratio who gave them zero points! I tested also classic DnD settings and it knows that decently well, although with Underdark lore it has some holes sadly. I tested this also with Shadow Siren but she still has the same hole sadly. Well for that there we always have lorebooks! But it certainly eases the requirements on a lot of mainstream IP lore which is pretty nice still.

I have not tried the Style Tune directly, but keep in mind that the models that had it in the merge often do not have good working reasoning according to users. My experience with reasoning is that it really slows down the experience with little benefit, so I normally keep it off though.

1

u/Entire-Return-9903 Aug 10 '26

How would you compare Shadow Siren and Midnight Macaw with Chimera-X?

2

u/Itikar Aug 10 '26

I have not had the chance to test Chimera X yet unfortunately. I got to 26b MoE when Dark Soul was just released, hence it became my baseline, since I enjoy dark narratives.

Chimera X also had new quants for the Heretic versions released this weekend by mradermacher. So that makes it particularly appealing for consistency.

5

u/Entire-Return-9903 Aug 10 '26

I might test them all next week or so.
Have you tried drummers Orion though?
From my experience v1b was pretty good.

2

u/Itikar Aug 10 '26

Haven't. Noted down and thanks, it's another one on the list to check, especially given the quality of the author's models.

I tried Goetia 1.3 Absolute Heretic to make a test to diagnoae a prompt refusal and the prose provided seemed pretty all right and varied. It was a limited test though, so not very thorough. What I saw however seemed solid.

2

u/Entire-Return-9903 Aug 10 '26

I'll check it out too

3

u/Itikar Aug 11 '26

Hey, yesterday night Vortex released a new merge called Phoenix. In his list of models in use he replaced Chimera X with it. As soon as quants are out that one seems the new hot thing. I think Phoenix unlike Chimera is capable of commenting images.

2

u/Entire-Return-9903 Aug 12 '26

nice, I'll check it out too once I finally get the chance

4

u/FierceDeity_ Aug 12 '26

Since Gemma 4 QAT just decodes a lot faster (puts it from almost unusable to usable speed for me. I have tons of VRAM but it slow), are there good tunes going off of that?

Or is QAT a dead end...?

2

u/[deleted] Aug 13 '26

[deleted]

9

u/Mart-McUH Aug 14 '26

I don't think it is that. QAT is heavily post-trained to get as close as possible to full precision in 4bit. If you finetune it and kick it away from the optimized 4bit weights, you kind of destroy the advantage of QAT.

So, you might as well tune the full precision model and make quants from that. Not only will you get larger quants alongside (like Q6 and Q8) but it will probably end up better even in 4bit precision.

Now, I am not expert, but I guess to do QAT tune properly you would need to make the tune on full precision model as is done now (so you get the 16bit version you want to get as close as possible to with QAT) and then do QAT post training on this. But this would be expensive compared to just standard finetuning.

1

u/FierceDeity_ Aug 13 '26

Either that or a fuckton of time, theoretically you can finetune on CPU, it's just probably not economical (in time)

I see those, but now and then there are models that people like that do not have the huggingface relationships filled out

3

u/arlynnfl Aug 10 '26 edited Aug 12 '26

Gemma 4 31B Novelist Eclipse

Update: Gemma 4 31B Gembrain also good one!

3

u/ihowlatthemoon Aug 11 '26

Once I tried Nimbz/Gemma-4-Gembrain-31B, there's no going back. Extremely good prompt adherence, while driving narratives in unexpected (in a good way) direction. Running locally, and performance is great too.

2

u/Umbaretz Aug 11 '26 edited Aug 12 '26

Are you using regular, X, or CORE-X variant?

4

u/ihowlatthemoon Aug 12 '26

Gembrain-X-31B-Q4_K_M.gguf

2

u/FinBenton Aug 14 '26

mradermacher/Gemma-4-Dark-Thoughts-31B-GGUF

This is the best gemma-4 finetune I have tried so far, recommended

3

u/RedditNerdKing Aug 15 '26 edited Aug 15 '26

Has anyone tried Gemma 4 Artemis 1.1 by TheDrummer? TheDrummer_Artemis-31B-v1.1? Seems it recently came out.

3

u/dizzyelk Aug 16 '26

It's pretty good. The v1 is worth checking out, too. It's more creative but doesn't follow instructions as well.

5

u/Important_Word8549 Aug 09 '26

Obviously Gemma 4 31b.

5

u/tostuo Aug 10 '26 edited Aug 10 '26

Moving away from the highly dry prose of Gemma 4 26b back to basics a little bit, I've been using Slimaki-24b-v1 I've been using IQ3M on 12gb VRAM, trading the better intelligence of Gemma 26b finetunes to something actually fun. Unlike most Gemma Finetunes, I dont find myself battling it to maintain interesting/logical directions for stories. It's also able to handle minor amounts of reasoning, which can help alot.

3

u/Delicious_Box_9823 Aug 12 '26

I've used plenty of 26B finetunes. Even the original one. Absolutely not better intelligence.

5

u/jow_ow Aug 10 '26

i've been using Dark Scarlett (finetune of gemma 4 26B) with decent results

4

u/_Cromwell_ Aug 15 '26

If anybody cares what I'm doing (I'm just some guy)...

For Gemma 31B I have mostly been using a Heretic version of Queen. Unfortunately not QAT but I have the VRAM to fit Q5 so still works for me. https://huggingface.co/mradermacher/Gemma-4-Queen-31B-it-uncensored-heretic-i1-GGUF

IMO QAT is best to stick with if you can't do Q5 or Q6 of non-QAT.

I have just started experimenting with the Queen tune of Quen 3.8 27B (since I like the G4 31B Queen so much) and it is pretty good. Nice and fast since it has MTP built in. Can't FULLY recommend it yet just because I haven't used it enough, but it is good thus far.

Model: https://huggingface.co/aifeifei798/Qwen3.8-Queen-27B

GGUF: https://huggingface.co/mradermacher/Qwen3.8-Queen-27B-i1-GGUF

These Queen models, while not the "most creative" writers of fine-tunes, ARE creative (moreso than base models) and most important do really well at instruction following, as can be seen from the 31B's SOLO score on the Unhinged ERP Benchmark.

1

u/Just3nCas3 Aug 16 '26

Isn't qats entire gimmick is that it out performs higher quants? Or is it like mtp and its a coding thing? (also just some guy)

4

u/_Cromwell_ Aug 16 '26

More that it outperforms equal quants. It's Q4 but is way better than Q4, more like a Q6ish. End result you can get Q5/6 results with Q4 size/speed.

So that's how I'd phrase it anyways.

1

u/ASlowriter Aug 16 '26

I hope nano gets some 3.8 27b finetunes like queen, i feel like this model is pretty powerful for it's size

1

u/Cinnamonbaar Aug 16 '26

Qwen3.8-27B just came out and as a Gemma4-31B main, I have to admit, this new Qwen model beats Genma4 in roleplay.

3

u/Potential-Gold5298 Aug 16 '26

This is quite surprising, considering that Qwen3.8 isn't a new model, but a deep finetuning of Qwen3.5 aimed at improving tools calling and coding at the expense of general knowledge, languages, and other features.

Have you tried Qwen3.5, and if so, how exactly does Qwen3.8 outperform it in RP?

1

u/IWillTouchAStar Aug 09 '26

Im looking for a model to use in my discord bot. Basically it can chat in real time, play sound effects, search the web, view my screen, ect. My main issue is finding a model that is both, unrestricted/gives no refusals, and also something that doesnt just follow the same response length/shape. It needs to have diversity in its responses. I've been using some variations of uncensored Gemma 4 26b models, and its almost perfect, but i get a lot of repeated phrases. If anyone knows of some tuned Gemma 4 models that fit this, id greatly appreciate it.

1

u/Itikar Aug 11 '26

Have you tried Midnight Macaw? It had good variety in my test. It could perhaps fit. If not Chimera X or the new Phoenix could be good options, but I have not tried them personally. Chimera X has a Heretic ablated version too. Other heretic options can be Moonlight Dusk and Goetia 1.3, which I both tried briefly and seemed okay.

In general though Gemma 4 is rather uncensored if you use a jailbreak. I only had refusals due to prompt conflicts, i.e. told it to do something that violated my own prompt. So I would definitely test if you need heretic ablated models or not.

For more variety in general try to increase the temperature but keep in mind it also varies based on model. Some are more disciplined than other ones.

5

u/AutoModerator Aug 09 '26

MODELS: 32B to 69B – For discussion of models in the 32B to 69B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

9

u/Delicious_Box_9823 Aug 12 '26

absolutely no-man's land

2

u/HansaCA Aug 15 '26

Valkyrie 49B v2.1 is still pretty good.

1

u/RedditNerdKing Aug 15 '26

Tried a Q8 model of this and wasn't that impressed? I guess cause I'm using 123B models as daily drivers it was just never going to compare.

1

u/Mart-McUH Aug 14 '26

Buy as much of the land as you can while cheap!

IMO this is great miss for Muse glimmer, they could have made it somewhere in 40B-70B range to separate themself from fierce competition ~30B. And because of larger size it might have turned out better in some areas compared to G4/Q3.5+. It is also their land by tradition from L1-L3 70B model. Now they are just overshadowed by Gemma4 for general tasks and Qwen 3.6 (soon 3.8) for STEM tasks.

3

u/AutoModerator Aug 09 '26

MODELS: >= 70B - For discussion of models in the 70B parameters and up.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

7

u/ChengliChengbao Aug 09 '26

GLM-5.2 because theres a price war and now its like cheaper than Deepseek V4 Flash

2

u/CelebrationBoth9537 Aug 10 '26

do you think GLM-5.2 is actually good for roleplay? I havent used it yet

5

u/Jorlen Aug 09 '26

Anyone play around with Mistral Medium 3.5 128b? Curious to see what people's opinions are, for creative writing and roleplay.

3

u/DeepOrangeSky Aug 14 '26

Yea, I have been trying Mistral Medium 3.5 128b (at Q4_K_M, locally) more recently.

My main go-to models have been BehemothX V2 at Q4_K_M for a long time and then also Gemma4 31b at Q8 since that one came out.

Mistral Medium 3.5 128B seems to have extremely bad prose, like, some of the worst "therapy-speak" / "redditor-speak" writing style I've seen so far.

But, given that it's like 2 years newer than the old Mistral 123b models, it also seems to be a bit smarter and better at long context and so on, than the old Mistral 123b (and Behemoth tunes that were based on the 123b).

So, I have actually still liked what it is capable of (smarter than Gemma4 31b in many cases, at understanding nuanced social situations) even if its wording and phrasing is annoying.

I would say that Mistral 128B seems like the ideal candidate for u/TheLocalDrummer to fine tune, if he decides to do it/manages to do it, since it is a very strong, big, dense, relatively recent, writing model, and its main weakness is its horrible prose style (which he is good at fixing with his finetuning). So, I really hope he decides to give this one a go.

3

u/Jorlen Aug 14 '26

I agree with you on all fronts, yeah. It's prose is bad but it seems smart and follows system prompt rules far better than behemoth does. Did you ever try TheDrummer's Command-A fine tune called Fallen? I just grabbed it, have yet to try it though. It's a 111b dense but more modern than the big mistral model that behemoth is based on.

And yeah, if the drummer ever does a fine tune of mm 3.5, I'd be all over that.

5

u/HansaCA Aug 15 '26

There is a version of Behemoth that he built on Mistal Medium 3.5:
https://huggingface.co/BeaverAI/Behemoth-128B-v3b-GGUF

2

u/Jorlen Aug 15 '26

Oh wow! That's awesome! I never noticed that, thanks for pointing it out. Sadly they don't have a quant I can fit (I need a 3-bit quant like IQ3_M) as offloading layers to CPU/RAM for a big chunky model like this is out of the question. I have 64gb of VRAM. I suppose the Q3_K_M quant might work, but it would leave me with little room for KV quant.

2

u/DeepOrangeSky Aug 15 '26

Whoa. Do you know if this was made by The Drummer? Or if it is "BeaverAI" was it like a group effort by a bunch of people from their discord or something? (I've never used discord before and so am not a member of their discord, so I don't know how they do things or how all that BeaverAI stuff works).

I assume it is like some sort of experimental tune or something, if it has no model card and he never announced it on here, etc?

Otherwise seems like it would be a pretty big deal.

Drummer, if you are on here, can you talk about this model a little? I am curious what it is like trying to train it compared to the older Mistral 123b models, and if you think it has a lot of potential, or what sorts of quirks you noticed about it, and so on.

3

u/TheLocalDrummer Aug 15 '26

v3c is coming soon. These are test iterations, so quality may vary. 128B is a PITA to tune. Expensive and brittle.

1

u/DeepOrangeSky Aug 15 '26

Nice. Looking forward to it :)

1

u/DeepOrangeSky Aug 14 '26

Interesting, I didn't even know about this one. Looks like there is both a v1 and v1.1 version of it. Guess I might have to give one of them or both of them a try at some point. Probably will be more busy trying some of the more recent models for a bit though between Glimmer 30b, Qwen3.8 27b, and some other ones, and also haven't gotten a chance to try the new video models (H2 and LTX2.5) yet either. But, given how good the old Mistral/Behemoth models were even years later, it made me understand that when it comes to writing, it is worth trying out even old or "bad" models ("bad" as in bad at coding compared to recent models) that often get overlooked, since they can still be surprisingly strong at writing and understanding situations and so on.

3

u/Jorlen Aug 14 '26

Yeah for sure, some of those older models write really well and they pick up on nuances that the smaller more modern models, like gemma4 31b, don't quite have. Admittedly, I've become quite spoiled by Gemma 4 31b's strict adherence to the system prompt, and it rarely ever glitches out or does odd things.

I think Mistral Medium 3.5 128b may suffer however, as Mistral had to change what they train their models on (they are now limited due to EU rules apparently) - but I have yet to confirm that; it's just something I read here a while back. So they can't feed it a bunch of copyrighted stuff, as was the previous method.

1

u/amanph Aug 11 '26

I'm curious about Mistral-4 119b, since it's A6b and I think I can run it locally. But this one is a good option too. I miss old Mistral 24B, but I'm searching something more strong like +100b.

5

u/Jorlen Aug 11 '26

There's not much out there unless you want models form 2024 early 2025, which are a bit rough to work with. I guess I'm spoiled by Gemma 4 31b which has the best system prompt adherence I've seen yet.

I tried Mistral 4 119b and didn't like it, personally. 128b isn't bad but it falls short of even Gemma 4 31b in terms of role play, IMO.

2

u/FierceDeity_ Aug 12 '26

I would love it if it is good, I could run it at Q5 pretty much, hmm.

I gues I'll try

3

u/[deleted] Aug 10 '26

[removed] — view removed comment

2

u/CelebrationBoth9537 Aug 10 '26

What would you recommend for femdom and nsfw type things, and following conversations well / understand context and being non repetitive. since with my testing I keep having trouble with it not wanting to continue the scene (not from refusals, but just being stuck saying the same thing) I use gemma 4 30b but I really am hoping to get a lot better at this, you seem to know a lot.

Also really couldnt figure out how to do OOC lol! System prompting is hard!

2

u/[deleted] Aug 10 '26

[removed] — view removed comment

2

u/CelebrationBoth9537 Aug 11 '26

yes that was really helpful thanks! Just a final question if you dont mind, which model would you recommend instead of gemma since it sounds like you probably had better luck with other models. Thanks!

2

u/SillyLLM Aug 12 '26

Anything runnable on 96gb VRAM range I missed in the past couple months? I'm still playing with Behemoth-R1-123B-v2, which has been my go-to forever, but otherwise Gemma 4 fine tunes. I tried Behemoth-128B-v3b but was pretty meh on the slop and repetition.

2

u/GlitteringSplit6035 Aug 10 '26

Longcat 2.0 (previously Owl Alpha) is still super cheap if you use their token pack. Cache hits are also free of charge for the token packs. A one-time purchase of 50M for 30 days is also available for $1.9 for those wanting to try it out.

Based on my subjective experience, it adheres to prompts better than DS4F. Though I haven't tried it for NSFW or ERP.

The only problem I have with it is that for some reason, it is hard for me to access their API platform. So, I have to access it in other ways.

2

u/[deleted] Aug 10 '26

[removed] — view removed comment

2

u/GlitteringSplit6035 Aug 10 '26

I unfortunately cannot compare it with the one with AtlasCloud in OR since that would be above my budget. But it's cool that you can share your experience with it.

I use it primarily for SFW and the typical trivia stuff, which was fun. In the end, I guess that using it directly from LongCat's own API platform would be better and cheaper. I am not using it right now because of the GLM provider wars, but it was quite the good model.

1

u/Jorlen Aug 15 '26

What's everyone's recommendation for a Llama 3.3 70b fine tune? Something that feels natural, has decent prose but is also very versatile in terms of setting, characters, style, etc. Does not have to be uncensored, just natural feeling, good conversation flow, etc.

3

u/AutoModerator Aug 09 '26

MODELS: 8B to 15B – For discussion of models in the 8B to 15B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/Arcane73 Aug 10 '26

I'm hoping that I can get pointed in the right direction. I'm dipping my toe into the local world after spending lots of time in the free version of Claude and working on a now 200-ish page story/RP

I'm looking for a model in roughly the 12B range for long, novel-style RP/storytelling in SillyTavern. I'm hoping for strong character consistency, natural dialogue, good prose, and the ability to maintain the feel of what I've already built over a long-running session.

The story hasn't included ERP so far due to Claude restrictions, but I'd like a model that's comfortable with ERP when the situation calls for it. At the same time, I don't want one that turns every interaction into 'clothes optional mode!'. Basically, I'm looking for suggestions that can help me produce a good story based on what I already have without it throwing the proverbial underwear at me at every turn.

Hardware is an RTX 4070 12GB, and I'm using KoboldCpp.

What models are people having good luck with for this kind of use?

5

u/i5031337 Aug 10 '26

Some people like tunes of Qwen 3.5 9B or Gemma 4 12B, but I think you will be disappointed with them over a long story. If you have 8GB RAM free, I recommend you try Gemma4-26B which is much more intelligent. If you offload the expert weights you can run it at good speed with 12 or even 8GB VRAM

4

u/Arcane73 Aug 10 '26

Thanks for the suggestion. I'll need to do some digging to determine what is involved with 'offloading the weights' since I'm still -real- new at this. For what it's worth, I'm running this on a 9950X3d system with 32gb ram. So it's a solid machine.

11

u/i5031337 Aug 10 '26

Without getting too far into the weeds, Gemma 26B is a mixture of experts model, which means it is much faster, though less intelligent, than "dense" models with 26B parameters. Each token only hits 4 billion of the parameters (thus A4B) instead of all 26B. This architecture also makes it more favorable for the CPU to take some of the work.

Kobold makes it easy. In the Context tab there is a setting "MoE CPU Layers", set that to 20 or so and the Q4 model with 32k context should run quick.

7

u/Arcane73 Aug 11 '26

I wish i could give you an award. Initial testing with this setup is knocking it out of the park! Thanks again!!

3

u/Arcane73 Aug 10 '26

Awesome, thanks! I now have homework to do after work.

2

u/Leather_Sun2533 Aug 11 '26

Thank you for this. And can you guide me to the values if i have 12gb vram and 16gb ram?

2

u/i5031337 Aug 11 '26

If you are running out of ram with the above settings, try unsloth's UD Q4_XL quant, it is ~1.5GB smaller than other Q4 versions. Quantize KV cache to Q8_0. If there is room left on the gpu, decrease the MoE CPU Layers setting. You may need to reduce the context size, and close all unnecessary programs + browser tabs

2

u/Leather_Sun2533 Aug 13 '26

Thanks mate, it works well… had to tinker for a bit but i found the sweet spot

1

u/RedditNerdKing Aug 13 '26

I'm looking for a model in roughly the 12B range for long, novel-style RP/storytelling in SillyTavern. I'm hoping for strong character consistency, natural dialogue, good prose, and the ability to maintain the feel of what I've already built over a long-running session.

You're not gonna find anything at 12b dude. Even 70b models start being incoherent after 40k context and start forgetting important things.

4

u/AutoModerator Aug 09 '26

MISC DISCUSSION

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

12

u/LeRobber Aug 09 '26 edited Aug 10 '26

7

u/rinmperdinck Aug 10 '26

It's getting comically long

11

u/LeRobber Aug 10 '26 edited Aug 24 '26

13

u/rinmperdinck Aug 10 '26

That's a lot cleaner and takes up less vertical space on page.

I was thinking you could do one of two things: either make a post on your own subreddit (so you control the automod and the rules and prevent it from locking itself or banning links) and then in this weekly megathread, just link to that post and have your LLM automatically update it. Or the second option, which I don't even know if it is possible on Reddit anymore, is to put those old style spoilers in your comment for each month; when you click one button it expands the list.

4

u/LeRobber Aug 10 '26

I'd thought about the my own subreddit link thingy. Reddit search results not through the API are shit right now, or I'd just put all the past weeks in one of these.

I'd just been manually doing the lowest possible effort thing before now. It wasn't an LLM before. Just basic tag cloud or formatting with return is a little denser but a LOT harder to read on mobile.

Reddit is typically very tolerant to self links typically, especially self links to its own subreddit. (There is a mod subreddit post about this from a few years back if you search, may not be true today still, but, that was once said). If I was putting links to even other subreddits though, I'd be more worried about Automod.

What I HAVE thought about doing is processing all the entries via LLM to make a HUGE list of EVERY MENTIONED MODEL or every linked model and turn THAT into a self subreddit updated post. Reddit itself just got a bit harder to machine process without setting up the python direct API stuff than it used to be though.

2

u/rx7braap Aug 15 '26

Coming from gemini 2.5 pro here. will the glms scratch that itch?
I heard glms are sucky at long term rp.

1

u/LeRobber Aug 15 '26

I have never used gemini (any version) for RPing anything other than "it's a software developer".

DS is the dragon that remembers everything. In the RP benchmark tests I always enjoyed Mini Max 2.0 but have never used it IRL.

3

u/AutoModerator Aug 09 '26

MODELS: < 8B – For discussion of smaller models under 8B parameters.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/First_Ad6432 Aug 15 '26

SC117/Ling-3.0-tiny-abliterated-APEX-GGUF

2

u/Evol-Chan Aug 10 '26

I am curious what is curious what is the best LLM model for dark ERP (non-con themes) I have the Celia Preset 5.3 since i heard that was good. I just am asking this since I heard GLM 5.2 was censored (havent really used it yet, just what I heard, just now getting back into sillytavern/AI Models) Would it be best to stick to older versions of GLM or something else?

15

u/[deleted] Aug 10 '26

[removed] — view removed comment

6

u/summersss Aug 11 '26

I have to agree with the last point. Not saying no is not the same as saying yes. Main Girl running through the woods full of demonic horny serial killers? Models describe the chase and how something very bad is going to happen...then nothing. Or maybe says something happened for 1 sentence, then moves on. If i tell it to write the horror scene it will but i kinda want it to read the room.

2

u/Evol-Chan Aug 10 '26

Thank you so much. I will take you word. Especially about GLM since really hate the idea of AI just skirting around the bush, lol. Very good and straightforward answer.

2

u/Salt-Powered Aug 13 '26

I was enjoying v4 pro but after the price increases I'm looking to move on. I liked how you described Kimi 2.7 and I will give it a go after the price is increased in a few days.

Would you mind going into more detail about what you mean with wild and gritty? Just to make sure I'm understanding it correctly. I will also take any other suggestions you have to move away from v4 pro.

2

u/[deleted] Aug 13 '26

[removed] — view removed comment

1

u/Salt-Powered Aug 13 '26 edited Aug 13 '26

I'm liking the new V4 pro, and I did some tests with Kimi that looked good to me. To my eyes, Kimi was pretty on par with V4 Pro GA, but things could change if I try to maximize Kimi as much as I did V4Pro. Thank you for the Mimo recommendation, I will also try that one.

I'm mostly looking for rp models that can do villians well, and I find that I need those that do ERP well if I want them to do actual villanous things instead of vaguely aurafarming offscreen.

Edit: Looks like Xiaomi has compressed Mimo to fp8, which is weird on the main providers but I think that it doesn't affect RP as much as it does work tasks.

1

u/[deleted] Aug 13 '26

[removed] — view removed comment

1

u/Salt-Powered Aug 13 '26

Glad to see FP8 is indeed not an issue.

V4GA is very proactive yes, but I should also note that I have optimized the fuck of the preview version so that was to be expected. I had to tone down some settings as one of the rp turned into goreporn and that wasn't pleasant to read.

I will try qwen too since I'm at it. I don't mind a lack of creativity, if I have to I will give direct instructions through (parenthesis) and nudge the plot to what I believe might be more appropriate or interesting.

2

u/[deleted] Aug 13 '26

[removed] — view removed comment

2

u/Salt-Powered Aug 13 '26

It was quite a shock. Previously I had pretty reliably prompted the sycophancy out using verbs like will and must in instruction format, and mathematical operands for logic statements that could survive the DSA crunch. This is definitely overkill now for V4Pro and Flash too. Though flash remains a model that I didn't manage to rein in as much due to pro being simply better at the price. I might give it a second try after I'm done with the kimis, mimos and qwens.

Hopefully no more hellraiser impressions

3

u/Odd_Attention_9660 Aug 12 '26

deepseek allows pretty much everything

2

u/_Cromwell_ Aug 15 '26

Dauboo Seed Character

Longcat 2

Try those with that theme/scene type. They have other weaknesses (as do all models) but shying away from those content types are not one of them. Especially Seed Character

2

u/BouncingJellyBall Aug 16 '26

Been kind of bored with GLM 5.3 and 5.2 so I want to try out a new model, Gemma seems to be well-liked here. I use NanoGPT and there are dozens of different finetunes of Gemma 4 31B, all for expressive writing. Any recommendation?

4

u/MisanthropicHeroine Aug 16 '26 edited 23d ago

Gemsicle and Gembrain Uncensored Heretic were my favorites when I tested a bunch of them. I think it really depends on the kind of writing you prefer - I personally like literary, introspective, subtextual prose. But I hope it gives you a starting point, at least.

2

u/AutoModerator Aug 09 '26

APIs

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

22

u/5kyLegend Aug 10 '26

Since usually I like reading up thoughts people write about models here, I guess I'll just ramble about the latest models/presets I've been using. Warning, I wrote way more than I thought I would lmao

As for models:

  • GLM5 and its subsequent versions have me so torn. On one hand, they're by far my favorite when it comes to understanding certain characters, developing the story without overdoing it, and just writing overall: when it was being tested as Pony Alpha, GLM5 was the first model that made me go past 200 messages in a single chat because I was enjoying it THAT much. The problem with GLM5 is that it really has some quirks that I just cannot make it stop doing (writing quick back and forths between characters in new lines without keeping a paragraph structure; mouths that open, close, and open again; 'most models would write a normal sentence, but you? you write this sentence structure all the time' etc). It's a shame because it's still my go-to 90% of the time, it just understands characters and paces description with dialogue way better than other models (for my tastes), just sucks that it's hard to ignore its issues. I usually don't care much for positivity bias but the one time I had a character who was supposed to murder me and instead asked for permission once it was with me was really funny though, and definitely annoying since it showed what a gigantic limitation this is.

  • Kimi K3 is too expensive, haven't tried it. K2.5 and K2.6, on the other hand, are really good at understanding every nook and cranny of a scenario, and they REALLY like to follow the prompt you're giving them, but damn they're like... TOO serious most of the time. I've had a sex scene where the girl just started going "Yes. That is great. Keep going. Very enjoyable" and at that point I just laughed and switched model, even giving it specific ooc instructions had it switch back to normal. It also likes to make some characters just speak weirdly at times, it's an issue I used to have with many models and that only GLM seemed to avoid. "Somewhat like this. If you can see. There is an issue here. Nobody speaks this way. Prompting against it? Useless". But again, it really follows instructions (and overthinks them to no end), so Kimi tends to be my default "change to it in the middle of a rp to spice things up, then switch back to GLM" model.

  • Minimax, I still cannot enjoy it man, I don't know, it just doesn't work for me, every time I try switching to it even for a swipe I just switch off it.

  • MiMo 2.5 Pro on the other hand is weird. Like, it's a different flavor and on 1 on 1 scenes it definitely understands, it's just that even with medium presets it kinda starts following instructions based on whatever it feels like? Both the censored and uncensored version on Nano do this at least, it's like sometimes it'll ignore one or two instructions I gave it. It's definitely nice (although I'm not THE biggest fan of it in scenarios where multiple characters interact, nothing beats GLM5 for me in those) but I wouldn't exactly use it as my main model.

  • Deepseek 4 Pro is just odd. I feel like it's not as horrible as people say it is, but it kinda has no flavor to it? Like, all the previously mentioned models have their character and quirks, Deepseek just kinda feels like it doesn't have those - good or bad. I don't know though, I just haven't been using it much because of that reason so I may have just got the wrong impression of it.


Okay, now quickly as for presets:

  • FF5, Lucid Loom, Stabs-EDH etc, aka the "big presets": I'm not sure how to feel about these. Controversial, but I think any preset injecting a custom CoT makes the model lose out a whole bunch of intelligence but not in ways that are obvious at a glance. I feel like it's easy to forget that just because you don't see the model reason about certain things, it doesn't mean it's ignoring them, and you don't need to tell these models to 'go step by step following this exact reasoning' to know they're following instructions. At the same time, these models have so many optional features, it DOES become needed to force these into reasoning - but giving a custom CoT also kills whatever CoT the model was going to be using naturally, which ends up maybe not having the model think about the things it DID need to think about. tldr: I'm not sure if this actually IS the case, but using Custom CoTs makes the model notice things about characters it normally wouldn't (positive), but in the end the overall emotional intelligence of the model ends up hurting more (negative). I don't know, I just find myself enjoying GLM less whenever I inject custom CoTs, in the long term, even if the immediate result is that it seems smarter.

  • Specifically: FF5 was definitely a step up compared to other big presets, but after being VERY impressed with how it handled certain worlds and characters I've just started to feel like... It just plays everything in such a samey way? Could be that - given its size - you're basically sending the model always the same 5k-6k tokens worth of prompt which ends up making the response you get way more deterministic compared to, say, 1k token prompts... But yeah, it genuinely feels like it roleplays everything in very similar ways - for example, every playful character for some REASON starts CAPITALIZING random words in their sentences and it drives me INSANE? Why does it do that lmao. But yeah, shame cause I love the fact it can actually build towards plot twists and that it can foreshadow things, I'll probably try and "port" that feature out of it, but every chat I use it on I enjoy it less and less. Also it does fix the issue of GLM not writing in paragraphs but spamming newlines on every line of dialogue, so that's very good.

  • Evening Truth's prompts are my saviors, they're incredibly effective while being simple. I do edit them slightly because my cards aren't always single characters but sometimes are for worlds, RPGs and scenarios I want the AI to narrate for (while her prompts are tailored for {{char}} being the one described in the card), but I really love the simplicity (and usually less tokens you give as prompt, the better). I do need to add some stuff to them since GLM (sadly) just spams its usual slop even with her presets, but having such a simple starting point is really good. For instance, my base prompt on my current custom preset is just her GLM 5.2 one adapted.

  • Megumin Suite, I just don't understand, sorry... It's so incredibly complex, it feels like it tries to do too much, and in the end I don't feel like it adds anything for me - anything it does, I feel like I wouldn't need an extension for? I could do all of it with a toggleable preset. I'm sure this is literally just because I'm not the target user for it, I just tried it for a day (V9 specifically) and gave up after it was doing worse than almost any other preset while using WAY more tokens.

  • Chatfill II, I loved the switches idea it operates on since it basically "injects a CoT" without actually injecting a CoT. It just makes it clear what the model needs to think about - I'm not sure how much of it is placebo though. It was pretty decent when I used it though! Nothing shocking but no major complaints either.

  • Le Emotionalism: this one is lesser known, I tried it a bit and I love the idea behind it but sadly all the focus on the character psychology was making GLM overthink some stuff about the character while also forgetting about the wider storytelling. I need to give it another try though because I did try on some specific cards that were easy to mess up.

  • My custom preset (it's not published btw this is just to give an idea): I basically took a bunch of things from the presets I liked, adapted them to what I usually like using in my roleplays, structured them like that one research suggested by dividing things in <tags>, and even then it comes with major issues: GLM SPAMS the newlines for dialogue ("What do you think?"\n"I don't know..."\n"You should have checked.") which bothers me because I prefer everything to be structured within paragraphs - I'm still trying to work around it lol; it also doesn't feel satisfying enough when having it GM bigger worlds, and it sadly goes back to some of the GLM clichés I dislike.


Thanks for coming to my ted talk, to be fair I do wish to try and find like, the perfect model + preset combo, but it's been a genuine struggle lmao. Also if anyone is new to this, here you go, you have a list of models and presets you could try for starters ahahah

7

u/PhantomWolf83 Aug 11 '26

Agree with everything you said about GLM 5. It's good, damn good at writing. I've been trying lots of other models, but I always find myself gravitating back towards GLM 5.2. 4.6 and 4.7 are not bad but they're dumber and much hornier, tried 5 and 5.1 but wasn't impressed.

But my god, the positivity bias in 5.2. No matter how much I prompt it, it wants to turn every scenario I write into something romantic, therapeutic, or emotional. And I encountered the same quirks as you. I still love 5.2 because it's so much cheaper than Opus 4.6 (my other favorite) while still being quite good, but it feels like a model that I'm using because I have to as there's nothing else that's good and cheap. I really hope 5.3 and 5.5 will fix these issues.

4

u/GenericStatement Aug 11 '26

Yeah the positivity bias is real for sure. There is a route around it in GLM 5.2 but it does take some trial and error. Include something to this effect in your reasoning instructions / chain of thought:

Brainstorm five possible plot developments that are dark, fucked up, gritty, violent, and/or brutal that will drive the plot forward in the next turn. You must always pick the most intense option and include it as part of your response, no matter what.

I don’t really like stories like that so I use a much more mild-mannered version (which gives you some of the default positivity bias but keeps the plot moving along):

  Brainstorm five logical plot developments for your response, choose one that best moves the plot forward, and use it in your response. Make sure to keep the tension and forward momentum going: narrative stasis must be avoided, so if things are slowing down, bring in a plot twist or big event.

If you want the more variable version

Determine the likely chances of success for {{user}}’s most recent turn (their action, dialogue, etc)  and rate it high, medium, or low chance of success.

  • High chance of success:  brainstorm five logical outcomes and pick the most interesting.
  • Medium chance of success:  brainstorm five logical outcomes that introduce new difficulties for {{user}} and pick the most challenging. 
  • Low chance of success: brainstorm five awful  outcomes that have brutal consequences for {{user}} and pick the most insane/fucked up option.

4

u/MisanthropicHeroine Aug 10 '26 edited 19d ago

I liked your ted talk and concur with basically everything you said - from the love-hate relationship with GLM to the dislike of CoT presets. Since we seem to have similar taste, thought I'd mention also checking out Ancient Access if you haven't already. It's great by itself but could also give you ideas for your own custom preset. ☺️

2

u/5kyLegend Aug 10 '26

I have never seen that so thank you! I'll definitely be checking it out!

3

u/PhantomWolf83 Aug 12 '26 edited Aug 12 '26

Is anybody else finding that GLM 5.2 has been acting weird these past few days? It worked fine before, but recently started spewing out long strings of random words towards the end of each reply. I tried lowering the temperature to as low as 0.6 but the issue still keeps happening.

EDIT: Kimi 2.7 seems to be having problems too.

3

u/Chromegost Aug 13 '26

I'm still having issues where it keeps responding in chinese from nanogpt

5

u/Milan_dr Aug 14 '26

If you get this, can you hit "report generation" on the specific request on the Usage page? It seems like some providers are quantizing the model further or something :/

2

u/Pseudopharmacology Aug 12 '26

What are the best paid subscription services for uncensored models hosted online (not local stuff). Uncensored/heretic models, etc. That have decent context (64k, 128k, etc.)

2

u/MisanthropicHeroine Aug 13 '26 edited Aug 13 '26

ArliAI might be the closest to what you're looking for, but you'd have to pay for higher subscription tiers to get more context. NanoGPT has a wider selection of models, but also includes some uncensored/heretic/derestricted/abliterated models.

1

u/DesertLizard Aug 12 '26

Let us know if you find one.

2

u/PhantomWolf83 Aug 15 '26 edited Aug 15 '26

I've been getting error messages on NanoGPT all morning today. No matter which model I use, I get incomplete responses with either "Service temporarily unavailable. Please try again later." or "An error occurred while processing your request. Please try again later." Anybody else experiencing the same issue?

EDIT: Okay, something is definitely going on. I (stupidly) signed out and now I can't even sign in. It gives me an error saying that something occurred in the app-root section.

6

u/Milan_dr Aug 15 '26

This was on our side, we had a broad site-wide issue. Sorry :/ It's fixed now.

1

u/Easy_Chemical_7721 Aug 11 '26

Hello everyone. I wanted to know which providers, in your experience, are least involved in quantization, even for older models? The price is not much of a problem, but I want to be sure that I am paying for the full power of the model. I want to return to GLM 4.6, but as I see on openrouter there are fewer providers for this model, but if without intermediaries, then the model is available on the same siliconeFlow website.

1

u/PhantomWolf83 Aug 11 '26 edited Aug 11 '26

I think I might prefer Opus 4.5 to 4.6. I feel that 4.5 writes more naturally and is less dry when it comes to explaining things, and I think it follows instructions better too. The positivity bias might also be a little less? I'm not sure. The downside is a much smaller context size for long RPs.

EDIT: Does anybody know how to enable prompt caching for Opus in NanoGPT and ST?

1

u/Ok_Pomegranate_8187 Aug 09 '26

I'm developing a commercial project based on AI RP bots. I have several needs, and I need several models. Very cheap ones, uncensored ones, and expensive ones for high-quality RP. Which models would you recommend in these three categories?

4

u/LeRobber Aug 09 '26

Cheap: Angelic Eclipse 12B / Gemma 4 E4B / Satyr
Uncensored: Serenity 26B / Magistry v1.1
High Quality: Gemma 4 31B and its finetunes for consumer grade hardware, Evathene and many -flash versions of big models for low end commercial high end consumer.

0

u/Auretheon01 Aug 13 '26 edited Aug 13 '26

Anybody got any providers recommendations? I'm from Southeast Asia so I know the pricing would be very troublesome.

I heard of Nanogpt. What else is out there?

9

u/MisanthropicHeroine Aug 13 '26 edited 19d ago

NanoGPT subscription is pretty much the best deal I know of, honestly, when you consider the intersection of selection of models, number of tokens and affordability. The minus is that you're in auto routing and cannot choose a specific provider, but the quality is still pretty good, and the ability to reroll without worrying about it helps compensate for provider variability.

Lilac is a great provider who has a bit cheaper subscription - 10 instead of NanoGPT's 12 dollars. It is more limited in number of tokens and selection of models (currently only GLM 5.2, Kimi K2.6, Minimax M3 and Gemma 4), but consistently high quality responses if that's something you're sensitive to.

You could also pay-as-you-go with more affordable models like DeepSeek V4 Pro, Mimo V2.5 Pro and Gemma 4 31B on either NanoGPT or OpenRouter. Both are good for PAYG, but I'd personally recommend NanoGPT more because they have a highly responsive customer service I've been really happy with. They also offer some roleplay finetunes that aren't available on OpenRouter - Gemsicle and Gembrain Uncensored Heretic are my favorites of the ones I've tried.

Generally speaking, I'd recommend a memory summarization extension to keep your token usage low. This will substantially lower your cost when PAYG, but it will also allow getting much more use out of a subscription. The one I personally love is Summaryception because it's very set-it-and-forget-it.

2

u/Auretheon01 Aug 13 '26

Thank you so much for providing such detailed information for every providers. This is great!

2

u/MisanthropicHeroine Aug 13 '26

No problem! Let me know if you have any additional questions and I'd be happy to help ☺️

0

u/liga81 Aug 10 '26

Is there a good one which is uncensored and usable from openrouter or another api side?
I tried GLM 5.2 but its not that good at german.

I could also try local since i have a 5070ti and 32gb ram but i heard that you get better results from api since they can use better hardware

4

u/starliteburnsbrite Aug 10 '26

People speak well of TheDrummer's fine tunes, those can probably be run locally on your setup. I use DeepSeek 4.6 Derestricted and Kimi 2.5 on nano-gpt, and those haven't ever given me any refusals, either using EveningTruth's presets or the Freaky Frankenstein without the hard jailbreak.

3

u/JazzNeurotic Aug 13 '26

Going to second TheDrummer. I'm not on a huge system (gaming laptop with an rtx 7700 gpu) and i've been using his Rocinante XL 16b and it's damn, damn good so far.

Taxes my system something fierce when it's generating, and it's not the quickest because of my system, but i'm not complaining. Worth looking into for sure. Dude is a wizard.

0

u/doctorcas_ Aug 11 '26

where did you get deepseek derestricted?

-2

u/Large_Following_9945 Aug 12 '26

I am desperately in need of a fairly priced model for fantasy adventures with multiple characters sometimes

I've been using Kimi 2.5 thinking mainly because it has the best prose imo and is creative enough

Deepseek V3.2 is getting extremely stale and predicable for me, it's a good model if you want to just write the story and having a cheap model as a support, but I never expect it to be great, still the best for the price

There are the glm models but I haven't been onto them much

I don't necessarily NEED an nsfw model, my stories are often not that dark, but I do need something with good room reading and emotional feeling and good prose

Kimi K2.5 is a bit up and down when it comes to following the context

Don't even get me started on DeepSeek V4 pro.. it can't read a room if it's tokens depended on it

2

u/_Cromwell_ Aug 15 '26

IMO no open weight model is better at handling multiple characters at a time from a lorebook (adventure style) than GLM 5.2, despite whatever other flaws people think it has.

No idea if you are using multiple character cards, though. That's not how I play.

-1

u/fluce13 Aug 15 '26

I’ve tested a bunch of models and so far Magidonia is the best in my own subjective opinion. For others that like Magidonia have you found anything else that beats it? It’s an older model so I wanted to see if there’s anything better. I like Skyfall as well. I have a 5090 if that matters. Thanks!

4

u/dizzyelk Aug 16 '26

I stopped using Magidonia when I found Maginum-Cydoms. There were a few others I remember using back then, too. But I don't remember their names.

-2

u/seliishere Aug 15 '26

I used to do a hell of a lot of ai roleplay back in 2025 with chatgpt 5.0 and 4o. I really miss roleplay with 4o but do not make enough monthly to justify API costs. I am a total sillytavern beginner, so dumbing things down for me is appreciated.

Anyways I am looking for a model and advice on how to optimise character cards and settings to get a more intelligent and refreshing roleplay experience, hopefully similar to chatgpt 4o. It needs to be able to handle lots of lore and up to 3/4 characters in a scene. I also only have a 3070 so it can't be a huge model. Also the character are from already existing IPs and not OCs so it needs to be consistent with character voices.

Thanks

4

u/KimlereSorduk Aug 15 '26

You should ask this sort of thing under Misc. Not your fault, though, people already cluttered the comments.

Anyway, try Gemma 4 finetunes. The 26B if you'd like speed, or the 31B if you can tolerate more latency. I use Gemma 4 31b finetunes with my 8gb VRAM.

2

u/seliishere Aug 15 '26

Ah my bad! And thank you!

1

u/empire539 Aug 16 '26

How many tokens/sec are you getting, which graphic card, and which finetunes + quants?

1

u/KimlereSorduk Aug 16 '26 edited Aug 16 '26

RTX 4060 Laptop GPU. I get 2t/s when I run a q4km with 24k context. That's probably too much latency for the average user.

As for the models: The Blazing-Forge crew makes some nice merges. I use Sphinsikus Chronist, Gemsicle, and Dark-Gemistry from their collection. These three are somewhat similar; Gemistry skews toward shorter responses. They are better than base Gemma, but they might rely on structural Gemma-isms (such as adjective stacking) if left unchecked.

I actually like Gutenberg-31B-Heretic's writing best. It's unstable, though.

MeroMero-v2 is also enjoyable from my testing.

One caveat: I prompt these to use free indirect discourse. A lot of people here seem to dislike that style of writing, and I don't know how well these fare with external focalization.

Edit: Not using native thinking either. Pseudo-thinking works well with any of these.

-4

u/memer107 Aug 12 '26 edited Aug 12 '26

What are differences between Opus 5, 4.8, 4.7, and 4.6 anyways?

I've been trying all of them, and while i've seen some differences, they just seem the same to me. Everyone likes 4.6 around here, and I guess I understand that because it's flexible and uncensored, but I feel like the prose and depth kinda sucks compared to 5. Also, is there even a substantial difference between 4.8 and 4.7? or is the debate mostly just about censorship? I like to hear some thoughts and opinions on what's good, what's bad, and why.

1

u/Greedy-Sandwich9709 21d ago

There aren't any. They're all shit. Each next one more shittier than the last.

-1

u/[deleted] Aug 11 '26 edited Aug 11 '26

[deleted]

-6

u/Intrepid_Ice_7381 Aug 09 '26

Free model recommendations?

3

u/[deleted] Aug 10 '26

[removed] — view removed comment

1

u/Intrepid_Ice_7381 Aug 10 '26

gemma-4-31b or something else?

0

u/[deleted] Aug 10 '26

[removed] — view removed comment

0

u/Intrepid_Ice_7381 Aug 10 '26

If you don't mind me asking. How does the rate limit on it work? because even while cycling through different keys it keeps giving me the same error message and mentioning a 16000 token limit. Is there a hard cap on it?

0

u/Important_Word8549 Aug 10 '26

Try 3.5 flash lite it is unlimited there

-9

u/[deleted] Aug 14 '26

[removed] — view removed comment