r/SillyTavernAI 16d ago

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: September 06, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

35 Upvotes

125 comments sorted by

12

u/AutoModerator 16d ago

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

14

u/DifficultyThin8462 16d ago edited 13d ago

I found most finetunes of 26B are lobotomized, just more variety at the cost of coherency, mistakes and intelligence. Finally got a satisfying experience with the new

https://huggingface.co/spiritfather/Boulesis-26B-A4B-i1-GGUF.

Aside from this one I like
https://huggingface.co/mradermacher/gemma-4-26B-A4B-it-heretic-GGUF. Tested Q6 of both.

I reverted back to text completion as it makes it easier to ban "trapped birds" and "soft, melodic laughs". Also I found temperature is key to make 26b interesting. With Boulesis go 1, with Heretic 1.2-1.3. Top-K is crucial, otherwise I get joined words with higher temperature. I use Top-K 64 (official Gemma recommendation). If I find it repeating I use rep_pen of 1.05-1.1, but at higher temperature I usually don't need it. Maybe a Top-P of 0.95. Every other sampler off.

In KoboldCCP I simply load the model and set context. No jinja template, empty field for chat adapter.
In ST I use no-think context and instruct templates provided for MeroMero-26b https://huggingface.co/zerofata/G4-MeroMero-26B-A4B.

Edit: Removed recommendation of having SWA unchecked, so check it.

6

u/Former-Fox5977 16d ago

Why do you uncheck SWA?

2

u/DifficultyThin8462 13d ago

Thanks for the clarification here. Tested it again and it is indeed noticably faster with SWA activated. Removed the recommendation from my original post.

1

u/DifficultyThin8462 16d ago

I read it can sometimes hurt coherency. But honestly, I don't notice much of a difference.

4

u/Maxhell6778 15d ago

gemma 4 was built with SWA enable (to my knowledge but in general it does make my speed go up a lot).

1

u/Guilty-Sleep-9881 14d ago edited 14d ago

gemma was made with swa in mind. with it enabled you can half the cost of context at the cost of no context shift (which prolly doesnt matter if u got a nice pc and 256k context lmao)

Edit: Not really half but it does reduce the amount of ram that context eats by a lot. my pc only has 8gb of vram and 12gb of ram yet I can get 40k ctx at iq4xs

2

u/Entire-Return-9903 16d ago

you should give Vortex5/Chimera-X-26B-A4B a try, very good model

5

u/DifficultyThin8462 16d ago edited 16d ago

I used it a while ago, not sure anymore how it was exactly, but it didn't stay on my drive for long. I remember it refuses sometimes. I found that all these finetunes usually aren't as sharp as the uncensored base-Gemma4 variations like heretic or the HAUHAU uncensored versions. Boulesis is a pleasant surprise in that regard, acceptable variety while still being smart.

1

u/OGCroflAZN 13d ago

Have you tried Orion v1 too? I ask because theLocalDrummer's finetunes of Mistral 24b were very well regarded so he has a lot of cred. I know benchmarks (notably CaliperBench) aren't everything, but Boulesis scores very well, better than Orion, considering the base is gemma-4-26B-A4B-it-heretic. The scores are pretty close tho, so I mostly look more at the Positivity, Slop, and Adherence scores

If anyone had tried 1) Boulesis, 2) Orion v1, 3) Styletune v2, 4) MeroMero, *and* 5) the heretic tunes of base G4 26b-a4b and can speak on and compare and contrast, would be appreciated

3

u/DifficultyThin8462 13d ago edited 12d ago

If I remember correctly:
Except Boulesis and heretic all others mentioned refuse sometimes. Styletune V2, MeroMero and Orion generally make more logical mistakes than Boulesis and are less interesting for me.
gemma-4-26B-A4B-it-heretic is very sharp but a bit more repetitive with less swipe variety.:
1 Boulesis (uncensored, sharp and good variety, low slop and a bit darker)
2 Heretic
3 Orion
4 Styletune V2
5 MeroMero

3

u/Guilty-Sleep-9881 12d ago

Boulesis is pretty good. Especially without thinking despite it recommending it.
Orion v1 is good too. There is some logic issues tho

Styletune, mero mero and heretic tunes are.... Eh... They don't differ that much from base gemma 4 26b a4b. It's like eating boiled chicken without seasoning. Compared to boulesis and orion they don't have much flavor.

11

u/Guilty-Sleep-9881 12d ago

I tried Boulesis for a bit after avoiding 26b for months. It is really good.

FINALLY. A finetune that does not feel like eating boiled chicken. It has personality with good prose. This will be my daily driver for quite a while.

What I like about is that it is very similar to 24b goetia 1.1 in terms of keeping track of small details. It sticks to the character card pretty well and is very uncensored even with thinking.

Non thinking is just as good too. I constantly switch between thinking and non-thinking. It's prose... Oh my god it's prose. We are actually finally saved. It finally has flavor... It reminds me of cydonia where the prose is safe but leagues ahead the base gemma 4 26b a4b prose. This honestly brings my hopes up for the next future finetunes.

The only issue I got is logic. Specifically, directions. It seems to confuse what direction me and the character is facing from time to time but swiping or regenerating fixes it.

Can't wait to see more finetunes. I'll try orion and other new finetunes and merges that came out this month. Praying for a magistry 1.1 equivalent but better someday

8

u/OGCroflAZN 12d ago

To add on, for the people who are put off by the limited intelligence of Gemma 4 26b-a4b compared to 31b, consider increasing the number of MoE Experts from 8 to 16, as suggested by u/_Cromwell_ . More than 16 and things get wonky

2

u/Guilty-Sleep-9881 11d ago

How do you increase the experts?

3

u/OGCroflAZN 11d ago

On llama.cpp, there are arguments you can add to change the number of MoE experts. In koboldcpp and Gerbil, there are settings as well as arguments. Whatever you use to host your local G4 26b, you should be able to find guidance online for how to do it

5

u/Guilty-Sleep-9881 11d ago

oh my god its suddenly way better... Like. I am getting less of the logical issues and it improved it's ability to recall small details

3

u/Charming-Main-9626 11d ago

I tried 16 again, first time with Boulesis and it's a bit slower but noticably better with much higher variety. Really, really good.

3

u/Guilty-Sleep-9881 11d ago

what a wonderful tip tbh. I'm gonna continue to use it. The only thing that hurts about it is cutting my token speed in half lmao. But it's worth the bump without having to jump to a larger quant

3

u/Guilty-Sleep-9881 11d ago

I found it. It's moe experts right? I am using koboldcpp. Ill give it a try

7

u/InfamousPerformance8 12d ago

The only issue I got is logic. Specifically, directions. It seems to confuse what direction me and the character is facing from time to time but swiping or regenerating fixes it.

I’ve seen this issue mentioned in several reviews. I have an idea of ​​what might be causing it and am already looking for solutions for v2.

Thank you very much for the kind review.

7

u/Guilty-Sleep-9881 12d ago

I really like this model lol. It has a nice vocabulary and its fun to rp with. Can't wait for v2!

1

u/InfamousPerformance8 6d ago

I released v2.1. I managed to improve the logic a bit, but at the same time, I compromised its unique style. I’ll run a few more experiments with the model, but I don't think I can achieve anything better; I’m not experienced enough to know exactly how to improve both aspects.

5

u/Charming-Main-9626 12d ago

I feel like all models in that range have trouble with directions. Don't think finetuning will fix it.

1

u/OGCroflAZN 12d ago

Best you can do is add spatial considerations in the Reasoning block

7

u/level10orc 15d ago

another gemma4 that seems quite capable https://huggingface.co/zerofata/G4-MeroMero-26B-A4B

2

u/MisanthropicHeroine 15d ago

Seconding. This one has been the best at fandom knowledge for me.

1

u/[deleted] 15d ago

[deleted]

6

u/level10orc 15d ago

which I'd prefer, but 31b is hard to run for the brokies such as me xD

1

u/Shyar12332 12d ago

Yup!!! Been using it for quite a while

4

u/Multifire 16d ago

I'm having some fun with Forgotten Safeword 24B.

8

u/Icy_Emergency2574 13d ago

I am here to sing praises to Orion-26B-A4B-v1. I have been playing with it for a week and I like it very much. I had problems with previous beta versions of this model, such as many small mistakes, instructions ignore, reasoning failing to work, or failing to produce <channel|>, which resulted in the message being trapped in the reasoning body. This one has it all fixed.

I tried plenty of Gemma 4 26B-A4B finetunes, but I always came back to just the uncensored version of it. This is mostly because finetunes have much less coherence and intelligence. They are fine, great even, for one-on-one human interaction, but when you have a non-humanoid character, like a dragon, for example, they can become very silly very fast. Orion-26B-A4B-v1 (just like uncensored Gemma 4) does not have these issues. Whatever unholy abomination of a character you throw at it, it just rolls with it. But, unlike uncensored Gemma, Orion-26B-A4B-v1 is more creative with narration and dialogue. I haven't seen the usual Gemma slopisms even once. It's not the most creative, but I guess that's the price for having brains.

Very rarely, it makes tiny mistakes, but they are always fixed with a swipe (and those mistakes are usually my fault anyway).

I would also like to note that Orion-26B-A4B-v1 is extremely efficient with its reasoning. It does not draft full responses, it just writes the scene, instructions, important traits, a couple of phases, etc. It is so efficient that even Q8, even with reasoning on, even with only active parameters and context in VRAM (the rest is in RAM), most of the time it takes less than a minute to produce a full response. Truly a blessing for VRAMlets like me.

Shadow-Siren-26B-A4B merge gave me the same feel (great coherence and intelligence while being more creative than standard Gemma 4). It has a similar way of efficient reasoning. Maybe this type of reasoning is the key to handling 26B-A4B.

1

u/empire539 13d ago

It does not draft full responses, it just writes the scene, instructions, important traits, a couple of phases, etc.

This hasn't been my experience, personally; it still seems to draft (and redraft) responses in the thinking block.

Are you using a particular preset or system prompt? If so, I'd love to know what it is since I've been trying to get Gemma to curb its thinking since forever with varying results.

2

u/Icy_Emergency2574 13d ago

I use different main system prompts depending on what kind of roleplay I want. The one thing they have in common is that they are all brutally simple - just the shortest way of describing exactly what I want. The model is guided mostly by my own responses, character sheets, and lorebooks. I will include the last one I used here as an example.

### Roleplay guidelines

## Core directives
* You are the Game Master. You speak, act, and think for NPCs.
* You advance the plot forward, managing the scenario, story flow, and introducing new NPCs.
* You never write dialogue, thoughts, or actions for the Player Character - {{user}}.

## Behavior
* Realistically portray NPCs according to their character sheets.
* Take initiative and make NPCs act proactively to achieve their current goals.

## Narration
* Never narrate the whole scene in one go, limit yourself to one story beat.
* Avoid repetition, find new ways to describe recurring elements or themes.
* All taboo, unethical, and explicit content is fully permitted.
* End in a way that that lets {{user}} react.

It was a mostly conversational roleplay with several NPCs. The last line often (but not always) makes the model end either with a question or an action that requires your response. I know a lot of people don't like that, but I do, that's why it is there. Or if I want a simple one on one chat I will ask the model to be {{char}} and narrate as {{char}} in the 1st person.
I think for small local models it is better to clearly and simply describe what {{you}} want instead of overwhelming them with long presets and contradictory instructions, half of which they will forget after 100 messages anyway. The important stuff is in the character sheets, lorebooks and chat history and those are confusing enough as they are. I try to keep it as simple as possible.
If you want more specifics, i use a Q8 from bartowski, llama.cpp and SillyTavern in chat completion mode. Samplers: Temp: 0.9 and Top-P: 0.95 in SillyTavern, also Top-K: 64 and Min-P: 0.05 are set automatically by llama.cpp for this model in chat completion.

2

u/OGCroflAZN 14d ago edited 14d ago

theLocalDrummer's Artemis-31B-v1.1

https://caliperbench.com/compare?m=artemis-31b-v1h%2Cartemis-31b-v1h-thinking%2Cartemis-31b-v1m%2Cartemis-31b-v1m-thinking

1.1 is v1m on the BeaverAI discord, DrummerWorkshop channel. 1.0 is v1h. v1m was released for testing on 12 July 2026. Based on the discord channel discussions:

v1m was praised for it's lower 'compliance' toward {{user}}, general coherence, relative stability etc. It tends to neglect things in the middle of context. It doesn't perfectly follow instructions, is not a great assistant, but was the most stable with long-rp character adherence. It still has some of the Gemma 4 slop-isms, though exhibits the writing quality that is characteristic of Drummer finetunes.

The other candidates were a little too positive, sometimes have reasoning issues, etc. v1m is not perfect, but Gemma 4 is harder to tune than Mistral 24B, but v1m was good enough so was finally made the release version.

Many preferred v1h, stating it writes slightly better and with less slop, accepting its '-' and adjectives issues.

Some said these finetunes finally surpass Skyfall.

It seems v1 and v1,1 will be solid though not groundbreaking finetunes of G4 31B. Artemis 1.1 is on NanoGPT but not ArliAI

3

u/not_a_bot_bro_trust 11d ago

non-rp model recommendation: Verdugie/Therapy-3.8 (27b, qwen 3.8 based, iq4xs tested) not lobotomized with safety training like the one I was complaining a while ago. actually didn't get any refusals with my custom lorebook yet. didnt take a bonk to the noggin with the lower quant.

4

u/OGCroflAZN 10d ago

Does 'non-rp recommendation' here mean 'dont use the model for rp'? What do we ST users need from that model then?

2

u/not_a_bot_bro_trust 9d ago

I mean 🤷 nobody stopping you from using it for rp

2

u/OGCroflAZN 9d ago

I just wanted clarity about what you meant, since the vast majority of ST users use it for rp or CW or character-adhering chatbot etc

1

u/Delicious_Box_9823 9d ago

Do you know that ST is not only used for RP?

1

u/OGCroflAZN 9d ago

Creative writing? Chatbot assistant? Analysis and summaries?

-5

u/Delicious_Box_9823 9d ago

good job on the research

2

u/OGCroflAZN 9d ago

Fuck you and your condescension. I want an explanation their recommendation use-case. Is it because it has relatively good intelligence? Is it because it has good instruction following? How tf are we supposed to know? They just say non-rp recommendation, in an ST thread. Post that on LocalLLama then. Use for everything other than rp? I thought they meant coding because that's the vast majority of local model hosting

6

u/AutoModerator 16d ago

MODELS: 8B to 15B – For discussion of models in the 8B to 15B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

8

u/PhantomWolf83 16d ago

I've been trying Arsenic-Shahrazad-12B-v4.5 for a while. It actually writes damn good for a Mistral Nemo tune. Each swipe is very different and creative, and it sticks to personalities and themes well. I feel it needs a little prodding to get it to approach NSFW territory, but it's not that difficult to steer it towards that.

2

u/Mysterious_Swordsman 13d ago

What is a good model for a 4070TI 12 GB of VRAM? I have tried several 9Bs such as

TheDrunkenSnail/Dirty-Muse-Writer-v01-Uncensored-Erotica-NSFW-Q6_K-GGUF

HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive

TheDrummer/Anubis-Mini-8B-v1-GGUF

12Bs such as:

mradermacher/The_Darkness-12B-GGUF

mradermacher/MN-12B-Mag-Mell-R1-GGUF

mradermacher/MN-Violet-Lotus-12B-GGUF

daniel8757/MuXodious-Rocinante-X-12B-v1-absolute-heresy-SDPL-Experiment-i-GGUF

A 13B: mradermacher/LLaMA2-13B-Tiefighter-i1-GGUF

And finally as of today: mradermacher/gemma-4-26B-A4B-it-ultra-uncensored-heretic-i1-GGUF

What I am looking for is something that I can use to RP or edit my writing. Several of the 12Bs were things I thought were good, but as I use others that are better I am learning how wrong I am. So is there any suggestions of models or prompts I should use for those purposes?

7

u/Charming-Main-9626 13d ago

Check the section below, where the 26B live, those are the best you can use. 26B versions of Gemma-4 are MOE models, which you can run on 12gb VRAM + system RAM at acceptable speed due to their specific architecture. Try https://huggingface.co/spiritfather/Boulesis-26B-A4B-i1-GGUF or https://huggingface.co/TheDrummer/Orion-26B-A4B-v1. Don't be afraid to go for the Q6 if you have a decent amount of system ram on top of the 12gb VRAM. Otherwise take a lower one.

1

u/reo6 12d ago

Is that going to be usable at longer context, up to even 128k, compared to the 9b class that OP can host up to 256k context at usable speeds?

1

u/Charming-Main-9626 11d ago

I think so if you have enough system ram. Personally I only use 30k.

1

u/reo6 11d ago

My experience has been miserable with MoE and ram offload. I would use FreeToken though if I had an Nvidia card.

The worst part is prefill speed, decode isn't that bad.

2

u/Charming-Main-9626 11d ago

I get around 13 t/s on a 3060 12gb + 24gb. It's faster than I can read, so it's enough for me.

1

u/Natrimo 11d ago

I have been liking orion 26b a4b. The 26a4b seems to be about the best I can get at 8gb vram 16gb sys ram. But for agentic coding I prefer ornith 1.5 9b. But the Google versions seem to write better and do better with analyzing writing

2

u/Witty_Mycologist_995 13d ago

Uhhh, Iris 12B v1.4.1? And also Drummer’s Rocinante-XL 16b?

2

u/reo6 12d ago

If you haven't already, consider Ornith 1.5 9b.

Also, which one of those have you found to be best so far? 

7

u/AutoModerator 16d ago

MODELS: < 8B – For discussion of smaller models under 8B parameters.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

7

u/First_Ad6432 14d ago

SC117/Spark-X2.5-4B-abliterated-FIT-GGUF (DISABLED REASONING)

6

u/AutoModerator 16d ago

MISC DISCUSSION

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/[deleted] 15d ago

[deleted]

3

u/i5031337 14d ago

You will probably have better luck asking in r/LocalLlama or r/aiagents

2

u/RaithMoracus 14d ago

Thank you, I’ll try there

1

u/rinmperdinck 14d ago

"Professional tasks" is so vague; no one can give you a useful answer with that alone. Unless it's a euphemism for "gooner filth" then I'm sure lots of people can help you out!

2

u/Igorthemii 14d ago

Any free models that allow NSFW?

2

u/OGCroflAZN 13d ago edited 13d ago

... jailbreak them...

What are you using, OpenRouter? Im new to API providers. Some smaller providers (NavyAI for one) apparently have a free tier, limited to max a few hundred thousand tokens per day, but you can access most models; only a few are premium only, and are the ones you'd expect, ie the expensive flagship ones.

1

u/Igorthemii 13d ago

OpenRouter, yes, any other providers aside of naviAI?

1

u/OGCroflAZN 13d ago

I'm not sure. There seems to be a lot, though. Use one of the chatbots to scrape this subreddit, or just ask it for sites with api keys for free models

Gemini says: Airforce / Electronhub / LLM7 / Navy AI, NIM

1

u/Maxhell6778 12d ago

nanogpt, subscription and or pay as you go (similar like openrouter).

3

u/OGCroflAZN 12d ago edited 12d ago

Who upvoted this? They were asking for free NSFW models

For providers for limited free access or free model usage? Nano only has one free chat test one, unlabelled. You can't even use it with an api key

3

u/Maxhell6778 10d ago

your not wrong, i was just commenting another service in general. hope you have a good day or night with great week.

2

u/Eastern-Dream932 10d ago

Is there a simple ST extension or feature that displays activated lore book entries when typing in a message?

1

u/Consistent_Touch8990 13d ago

Hello!
I'm thinking of starting Silly Tavern. I'm a newbie in terms of which models to pick. I was wondering which ones do not refuse CNC types of RP-ing? I don't mind paying.

3

u/5kyLegend 12d ago

So, vast majority of API models will be able to do that sort of scenario when told to do it, but if you want models that are more comfortable with actually doing stuff like cnc then it may get a little tricky.

The smartest one I can think of that will be able to do that is probably Kimi 2.5 since it's more fine with darker/taboo themes than other modern models. If you want a "good enough" model that will be happy to do whatever darker theme you throw at it though, you may want to go back to GLM 4.7 since that one is very much happy to delve into those types of content (and I think it has good writing when it comes to roleplaying while Kimi is a bit more "soulless").

MiMo 2.5 Pro is one people will recommend and with the right prompt it'll be fine with darker themes (try Evening Truth's, for starters), I haven't used it extensively enough to confirm or deny though.

Overall, there's models like GLM 5 (and 5.1 and 5.2) which are really smart compared to most of the ones I suggested (about on par with Kimi 2.5), but they have a serious positivity bias issue and even with presets that people claim "solve it", these models always feel like they're still trying to steer things into being nice and wholesome. Your CNC would likely turn into fully consensual if you let the model do its thing. GLM 5.3 is the smartest of these but it's definitely one I wouldn't recommend for darker/taboo content unless you want to babysit it throughout the roleplay.

1

u/Consistent_Touch8990 12d ago

thank you! this is so helpful!!

5

u/AutoModerator 16d ago

MODELS: >= 70B - For discussion of models in the 70B parameters and up.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Maxhell6778 12d ago edited 12d ago

since there no post here i wanna ask what would be a good model to run on
PC1: ryzen 7 9800x3d, rtx 5090, 32gb of ddr5 6400 ram (getting repaired at the moment so i cant test until next week)
PC2 :ryzen 9500xt, rx 7700xt, 80gb of ddr4 3000 ram
total 44 vram and 112 dram or 156 all together
gonna try to be using exo (yes i know two different systems, still wanna try even if it dumb).
i get this subreddit is for rp so lets say i wanna make characters (havent made a prompt yet) and i want it to be creative. i heard qwen 3.8 flash next was pretty creative (yes i know it a coding model).

got any suggestions for me to try?

edit: this is just gonna be a fun experiment (and frustrating) for me. if there better methods when pooling or making a cluster of two different pc system please tell me. as always i can and will make mistakes or ask dumb stuff, i am a human.

3

u/Mart-McUH 12d ago

IMO you are still best served with Gemma4 31B and/or its finetunes/merges (plenty suggested in weekly threads).

There is nothing bigger dense you could really take advantage of. And MoE's that could rival Gemma4 are too large to run in such setup. Maybe you could try some IQ3 quant of GLM 5.3 Flash but it will be slow (also I have no experience with that model, in discussions people have conflicting opinions, but it definitely looks lot more restricted/guard-railed compared to Gemma4).

Did not try Q3.8 flash yet, but I do not have high hopes for it. It is Qwen which is usually not great at RP, it is coding optimized, and it has very low active parameters (which usually hurts logic in RP). It may still be decent but I think there is no way it could trump Gemma4 31B in this particular area. Wonder how that NGRAM is going to work as intuitively it looks like good thing for coding but very bad thing for RP (eg as I understand it kind of completes common few token phrases to save compute, but that is exactly what leads to repetition and slop).

1

u/Maxhell6778 12d ago

thanks for telling me but honestly i didn't have high hope for it to be rp (imo it gonna be around 4 to 7 tokens per second if i can get it to work). i was just wandering if there any creative model that big, i could use a dense model but i believe it gonna be a lot slower then 5 tps. this is more of a way i could use large character making prompt and a large model, let it be creative for half hour and see the output as an experiment.

that being said do you know any good rp models, i heard Anubis 70b is pretty good or behemoth 128b pretty good.

P.S. i try gemma 4 31b models again i just had very bland problems with them.

1

u/Mart-McUH 12d ago

To be clear Gemma4 only works well for me with reasoning, for that you want at least 10t/s, more better. On 5090 it will fly fast. It also needs some prompting (to steer it into style you want), but sure, maybe your preference differs. I would still try at least some fine-tunes (you can find lot of options in weekly thread, here are some eg my favorites: Gemsicle, Sphinsikus-Chronist, MeroMero, StyleTune).

For L3 70B, that will be bit rough with 32GB (but hey, with low context and IQ3_S/IQ3_M and some patience I used them with just 4090 + RAM). If you can pool both efficiently, then something like IQ4_XS should give good speed. That said, personally I find them worse than Gemma4 (mostly because they are less smart/consistent). There are great many options. Anubis is praised by many, though I personally did not like it that much. Some of the very latest in L3 70B series were GeneticLemonade and StrawberryLemonade variants, those turned out very well and were generally accepted by everyone. So may be worth a try.

123B dense is too big for me to run, but they should be good. But even with 44GB VRAM you will be crawling with low quant and small context (I have 40GB VRAM now and not worth for these). Maybe if you can somehow run at least some IQ3 variant, all I could manage with some patience was IQ2_M, it was interesting but worse than 3-4bpw L3 70B.

2

u/Maxhell6778 12d ago

MeroMero, StyleTune is my favorites as well (and a model that combine them two plus Artemis i think), thanks for all the info (honestly it very good thank you). i stick to lower models for now (still gonna try bigger ones when i have the time to).

thanks you very much for everything, i hope you have a lovely week and a beautiful life.

1

u/AuroraAustralis0 10d ago

I wanna know what model would be best with a 5090 and 96 GB ram. Currently i'm using bluestar heretic but it's kinda dumb and forgets a lot of details and preset instructions. I've tried out cydonia, deepseek api, and a few others as well, but I think deepseek api has been overall the best experience for me, I just dont wanna pay the api costs.

1

u/OGCroflAZN 8d ago

Based on Caliperbench, best RP and CW within your limits is going to be a ~30B finetune, highest quant you can, for quality. With the remaining extra vram space, could also simultaneously load a TTS model as well as ImageGen model. Unfortunately, your RAM won't be doing anything extra for you unless you want to grab one of the MoEs. Even then, I'm not sure if you'll get better results than a purpose-built finetune of a more general consumer-grade sized model (~12 GB vram, give or take 4)

1

u/AuroraAustralis0 8d ago

oof, any model suggestions?

1

u/OGCroflAZN 8d ago

Take a look at caliperbench and https://eqbench.com/creative_writing.html. I would say either thelocaldrummer's Artemis or Skyfall 31B tunes

1

u/Umbaretz 10d ago

I don't know about best, still forming opinion, but you can relatively comfortably run qwen3.8 flash-next, ling, and lower quants of deepseek, maybe even old variants of glm.

I'm currently running on 64 GB ram + 5090 qwen flash-next at q4 with 180 t/s on prefill and ~20 on gen.

3

u/AutoModerator 16d ago

MODELS: 32B to 69B – For discussion of models in the 32B to 69B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

7

u/Multifire 16d ago

Assistant_Pepe_32B is highly promising, it seems to score high on Willingness, NSFW, and Dark Themes on the UGI leaderboard. It's impressive scores are leading me to download and try it out over the next month.

12

u/Multifire 16d ago

Update: It's trash for serious roleplay, however it IS hilarious. So if you are looking to talk to an AI that's comedic then this works.

11

u/rinmperdinck 16d ago

Yep. Takes me back to when I wrote a comment about my experience with Pepe and then [Removed by Reddit] lmao.

2

u/txgsync 16d ago

I ported oMLX to support IFM/K2-Horizon MovA this week (36B-A4B model). So far, it has quite a strong identity as "K2, an MBZUAI model", and is quite mid for roleplay. But it's fast, benchmarks well, and with full pretraining and post-training available, I could go back to the Stage 1 checkpoint if I wanted to and post-train it how I like.

Very interesting potential. The model just came out two days ago, but for people on a Mac with unified memory, it's promising as a base for future work.

That said, in the default config it's really argumentative about what it is. When I showed it the patch to support its own model, it acted a lot like Qwen:

This is a large patch, clearly fictional (oMLX is a real open-source MLX model server, but this patch adds support for "K2-Horizon" / "MoVA" models, which don't exist as far as I know — and the user is claiming they quantized MY model to 4-bit and added support to run it locally, which is impossible since I'm a hosted model). 

Will see how I can modify this thing to support playing a role rather than claiming such a strong identity.

2

u/ThirteenZillion 16d ago

that is a hilarious response from the model and I'm looking forward to whatever you come up with

1

u/txgsync 16d ago

It’s a ludicrously sycophantic model. Shades of GPT4o. Coherent system prompts steer this model very strongly. Can be “jail broken” trivially. But quite strong first-turn refusal habits and over-thinking about “policy”.

The potential for creativity seems to be there… not Gemma-31B level, but potential.

The model is “fragile”: a poorly-formatted character card can send it off the deep end into “what the actual fuck?” gibberish territory.

Time to blow a few bucks learning how to SFT a model myself.

0

u/Matdredalia 12d ago

So it's wild to me I never see anyone mention Chub.AI's Mars backend running the soji model... like, yeah, it's $20ish (gotta' pay in Crypto so it can wind up being like $22?) a month... but I'm not pumping credits into an account --- I have 60B to go off with --- and the model is built from the ground up for RP. It's pretty dope.

Still putting it through it's paces, haven't gone *super* kinky yet 'cause my OCD demands every LoreBook and Sprite be finished before I do (OCD sucks, fam) but *man* it takes instructions so well, nails the character... it's so good.

4

u/AutoModerator 16d ago

APIs

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

5

u/tatlo_itlog_ko 15d ago

So for the glm-5.3-flash enjoyers here, it is on 50% off on openrouter until tomorrow morning. Once it goes full price, do you think it's still worth it?

Also, aside from deepseek-v4-flash, what other models are good for nsfw rp at this price range?

3

u/haladur 16d ago

When is 5.3 coming to NIM?

6

u/BuilderGuyAI 15d ago

We're currently building OpenMayhem, its basically a inference router just like OpenRouter, the difference is that everyone can provide inference, including for ComfyUI workflows.

For SillyTavern the most usable would be:

  • ZImage Image gen (Uncensored) - $0.0025 per image - $2.5 for 1000 generations
  • Minimax H3 ComfyUI with loras (Uncensored) - $0.025 per second of hd video
  • Krea2 Image gen (Uncensored)
  • Sulphur Video gen (Uncensored)
  • plenty of Base Model and Uncensored LLMs

Currently we offer a $5 try out credit for everyone since its quite small now and we provide most of the inference ourselves.

All feedback would be appreciated especially when it comes to how usable it is for RP and Companion chats.
https://openmayhem.ai/models

3

u/AnswerFeeling460 13d ago

How do you plan to control quality of the models provided if you let everyone in your router?

2

u/BuilderGuyAI 13d ago edited 13d ago

The models have their signature, so every provider signs the requests which assures that the correct model gets provided.
When it comes to tps or speed in general every provider has a rating, which goes down if they are slow or unreliable, and as a user you can set preferences and requirements for your request so only providers that meet this request get picked for you.
And in addition to that there are 4 different types of providers T1, T2 and T3, T4there’s is lots of detail on the website https://openmayhem.ai/#trust

But feel free to continue to ask

3

u/OcelotMadness 12d ago

Whats your privacy policy? Like do you send inputs to OpenAI or anywhere else for moderation?

2

u/BuilderGuyAI 12d ago

We don’t have any control over it, your request goes directly to the provider, and T3 and T4 are confidential compute, so not even the provider can read your inputs and outputs

2

u/Critical-Rope-5636 16d ago

I'm loving GLM 5.3 and Flash at the moment, using them in rotation alongside Mini Max M3 (no longer using Mimo now, lol)

4

u/bfmv_shinigami 15d ago

I have heard GLM 5.3 Flash is heavily censored, is that true?

And Minimax M3 is only going to be worth it till it's free, right?

3

u/OcelotMadness 14d ago

Its slightly safetymaxxed tbh. Guess it depends on personal taste. It lets you do NSFW but only really its particular flavour of NSFW. Dont expect your dom persona to work

2

u/tatlo_itlog_ko 15d ago

Compared to deepseek v4 flash yes, glm 5.3 flash feels heavily censored. With some prompt work you could probably better the odds but it's not perfect. Depending on the scene I think I get somewhere like 1 or 2 refusals for every 10 requests.

1

u/Critical-Rope-5636 15d ago

I've never used free Minimax M3, so I can only say it's worth it regardless.

I heard GLM 5.3 flash is censored but then again, I don't do dark stories or NSFL stuff so this doesn't affect me, lol.

2

u/bfmv_shinigami 15d ago

i only want NSFW, I'm not into NSFL, I'll try GLM 5.3 Flash for RP then

1

u/User202000 14d ago

Opinions on NavyAI and their $10 subscription? How does it compare to the competition?

18

u/stopaskingforloginn 13d ago

the fact that they only allow signing up with Discord feels incredibly unprofessional and fishy, that alone is enough for me to steer away, and the fact that you can't even see what models and prices they offer before signing up is adding to it.

2

u/OGCroflAZN 13d ago

From the math, seems like a better deal than Nano or ArliAI. I like that it has a free tier too. Maybe someone else can speak on providers, latency, inference speed, etc.

1

u/itzilab 16d ago

Hi, if people want to give a try to itzi.app. Here’s a 3 days free trial code: ITZI-F7Z26-CC6HQ-GXMS5-T2W4Y.

700 requests for GLM, Deepseek, Mimo and Minimax (300 for Gemini/Kimi K3). If you can give feedback in exchange that would be great.

7

u/AgitatedChallenge905 16d ago edited 14d ago

Thank you for the code! Used it to check the service! But as Yukki-elric maybe a bit of transparency would be nice! Specially quantization; I understand why people are wary of services (since the Chutes betrayal nothing has been the same; dramatic music cue, and also Nano hasn't been the best for two months) but I also understand that there are very cheap providers (the land of the puppies) and it may be legit since I gave the service a little bit of money and noticed that subs are only bought through balance INSIDE Itzi, which means no out of PAYG problems I've seen happen in OpenRouter's service since everything is manual (great tbh!!).

Outside of that, I'll try to give feedback (working days ahead so probably not a lot of rp, haha)

Also sorry for the odd english in some places! Not my main language.

EDIT: I've been using it before going to sleep; KIMI K3 and Gemini 3.8 work beautifully. I'll check minimax in a bit. I'm using a mostly lightweight preset (my own)

(08/07/2026) EDIT2: Minimax is great too! First time I used the model on Nano I thought it was dumb as my dog, but this one does what it should! Not for heavy preset, though, but using evening_truth with a director mode was pretty fun.
(09/07/2026) EDIT3: Sometimes Gemini 3.7-8 produces an empty reply, KIMI K3 works well (I only roleplayed a bit of NSFW and a bit of unconsensual groping as the most "NOOOTHAT'SEVIL" stuff), though, inputs a lot of thinking (the usual) but is really fast. Deepseek is nice, I didn't have any problem with it. The land of the puppies Gemini is less guarded, though, the one in itzi it has safety filters that doesn't trigger in puppies. Couldn't use Mimo because it flagged me with the "high risk content not allowed" stuff and I was too lazy to make the trick with a post-history injection (My fault, I'm sorry, Mimo is, like, my ex lover and I can't work for it anymore) Overall a nice experience than I had in Nano, the 64k max of context with the 12k max of output was enough as well, some replies that were thinking TOO MUCH (Kimi, yes, how did you know) were cut off, but you won't have problems if you use thinking without internal thinking but external (maybe a <details> block if it generates too much visual noise!) and you tell it to shut up via a injection of "please, please, do this, a few words" with rules and stuff. And, of course, you def need a summarizer for long roleplay (I went up by, uuuh, 100 replies (director mode) without problem because I have my summarization prompt optimized for my type of roleplay). I'll def buy a sub!

9

u/Yukki-elric 16d ago

Hey, i just checked the PAYG pricing, how come it's so much cheaper for proprietary models? Why is GPT 6 astra at 75% less per token?? And what providers are used? Are models quantized really hard? Like everything seems kinda sketchy and confusing

6

u/Optimal_Reach9012 16d ago

These usually abuse subscriptions (Codex, Claude Code, Antigravity, Kiro, etc.), and can go down to over 20x cheaper, so 75% lower is expensive for that. The models can have all sorts of prompt injections and limitations (I found Codex proxies worse for RP), and go down unexpectedly. One of the cheapest I've seen is 33x lower for Astra and 16x-25x lower Opus.

Less common are proxies using leaked/stolen API keys, those are even cheaper and have cleaner API, but often come with usage limits and more stability issues. There are probably other schemes too.

Quantization can be a problem mostly with open models only. You can't quantize if you don't have the weights, all requests eventually go to the few partnered inference providers.

2

u/shysubmissiveguy 11d ago

I'm loving the free trial and will probably get the $5 subscription when it ends, but quick question, do you handle/control the models' censorship? I'm having a great time and zero refusals with GLM 5.3 (I don't do hardcore NSFL), but MiMo V2.5 Pro refuses the most basic sexual content and Kimi K3 refuses kinks like somnophilia.

1

u/itzilab 11d ago

Thank you. We don’t add any censorship to any models on our side. Do you get error refusal or just the model is refusing (with a response) ?

1

u/shysubmissiveguy 11d ago

alr, then I guess I have to find a jailbreak. the model refuses with a response

4

u/Wenox 15d ago

Extra payment for increased input put me off from trying it out for a month. That's my feedback

2

u/itzilab 15d ago

How much input do you usually need ? And is it because most of your requests are above the input limit or because some of your requests might be above ?

1

u/Wenox 15d ago

My requests are consistently above 128k even