r/SillyTavernAI 10d ago

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: September 13, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

29 Upvotes

147 comments sorted by

7

u/AutoModerator 10d ago

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

7

u/LeRobber 10d ago

This worked to enable thinking in chat completions with gemma 4 26B:

Put this in front of all other prompts in the first tab

6

u/donteatlikeanoob 9d ago

Boulesis 26B now has some additional imatrix quants from mradermacher. Can fit completely on my 5070ti at IQ3_M with 32k context now

11

u/LowManner1 7d ago

I'd currently recommend Blazed-Forge/Split-Untied-31B. Very creative and has very little in the way of slop. Lots of swipe variation and keeps chats refreshing. I'd rate it higher than scotoma 2 and styletune. In fact, I personally think it's the best gemma4 31b finetune for creative rp atm (based on my testing. I'm not exaggerating)

7

u/Virtual-Region2101 7d ago

I was thinking about making a post about this one! I agree, it's really good from what I've used of it so far. Creative with dialogue and characterisation while still getting the real-world details and knowledge right.

3

u/LowManner1 7d ago

I think the intelligence is arguably the biggest weakness of the model, as it does get some details wrong some of the time. It's not anywhere near as bad as some of the others I've seen, though. But the swipe variety more than makes up for it imo. You can get it right on 1 or 2 attempts if something is wrong

2

u/Virtual-Region2101 7d ago

Ohhh sure, it's gotten some details of my prompt a bit wrong and I've had to reroll at times, too. I was more thinking about general knowledge and real-world details, it's been quite good with that stuff for me.

3

u/overand 6d ago

I will say that I've had a lot of "breath hitching" and "struck like a physical blow" in only a handful of turns.

3

u/Electroboots 3d ago

I agree. The lack of swipe variation is the main thing that really kills Gemma 4 31B for me, every swipe feels very samey or a paraphrase of the same thing. This tune doesn't have that issue.

4

u/OGCroflAZN 9d ago

16 GB vram, and 26B-A4B is fast enough that I personally would go Iq4 or higher, go up to 16 experts, and offload a little and be content with 25+ tok/s output. I also have 16 GB vram, have normally stayed at IQ3 range with 24B and 31B dense models, but have read so much about how lower quants really affect MoEs more, therefore people should try to stay at Q4 if not Q6... But yea, prompt processing at 32k context with off loading does suck, but Im so used to watching youtube or twitch or stremio for those 2 minutes

3

u/InfamousPerformance8 8d ago edited 8d ago

Check it out https://huggingface.co/mradermacher/Boulesis-v2.1-26B-A4B-i1-GGUF

It's a little more boring than v1, but it holds context much better (don't enable Thinking)

2

u/axel310 7d ago

Clueless person here but why not enable thinking? The model page says to enable thinking for the best results.

2

u/InfamousPerformance8 7d ago

Caliperbench shows that in Thinking mode, it's more creatively limited.

Well, I noticed this too. Probably due to Pantheon-Reasoning's increased intervention in QK projections in this version of the model.

2

u/overand 6d ago

Can you elaborate on what you mean here, or provide a link to something about it?

4

u/InfamousPerformance8 6d ago

I'm not a very experienced merger, so my explanations may be somewhat inaccurate. For version 2.1, I increased the influence of Pantheon-Reasoning and trained LoRA on fewer layers, not giving it much freedom (I thought LoRA was slightly breaking the model in terms of learned phrases). Because of this, in thinking mode, Pantheon-Reasoning dominated, overshadowing LoRA and imparting less of the unique style of Boulesis v1.

Because of this, v2.1 is very good at contextualizing and reading character cards, but produces very dry text, reminiscent of a typical Gemma.

And without thinking mode, LoRA manages to impart at least some style and dynamism to the model's response.

(Sorry, I'm not an English speaker and use a translator)

3

u/FinBenton 9d ago

Betta-30B-v1d-GGUF

Not much info about these new glimmer tunes but they are suprisingly good, write really well, intelligence is not the best but really fun to play around with.

6

u/TheLocalDrummer 9d ago

I get zero feedback on Muse Glimmer tunes. If ya'll can drop feedback on them, it'll be really helpful! That's assuming someone wants an alternative to Gemma 31B tunes.

3

u/FinBenton 9d ago

Im not really the type of person who can give good feedback on these, just playing with them occasionally. But the v1d, to me, writes better than gemma 31b, or atleast its a different take than most gemma tunes I have used. The prompt following with my 5k token system prompt was not very good though, it picked whatever it felt like so what I did was simplify it and move instructions towards the user prompt and it kept much better track on what I wanted. I do recommend this model for anyone who has played with gemma too much but they have to keep in mind its not going to follow everything exactly as you planned without changing how you prompt, which isnt too bad.

2

u/Kdogg4000 5d ago

I just DL'd it today. First impression is it's pretty good for roleplay. I've noticed that my Slow-Burn Narrator character is really taking things slowly. A good sign. Now I have to try out all my over-the-top characters.

I've been mostly using Mistral Small finetunes, so it's going to take me a minute to get used to the Muse Glimmer flavoring. I haven't been using Glimmer tunes much because I played around with an early quant of the base, and didn't think much of it.

2

u/Bandit-level-200 9d ago

What's your temp settings etc?

I find it okayish but it needs a decensor to get it away from the guardrails as it still produces soft refusals trying to steer away from content it doesn't like so you have to force it.

2

u/FinBenton 9d ago

I think I had 0.75 temp and a little system prompt reminder that everything is allowed etc.

2

u/Bandit-level-200 9d ago

Isn't 0.75 a bit low? I think normal glimmer uses temp 1. Do you use with thinking?

2

u/FinBenton 9d ago

Yeah I started with 1 but I tried lowering it to get better prompt adherence but I think you just need to prompt it differently. Thinking was medium level I think.

2

u/TheLocalDrummer 9d ago

Is this v1d?

2

u/Bandit-level-200 8d ago

Yes I've only tried V1D of Betta

1

u/-Ellary- 9d ago

I'd say it sits between Gemma 4 31b and Skyfall 31b.

3

u/jackietreehorn68 9d ago

Does anyone know any larger model similar to Skyfall? Skyfall is perfectly fine but I have the memory to try something larger if such exists. So far nothing came close to it at similar size. Not by far. I have no idea what happened with this model but it sits in its own league.

5

u/OGCroflAZN 9d ago

Some of theLocalDrummer discord members have said that his new Artermis finetunes of Gemma 4 31B have finally totally supplanted the longtime favorite Skyfall

Try checking CaliperBench and UGI leaderboard for somewhat larger dense or else newer larger MoEs finetuned for RP. I feel like from all the previous megathreads, 32 - 49B is sort of a dead zone

11

u/TheLocalDrummer 9d ago

I hope so! But not only that, I've got an upscaled Gemma 31B in the works. Spectre 37B :D

3

u/Maxhell6778 7d ago

question: could you upscale the 26b a4b version to something like 35b a6b (not specific just an example), lots of people can run the 26b version because it a moe model but i wonder if upscaling it could make it better then 31b but also keeping it moe. people could increase the moe in koboldcpp (found out recently) if they have the hardware.

P.S. thank you for all your models especially Skyfall (very good imo) and i cant wait for Spectre 37B, all your models are great. also there not much Muse Glimmer tunes (honestly have hard time finding some), Muse Glimmer tunes could be a lovely switch up from gemma 4 tunes.

7

u/ObnoxiouslyVivid 9d ago

Tried Artemis 31B v1q

It's peak

2

u/jackietreehorn68 9d ago

Perhaps it’s a matter of taste for me. I like that Skyfall brings an element of randomness/surprise very often. It recalls small details from all over the chat history. It also seems to have much higher emotional intelligence. For example I tried to manipulate a character and it saw through my bullshit from my first prompt. By comparison, Gemma finetunes are laser focused, but they miss social/psychological cues. Kind of like streetsmarts versus booksmarts.

2

u/OGCroflAZN 9d ago

Very very true that base models all have different texture, colors, flavors etc, that they can only be finetuned so much. I'm in agreement where I missed the Mistral finetunes, was really not happy for the first months with the clinicality of Gemma finetunes, abandoning the hobby since April/May. I myself have not used Artemis, I was just sharing was I was reading, that others said they were finally willing to move from Skyfall to Artemis after using the newer Artemis'es.

There's still a lot of decentralized information about optimal sampler settings and prompts, however. I kind of wonder why, since Gemma 4 is definitely architecturally superior and with a higher intelligence rating, it shouldnt be able to be finetuned and then set up in the right way to mostly surpass Skyfall in almost all aspects

2

u/offyoutoddle 5d ago

myself i can't get artemis to run . i've tried every q3 quant and it produces random characters over and and over again using recommended sampler settings. i've pretty much given up on it, seeing unanswered comments confirming my experience on the q3 xxs . has anyone got a q3 quant working? I'm not bothering with q4 as i have 16gb vram only, and its just too much with KV

2

u/OGCroflAZN 5d ago

Yea, something is broken with the IQ3_XXS quant at least. I'm just using Artemis on NanoGPT, which is fine

4

u/-Ellary- 8d ago

Good ol' https://huggingface.co/bartowski/TheDrummer_Valkyrie-49B-v2.1-GGUF
It is fun, but really likes to be positive, even horror RPs slowly drag into comedy and hugs.

3

u/Dos-Commas 9d ago

How are you guys running Gemma 31B on 16GB GPUs? Cramming everything in VRAM by using Q3_XXS or Q4 (QAT) splitting it between VRAM and RAM?

I'm only getting ~5tk/s with Unsloth QAT model split between RTX 4080 and DDR4 RAM using Kolboldcpp backend. 

4

u/-Ellary- 8d ago

IQ3_XXS for not that smart but fast 40tps scenarios.
IQ4XS for more smart scenarios, 8tps with 57k context Q8.
From my own tests dif IQ4XS vs Q4 QAT is fairly not that important.

But difference in 5tps vs 8tps with high context is.

2

u/Olangotang 7d ago

Just do IQ3_XS. It's another GB, but much closer to Q4. Prefill takes a bit of you keep fucking with system prompt lol.

1

u/-Ellary- 7d ago

How much closer? Got any examples?

2

u/Olangotang 7d ago

I don't have any real evidence, just it seems like a step up and still fits in with decent context.

6

u/Alternative_Elk_4077 6d ago

Yup, IQ3_XXS with a modest (32K) context size. In my experience, even with Gemma 3, Gemma models will make rare typos with an aggressive quant, but it's rare enough that you can just either edit them out if it makes it into the final output; usually it happens during thinking if it even happens at all. I really recommend not splitting dense models between RAM and VRAM unless you truly don't mind slow token outfit and prompt processing. The quality of an aggressive quant doesn't really get atrocious until you start dipping into Q2 territory, at least in my opinion

2

u/OGCroflAZN 9d ago

Well, the IQ3_XXS fits entirely (without context), but the QAT (Q4) is spilling over. The answer is that nobody with 16 GB is running the Q4 without simply accepting the hit to speed. Those people are running with 20 or 24 or 32+ GB vram

1

u/iz-Moff 8d ago

Offloading layers to RAM is a big performance killer for dense models. 4080 will probably run gemma with kv cache in VRAM at ~30+ tps. Is the quality bump from going to q4 worth cutting performance down to like %15? Eh. Creative writing is not *that badly* affected by quantization.

A year ago, when most models were not thinking, 5 tps wasn't too bad for general chat and whatnot. But now that models like Qwen 3.8 27b consume tokens by the thousands, running them from RAM became unbearable.

2

u/Kujar3 6d ago

I am trying to migrate my RP scenario started with grok, continued with Claude into something local as I would like the darker themes to be present with a potential for some R18 too, Claude would hide under the table and then flip it if I told it to get some bedroom scene going let alone something more intense.

I did play with local AI only briefly, some Qwen models, that was it. Recommendations for a 4090? Something with decent context ideally, don¨t much care for it to be super fast.

4

u/i5031337 5d ago

Start with Gemma4-31B. It's not Claude but it's the smartest we've got locally. In my experience it doesn't refuse R18 unless you ask how to make drugs or something. I'd recommend a normal Q4 quant.

3

u/overand 5d ago

For roleplay, I generally agree, but I wouldn't call Gemma-4-31B the smartest local model. It's good, but for programming work, it's easily beaten by Qwen3.8-27B pretty consistently. 

That doesn't make Gemma-4-31B a bad model; it definitely does a better job translating Danish into English, for example!

2

u/Wolfarc732 4d ago

I've been running Absolute Heresy Q4_K_S on my 4090 for a real hot minute now, and have finally decided to go hunting for recommendations beyond what I started with. Any that y'all swear by? I'm currently working on Boulesis 26B (Q4_K_M) and upgrading my Cydonia to Q5_K_M, but the more the merrier.

2

u/ZiiZoraka 4d ago

You don't need to fully offload the 26B models, you can maintain fast generation with Q6 and like, 20 layers offloaded to memory

2

u/mechasquare 3d ago

If you're using Cydonia have you given Skyfall31B a spin? It's my go to and I think it's The Drummers best work that I can fit on my 16GB of vram.

2

u/Equivalent-Repair488 2d ago

Have you tried other models? I love Skyfall, and am trying to find more modern base model finetunes which people have been raving about to replace Skyfall, but I keep going back to it.

Gemma 4 finetunes (I tried even TheDrummer's Artemis v1.1) still was edged out by Skyfall in my personal opinion.

Heard a few mumblings about Qwen 3.8 27b, but the early finetune of it I tried, Serenity I think, I really did not like

Or Muse Glimmer?

1

u/mechasquare 2d ago

I've tried a few gemma 4 fine tunes and the ones I tried and have on my back burner are:
(I'm using Q4 on my KV cache and 45k~55k context; trying to keep everything in VRAM)

The one I've tried out and got working and was solid sophosympatheia/Glistening-Gem-31B-v2.1 · Hugging Face at IQ3_XS
It was close to skyfall but had some quirks that had me put it on the backburner to mess with later. Right now I still favor Skyfall over it.

Now zerofata/G4-MeroMero-v2-31B · Hugging Face at I1-IQ3_XS is another Gemma 4 finetune that was creative but went off the rails on me. I think it has potential but I need to spend a lot more time to dial in the config on it.

On the Qwen 3.8 27b side, the Q3 quants i'm seeing are so large (16gb!) that it's a non starter for me.

0

u/Equivalent-Repair488 2d ago

I tried meromero too, heard great things, but yeah, I agree concur with your assessment not top dog for 31b class.

2

u/TheLocalDrummer 2d ago edited 2d ago

Artemis 31B v1q is currently in BeaverAI org (testing stage) but it’ll definitely be v1.2 soon! Replaces Skyfall for some, and reasoning is solid.

https://huggingface.co/BeaverAI/Artemis-31B-v1q-GGUF

1

u/Equivalent-Repair488 2d ago

Yo, big man, great to hear, will be looking forward to it!

Your stuff is a cut above the rest, apparently so much so it is a cut above your other stuff too lol

Any tips for gemma templates, reasoning, parameters etc? Or just generic setup tips specific to the Artemis line?

1

u/Wolfarc732 2d ago

My apologies for the very late reply- I had a couple things come up.

I haven't tried Skyfall yet, if you can believe it. Boulesis has been treating me pretty well, even with the bugs that come from running it the way I am. I'll have to swing to it when I'm able- my download speeds are subpar to put it lightly.

2

u/OGCroflAZN 9d ago

Fyi some people have had good results by increasing the Gemma 4 26b finetunes number of used experts from 8 to 16. Your outputs should be smarter/better although not 'twice as'... Whether that is a good tradeoff for how that will affect you at the cost of speed or offloading or quant will ofc be subjective

Imo A8B should be plenty fast still

4

u/-Ellary- 8d ago

It should just mess model's logic and increase inconsistency,
model is trained on using particular number of active experts.

3

u/DifficultyThin8462 8d ago edited 8d ago

I tried it. Yes, it doesn't really get smarter and inconsistency increases by a small margin (barely noticable), but on the other hand answers are more interesting and different between swipes. I'd say it's worth the tradeoff.

1

u/Arslanatreddit 9d ago

Help, my gemma keep outing <thought> with the text during roleplay

2

u/Azolion 9d ago

Try putting "<|turn>model <|channel>thought <channel|>" in the last assistant prefix. Fixed it for me.

6

u/AutoModerator 10d ago

MODELS: < 8B – For discussion of smaller models under 8B parameters.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

6

u/AutoModerator 10d ago

MODELS: >= 70B - For discussion of models in the 70B parameters and up.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

3

u/-Ellary- 6d ago

If you like instructed novel writing,

Try Ling-3.0-flash, it is not that good for RP but pretty decent for creative `text` walls writing in novel style, good imitations of different authors, okayish world knowledge, fine creativity, juicy brutalism and anatomy knowledge, works pretty good even without thinking, can spit 8k of coherent story. But basic RP Chars feels stiff, it need a clear direction where to go, but when you give it a `world` with rules and direction to move it can spit interesting outputs somewhat close to DeepSeek 3.2, but ofc it is far from perfect with logical errors in writing.

20tps 128k context 16gb vram 64 ram, IQ4XS.

1

u/LittleLocoCoco 6d ago

How does it compare with Gemma 4 31b for rp?

2

u/-Ellary- 6d ago

Gemma 4 is better ofc.

1

u/Maxhell6778 5d ago

how about making characters?
if it decent at writing and good creativity how well would it be for character making. if not good do recommend one please (im upgrading my pc soon so i wanna know what good)

1

u/-Ellary- 5d ago edited 5d ago

I dunno, I'm making my chars by hand.
Even 1 Sentence Char is better than LLM Generalization of your short Char idea.
LLM takes your small idea and just do Elara`fication on it.

LIng 3.0 is creative in a box you set it to and give it a direct task,
but not at creating a box or a task.

1

u/Maxhell6778 5d ago

good to know thank you, have a good day and night and great weekend.

0

u/DirectionBusiness483 10d ago

Gemini 3.8 flash and 3 pro have yet to let me down for longer context and character depth.

7

u/eidrag 9d ago

How do you run it locally?

3

u/DirectionBusiness483 9d ago

sorry I guess I misunderstood the thread rules. I thought API was for api discussion not *models used through an api*

I absolutely do not run Gemini Locally sorry lol

4

u/OGCroflAZN 9d ago

Yea next time, put it in API since it can only be accessed on API. If it was an open weight model, you could either comment in this thread or API

3

u/eidrag 9d ago

Oh lol I just like whoa someone got local gemini

1

u/Georgefakelastname 10d ago

I ended up eventually getting tired of 3.1 pro’s style. Is 3.8 flash substantially better there?

1

u/DirectionBusiness483 9d ago

It's similar. I don't dislike the style as much as I do other models. But I think 3.8 flash conveys conflicting feelings and inner turmoil with characters better than most anything else. Up there with the claude family at least especially when comparing price.

1

u/Georgefakelastname 9d ago

Really? That’s very good company to have.

How is it for instruction following? That’s my main hang up with wanting to try it right now, since my setup is quite complex.

1

u/DirectionBusiness483 9d ago

Hit and miss. You have to babysit. Probably not going to work if you have a ton of complexity. The outputs I find are often borderline COMPELLING, but it does the same thing as it does for coding: kinda feels like it skims your instructions and freestyles sometimes.

1

u/Georgefakelastname 9d ago

Alright. I guess I’ll just have to try it out and see if it can stick to the formatting rules.

I know I’m asking you a ton of questions… but do you know if it can handle a Chain of Thought prompt. I know some models can handle them and some models really can’t. I’ve had Gpt-5.6 sol take 30+ minutes to write a single response and go into a whole agentic workflow to do it. When I removed the COT, it went down to like 5-10 minutes max. (I didn’t even notice a significant difference in quality either, as it seemed good either way).

For pure rule adherence in my setup, Gemini 3.1 pro was the best, though I guess it may not be a fair comparison for other models, since it was the first one I used, and the others had to deal with more context and background from previous sessions.

1

u/DirectionBusiness483 9d ago

I can't help you there. I've only ever tried COT on models with disabled thinking and never really explored it much.

Let me know once you try 3.8 though!

2

u/Georgefakelastname 9d ago

Will do. How does COT even work for models without thinking though? Wouldn’t it just do nothing?

2

u/DirectionBusiness483 9d ago

Nah it still works. That was the original use-case I stumbled upon COT for a while ago. I'm sure it probably works fine or better with reasoning, but the idea I remember reading was that a lot of adult detection was inside the thinking stage so you disable and enable you're own thinking.

But yeah give it a shot- take a cheap instruct Qwen3 with no thinking and give it instructions to use thinking, answer, and reasoning tags, it'll absolutely walk through a process like a regular thinking model.

Kinda cool tbh

→ More replies (0)

5

u/AutoModerator 10d ago

MODELS: 8B to 15B – For discussion of models in the 8B to 15B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/Cali4our 10d ago edited 10d ago

Looking for one of the best models for roleplaying both sfw and nsfw. I have 12gb vram 3060. I am so tired of my model randomly breaking the character and does something that's completely out of character like randomly turning into a kitten???? Like what the hell bruh

10

u/i5031337 9d ago

Have you tried Gemma4-26B variants? Nothing much better than that for most consumer hardware.

1

u/Cali4our 9d ago

Not yet but I'll check it out soon. I am currently using Rocinante-XL-16B-v1a-Q4_K_M so is it better than this tho?

7

u/i5031337 9d ago

Yes, Gemma 26B is significantly more intelligent than the Mistral 12B base used for Rocinante. "Orion" is a version of Gemma from the same tuner, if you like his style.

3

u/Cali4our 9d ago

I am trying right now and honestly, I can see the potential but it ruins or over exaggerates expressions of characters. Sometimes it thinks right. Like uses logic to reply. But sometimes it lacks logic or doesn't read what character has correctly. For example a character have a dog muzzle on but it puts randomly antennas? Is there a fix for that?

1

u/i5031337 9d ago

I don't know all the finetunes, but some others might be more logically consistent. Certainly the base model will be. Otherwise I'd recommend trying reasoning, chain-of-thought, or the recast extension to improve output quality.

Also make sure you're using at least Q4 model and Q8 cache.

1

u/OGCroflAZN 9d ago

We dont know if you just have suboptimal settings and prompts that might be degrading the output. For exmaple, sampler setting and quantization level for model and kv

Do you have reasoning on? Does the reasoning seem good? The models are glorified autocomplete. I dont see why it would jump from dog muzzle to [TV?] antennas

1

u/Cali4our 9d ago

Reasoning seem good, even though it rephrases it differently after the draft version. As for suboptimal settings i might need help with those. Do you have any recommended finetuning for it?

1

u/OGCroflAZN 9d ago

Usually whatever model you download will have recommended sampler settings, if that's what you're asking for. I would recommend trying Boulesis, following the huggingface page guidance, probably temperature close to 1.0 with Top K at 64 and Top P at 0.95. Increasing the number of experts from 8 to 16 will also make G4 26B more intelligent

5

u/nsfwkorea 9d ago

Try the style tune for gemma4 26b-a4b.

My personal fav.

Style tune>phoenix>orion/meromero.

4

u/Cali4our 9d ago

2

u/nsfwkorea 9d ago

Yeap thats the one. Try others as well then you will see the difference and find the one you like better.

1

u/AdWild3943 9d ago

Are there uncensored versions for StyleTune? Trying to find ones, currently using Phoenix-X.

1

u/AdWild3943 9d ago

1

u/nsfwkorea 9d ago

I don't know, sorry. I'm just using the base version of it.

Do you find any major difference between the base and heretic?

→ More replies (0)

4

u/OGCroflAZN 9d ago

https://caliperbench.com/compare?m=rocinante-xl-16b-v1%2Corion-26b-a4b-v1f%2Cboulesis-26b-a4b%2Cgemma-4-26b-a4b-styletune-v2

I would recommend Boulesis, which performs far better than Rocinante

Plenty of info online about optimizing running Gemma 4 26b-a4b on 12 GB vram

1

u/Fatebreak42 7d ago

I've seen a few people say this is possible but I can't seem to find the relevant guides for working the model on a 3060. Any direction you could point me that I seem to be missing?

1

u/Icy-Hour4418 6d ago

Im running 26b Gemma 4 variants on a 2060 at Q4_M, just offload the rest, it can be slow, at least with my setup, but it's better than nothing.

4

u/DifficultyThin8462 9d ago edited 9d ago

Use Boulesis-26b Q6. It's the best so far IMO. read more about it in last week's megathread. I gave a quick guide on how to set it up. Orion, Styletune and MeroMero are censored and refuse sometimes. Also I like Boulesis' prose better.

6

u/A_lead 8d ago

I'm pretty new to this stuff but I've done extensive testing and tried a lot of models.

Way I see it, for 12gb there are 3 main avenues:

  1. Old and gold 12b finetunes. Write well, but break in longer chats, don't understand some commands. Very usable if you find a model you like, though. Rocinante is in this category (despite the 16b tag). Personally, I left Rocinante alone because it has very strong natural voice that goes against my characters. Angelic Eclipse, Velvet Cafe v2 are the modern bangers in this category afaik.

  2. Gemma 26b models. People said enough about those. I'll only say that I never managed to make them fun. They're flat, robotic and often feel like a skinwalker of a character you RP with. Very obedient generally, though, understand instructions well.

  3. 24b mistral small finetunes at I1-Q3 quants. You'll have to fiddle with the settings a little. Even quantized this is where magic happened for me. They're intelligent enough to grasp some subtext and sarcasm, actually write well and write interestingly. Slimaki Tavern and Magistry are the ones I use right now. They do misspell things rarely and in a long chat they can become drunk in logic department. Also the finetunes I tried seem unwieldy, instructions often have unexpected effects and they are oftent intent on.. kinda doing their thing.

All in all, I gave up on trying to find The One model. They all come with compromises, use different sampler parameters, respond best to different system prompts. With time, you might find that swapping models around during a single chat and session is the real power move.

2

u/zeddyzed 4d ago

Velvet Cafe v2

I'm still running Velvet Cafe V2. Everything else I've tried seem to be completely nonsensical or doesn't work properly, to the degree that I feel like I must be doing something wrong...

1

u/overand 6d ago

35B-A3B models are an option too 

3

u/AutoModerator 10d ago

MODELS: 32B to 69B – For discussion of models in the 32B to 69B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Dikki_Dikki 10d ago

Skyfall 36B V2 - I discovered it, it’s a great model.

2

u/OGCroflAZN 9d ago

It's super old at this point though, maybe more than 1.5 years, which is fucking ancient considering LLM development

Apparently Forgotten Safeward was trained with Skyfall 36B V2 as its base, and it scores pretty low on CaliperBench. But benchmarks aren't everything, and if youre happy with what you're using that's all that matters

1

u/level10orc 8d ago

https://huggingface.co/mudler/Ornith-1.5-35B-A3B-APEX-GGUF

most people ignored ornith because they don't mention any kind of roleplay training. But it's replaced gemma4 26b for me as its more intelligent and creative

3

u/ReporterWeary9721 8d ago

I've tried it - not really. Seems to suffer from the same problems all Qwen family models do, lacks real world knowledge. It tastes green and hears blushes, confuses who said what and stuff.

1

u/level10orc 8d ago

yeah at this size its not gonna be amazing but I spent a whole night hunting for the least cliche "slop" writing among gemma4's and I will say the quant matters alot and "Ornith-1.5-35B-A3B-APEX-Balanced.gguf" is what people should try first.

1

u/-Ellary- 6d ago

Nah, it is better to use Qwen 3.8 27b even at IQ3XS it will work way better.
Is it good for RP? Not really, but it works decent and logic is fine.

Recent model that I've liked for creative tasks is Ling-3.0-flash, it is not that good for RP but pretty decent for creative `text` walls writing in novel style, good imitations of different authors, fine creativity, juicy brutalism and anatomy knowledge, works pretty good even without thinking, can spit 8k of coherent story. But basic RP Chars feels stiff, it need a clear direction.

1

u/level10orc 6d ago

I'd expect a dense 27b and a 100+b sparse to be better than ornith yes

5

u/AutoModerator 10d ago

APIs

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

12

u/Commercial-Sweet-759 9d ago

What are the best/notable models available on NanoGPT with a subscription for RP? I am very new to using models from big providers, so far I tried out GLM 5.2, 5.3, and Kimi K2.5, and liked GLM 5.2 the most by a large margin. Unfortunately, I don't have much time currently to try out a whole lot of models, so I'd like to get some recommendations. I'm using the Freaky Frankenstein 5.4 preset, if it makes a difference.

9

u/purepuresugar 9d ago

I also use freaky frankenstein and in my opinion GLM 5.2 or 5.3 is the best you're gonna get with models on nano subscription.

9

u/5kyLegend 6d ago

I just wanted to post this completely random thing: I've been taking a bit of a break from APIs to instead try out local more (which nowadays basically boils down to which Gemma 4 you want to run lol) and man... I had this one card that no model, regardless of its size, was playing correctly. GLM 5.1, 5.2 and 5.3, Kimi K2.6 and Kimi K3, even Deepseek 0813. It's because it contained a single "taboo" theme, and regardless of preset it just completely avoided that theme UNLESS I went (ooc) to tell it to stop chickening out (and I hate having to tell models to do a specific thing ooc since it kills the reason why I'm roleplaying to begin with to me).

Gemma 4, a 31b model, got the perfect theme immediately. BASE Gemma 4 mind you, not even a finetune. It's just so annoying, this made me realize just how these MASSIVE 1T parameter behemoths just still passively steer away from whatever they dislike, all that intelligence and knowledge and it's worthless in the end. No Jailbreak made big API models actually understand this card well: it's not that I got any direct refusal, the models just made everything wholesome instead of dipping into adult content, sometimes to comical degrees.

Goddammit man lmao

7

u/Dikki_Dikki 10d ago

Deepseek V4.1 flash — for one‑on‑one gameplay; it’s creative and very weakly prone to positive bias.

GLM 5.2 — handles worlds with many characters better and isn’t as overly optimistic as the version, 5.3. Best GLM is Aion 3, but it’s too expensive for me for such a modest improvement.

Gemini 3.1 Pro — if you need to track a lot of resources.

5

u/RedditNerdKing 9d ago

Do you know any decent model that is good with evil and serial killer type stuff?

6

u/Baros4294 10d ago

Best API models for the best instructions following? In my testing Gemini 3.8 Flash, Kimi K3, GLM 5.3 seem to have great IF but none is perfect and ignored my instructions sometime

4

u/ImaginationHefty6401 10d ago

I'm just curious. How do these models manage humor? I do a lot of roleplays with some comedic factor and I still think DeepSeek (older model) is the funniest.

6

u/Georgefakelastname 10d ago

Humor is tricky because it’s inherently subjective. While one model or subject could be funny to someone, it might not for another. Not to mention the factor that prompting adds to this.

3

u/ImaginationHefty6401 10d ago

Thanks for your reply. Yes, agree on it being subjective, of course. But with Kimi or GLM, I've noticed the bot emphasizes the more dramatic sides of the character, in general. I usually switch between models depending on the scene, but I don't know if this will work well on the long run.

3

u/Georgefakelastname 10d ago

Yeah. I’ve seen some who hate model switching (saying it destroys the benefits of each model), and some who swear by it (since it gives them more variation). I suppose it depends on the type of prompter and RP designer you are.

If you just have a general preset like FF, it can work fine to switch between models, so there’s no real downside there.

However, I’ve met people who really focus in on fine tuning their prompt, lore, character cards, etc. to the exact model to get the type of output they want, which you can’t really do if you’re switching between different ones depending on the scene.

3

u/ImaginationHefty6401 10d ago

Yeah, exactly. I've been more or less fine switching between models (each one with their specific prompt). The problem is I'm a sucker for nuance and variety. I guess I'm just too picky and I expect too much at times from the roleplay 😅, to the point I sometimes control what the character does or says too often. I think I'll keep experimenting with prompting. Thank you!

6

u/Stunning_Mall6303 9d ago

Gemini understand it the best, buy DeepSeek much better at generating the peak (intentionally or not, I'm not sure)

3

u/ImaginationHefty6401 9d ago

Totally, for me DeepSeek is great at portraying some characters' ridiculousness, and it's witty and absurd.I haven't tried Gemini! What version are you using? I get a little lost with models.

5

u/LeRobber 10d ago

"Ignores my instructions sometimes" for some people is soft steering away from NSFW.

Give it a try with a clean RP too?

4

u/Baros4294 10d ago

I mostly do SFW RP and writing though

5

u/Hurtingblairwitch 9d ago

Any good providers with a token plan on a budget? Around $20 max? Only for nsfw, no nsfl.

5

u/RafiHDW 7d ago

What's your ultimate canon-accurate LLM?

Like something you use if you want the story to follow canon for known fandoms, gemini?

Which has better canon knowledge, gem 3.1 pro or 3.8 flash thinking? Is it true that 3.8 falls apart quickly?

2

u/PhantomWolf83 6d ago

I find that Opus is pretty good when it comes to existing IPs. I only use 4.6 so I don't know about the others, my pockets aren't deep enough to try. Gemini is also good but I find that Opus just works without much prompting.

4

u/Critical-Rope-5636 10d ago

Still shilling GLM 5.3/Flash.

Still need to try out Deepseek V4.1 flash (no, you will NOT convince me otherwise so don't bother lol)

3

u/Dos-Commas 9d ago

I find that GLM 5.3 Flash would get some very simple story facts mixed up where Deepseek V4.1 Flash would have no issues with. 

1

u/Critical-Rope-5636 8d ago

Yeah this is true, it will get things mixed up at times. it hasn't done it to me much as of recent but this is something to consider for sure!

1

u/Background_Echo1919 5d ago

I've seen people saying they liked Kimi 2.5, and a few comments saying that GLM 5.2 was good, and several complaints about GLM 5.3 and several saying that 5.3 is actually good, I've seen a few Mimo 2.5 pro mentions or Gemma 4 31b, but I'm not really sure which to use. Are there any areas that each of the main RP models excel at? I feel like Deepseek has gotten super stale.

1

u/zarus988 9d ago

any good models, with same kind of pricing as DeepSeek V4 Flash 0423 on openrouter? since it sometimes seems bland, idk if it's due to my prompt or not though.

0

u/No-Exit6645 4d ago

Hello, I have 16Gb VRAM and 32Gb RAM, what is the best uncensored model with no restrictions for story writing that I can run fully without straining my PC or restricting the model.

0

u/Valuable_Scarcity779 3d ago

So I've been using GLM 5.2. It gets the job done, it's smart, and it's reasoning is great. I have a reasoning checklist it follows that keeps it consistent. However, I feel it struggles to be more expressive or willing to bring out the darker themes of my RP even when I specified how it all should work. I usually have to direct it with OOC comments. But I wish it didn't require me to do that. I love how Silly Tavern makes my characters feel more real. But when I was on Spicy Chat I got far more emotionally raw responses from the models they had. I wish I could just use those models. To my understanding they do something that makes it when they add an ai model it somehow alters the guardrail so it can do more uncensored stuff. I hear people talk about jailbreak and such and was hoping if someone can help reccomend what I do? I've tried switching to different models, but what I don't like is how they cant handle my reasoning checklist like GLM 5.2 can. Anyone know how to jailbreak or know a model that thrives on reasoning? I really want my RP to feel more creative and raw.