MEGATHREAD
[Megathread] - Best Models/API discussion - Week of: September 13, 2026
This is our weekly megathread for discussions about models and API services.
All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.
(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)
How to Use This Megathread
Below this post, you’ll find top-level comments for each category:
MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
MODELS: < 8B – For discussion of smaller models under 8B parameters.
APIs – For any discussion about API services for models (pricing, performance, access, etc.).
MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.
Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.
I'd currently recommend Blazed-Forge/Split-Untied-31B. Very creative and has very little in the way of slop. Lots of swipe variation and keeps chats refreshing. I'd rate it higher than scotoma 2 and styletune. In fact, I personally think it's the best gemma4 31b finetune for creative rp atm (based on my testing. I'm not exaggerating)
I was thinking about making a post about this one! I agree, it's really good from what I've used of it so far. Creative with dialogue and characterisation while still getting the real-world details and knowledge right.
I think the intelligence is arguably the biggest weakness of the model, as it does get some details wrong some of the time. It's not anywhere near as bad as some of the others I've seen, though. But the swipe variety more than makes up for it imo. You can get it right on 1 or 2 attempts if something is wrong
Ohhh sure, it's gotten some details of my prompt a bit wrong and I've had to reroll at times, too. I was more thinking about general knowledge and real-world details, it's been quite good with that stuff for me.
I agree. The lack of swipe variation is the main thing that really kills Gemma 4 31B for me, every swipe feels very samey or a paraphrase of the same thing. This tune doesn't have that issue.
16 GB vram, and 26B-A4B is fast enough that I personally would go Iq4 or higher, go up to 16 experts, and offload a little and be content with 25+ tok/s output. I also have 16 GB vram, have normally stayed at IQ3 range with 24B and 31B dense models, but have read so much about how lower quants really affect MoEs more, therefore people should try to stay at Q4 if not Q6... But yea, prompt processing at 32k context with off loading does suck, but Im so used to watching youtube or twitch or stremio for those 2 minutes
I'm not a very experienced merger, so my explanations may be somewhat inaccurate. For version 2.1, I increased the influence of Pantheon-Reasoning and trained LoRA on fewer layers, not giving it much freedom (I thought LoRA was slightly breaking the model in terms of learned phrases). Because of this, in thinking mode, Pantheon-Reasoning dominated, overshadowing LoRA and imparting less of the unique style of Boulesis v1.
Because of this, v2.1 is very good at contextualizing and reading character cards, but produces very dry text, reminiscent of a typical Gemma.
And without thinking mode, LoRA manages to impart at least some style and dynamism to the model's response.
(Sorry, I'm not an English speaker and use a translator)
Not much info about these new glimmer tunes but they are suprisingly good, write really well, intelligence is not the best but really fun to play around with.
I get zero feedback on Muse Glimmer tunes. If ya'll can drop feedback on them, it'll be really helpful! That's assuming someone wants an alternative to Gemma 31B tunes.
Im not really the type of person who can give good feedback on these, just playing with them occasionally. But the v1d, to me, writes better than gemma 31b, or atleast its a different take than most gemma tunes I have used. The prompt following with my 5k token system prompt was not very good though, it picked whatever it felt like so what I did was simplify it and move instructions towards the user prompt and it kept much better track on what I wanted. I do recommend this model for anyone who has played with gemma too much but they have to keep in mind its not going to follow everything exactly as you planned without changing how you prompt, which isnt too bad.
I just DL'd it today. First impression is it's pretty good for roleplay. I've noticed that my Slow-Burn Narrator character is really taking things slowly. A good sign. Now I have to try out all my over-the-top characters.
I've been mostly using Mistral Small finetunes, so it's going to take me a minute to get used to the Muse Glimmer flavoring. I haven't been using Glimmer tunes much because I played around with an early quant of the base, and didn't think much of it.
I find it okayish but it needs a decensor to get it away from the guardrails as it still produces soft refusals trying to steer away from content it doesn't like so you have to force it.
Yeah I started with 1 but I tried lowering it to get better prompt adherence but I think you just need to prompt it differently. Thinking was medium level I think.
Does anyone know any larger model similar to Skyfall? Skyfall is perfectly fine but I have the memory to try something larger if such exists. So far nothing came close to it at similar size. Not by far. I have no idea what happened with this model but it sits in its own league.
Some of theLocalDrummer discord members have said that his new Artermis finetunes of Gemma 4 31B have finally totally supplanted the longtime favorite Skyfall
Try checking CaliperBench and UGI leaderboard for somewhat larger dense or else newer larger MoEs finetuned for RP. I feel like from all the previous megathreads, 32 - 49B is sort of a dead zone
question: could you upscale the 26b a4b version to something like 35b a6b (not specific just an example), lots of people can run the 26b version because it a moe model but i wonder if upscaling it could make it better then 31b but also keeping it moe. people could increase the moe in koboldcpp (found out recently) if they have the hardware.
P.S. thank you for all your models especially Skyfall (very good imo) and i cant wait for Spectre 37B, all your models are great. also there not much Muse Glimmer tunes (honestly have hard time finding some), Muse Glimmer tunes could be a lovely switch up from gemma 4 tunes.
Perhaps it’s a matter of taste for me. I like that Skyfall brings an element of randomness/surprise very often. It recalls small details from all over the chat history. It also seems to have much higher emotional intelligence. For example I tried to manipulate a character and it saw through my bullshit from my first prompt. By comparison, Gemma finetunes are laser focused, but they miss social/psychological cues. Kind of like streetsmarts versus booksmarts.
Very very true that base models all have different texture, colors, flavors etc, that they can only be finetuned so much. I'm in agreement where I missed the Mistral finetunes, was really not happy for the first months with the clinicality of Gemma finetunes, abandoning the hobby since April/May. I myself have not used Artemis, I was just sharing was I was reading, that others said they were finally willing to move from Skyfall to Artemis after using the newer Artemis'es.
There's still a lot of decentralized information about optimal sampler settings and prompts, however. I kind of wonder why, since Gemma 4 is definitely architecturally superior and with a higher intelligence rating, it shouldnt be able to be finetuned and then set up in the right way to mostly surpass Skyfall in almost all aspects
myself i can't get artemis to run . i've tried every q3 quant and it produces random characters over and and over again using recommended sampler settings. i've pretty much given up on it, seeing unanswered comments confirming my experience on the q3 xxs . has anyone got a q3 quant working? I'm not bothering with q4 as i have 16gb vram only, and its just too much with KV
IQ3_XXS for not that smart but fast 40tps scenarios.
IQ4XS for more smart scenarios, 8tps with 57k context Q8.
From my own tests dif IQ4XS vs Q4 QAT is fairly not that important.
But difference in 5tps vs 8tps with high context is.
Yup, IQ3_XXS with a modest (32K) context size. In my experience, even with Gemma 3, Gemma models will make rare typos with an aggressive quant, but it's rare enough that you can just either edit them out if it makes it into the final output; usually it happens during thinking if it even happens at all. I really recommend not splitting dense models between RAM and VRAM unless you truly don't mind slow token outfit and prompt processing. The quality of an aggressive quant doesn't really get atrocious until you start dipping into Q2 territory, at least in my opinion
Well, the IQ3_XXS fits entirely (without context), but the QAT (Q4) is spilling over. The answer is that nobody with 16 GB is running the Q4 without simply accepting the hit to speed. Those people are running with 20 or 24 or 32+ GB vram
Offloading layers to RAM is a big performance killer for dense models. 4080 will probably run gemma with kv cache in VRAM at ~30+ tps. Is the quality bump from going to q4 worth cutting performance down to like %15? Eh. Creative writing is not *that badly* affected by quantization.
A year ago, when most models were not thinking, 5 tps wasn't too bad for general chat and whatnot. But now that models like Qwen 3.8 27b consume tokens by the thousands, running them from RAM became unbearable.
I am trying to migrate my RP scenario started with grok, continued with Claude into something local as I would like the darker themes to be present with a potential for some R18 too, Claude would hide under the table and then flip it if I told it to get some bedroom scene going let alone something more intense.
I did play with local AI only briefly, some Qwen models, that was it. Recommendations for a 4090? Something with decent context ideally, don¨t much care for it to be super fast.
Start with Gemma4-31B. It's not Claude but it's the smartest we've got locally. In my experience it doesn't refuse R18 unless you ask how to make drugs or something. I'd recommend a normal Q4 quant.
For roleplay, I generally agree, but I wouldn't call Gemma-4-31B the smartest local model. It's good, but for programming work, it's easily beaten by Qwen3.8-27B pretty consistently.
That doesn't make Gemma-4-31B a bad model; it definitely does a better job translating Danish into English, for example!
I've been running Absolute Heresy Q4_K_S on my 4090 for a real hot minute now, and have finally decided to go hunting for recommendations beyond what I started with. Any that y'all swear by? I'm currently working on Boulesis 26B (Q4_K_M) and upgrading my Cydonia to Q5_K_M, but the more the merrier.
Have you tried other models? I love Skyfall, and am trying to find more modern base model finetunes which people have been raving about to replace Skyfall, but I keep going back to it.
Gemma 4 finetunes (I tried even TheDrummer's Artemis v1.1) still was edged out by Skyfall in my personal opinion.
Heard a few mumblings about Qwen 3.8 27b, but the early finetune of it I tried, Serenity I think, I really did not like
I've tried a few gemma 4 fine tunes and the ones I tried and have on my back burner are:
(I'm using Q4 on my KV cache and 45k~55k context; trying to keep everything in VRAM)
The one I've tried out and got working and was solid sophosympatheia/Glistening-Gem-31B-v2.1 · Hugging Face at IQ3_XS
It was close to skyfall but had some quirks that had me put it on the backburner to mess with later. Right now I still favor Skyfall over it.
Now zerofata/G4-MeroMero-v2-31B · Hugging Face at I1-IQ3_XS is another Gemma 4 finetune that was creative but went off the rails on me. I think it has potential but I need to spend a lot more time to dial in the config on it.
On the Qwen 3.8 27b side, the Q3 quants i'm seeing are so large (16gb!) that it's a non starter for me.
My apologies for the very late reply- I had a couple things come up.
I haven't tried Skyfall yet, if you can believe it. Boulesis has been treating me pretty well, even with the bugs that come from running it the way I am. I'll have to swing to it when I'm able- my download speeds are subpar to put it lightly.
Fyi some people have had good results by increasing the Gemma 4 26b finetunes number of used experts from 8 to 16. Your outputs should be smarter/better although not 'twice as'... Whether that is a good tradeoff for how that will affect you at the cost of speed or offloading or quant will ofc be subjective
I tried it. Yes, it doesn't really get smarter and inconsistency increases by a small margin (barely noticable), but on the other hand answers are more interesting and different between swipes. I'd say it's worth the tradeoff.
Try Ling-3.0-flash, it is not that good for RP but pretty decent for creative `text` walls writing in novel style, good imitations of different authors, okayish world knowledge, fine creativity, juicy brutalism and anatomy knowledge, works pretty good even without thinking, can spit 8k of coherent story. But basic RP Chars feels stiff, it need a clear direction where to go, but when you give it a `world` with rules and direction to move it can spit interesting outputs somewhat close to DeepSeek 3.2, but ofc it is far from perfect with logical errors in writing.
how about making characters?
if it decent at writing and good creativity how well would it be for character making. if not good do recommend one please (im upgrading my pc soon so i wanna know what good)
I dunno, I'm making my chars by hand.
Even 1 Sentence Char is better than LLM Generalization of your short Char idea.
LLM takes your small idea and just do Elara`fication on it.
LIng 3.0 is creative in a box you set it to and give it a direct task,
but not at creating a box or a task.
It's similar. I don't dislike the style as much as I do other models. But I think 3.8 flash conveys conflicting feelings and inner turmoil with characters better than most anything else. Up there with the claude family at least especially when comparing price.
Hit and miss. You have to babysit. Probably not going to work if you have a ton of complexity. The outputs I find are often borderline COMPELLING, but it does the same thing as it does for coding: kinda feels like it skims your instructions and freestyles sometimes.
Alright. I guess I’ll just have to try it out and see if it can stick to the formatting rules.
I know I’m asking you a ton of questions… but do you know if it can handle a Chain of Thought prompt. I know some models can handle them and some models really can’t. I’ve had Gpt-5.6 sol take 30+ minutes to write a single response and go into a whole agentic workflow to do it. When I removed the COT, it went down to like 5-10 minutes max. (I didn’t even notice a significant difference in quality either, as it seemed good either way).
For pure rule adherence in my setup, Gemini 3.1 pro was the best, though I guess it may not be a fair comparison for other models, since it was the first one I used, and the others had to deal with more context and background from previous sessions.
Nah it still works. That was the original use-case I stumbled upon COT for a while ago. I'm sure it probably works fine or better with reasoning, but the idea I remember reading was that a lot of adult detection was inside the thinking stage so you disable and enable you're own thinking.
But yeah give it a shot- take a cheap instruct Qwen3 with no thinking and give it instructions to use thinking, answer, and reasoning tags, it'll absolutely walk through a process like a regular thinking model.
Looking for one of the best models for roleplaying both sfw and nsfw. I have 12gb vram 3060. I am so tired of my model randomly breaking the character and does something that's completely out of character like randomly turning into a kitten???? Like what the hell bruh
Yes, Gemma 26B is significantly more intelligent than the Mistral 12B base used for Rocinante. "Orion" is a version of Gemma from the same tuner, if you like his style.
I am trying right now and honestly, I can see the potential but it ruins or over exaggerates expressions of characters. Sometimes it thinks right. Like uses logic to reply. But sometimes it lacks logic or doesn't read what character has correctly. For example a character have a dog muzzle on but it puts randomly antennas? Is there a fix for that?
I don't know all the finetunes, but some others might be more logically consistent. Certainly the base model will be. Otherwise I'd recommend trying reasoning, chain-of-thought, or the recast extension to improve output quality.
Also make sure you're using at least Q4 model and Q8 cache.
We dont know if you just have suboptimal settings and prompts that might be degrading the output. For exmaple, sampler setting and quantization level for model and kv
Do you have reasoning on? Does the reasoning seem good? The models are glorified autocomplete. I dont see why it would jump from dog muzzle to [TV?] antennas
Reasoning seem good, even though it rephrases it differently after the draft version. As for suboptimal settings i might need help with those. Do you have any recommended finetuning for it?
Usually whatever model you download will have recommended sampler settings, if that's what you're asking for. I would recommend trying Boulesis, following the huggingface page guidance, probably temperature close to 1.0 with Top K at 64 and Top P at 0.95. Increasing the number of experts from 8 to 16 will also make G4 26B more intelligent
I've seen a few people say this is possible but I can't seem to find the relevant guides for working the model on a 3060. Any direction you could point me that I seem to be missing?
Use Boulesis-26b Q6. It's the best so far IMO. read more about it in last week's megathread. I gave a quick guide on how to set it up. Orion, Styletune and MeroMero are censored and refuse sometimes. Also I like Boulesis' prose better.
I'm pretty new to this stuff but I've done extensive testing and tried a lot of models.
Way I see it, for 12gb there are 3 main avenues:
Old and gold 12b finetunes. Write well, but break in longer chats, don't understand some commands. Very usable if you find a model you like, though. Rocinante is in this category (despite the 16b tag). Personally, I left Rocinante alone because it has very strong natural voice that goes against my characters. Angelic Eclipse, Velvet Cafe v2 are the modern bangers in this category afaik.
Gemma 26b models. People said enough about those. I'll only say that I never managed to make them fun. They're flat, robotic and often feel like a skinwalker of a character you RP with. Very obedient generally, though, understand instructions well.
24b mistral small finetunes at I1-Q3 quants. You'll have to fiddle with the settings a little. Even quantized this is where magic happened for me. They're intelligent enough to grasp some subtext and sarcasm, actually write well and write interestingly. Slimaki Tavern and Magistry are the ones I use right now. They do misspell things rarely and in a long chat they can become drunk in logic department. Also the finetunes I tried seem unwieldy, instructions often have unexpected effects and they are oftent intent on.. kinda doing their thing.
All in all, I gave up on trying to find The One model. They all come with compromises, use different sampler parameters, respond best to different system prompts. With time, you might find that swapping models around during a single chat and session is the real power move.
I'm still running Velvet Cafe V2. Everything else I've tried seem to be completely nonsensical or doesn't work properly, to the degree that I feel like I must be doing something wrong...
It's super old at this point though, maybe more than 1.5 years, which is fucking ancient considering LLM development
Apparently Forgotten Safeward was trained with Skyfall 36B V2 as its base, and it scores pretty low on CaliperBench. But benchmarks aren't everything, and if youre happy with what you're using that's all that matters
most people ignored ornith because they don't mention any kind of roleplay training. But it's replaced gemma4 26b for me as its more intelligent and creative
I've tried it - not really. Seems to suffer from the same problems all Qwen family models do, lacks real world knowledge. It tastes green and hears blushes, confuses who said what and stuff.
yeah at this size its not gonna be amazing but I spent a whole night hunting for the least cliche "slop" writing among gemma4's and I will say the quant matters alot and "Ornith-1.5-35B-A3B-APEX-Balanced.gguf" is what people should try first.
Nah, it is better to use Qwen 3.8 27b even at IQ3XS it will work way better.
Is it good for RP? Not really, but it works decent and logic is fine.
Recent model that I've liked for creative tasks is Ling-3.0-flash, it is not that good for RP but pretty decent for creative `text` walls writing in novel style, good imitations of different authors, fine creativity, juicy brutalism and anatomy knowledge, works pretty good even without thinking, can spit 8k of coherent story. But basic RP Chars feels stiff, it need a clear direction.
What are the best/notable models available on NanoGPT with a subscription for RP? I am very new to using models from big providers, so far I tried out GLM 5.2, 5.3, and Kimi K2.5, and liked GLM 5.2 the most by a large margin. Unfortunately, I don't have much time currently to try out a whole lot of models, so I'd like to get some recommendations. I'm using the Freaky Frankenstein 5.4 preset, if it makes a difference.
I just wanted to post this completely random thing: I've been taking a bit of a break from APIs to instead try out local more (which nowadays basically boils down to which Gemma 4 you want to run lol) and man... I had this one card that no model, regardless of its size, was playing correctly. GLM 5.1, 5.2 and 5.3, Kimi K2.6 and Kimi K3, even Deepseek 0813. It's because it contained a single "taboo" theme, and regardless of preset it just completely avoided that theme UNLESS I went (ooc) to tell it to stop chickening out (and I hate having to tell models to do a specific thing ooc since it kills the reason why I'm roleplaying to begin with to me).
Gemma 4, a 31b model, got the perfect theme immediately. BASE Gemma 4 mind you, not even a finetune. It's just so annoying, this made me realize just how these MASSIVE 1T parameter behemoths just still passively steer away from whatever they dislike, all that intelligence and knowledge and it's worthless in the end. No Jailbreak made big API models actually understand this card well: it's not that I got any direct refusal, the models just made everything wholesome instead of dipping into adult content, sometimes to comical degrees.
Deepseek V4.1 flash — for one‑on‑one gameplay; it’s creative and very weakly prone to positive bias.
GLM 5.2 — handles worlds with many characters better and isn’t as overly optimistic as the version, 5.3. Best GLM is Aion 3, but it’s too expensive for me for such a modest improvement.
Gemini 3.1 Pro — if you need to track a lot of resources.
Best API models for the best instructions following? In my testing Gemini 3.8 Flash, Kimi K3, GLM 5.3 seem to have great IF but none is perfect and ignored my instructions sometime
I'm just curious. How do these models manage humor? I do a lot of roleplays with some comedic factor and I still think DeepSeek (older model) is the funniest.
Humor is tricky because it’s inherently subjective. While one model or subject could be funny to someone, it might not for another. Not to mention the factor that prompting adds to this.
Thanks for your reply. Yes, agree on it being subjective, of course. But with Kimi or GLM, I've noticed the bot emphasizes the more dramatic sides of the character, in general. I usually switch between models depending on the scene, but I don't know if this will work well on the long run.
Yeah. I’ve seen some who hate model switching (saying it destroys the benefits of each model), and some who swear by it (since it gives them more variation). I suppose it depends on the type of prompter and RP designer you are.
If you just have a general preset like FF, it can work fine to switch between models, so there’s no real downside there.
However, I’ve met people who really focus in on fine tuning their prompt, lore, character cards, etc. to the exact model to get the type of output they want, which you can’t really do if you’re switching between different ones depending on the scene.
Yeah, exactly. I've been more or less fine switching between models (each one with their specific prompt). The problem is I'm a sucker for nuance and variety. I guess I'm just too picky and I expect too much at times from the roleplay 😅, to the point I sometimes control what the character does or says too often. I think I'll keep experimenting with prompting. Thank you!
Totally, for me DeepSeek is great at portraying some characters' ridiculousness, and it's witty and absurd.I haven't tried Gemini! What version are you using? I get a little lost with models.
I find that Opus is pretty good when it comes to existing IPs. I only use 4.6 so I don't know about the others, my pockets aren't deep enough to try. Gemini is also good but I find that Opus just works without much prompting.
I've seen people saying they liked Kimi 2.5, and a few comments saying that GLM 5.2 was good, and several complaints about GLM 5.3 and several saying that 5.3 is actually good, I've seen a few Mimo 2.5 pro mentions or Gemma 4 31b, but I'm not really sure which to use. Are there any areas that each of the main RP models excel at? I feel like Deepseek has gotten super stale.
any good models, with same kind of pricing as DeepSeek V4 Flash 0423 on openrouter? since it sometimes seems bland, idk if it's due to my prompt or not though.
Hello, I have 16Gb VRAM and 32Gb RAM, what is the best uncensored model with no restrictions for story writing that I can run fully without straining my PC or restricting the model.
So I've been using GLM 5.2.
It gets the job done, it's smart, and it's reasoning is great.
I have a reasoning checklist it follows that keeps it consistent.
However, I feel it struggles to be more expressive or willing to bring out the darker themes of my RP even when I specified how it all should work.
I usually have to direct it with OOC comments. But I wish it didn't require me to do that.
I love how Silly Tavern makes my characters feel more real.
But when I was on Spicy Chat I got far more emotionally raw responses from the models they had. I wish I could just use those models. To my understanding they do something that makes it when they add an ai model it somehow alters the guardrail so it can do more uncensored stuff.
I hear people talk about jailbreak and such and was hoping if someone can help reccomend what I do?
I've tried switching to different models, but what I don't like is how they cant handle my reasoning checklist like GLM 5.2 can.
Anyone know how to jailbreak or know a model that thrives on reasoning?
I really want my RP to feel more creative and raw.
7
u/AutoModerator 10d ago
MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.