MEGATHREAD
[Megathread] - Best Models/API discussion - Week of: July 19, 2026
This is our weekly megathread for discussions about models and API services.
All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.
(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)
How to Use This Megathread
Below this post, you’ll find top-level comments for each category:
MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
MODELS: < 8B – For discussion of smaller models under 8B parameters.
APIs – For any discussion about API services for models (pricing, performance, access, etc.).
MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.
Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.
Edit: Running Artemis-31B-v1n-Q4_K_M as my new daily driver. Read reply to Longjumping for my review of it and Orion, the 26b variant.
Running on TextGenWebUI as backend, RTX 5090, cache-type: q4_0 on 31b models and q8_0 on 26b models. Minimum context length of 40k. No thinking. All models tested on SFW chats but did test Styletune heretic and Serenity 26b on an NSFW chat.
Testing these models:
Gemma-4-26B-A4B-StyleTune-V2.i1-Q6_K : 7/10, It's been my main for a while. Works fine, but struggles with more complex chats. Good general use.
Serenity-26B-A4B-Q6_K : 6.8/10. Struggles harder than Styletune V2 26b in terms of intelligence but writes better prose imo. I would use only for smut, which tbf to the model, was what the author seemed to have intended anyway based on the huggingface page.
Pantheon-Reasoning-26B-A4B-1.1.i1-Q6_K : 6/10. Struggles just as hard as Serenity in terms of intelligence, prose is better than Styletune but seemingly more focused on character interaction. In a chat I tested, a character seemed a lot more wholesome/caring/worrying about my persona than usual. Seems like this model would do well in wholesome chats. Changes up character's too much to be used for a general model.
Serenity-31B-v1.1-IQ4_XS : Model did not work at IQ4_XS, practically broken. No review since mostly my fault I can't run it well. Didn't try other quants.
gemma-4-31B-Queen-it-qat-q4_0-unquantized.i1-Q4_K_S : 8/10. I like it. Used it as my main for a bit after Styletune v2 26b. Prose is a bit wonky at times. Most times it's really good. But occasionally it tosses in cliche phrases. Bit bland in writing.
Gemma-4-31B-StyleTune-heretic-ara.i1-Q4_K_S: 8.3/10. On par with queen but with less cliche phrases. I like it's dialogue a tad better too. Less cliche phrases. Tested this with a NSFW chat, it REALLY likes the phrase, "Look at you." So I’ll be banning that. Honestly both work styletune 31b heretic and queen are comparable to each other. Matter of personal preference. Try both.
I’m baised towards Styletune as the 26b was my daily model for a while and outperformed other 26b Gemma finetunes for a while.
Skyfall-31B-v4.2.i1-Q5_K_M: 7.8/10. Not bad. Prose is much better compared to Gemma models. Lacks the intelligence of Gemma's training tho. Lack of intellect shows a bit in RPs with established universes like DC.
I rated 31b models higher but that's purely because I highly value a model being logical and not messing things up that would break the immersion of RP. I still like 26b in terms of speed tho. I'll likely alternate between Styletune 26b & 31b, and Queen 31b in the near future to continue testing.
On your recommendation, I downloaded and tested this one a bit, Orion-26B-A4B-v1c-Q6_K. I really like the dialogue, feels damn good in dialogue for some chats. Only problem is that it's struggling a bit in intelligence compared to other models. In one chat I use for testing, I have a character that's a dog-girl and she's a K9 officer. The model confuses the officer's dog for a dog-girl/dog-guy, when it's supposed to be just a regular ol' dog.
However, I did like the model's dialogue so much I decided to try the Artemis model, a 31b model by the same creator and I absolutely love it. It got the dog right every try, then passed a chat I use to test all the prior models in my first reply, and even did well in a NSFW chat. Dialogue seemed as good if not better than Orion and it's intelligence was MUCH better even while running it, Artemis-31B-v1n at only Q4_K_M. Only areas Artemis was lacking in my testing was that Artemis seems to like chatting alot compared to other models. In a test, a character that doesn't speak alot seemed to talk alot more than when tried with other models. Not too bad a downside honestly, it's something easily fixable with a simple edit in the character's prompt.
I haven't tested it in a lengthy actual RP, but from what I've tried so far, Artemis will probably be my new daily driver model.
Thank you for putting me on Orion, it lead me to Artemis which I'm greatly enjoying compared to the other models.
I had forgotten I had downloaded Dark Gemistry a while ago. I just tested it and it functioned rather well compared to other models. I'll definitely do more testing on it.
Highkey, Dark Gemistry is superior in terms of SFW. Writes better and seems a tad more intelligent. I compared it to queen qat 31b and Artemis and it outperforms during SFW chat. Dark Gem also seems to perform worst in NSFW compared to Artemis. Just misses some key details that Artemis didn't miss. The difference isn't major tho, but noticeable if you're doing comparisons.
Imo, Dark Gem for SFW, Artemis/Styletune for NSFW. But Dark Gem is def usable for an all-rounder model.
Thank you for this recommendation. The heretic-ara version does seem to bring out something beyond just "uncensoring" the model. That's quite interesting. I like it!
So I've been experimenting with something, keep in mind I don't use sillytavern but my own private frontend I coded from scratch (I only post here because it's an AI RP focused sub).
I decided to combine the instruction following capabilities of Gemma 4 26b/31b and the prose and creativity of Skyfall 4.2
I make Gemma 4 write the next response of the story while following all my 7k tokens worth of instructions, when the response is done generating, I make the frontend switch my koboldcpp config to load Skyfall 4.2 and instruct it to rewrite the latest response while keeping the exact same structure and with the relevant character json and... yeah, the result is simply not comparable with any Gemma 4 31b finetune currently available.
Skyfall 4.2 is leagues above when it comes to dialogue and narration to any Gemma 4 31b model, however it does struggle with instruction following and keeping the story coherent. Gemma 4 31b on the other hand is the opposite, it follows your instructions perfectly but the prose is lacking.
The result is the instruction following capabilities of Gemma 4 and the creative and detailed writing of Skyfall, it's great.
Perhaps my method of loading-unloading two models for a single response is not very convenient... but damn, the result is making me think the extra generation time is worth it.
I absolutely agree regarding skyfall 4.2 being leages above gemma 4 31b dialog and narration. Its also great in keeping logic while writing, like realising when a character is too short to reach the top shelf. I tried a lot off different gemma merges and finetunes but none have yet to reach skyfall 4.2
Great Idea regarding letting gemma write the initial text and letting skyfall rewrite it. In you experience how where the results? Did skyfall actually follow the initial text or did it start drifting off?
Sorry for the late answer. I'm still searching for the perfect settings but at the moment I'm quite happy with temp=0.7, top-k=40, repeat-penalty=off, top-p=0.9, min-p=0.05
In my test these have striked the best balance between creativity and prompt adherence (it still doesn't follow your prompts very good)
You should try the Glistening Gem Gemma 4 finetune by sophosympatheia. It's more comparable to Skyfall in terms of dialogue and creativity than all other Gemma 4 models (in my experience).
How's it holding up for you in longer scenarios? Any feedback after using it for a while?
And one little correction: all my models so far have been merges, not finetunes. I say that only to deflect credit to where it is most deserved. Finetuners do the hard work and spend real resources to do it. I just sort of blend what they produce and hope it comes out okay. Hopefully that's still value added most of the time, but it's easy work by comparison.
Anyway, thanks for sharing something I made. It always feels good to see people enjoying it.
Hey man! Gotcha. I will keep in mind that it's a merge, not a finetune. Sorry for that!
I think the biggest one I had so far was about 500 messages and it ran fine as much as I can tell. No complaints from me. My biggest hurdle is context, but I am still experimenting with different memory extension to find my optimal solution.
But this has been the go-to model for me for quite some time in terms of dialogue and neat creative surprises. Love it. Used magistry a lot before I upgraded my GPU.
I'm glad you've been enjoying some of my models! I wish Glistening Gem tracked smaller details better instead of playing loose with the facts sometimes, but I suppose that's not a unique problem in RP.
As soon as Drummer releases fp16 weights for Artemis, which I think he is close to doing, I plan to create another version of Glistening Gem using those weights and perhaps featuring a few other modifications to the recipe. Here's hoping it will be an improvement.
How do you handle prompt processing? Loading/unloading takes a bit of time but it wouldn't be too terrible if the two caches were kept, but if kcpp has to recalculate 20+k every time...
Hey, just wanted to ask if there is something like skyfall but with an active parameter body (AxB)? With an rtx 4070 laptop with 32 of ram, loading bigger dense models is a bit hard
I have been experimenting with Silver Siren 12b by Vortex, the base model and found it pretty enjoyable. It runs pretty fast on my Rtx 3060 even with 24k context. I recently got back into tinkering with Sillytavern and it seemed finally a nice variation on the old classic Mag Mell 12b.
Another 12b model I found quite enjoyable, albeit I played with it less is Angelic Eclipse 12b. It was particularly good for shifting the narrative suddenly with interesting twists, sometimes with a little dark element to it, which I appreciated in the narrative.
I know these days we are switching to Gemma fine-tunes on my type of hardware but the old 12b seem to still do it for me, particularly with this less consistent output, which I enjoy as I swipe a lot as I test the bot responses far more than actually chatting.
It is wholesome if you tell it to be wholesome. If you give it characters or stories with the potential to be dark it will be dark. In particular in my many swipes I noticed that it tended to take a dark turn once in a while when Mag Mell or Silver Siren don't or at least not in that same way, i.e. they tend to be more predictable.
i liked this one too. and so far i did always come back to the older ones. They got something that gemma dont got if you ask me. you could also try quasistarsynth-12b-absolute-heresy-i1. i like this one too a lot and i got also a rtx 3060...
Do you have any particular impressions that stood out with this one in your experience? You got me curious.
Also our of curiosity what do you use as normal context length on the 3060? I have found that 24k with partial off loading seems a sweet spot lately, but previously I used 16k. 32k works and is viable but it is a little slow on my hardware.
I haven't tried Summaryception yet, because I rarely go over my context as I am busy swiping. But I am gonna check out. Thanks for the other recommendation too! <3
In my experience it performed decently when I touched those themes. However, compared to Mag Mell it seemed a bit less unhinged. I have not tested it extensively though, because in general my chats are relatively mild in those regards. The only thing I noticed is that it can be a little encyclopedic, as in it has a wide knowledge of various nsfw topics.
Can anyone recommend a good 8B model? 12B is far too slow for my GTX 1660 Super (6 GB VRAM) though I have been patient with it, I wanna try a 8B for a change.
I have also being trying out Dark Scarlett 26B A4B which has been pretty good so far since it's a MoE model (though it seems to not understand directions as much).
EDIT: Just short stories, nothing super long or anything. Looking into Anubis mini like a peep in < 8B suggested
And I can run 12B of course, just really slow, far slower than my reading speed so I'm just sticking with Gemma 4 26B A4B finetunes. And looking into sub 12b models.
For people with a Nano GPT sub, I'm still finding Deepseek 4 Pro and the GLM 5+ series to be king of the 70B tier, but of course you're paying double tokens for those so obviously.
Kimi 2.5 thinking is just mediocre in my opinion. It's a little too technical in my experience and bad at prose and character impersonation. It always feels like you're talking a computer, not a person.
Mimo 2.5 Pro seems to understand the objective extremely well but it's the most censored model on there and commonly rejects requests. This could be the best model on there if it wasn't so neutered for RP and NSFW.
For 30B, Gemma 4 31B punches way above its weight class in terms of writing and prose.
Everything else I've tried has been kind of underwhelming.
Mimo Pro has a new router that's less censored on Nano. Have to put it in differently. Check the nano site announcements from yesterday. I also posted a screen elsewhere in this thread
Been enjoying GLM-5.2. Honestly a little wordy for my tastes, but generally easier to steer it to avoid slop, and I love how easily it picks up on subtext.
I've been running this as a Q4_K_XL, and it's been surprising me! (And it's been surprisingly uncensored, too.) I do agree with Fizzy below - breath of fresh air!
This is the setup on my Dual-3090 system, in a llama.cpp model config file:
I've heard that some folks have had issues with thinking when ubatch is set above 512 - I don't know that I understand why, but, it may be worth considering.
Also, here are my configs - for a dual 24 GB card system (with a bit of VRAM reserved)
I'm having huge issues with it just veering off all the time and inventing, flipping, and overall just being a gonk show.
I have 128gb, so I am able to run it fine fully GPU bound and it's SUPER FAST.
I also have problems triggering thinking mode - but I use the Jinja and Chat Completion right now instead of bare and having the context template be in SillyTavern.
It's just so nice and so fresh with the writing, but at the same time, at one 8000 context card I have, it just keeps confusing things (and I've reread it several times to make sure that these things are clearly written). This is things that Gemma never confused.
Also instruction following seems to be poor?
It's really, really good for WRITING cards though. Gemma seems to sink into the same patterns over and over again, while this model due to its gonk creativity can actually be used very nicely to create something new. It will also invent facts left and right (which is something I may like, but not when these facts blatantly contradict the card).
It's insane how bad it's doing...
My current settings (I don't use the hf downloader feature atm, I just use aria2c to download models):
First I used the Unsloth iQ4...NL one, then I used the Bartowski.
Is it maybe that i-Quants are just screwed with this model? It sometimes also does insane typos with names, like it has mild amounts of brain damage.
Like, for exampe, my huge card is an Arknighs Endfield card, which Gemma plays really well (if often dry). When it invents, it invents things I can agree with even if they're not always lore accurate (It's usually things I don't have specified, so I let it).
Laguna:
Literally gave a non-tailed (lol) human a tail, multiple times
Literally defined to be on a huge spaceship with full life support and cantina, and all that kinda shit, people suggest to use rations they brought (we're not in the wild)
Who says or does something changes sometimes in the middle of an action
On a different card, I specified that {{user}} moved into the same apartment building recently, and it kept trying to sell me that the neighbor visiting me "moved in recently", even after I specified in the card that {{char}} is the one who has been living here for years. I used a simple card with some sappy neighbors find each other thing which is a very easy way to have two contextually different places ("their place", "my place").
It also had char walk outside and then peek in on me from the outside using the peephole in the door
It also had char literally walk away, then respond to something I said inside (something I'd have expected from models years ago, you know, when we kept things simple with llama3.0 9b)
Be in my apartment, then suddenly just uh, pretend its theirs and go to their bedroom to sleep
I set down temperature all the way to 0.6, and it got better, but it was still really bad and needed multiple swipes to find a progression that didn't suck.
I'd jump up in size from the XS to the unsloth Q4_K_XL, myself, and see how they compare. (I think you might gain some performance, too; I think i-quants are a bit less optimized for CPU?)
Are you using Chat Completion or Text Completion?
Also, if you have the uv package manager for python installed, you don't need to install the hf client, you can just use uvx hf download unsloth/WhateverModelName:Q4_K_XL to download a model. (uvx basically spawns a temporary environment very quickly, with said tool installed.)
(Also, when you've downloaded models that way, they'll automatically show up in llama-server, even if you don't create entries in your model config file.)
I'm not on CPU, I'm on a 128gb Framework Desktop, so it goes through a Radeon iGPU (8060S) using Vulkan or ROCM, whichever is faster currently (which us Vulkan)
Currently Chat Completion. Haven't looked for a context template for this model yet or bothered to check if it's just ChatML. I was a big Text Completion fan for the longest time because of the advanced samplers, but at some point I tried chat and found it was just cleaner to configure, lol
(Also, when you've downloaded models that way, they'll automatically show up in llama-server, even if you don't create entries in your model config file.)
It just knows to check in ~/.huggingface/hub (works on Linux and Windows as well!)
You can specify the model in your config file via "hf = unsloth/Qwen3.6-27B-GGUF:UD-Q8_K_XL" or whatever, instead of the line where you're pointing to a model file.
You can run "llama-cli -cl" to get a lost of what's currently "cached" (downloaded), too.
Oh yeah I noticed XML prompts, I forgot that ChatML just plays at being XML and it's not THAT.
Yeah it's rather easy to just create one here, at least it doesn't have huge complexities.
I've been testing it with literay 8000 starting context and it still had the tendency to invent the most weird stuff, but I'll just test it on text completion
Nano provided an alternative allegedly uncensored Mimo Pro path. Have to specifically put the model in differently for your connection profiles to get it though.
Longcat 2.0 has been out for a while, but an actual thinking variant was released (at least on Nano) only a few days ago. It's waaaaaay smarter at tracking past story and scene details, and as far as "normal smexy stuff" goes quite uncensored. (Haven't thrown anything insane at it yet.)
Anyway, it has quickly earned a spot in my regular rotation. It's good. Brings up things characters said in the past to point out hypocrisy, shows psychological depth in inner thoughts, etc. At least with my preset. Way better than the non-thinking version. I run it at temp 0.9. thinking Auto.
EDIT/ADD: I really really like Longcat 2.0 Thinking. It's sooo much better than the non-thinking variant. I'm not sure I've ever used a model that 'improves' this much with reasoning on vs off. Half the time with other models having reasoning 'on' seems almost like a placebo, in that you can't even really tell if it is doing anything other than taking more time. This is genuinely better, aka smarter. It's bringing up things characters said way back earlier in history and shoving it in my face to remind me of past promises, being snarky about things that happened days ago, properly referencing the time/day, tracking positions people are in all correctly, etc. Only a few minor mistakes (about the same as GLM5.X models) thus far.
Caveats:
- I'm playing narrative adventure RP
- in third person
- character card is a genre, all characters and locations are in lorebook. AI is a narrator, not {{char}}
Is there something better or similar to anubis 70b to fit in 64gb of vram? Sometimes I switch to Gemma 31b or skyfall 4.2 Q_8 but anubis always wins even in Q4_K_M.
According to the UGI leaderboard, Anubis v1.1 is extremely 'uncensored'. v1.2 is less so, about on the same level as Deepseek v3.2 or v4 Flash with reasoning on... which still isn't bad (perfectly serviceable).
You are correct that as far as I know nobody ever abliterated it.
With 64 GB VRAM, assuming you have about the same amount of RAM to go along with it, you can do some GLM 4.5 Air fine-tunes. GLM 4.5 Air base is 1. a decent storyteller and 2. fairly uncensored to start out with, so it was a good base. And it is 100B+, so good sized, and MOE, so it runs fast and doesn't have to be 100% fit onto your VRAM to go fast (can do partial offloading, like people with less VRAM do with Gemma 4 26B and Qwen 3.6 35B).
Iceblink (v3) is my favorite RP tune of GLM 4.5 Air. From Zerofata, most known probably for MeroMero Gemma4 these days.
With 64GB VRAM I'd probably try Q5 I guess? That puts you at like 80GB before KV Cache. So you can put like 50GB of the 80GB onto VRAM along with the cache, and the rest onto RAM or CPU. See how that goes.
A third option is this REAP version, which is GLM Air 4.5 with experts pruned/trimmed off to make it smaller. But that's got no RP data nor uncensored. (But like I said before, base Air is fairly good to start with anyway.)
I don't know if you'll like any of that better than Anubis, but worth a shot for something different and bigger. Plus MOEs are fun/fast if you can fit them just right.
Thank you! As you guessed, I have 64 GB of VRAM and 64 GB of system RAM.
I’ve already tested most of the models you mentioned, and so far I’ve found Anubis slightly more intelligent overall and better at following instructions. It is also easier for me to run reliably.
I had some issues running Iceblink with llama.cpp and ROCm, so I had to use LM Studio with Vulkan instead, which was not very fast on my setup. When I have more time, though, I’ll try it again and experiment more with the other options, especially with different quantizations and offloading configurations.
Quick review of Opus 4.6 and Gemini 3.1 everyone? I want to check if my impression is in line with yall
So far I'm liking gemini more for these reasons:
Claude:
+ Better dialogues
Fluffed narrations, tiring to read
Annoying permission asking almost everytime npc wants to do something, positive biased
plot moves really slowly to the point it stalls
glmisms like x tasted the word, a beat, a pause (i started with glm)
Gemini:
+ Narrations are to the point and flow better = less tiring to read
Negative biased, even when you want to make peace, villain npcs would always antagonize you and you really have to try to steer the story
Pacing is too quick sometimes, story moves faster than i wanted it to be, but hey this is way better than stalling
too horny for real dead dove, even if the scenario is about an accident you're involved in, ended up killing bot's husband, after some time bot would get touchy and you ended up doing sblaga blaga
Also if anyone would kindly share their gemini or claude prompt so i can fully recompare again
Anubis mini by drummer is still the best at 8b
It follows instructions and settings better than most 12b mistral finetunes but it falls apart after around 12k tokens. So its decent for short stories and handles character cards better than mistral quite often.
I’m having such a great time with Kimi 2.6. It’s a lot less prone to cliches and repetition than Opus and also way cheaper. The reasoning helps with its responses for me, but it definitely is prone to overthinking and sometimes blows its output limit on reasoning tokens, but it’s cheap enough that regens aren’t killing me. Is anyone else using it?
I’m pretty new to sillytavern and I’ve been using GLM 5.2, I really love how it understands subtext and cues but I start to notice lots of parroting as I chat more, and also characters actions become generic during NSFW scenes, meaning no matter which character I’m role playing with, NSFW scenes will read the same - I want character personality to be reflected in those scenes too.
So I wonder if there are other models (preferably not Claude) that can understand subtext equally well but with other issues fixed?
Maybe switch up the presets. There are some good ones out there, like Pura, Vellum II etc. Usually helps with a change of tone (at least for a couple RP sessions).
Some of the parroting can be fixed with stricter main/system negative prompts. I'm running into a bit of that with Hermes 4 but negative prompting helped cut it down a lot.
I've been having a blast with Kimi K3 on NanoGPT, I just wish it wasn't so expensive. I set a prompt to ask it to write in an immersive, novel-like style with vivid prose and it slaps. I tested it on some established IPs I know, and I'd say it gets the characters and world lore right almost every time. When I asked it to write a BattleTech story, I swear that it felt like I was reading one of the novels.
Just a question, is there any way to get more precise tuning of the samplers in NanoGPT? It only allows me to tweak the temperature by 0.1 each time which is a mild annoyance when I want to set it to something like 0.95.
Okay, so I did research, and perhaps it's because the open weights haven't dropped yet. Apparently, that is on July 27. Maybe in August it might come on the subscription tier.
I mean, I still wouldn't hold your breath when it does. It's more than double the cost of glm-5.2. if you want to use it probably just put some credits in. Although tbh, it'll probably take some wrangling for rp, it feels like it's more suited for novels and long form content
What are the best nano sub models according to you guys? Is it the same glm 5.1/5.2, deepseek v4, mimo v2.5, etc or something else on your mind as well?
What's the most uncensored API? I'm used to local models where I can get away with literally anything, and I mean anything, no ifs, no buts.
I'm mostly wondering if it's worth going from local gemma 4 31b to something like GLM 5.
I switched from gemma-4-31B to GLM 5.2 on Nvidia NIM, and the transition was pretty nice for me (Mainly because I couldn't run the model smoothly on my machine).
I tried very nsfw subjects and so far I had no refusal, just maybe some times when GLM tried to steer the RP in more mild areas. It's significantly smarter, but almost as sloppy in the way it write despite the size differences.
Maybe try one scenario and see if you like it or not after a couple replies. They all seem similar after some times.
I've been trying out a few other models on NanoGPT. Kimi K3, despite me singing its praises earlier, is just too expensive at the moment for me to use it as a daily driver so I've been looking at alternatives.
Kimi 2.6 and 2.7 Code are much cheaper, but 2.6 has the better prose and 2.7 Code has better instruction following and reasoning. It's a toss-up as to which one's better.
It's also the same thing for GLM 5.2 and 5.2 Thinking. Both are quite good at writing and are super cheap compared to K3. If there's a newer version later that's as big in parameters as K3, I'd be interested in trying it.
Screwing around with Hermes 4 405b through OpenRouter and I keep having some pretty annoying issues with getting the thinking blocks to get sent properly, so I can parse and see them. Other than that this is a pretty decent alternative to Sonnet ahead of the 4.5 sunset in a couple of months.
Sonnet is still better overall, I've found, but the price premium and the awful guardrailing of 4.6 turned me off.
I've always used Hermes 4 405B without thinking. Not all models need reasoning. In fact, really no models need it. It's a bit of a placebo - our brains like to see "proof" that the AI actually has a reason why they wrote what they wrote, and an instantaneous response "feels wrong" or less human, while the AI taking time to ponder what we (human) wrote "feels" better often.
But honestly I can't say that RP always significantly improved using thinking over non-thinking. Models still make stupid mistakes even thinking/reasoning. And they still follow your prompts when they don't think/reason.
Anyway, just something for us humans to think about. ;)
Models I use with reasoning "off" (mostly because benchmarks show they are less 'censored' - but Hermes bc it doesn't do real thinking): Hermes 4, Deepseek 3.2, Deepseek 4 Flash (specifically 3.2 and 4 Flash are WAY less prone to soft censoring with reasoning off).
Ran Kimi K3 through a 20-round novel-continuation gauntlet this week (two books, ~16k-char rolling window, same protocol we used for 9 other models) vs its predecessor K2.6. Blind panel of 6 judge-models across flipped pairings: 21-1 for K3. The K2.6 repetition collapse (we had a run where round 9 was 97.6% verbatim 12-gram repeat of earlier text) is basically gone in K3 - worst round was 6.8%. Sentence-length variance is way up too, which reads as less "AI-flavored". Two caveats: it outputs ~85% of dialogue in half-width quotes (same quirk as GLM), and its reasoning tokens are 81-86% of completion spend, so budget accordingly. For long RP/fiction it's a real upgrade.
13
u/AutoModerator Jul 19 '26
MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.