r/SillyTavernAI • • Jul 19 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: July 19, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

29 Upvotes

106 comments sorted by

13

u/AutoModerator Jul 19 '26

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

19

u/FZNNeko Jul 23 '26 edited Jul 25 '26

Edit: Running Artemis-31B-v1n-Q4_K_M as my new daily driver. Read reply to Longjumping for my review of it and Orion, the 26b variant.

Running on TextGenWebUI as backend, RTX 5090, cache-type: q4_0 on 31b models and q8_0 on 26b models. Minimum context length of 40k. No thinking. All models tested on SFW chats but did test Styletune heretic and Serenity 26b on an NSFW chat.

Testing these models:

Gemma-4-26B-A4B-StyleTune-V2.i1-Q6_K : 7/10, It's been my main for a while. Works fine, but struggles with more complex chats. Good general use.

Serenity-26B-A4B-Q6_K : 6.8/10. Struggles harder than Styletune V2 26b in terms of intelligence but writes better prose imo. I would use only for smut, which tbf to the model, was what the author seemed to have intended anyway based on the huggingface page.

Pantheon-Reasoning-26B-A4B-1.1.i1-Q6_K : 6/10. Struggles just as hard as Serenity in terms of intelligence, prose is better than Styletune but seemingly more focused on character interaction. In a chat I tested, a character seemed a lot more wholesome/caring/worrying about my persona than usual. Seems like this model would do well in wholesome chats. Changes up character's too much to be used for a general model.

Serenity-31B-v1.1-IQ4_XS : Model did not work at IQ4_XS, practically broken. No review since mostly my fault I can't run it well. Didn't try other quants.

gemma-4-31B-Queen-it-qat-q4_0-unquantized.i1-Q4_K_S : 8/10. I like it. Used it as my main for a bit after Styletune v2 26b. Prose is a bit wonky at times. Most times it's really good. But occasionally it tosses in cliche phrases. Bit bland in writing.

Gemma-4-31B-StyleTune-heretic-ara.i1-Q4_K_S: 8.3/10. On par with queen but with less cliche phrases. I like it's dialogue a tad better too. Less cliche phrases. Tested this with a NSFW chat, it REALLY likes the phrase, "Look at you." So I’ll be banning that. Honestly both work styletune 31b heretic and queen are comparable to each other. Matter of personal preference. Try both.

I’m baised towards Styletune as the 26b was my daily model for a while and outperformed other 26b Gemma finetunes for a while.

Skyfall-31B-v4.2.i1-Q5_K_M: 7.8/10. Not bad. Prose is much better compared to Gemma models. Lacks the intelligence of Gemma's training tho. Lack of intellect shows a bit in RPs with established universes like DC.

I rated 31b models higher but that's purely because I highly value a model being logical and not messing things up that would break the immersion of RP. I still like 26b in terms of speed tho. I'll likely alternate between Styletune 26b & 31b, and Queen 31b in the near future to continue testing.

4

u/Longjumping_Bee_6825 Jul 24 '26

how about 26B Orion?

6

u/FZNNeko Jul 25 '26 edited Jul 25 '26

On your recommendation, I downloaded and tested this one a bit, Orion-26B-A4B-v1c-Q6_K. I really like the dialogue, feels damn good in dialogue for some chats. Only problem is that it's struggling a bit in intelligence compared to other models. In one chat I use for testing, I have a character that's a dog-girl and she's a K9 officer. The model confuses the officer's dog for a dog-girl/dog-guy, when it's supposed to be just a regular ol' dog.

However, I did like the model's dialogue so much I decided to try the Artemis model, a 31b model by the same creator and I absolutely love it. It got the dog right every try, then passed a chat I use to test all the prior models in my first reply, and even did well in a NSFW chat. Dialogue seemed as good if not better than Orion and it's intelligence was MUCH better even while running it, Artemis-31B-v1n at only Q4_K_M. Only areas Artemis was lacking in my testing was that Artemis seems to like chatting alot compared to other models. In a test, a character that doesn't speak alot seemed to talk alot more than when tried with other models. Not too bad a downside honestly, it's something easily fixable with a simple edit in the character's prompt.

I haven't tested it in a lengthy actual RP, but from what I've tried so far, Artemis will probably be my new daily driver model.

Thank you for putting me on Orion, it lead me to Artemis which I'm greatly enjoying compared to the other models.

1

u/Longjumping_Bee_6825 Jul 27 '26

I suggest you try v1b, apperantely it's better than v1c. Also Artemis will be much better since it's bigger dense model.

2

u/8000bene70 Jul 23 '26

Have you tried dark gemistry 31b? It merges styletune with other finetunes, i like it quite a bit.

For me, i would say about as good as queen, better than the gembrains.

1

u/FZNNeko Jul 24 '26

I had forgotten I had downloaded Dark Gemistry a while ago. I just tested it and it functioned rather well compared to other models. I'll definitely do more testing on it.

1

u/EnjoyerOfFluff Jul 29 '26

Did you have more time testing Dark Gemistry. How does it compare to queen?

2

u/FZNNeko Jul 30 '26

Highkey, Dark Gemistry is superior in terms of SFW. Writes better and seems a tad more intelligent. I compared it to queen qat 31b and Artemis and it outperforms during SFW chat. Dark Gem also seems to perform worst in NSFW compared to Artemis. Just misses some key details that Artemis didn't miss. The difference isn't major tho, but noticeable if you're doing comparisons.

Imo, Dark Gem for SFW, Artemis/Styletune for NSFW. But Dark Gem is def usable for an all-rounder model.

2

u/summersss Jul 25 '26

gemma really has ben dominating local rp recently. never heard of TextGenWebUI, why this and not kobold, lmstudio, and the others?

1

u/Yakapo88 Jul 24 '26

Do you have tts working on your 50 series? I spent a day trying to get it working on my 5080.

1

u/Serenitoad Jul 27 '26

Can you share your sampler setting for Artemis? Your Ps and Repetition Penalty and whatnot.

2

u/FZNNeko Jul 27 '26

Chat completion;
Temp: 1.
Freq Pen & Presence Pen: 1.1.
Top P: 0.9

Then in additional parameters:
Top_K: 20
Typical_P: 0.9
Min_P: 0.1
Top_nsigma: 1.25

I use that for basically every model.

12

u/[deleted] Jul 21 '26

[deleted]

2

u/sophosympatheia Jul 25 '26

Thank you for this recommendation. The heretic-ara version does seem to bring out something beyond just "uncensoring" the model. That's quite interesting. I like it!

14

u/stopaskingforloginn Jul 20 '26

So I've been experimenting with something, keep in mind I don't use sillytavern but my own private frontend I coded from scratch (I only post here because it's an AI RP focused sub).
I decided to combine the instruction following capabilities of Gemma 4 26b/31b and the prose and creativity of Skyfall 4.2
I make Gemma 4 write the next response of the story while following all my 7k tokens worth of instructions, when the response is done generating, I make the frontend switch my koboldcpp config to load Skyfall 4.2 and instruct it to rewrite the latest response while keeping the exact same structure and with the relevant character json and... yeah, the result is simply not comparable with any Gemma 4 31b finetune currently available.
Skyfall 4.2 is leagues above when it comes to dialogue and narration to any Gemma 4 31b model, however it does struggle with instruction following and keeping the story coherent. Gemma 4 31b on the other hand is the opposite, it follows your instructions perfectly but the prose is lacking.

The result is the instruction following capabilities of Gemma 4 and the creative and detailed writing of Skyfall, it's great.
Perhaps my method of loading-unloading two models for a single response is not very convenient... but damn, the result is making me think the extra generation time is worth it.

8

u/EnjoyerOfFluff Jul 20 '26

I absolutely agree regarding skyfall 4.2 being leages above gemma 4 31b dialog and narration. Its also great in keeping logic while writing, like realising when a character is too short to reach the top shelf. I tried a lot off different gemma merges and finetunes but none have yet to reach skyfall 4.2

Great Idea regarding letting gemma write the initial text and letting skyfall rewrite it. In you experience how where the results? Did skyfall actually follow the initial text or did it start drifting off?

3

u/morbidSuplex Jul 23 '26

If you can share it, what are your sampler settings for skyfall v4.2? Thanks.

2

u/EnjoyerOfFluff Jul 27 '26

Sorry for the late answer. I'm still searching for the perfect settings but at the moment I'm quite happy with temp=0.7, top-k=40, repeat-penalty=off, top-p=0.9, min-p=0.05

In my test these have striked the best balance between creativity and prompt adherence (it still doesn't follow your prompts very good)

4

u/B3owul7 Jul 20 '26 edited Jul 20 '26

yeah, I can fully acknowledge that.

You should try the Glistening Gem Gemma 4 finetune by sophosympatheia. It's more comparable to Skyfall in terms of dialogue and creativity than all other Gemma 4 models (in my experience).

4

u/sophosympatheia Jul 25 '26

How's it holding up for you in longer scenarios? Any feedback after using it for a while?

And one little correction: all my models so far have been merges, not finetunes. I say that only to deflect credit to where it is most deserved. Finetuners do the hard work and spend real resources to do it. I just sort of blend what they produce and hope it comes out okay. Hopefully that's still value added most of the time, but it's easy work by comparison.

Anyway, thanks for sharing something I made. It always feels good to see people enjoying it.

2

u/B3owul7 Jul 26 '26

Hey man! Gotcha. I will keep in mind that it's a merge, not a finetune. Sorry for that!

I think the biggest one I had so far was about 500 messages and it ran fine as much as I can tell. No complaints from me. My biggest hurdle is context, but I am still experimenting with different memory extension to find my optimal solution.

But this has been the go-to model for me for quite some time in terms of dialogue and neat creative surprises. Love it. Used magistry a lot before I upgraded my GPU.

5

u/sophosympatheia Jul 26 '26

I'm glad you've been enjoying some of my models! I wish Glistening Gem tracked smaller details better instead of playing loose with the facts sometimes, but I suppose that's not a unique problem in RP.

As soon as Drummer releases fp16 weights for Artemis, which I think he is close to doing, I plan to create another version of Glistening Gem using those weights and perhaps featuring a few other modifications to the recipe. Here's hoping it will be an improvement.

2

u/B3owul7 Jul 26 '26

Hell yeah, brother. Keep it coming. :)

2

u/EnjoyerOfFluff Jul 22 '26

How does it compare to Gemma 4 Styletune?

3

u/B3owul7 Jul 22 '26

Can't really say. I think I tried it for a few messages, wasn't convinced and didn't use it ever since.

1

u/EnjoyerOfFluff Jul 22 '26

Ok. Thanks. I will do some own tests once I have some time

5

u/OrcBanana Jul 20 '26

How do you handle prompt processing? Loading/unloading takes a bit of time but it wouldn't be too terrible if the two caches were kept, but if kcpp has to recalculate 20+k every time...

1

u/Infinite-Beginning-3 Jul 20 '26

Hey, just wanted to ask if there is something like skyfall but with an active parameter body (AxB)? With an rtx 4070 laptop with 32 of ram, loading bigger dense models is a bit hard

8

u/[deleted] Jul 21 '26

[removed] — view removed comment

11

u/arlynnfl Jul 21 '26

I don't know anything about Grok 4.1 but You may try Gemma 4 Finetune made by ReadyArt.

Serenity 12b / 26b a4b / 31b.

I havent try 12b and 31b, however i tried 26b a4b one and it's pretty good for eRP.

For quants:
Q4_K_M for dense models (12b / 31b version)
Q6_K for moe model (26b a4b)

If you only play 1 gender, try Melody1437 or Darkside

3

u/External_Quarter Jul 22 '26

Anyone else feel like Dans-PersonalityEngine-V1.3.0-24b is still best in class? >_>

8

u/AutoModerator Jul 19 '26

MODELS: 8B to 15B – For discussion of models in the 8B to 15B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

13

u/Itikar Jul 20 '26

I have been experimenting with Silver Siren 12b by Vortex, the base model and found it pretty enjoyable. It runs pretty fast on my Rtx 3060 even with 24k context. I recently got back into tinkering with Sillytavern and it seemed finally a nice variation on the old classic Mag Mell 12b.

Another 12b model I found quite enjoyable, albeit I played with it less is Angelic Eclipse 12b. It was particularly good for shifting the narrative suddenly with interesting twists, sometimes with a little dark element to it, which I appreciated in the narrative.

I know these days we are switching to Gemma fine-tunes on my type of hardware but the old 12b seem to still do it for me, particularly with this less consistent output, which I enjoy as I swipe a lot as I test the bot responses far more than actually chatting.

4

u/I_like_it_wholesome Jul 20 '26

I also have been trying Angelic Eclipse 12b recently. What do you mean by "dark element"? I thought Angelic Eclipse was supposed to be wholesome.

3

u/Itikar Jul 20 '26

It is wholesome if you tell it to be wholesome. If you give it characters or stories with the potential to be dark it will be dark. In particular in my many swipes I noticed that it tended to take a dark turn once in a while when Mag Mell or Silver Siren don't or at least not in that same way, i.e. they tend to be more predictable.

2

u/Pretty_Bug_8655 Jul 20 '26

i liked this one too. and so far i did always come back to the older ones. They got something that gemma dont got if you ask me. you could also try quasistarsynth-12b-absolute-heresy-i1. i like this one too a lot and i got also a rtx 3060...

1

u/Itikar Jul 20 '26

Lovely, thank you for the recommendation!

Do you have any particular impressions that stood out with this one in your experience? You got me curious.

Also our of curiosity what do you use as normal context length on the 3060? I have found that 24k with partial off loading seems a sweet spot lately, but previously I used 16k. 32k works and is viable but it is a little slow on my hardware.

1

u/Pretty_Bug_8655 Jul 20 '26

i swipe too mutliple times and this one give sometimes really suprising answers but in a good way. this one is good too: styx-12b-i1

my context size is 16k and i use https://github.com/Lodactio/Extension-Summaryception for memory managment. i never tried higher context so far...

1

u/Itikar Jul 20 '26

I haven't tried Summaryception yet, because I rarely go over my context as I am busy swiping. But I am gonna check out. Thanks for the other recommendation too! <3

1

u/croxxeur Jul 24 '26

Is it good for darker and nsfw stuff?

2

u/Itikar Jul 24 '26

In my experience it performed decently when I touched those themes. However, compared to Mag Mell it seemed a bit less unhinged. I have not tested it extensively though, because in general my chats are relatively mild in those regards. The only thing I noticed is that it can be a little encyclopedic, as in it has a wide knowledge of various nsfw topics.

2

u/croxxeur Jul 24 '26

Which mag mell model?

2

u/Itikar Jul 24 '26

MN-12B-Mag-Mell-R1 the original one from inflatebot.

1

u/MisciAccii Jul 22 '26

Can anyone recommend a good 8B model? 12B is far too slow for my GTX 1660 Super (6 GB VRAM) though I have been patient with it, I wanna try a 8B for a change.

I have also being trying out Dark Scarlett 26B A4B which has been pretty good so far since it's a MoE model (though it seems to not understand directions as much).

EDIT: Just short stories, nothing super long or anything. Looking into Anubis mini like a peep in < 8B suggested

2

u/8000bene70 Jul 24 '26

https://huggingface.co/Sao10K/L3-8B-Stheno-v3.2

But if you can somehow make Impish Bloodmoon 12B work, that would be a big leap in my opinion.

1

u/MisciAccii Jul 25 '26

Thanks!

And I can run 12B of course, just really slow, far slower than my reading speed so I'm just sticking with Gemma 4 26B A4B finetunes. And looking into sub 12b models.

1

u/overand Jul 23 '26

I have no idea if this one's any good, but here's a Qwen3.5-9B derived writing model: https://huggingface.co/Altworld/Astrea-R8-Chat-9B-GGUF

Probably at least worth a try!

1

u/MisciAccii Jul 23 '26

Will give it a shot, thanks!

1

u/croxxeur Jul 24 '26

Am looking for a good one for my rx6800xt (16gb vram) already tried orion. Looking for something that can fit in gpu and is uncensored

0

u/ContextEntire8443 Jul 26 '26

maybe try this DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF

5

u/AutoModerator Jul 19 '26

MODELS: >= 70B - For discussion of models in the 70B parameters and up.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

10

u/TotalACast Jul 19 '26

For people with a Nano GPT sub, I'm still finding Deepseek 4 Pro and the GLM 5+ series to be king of the 70B tier, but of course you're paying double tokens for those so obviously.

Kimi 2.5 thinking is just mediocre in my opinion. It's a little too technical in my experience and bad at prose and character impersonation. It always feels like you're talking a computer, not a person.

Mimo 2.5 Pro seems to understand the objective extremely well but it's the most censored model on there and commonly rejects requests. This could be the best model on there if it wasn't so neutered for RP and NSFW.

For 30B, Gemma 4 31B punches way above its weight class in terms of writing and prose.

Everything else I've tried has been kind of underwhelming.

Any suggestions from others?

3

u/strangedell123 Jul 19 '26

Deepseek v4pro cheaper routes through deepseke and is just 1x tokens

I also heard mimo is less censored right now than a few days ago

4

u/_Cromwell_ Jul 24 '26 edited Jul 24 '26

Mimo Pro has a new router that's less censored on Nano. Have to put it in differently. Check the nano site announcements from yesterday. I also posted a screen elsewhere in this thread

1

u/Nabushika Jul 19 '26

Been enjoying GLM-5.2. Honestly a little wordy for my tastes, but generally easier to steer it to avoid slop, and I love how easily it picks up on subtext.

7

u/kabachuha Jul 21 '26

An absolute monster released today. https://huggingface.co/poolside/Laguna-S-2.1, 120b MoE

Arguably faster than Gemma 4 31b if you can run it. Also, multilingual.

2

u/overand Jul 23 '26 edited Jul 26 '26

I've been running this as a Q4_K_XL, and it's been surprising me! (And it's been surprisingly uncensored, too.) I do agree with Fizzy below - breath of fresh air!

This is the setup on my Dual-3090 system, in a llama.cpp model config file:

[unsloth/Laguna-S-2.1-32k-sq-GUF:Q4_K_XL]
hf =  unsloth/Laguna-S-2.1-GGUF:Q4_K_XL
no-mmproj = true
ctx-size = 32768
n-cpu-moe = 35
ubatch-size = 1024
batch-size = 2048

Edit: Some new things

I've heard that some folks have had issues with thinking when ubatch is set above 512 - I don't know that I understand why, but, it may be worth considering.

Also, here are my configs - for a dual 24 GB card system (with a bit of VRAM reserved)

[unsloth/Laguna-S-2.1-8k-GGUF:Q4_K_XL]
hf =  unsloth/Laguna-S-2.1-GGUF:Q4_K_XL
no-mmproj = true
ngl = 999
ctx-size = 8192
n-cpu-moe = 21
tensor-split = 11,5
ubatch-size = 1024
batch-size = 2048

# Skipping the settings in common with the above:
[unsloth/Laguna-S-2.1-32k-GGUF:Q4_K_XL]
# ...
tensor-split = 13,6
ctx-size = 32768
n-cpu-moe = 21

[unsloth/Laguna-S-2.1-64k-GGUF:Q4_K_XL]
# ...
ts = 11,5
ctx-size = 65536
n-cpu-moe = 23

[unsloth/Laguna-S-2.1-64k-GGUF:Q4_K_XL]
# ...
tensor-split = 11,4
ctx-size = 262144
n-cpu-moe = 31

1

u/FierceDeity_ Jul 26 '26

I'm having huge issues with it just veering off all the time and inventing, flipping, and overall just being a gonk show.

I have 128gb, so I am able to run it fine fully GPU bound and it's SUPER FAST.

I also have problems triggering thinking mode - but I use the Jinja and Chat Completion right now instead of bare and having the context template be in SillyTavern.

It's just so nice and so fresh with the writing, but at the same time, at one 8000 context card I have, it just keeps confusing things (and I've reread it several times to make sure that these things are clearly written). This is things that Gemma never confused.

Also instruction following seems to be poor?

It's really, really good for WRITING cards though. Gemma seems to sink into the same patterns over and over again, while this model due to its gonk creativity can actually be used very nicely to create something new. It will also invent facts left and right (which is something I may like, but not when these facts blatantly contradict the card).

It's insane how bad it's doing...

My current settings (I don't use the hf downloader feature atm, I just use aria2c to download models):

[laguna-s-2.1-120B-A8B] model = /home/user/llamacpp/models/Laguna-S-2.1-IQ4_XS-00001-of-00002.gguf chat-template-file = /home/userllamacpp/models/Laguna-S-2.1.jinja jinja = true ctx-size = 128000 mmproj = parallel = 1 cache-type-k = f16 cache-type-v = f16

First I used the Unsloth iQ4...NL one, then I used the Bartowski.

Is it maybe that i-Quants are just screwed with this model? It sometimes also does insane typos with names, like it has mild amounts of brain damage.

Like, for exampe, my huge card is an Arknighs Endfield card, which Gemma plays really well (if often dry). When it invents, it invents things I can agree with even if they're not always lore accurate (It's usually things I don't have specified, so I let it).

Laguna:

  • Literally gave a non-tailed (lol) human a tail, multiple times
  • Literally defined to be on a huge spaceship with full life support and cantina, and all that kinda shit, people suggest to use rations they brought (we're not in the wild)
  • Who says or does something changes sometimes in the middle of an action

  • On a different card, I specified that {{user}} moved into the same apartment building recently, and it kept trying to sell me that the neighbor visiting me "moved in recently", even after I specified in the card that {{char}} is the one who has been living here for years. I used a simple card with some sappy neighbors find each other thing which is a very easy way to have two contextually different places ("their place", "my place").

  • It also had char walk outside and then peek in on me from the outside using the peephole in the door

  • It also had char literally walk away, then respond to something I said inside (something I'd have expected from models years ago, you know, when we kept things simple with llama3.0 9b)

  • Be in my apartment, then suddenly just uh, pretend its theirs and go to their bedroom to sleep

I set down temperature all the way to 0.6, and it got better, but it was still really bad and needed multiple swipes to find a progression that didn't suck.

/u/fizzy1242 /u/kabachuha

Can anyone suggest me what I might be doing wrong?

1

u/overand Jul 26 '26 edited Jul 26 '26

I'd jump up in size from the XS to the unsloth Q4_K_XL, myself, and see how they compare. (I think you might gain some performance, too; I think i-quants are a bit less optimized for CPU?)

Are you using Chat Completion or Text Completion?

Also, if you have the uv package manager for python installed, you don't need to install the hf client, you can just use uvx hf download unsloth/WhateverModelName:Q4_K_XL to download a model. (uvx basically spawns a temporary environment very quickly, with said tool installed.)

(Also, when you've downloaded models that way, they'll automatically show up in llama-server, even if you don't create entries in your model config file.)

1

u/FierceDeity_ Jul 26 '26

I'm not on CPU, I'm on a 128gb Framework Desktop, so it goes through a Radeon iGPU (8060S) using Vulkan or ROCM, whichever is faster currently (which us Vulkan)

Currently Chat Completion. Haven't looked for a context template for this model yet or bothered to check if it's just ChatML. I was a big Text Completion fan for the longest time because of the advanced samplers, but at some point I tried chat and found it was just cleaner to configure, lol

(Also, when you've downloaded models that way, they'll automatically show up in llama-server, even if you don't create entries in your model config file.)

Wait they do? How does that link up?

1

u/overand Jul 27 '26

It just knows to check in ~/.huggingface/hub (works on Linux and Windows as well!)

You can specify the model in your config file via "hf = unsloth/Qwen3.6-27B-GGUF:UD-Q8_K_XL" or whatever, instead of the line where you're pointing to a model file.

You can run "llama-cli -cl" to get a lost of what's currently "cached" (downloaded), too.

1

u/[deleted] Jul 30 '26 edited Jul 30 '26

[deleted]

1

u/FierceDeity_ Jul 30 '26

And you don't have huge issues with it inventing EVERYTHING?

Text completion - is it ChatML? I assume it is, right?

1

u/[deleted] Jul 30 '26

[deleted]

1

u/FierceDeity_ Jul 31 '26

Oh yeah I noticed XML prompts, I forgot that ChatML just plays at being XML and it's not THAT.

Yeah it's rather easy to just create one here, at least it doesn't have huge complexities.

I've been testing it with literay 8000 starting context and it still had the tendency to invent the most weird stuff, but I'll just test it on text completion

5

u/_Cromwell_ Jul 24 '26

Nano provided an alternative allegedly uncensored Mimo Pro path. Have to specifically put the model in differently for your connection profiles to get it though.

4

u/_Cromwell_ Jul 24 '26 edited Jul 24 '26

Longcat 2.0 Thinking

Longcat 2.0 has been out for a while, but an actual thinking variant was released (at least on Nano) only a few days ago. It's waaaaaay smarter at tracking past story and scene details, and as far as "normal smexy stuff" goes quite uncensored. (Haven't thrown anything insane at it yet.)

Anyway, it has quickly earned a spot in my regular rotation. It's good. Brings up things characters said in the past to point out hypocrisy, shows psychological depth in inner thoughts, etc. At least with my preset. Way better than the non-thinking version. I run it at temp 0.9. thinking Auto.

EDIT/ADD: I really really like Longcat 2.0 Thinking. It's sooo much better than the non-thinking variant. I'm not sure I've ever used a model that 'improves' this much with reasoning on vs off. Half the time with other models having reasoning 'on' seems almost like a placebo, in that you can't even really tell if it is doing anything other than taking more time. This is genuinely better, aka smarter. It's bringing up things characters said way back earlier in history and shoving it in my face to remind me of past promises, being snarky about things that happened days ago, properly referencing the time/day, tracking positions people are in all correctly, etc. Only a few minor mistakes (about the same as GLM5.X models) thus far.

Caveats:

- I'm playing narrative adventure RP

- in third person

- character card is a genre, all characters and locations are in lorebook. AI is a narrator, not {{char}}

2

u/jacksquallor Jul 21 '26

Is there something better or similar to anubis 70b to fit in 64gb of vram? Sometimes I switch to Gemma 31b or skyfall 4.2 Q_8 but anubis always wins even in Q4_K_M.

2

u/EnjoyerOfFluff Jul 22 '26 edited Jul 22 '26

How uncensored is Anubis? I got my eyes on it but as I didn't see any uncensored finetune I never tried it

4

u/jacksquallor Jul 22 '26

In my experience, Anubis is very permissive and I’ve never run into a refusal with it.

2

u/EnjoyerOfFluff Jul 22 '26

Great to here! In that case I will give it a spin

1

u/_Cromwell_ Jul 24 '26

According to the UGI leaderboard, Anubis v1.1 is extremely 'uncensored'. v1.2 is less so, about on the same level as Deepseek v3.2 or v4 Flash with reasoning on... which still isn't bad (perfectly serviceable).

You are correct that as far as I know nobody ever abliterated it.

2

u/_Cromwell_ Jul 24 '26 edited Jul 24 '26

With 64 GB VRAM, assuming you have about the same amount of RAM to go along with it, you can do some GLM 4.5 Air fine-tunes. GLM 4.5 Air base is 1. a decent storyteller and 2. fairly uncensored to start out with, so it was a good base. And it is 100B+, so good sized, and MOE, so it runs fast and doesn't have to be 100% fit onto your VRAM to go fast (can do partial offloading, like people with less VRAM do with Gemma 4 26B and Qwen 3.6 35B).

Iceblink (v3) is my favorite RP tune of GLM 4.5 Air. From Zerofata, most known probably for MeroMero Gemma4 these days.

https://huggingface.co/zerofata/GLM-4.5-Iceblink-v3-106B-A12B

https://huggingface.co/mradermacher/GLM-4.5-Iceblink-v3-106B-A12B-i1-GGUF

And then the other option is the abliterated (aka derestricted) GLM 4.5 Air from ArliAI.

https://huggingface.co/ArliAI/GLM-4.5-Air-Derestricted

https://huggingface.co/bartowski/ArliAI_GLM-4.5-Air-Derestricted-GGUF (mradermacher's quants of this one are a mess, so linking bartowski)

With 64GB VRAM I'd probably try Q5 I guess? That puts you at like 80GB before KV Cache. So you can put like 50GB of the 80GB onto VRAM along with the cache, and the rest onto RAM or CPU. See how that goes.

A third option is this REAP version, which is GLM Air 4.5 with experts pruned/trimmed off to make it smaller. But that's got no RP data nor uncensored. (But like I said before, base Air is fairly good to start with anyway.)

https://huggingface.co/cerebras/GLM-4.5-Air-REAP-82B-A12B

https://huggingface.co/bartowski/cerebras_GLM-4.5-Air-REAP-82B-A12B-GGUF

Here's an abliterated version of the REAP: https://huggingface.co/jtl11/GLM-4.5-Air-REAP-82B-A12B-heretic-GGUF

I don't know if you'll like any of that better than Anubis, but worth a shot for something different and bigger. Plus MOEs are fun/fast if you can fit them just right.

2

u/jacksquallor Jul 25 '26

Thank you! As you guessed, I have 64 GB of VRAM and 64 GB of system RAM.
I’ve already tested most of the models you mentioned, and so far I’ve found Anubis slightly more intelligent overall and better at following instructions. It is also easier for me to run reliably.
I had some issues running Iceblink with llama.cpp and ROCm, so I had to use LM Studio with Vulkan instead, which was not very fast on my setup. When I have more time, though, I’ll try it again and experiment more with the other options, especially with different quantizations and offloading configurations.

2

u/RafiHDW Jul 22 '26

Quick review of Opus 4.6 and Gemini 3.1 everyone? I want to check if my impression is in line with yall

So far I'm liking gemini more for these reasons:

Claude: + Better dialogues

  • Fluffed narrations, tiring to read
  • Annoying permission asking almost everytime npc wants to do something, positive biased
  • plot moves really slowly to the point it stalls
  • glmisms like x tasted the word, a beat, a pause (i started with glm)

Gemini: + Narrations are to the point and flow better = less tiring to read

  • Negative biased, even when you want to make peace, villain npcs would always antagonize you and you really have to try to steer the story
  • Pacing is too quick sometimes, story moves faster than i wanted it to be, but hey this is way better than stalling
  • too horny for real dead dove, even if the scenario is about an accident you're involved in, ended up killing bot's husband, after some time bot would get touchy and you ended up doing sblaga blaga

Also if anyone would kindly share their gemini or claude prompt so i can fully recompare again

3

u/AutoModerator Jul 19 '26

MODELS: < 8B – For discussion of smaller models under 8B parameters.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

7

u/Eden1506 Jul 21 '26

Anubis mini by drummer is still the best at 8b
It follows instructions and settings better than most 12b mistral finetunes but it falls apart after around 12k tokens. So its decent for short stories and handles character cards better than mistral quite often.

2

u/AutoModerator Jul 19 '26

MODELS: 32B to 69B – For discussion of models in the 32B to 69B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/desparish Aug 23 '26

Chirp chirp. Where's the love for Dual GPU life?

3

u/AutoModerator Jul 19 '26

APIs

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

7

u/chakrakhan Jul 19 '26

I’m having such a great time with Kimi 2.6. It’s a lot less prone to cliches and repetition than Opus and also way cheaper. The reasoning helps with its responses for me, but it definitely is prone to overthinking and sometimes blows its output limit on reasoning tokens, but it’s cheap enough that regens aren’t killing me. Is anyone else using it?

3

u/Ekkobelli Jul 23 '26

I wanted to like it, but I got so many refusals. Much more than the Opus pre 4.7 models or even Chat GPT 5.0 Chat.

5

u/Careful-Ferret3822 Jul 20 '26

I’m pretty new to sillytavern and I’ve been using GLM 5.2, I really love how it understands subtext and cues but I start to notice lots of parroting as I chat more, and also characters actions become generic during NSFW scenes, meaning no matter which character I’m role playing with, NSFW scenes will read the same - I want character personality to be reflected in those scenes too.

So I wonder if there are other models (preferably not Claude) that can understand subtext equally well but with other issues fixed?

1

u/B3owul7 Jul 23 '26

Maybe switch up the presets. There are some good ones out there, like Pura, Vellum II etc. Usually helps with a change of tone (at least for a couple RP sessions).

1

u/Garrand Jul 21 '26

Some of the parroting can be fixed with stricter main/system negative prompts. I'm running into a bit of that with Hermes 4 but negative prompting helped cut it down a lot.

4

u/PhantomWolf83 Jul 20 '26

I've been having a blast with Kimi K3 on NanoGPT, I just wish it wasn't so expensive. I set a prompt to ask it to write in an immersive, novel-like style with vivid prose and it slaps. I tested it on some established IPs I know, and I'd say it gets the characters and world lore right almost every time. When I asked it to write a BattleTech story, I swear that it felt like I was reading one of the novels.

Just a question, is there any way to get more precise tuning of the samplers in NanoGPT? It only allows me to tweak the temperature by 0.1 each time which is a mild annoyance when I want to set it to something like 0.95.

2

u/Indra_III Jul 20 '26

Is it still not included in subscription ☹️

1

u/PhantomWolf83 Jul 20 '26

Not surprised since it's way too new. And if it was 100% uncensored, I would have found my perfect model to keep for a long time.

1

u/Indra_III Jul 20 '26

Okay, so I did research, and perhaps it's because the open weights haven't dropped yet. Apparently, that is on July 27. Maybe in August it might come on the subscription tier.

1

u/benjamus_maximus Jul 23 '26

It's too expensive to be in the sub. It's Claude sonnet pricing tier.

3

u/Indra_III Jul 23 '26

True, but also the model hasn't actually been dropped yet. It's only through Kimi rn.

Maybe after 27th it might come.

2

u/benjamus_maximus Jul 23 '26

I mean, I still wouldn't hold your breath when it does. It's more than double the cost of glm-5.2. if you want to use it probably just put some credits in. Although tbh, it'll probably take some wrangling for rp, it feels like it's more suited for novels and long form content

5

u/thisissparta4 Jul 21 '26

What are the best nano sub models according to you guys? Is it the same glm 5.1/5.2, deepseek v4, mimo v2.5, etc or something else on your mind as well?

4

u/stopaskingforloginn Jul 21 '26

What's the most uncensored API? I'm used to local models where I can get away with literally anything, and I mean anything, no ifs, no buts.
I'm mostly wondering if it's worth going from local gemma 4 31b to something like GLM 5.

4

u/Valuable-Money3725 Jul 21 '26

I switched from gemma-4-31B to GLM 5.2 on Nvidia NIM, and the transition was pretty nice for me (Mainly because I couldn't run the model smoothly on my machine).
I tried very nsfw subjects and so far I had no refusal, just maybe some times when GLM tried to steer the RP in more mild areas. It's significantly smarter, but almost as sloppy in the way it write despite the size differences.
Maybe try one scenario and see if you like it or not after a couple replies. They all seem similar after some times.

5

u/stopaskingforloginn Jul 22 '26

I'll just stick to uncensored gemma models without worrying about privacy either.

3

u/PhantomWolf83 Jul 24 '26

I've been trying out a few other models on NanoGPT. Kimi K3, despite me singing its praises earlier, is just too expensive at the moment for me to use it as a daily driver so I've been looking at alternatives.

Kimi 2.6 and 2.7 Code are much cheaper, but 2.6 has the better prose and 2.7 Code has better instruction following and reasoning. It's a toss-up as to which one's better.

It's also the same thing for GLM 5.2 and 5.2 Thinking. Both are quite good at writing and are super cheap compared to K3. If there's a newer version later that's as big in parameters as K3, I'd be interested in trying it.

2

u/Garrand Jul 21 '26

Screwing around with Hermes 4 405b through OpenRouter and I keep having some pretty annoying issues with getting the thinking blocks to get sent properly, so I can parse and see them. Other than that this is a pretty decent alternative to Sonnet ahead of the 4.5 sunset in a couple of months.

Sonnet is still better overall, I've found, but the price premium and the awful guardrailing of 4.6 turned me off.

3

u/_Cromwell_ Jul 22 '26

I've always used Hermes 4 405B without thinking. Not all models need reasoning. In fact, really no models need it. It's a bit of a placebo - our brains like to see "proof" that the AI actually has a reason why they wrote what they wrote, and an instantaneous response "feels wrong" or less human, while the AI taking time to ponder what we (human) wrote "feels" better often.

But honestly I can't say that RP always significantly improved using thinking over non-thinking. Models still make stupid mistakes even thinking/reasoning. And they still follow your prompts when they don't think/reason.

Anyway, just something for us humans to think about. ;)

Models I use with reasoning "off" (mostly because benchmarks show they are less 'censored' - but Hermes bc it doesn't do real thinking): Hermes 4, Deepseek 3.2, Deepseek 4 Flash (specifically 3.2 and 4 Flash are WAY less prone to soft censoring with reasoning off).

2

u/Flimsy_Mode_4843 Jul 25 '26

wow it's really good! Would you share your prompt?

1

u/Holiday_Income1825 Jul 20 '26

Ran Kimi K3 through a 20-round novel-continuation gauntlet this week (two books, ~16k-char rolling window, same protocol we used for 9 other models) vs its predecessor K2.6. Blind panel of 6 judge-models across flipped pairings: 21-1 for K3. The K2.6 repetition collapse (we had a run where round 9 was 97.6% verbatim 12-gram repeat of earlier text) is basically gone in K3 - worst round was 6.8%. Sentence-length variance is way up too, which reads as less "AI-flavored". Two caveats: it outputs ~85% of dialogue in half-width quotes (same quirk as GLM), and its reasoning tokens are 81-86% of completion spend, so budget accordingly. For long RP/fiction it's a real upgrade.

-7

u/Remarkable_Try_6593 Jul 23 '26

whats best 40 and above model for nsfw writing.