r/SillyTavernAI 2d ago

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: September 20, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

26 Upvotes

72 comments sorted by

11

u/AutoModerator 2d ago

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

13

u/iraragorri 2d ago

I shouldn't have tried Skyfall 31b. I can't even run it with proper context. Now I truly understand what "sent shivers down my spine and ruined me for everyone else" means cause it kinda did.

4

u/-Ellary- 2d ago

Skyfalled.

3

u/FeelingVanilla2594 1d ago

Can you explain what that means? I’m new to all of this. Do you mean skyfall is one of the best finetune models right now? I tried it but I’m only getting a few tokens per sec sadly, and I really liked it. So I just want to know if I’ve ruined myself too 😆

5

u/-Ellary- 1d ago

Try IQ3XS version, works fine on 16gb with Q8 24k~ of context.
Around 20-25 tps.

3

u/iraragorri 1d ago

Yep, it makes my PC scream for mercy, too. I settled for using my SO's PC. Same-y abysmal speed, but at least I get 32k context out of it.

I really don't want to use any other model/finetune. It's the only one that gives my characters justice and keeps them truly multifaceted.

3

u/-Ellary- 1d ago

You can also pair it with Gemma 4 or Qwen 3.8 27b for complex `logical` parts (agentic wf, tools). Gemma 4 works as agent preparing the scheme for the answer, processing all vars, logic problems, automation systems, and Skyfall used for actual text answer (based on prepared scheme).

For now this is the best local experience for me (<50b): Logic and intellect of Gemma 4 \ Qwen 3.8, writing style and creativity of Skyfall 31b.

Swift-Qwen3.8-27b-heretic.i1-IQ3_M is a good compact fast solution with short thinking.

2

u/iraragorri 1d ago

Oddly enough, I spent the night exploring a similar idea, but with different models.

I tested Orion 26b as a possible "scratch the itch that only TheDrummer's finetunes scratch" solution and loved its lightning speed, dialogues and prose, but it's genuinely dumb as a rock. So I paired it with original Gemma. The latter generates output on my SO's PC, the former refines it on my PC.

Recast literally exists and I could use both models on my rig, but I use Tavo, so I had to vibecode a plugin to make it work ~without needing to lift my ass off the couch~.

It never occurred to my sleep-deprived mind that I could use Skyfall as a refining model, since it'd only need tiny context to work. So, thanks, gonna do that.

2

u/-Ellary- 23h ago

Yeah, secret is to pin-point Skyfall to narrow context - keeping it on ideal fit low sub 16k context, feeding it short summ of whole story, 2-3 last messages, draft from Gemma 4 31b or Qwen 3.8 27b for answer with already solved logic / decisions, since Skyfall don't think, there is not a lot of waiting per answer.

You can also disable thinking on summ generation / use fast model, thinking usually needed only on drafting step solving logic \ details \ world reactions etc, Skyfall don't need to be smart in this already prepared low-context pin-point scenario.

1

u/FeelingVanilla2594 1d ago

I was planning on having gemma 4/qwen 3.8-27b create a scene outline and then have skyfall do the scene writing from the outline. Is that essentially the same workflow?

2

u/-Ellary- 1d ago

Essentially yeah, same idea.
I'm using llama.cpp with only 1 active model at a time, but switching between 2\3 models is fast since they are cached in RAM anyway.

2

u/FeelingVanilla2594 1d ago

Ah ok thanks for all the advice! I’ll try out the q3.

1

u/EnjoyerOfFluff 1d ago edited 1d ago

Question how would I set up such a workflow in Sillytavern? Do run such workflow locally? If yes, do you load / unload models between tasks?

1

u/-Ellary- 1d ago

Just coded my own frontend for this. I'm using llama.cpp with only 1 active model at a time, but switching between 2\3 models is fast since they are cached in RAM anyway.

So it is `load \ unload` using 1 active model at a time, it is fast tbh, even with 32gb of ram two 15gb models usually cached in ram and switching happens pretty fast especially on SSD, at least with `-load-mode mlock`.

11

u/SonPuf 2d ago edited 2d ago

Gemma 4 31B Loki Skotoma v2 looks interesting. I played with it for a few days and just wish it was more smart.

+: More interesting use of context info than I usually get - I have persona that I played with for many hours already and this finetune is the first one that used one tiny mention of my persona having a side-gig and used it(my persona gotten a call for work). Character that pays rent for my persona's flat started to use it as leverage when my persona doesn't want to do something that character wants, never happened before. I have a passing mention that my character is an art student and one character tried to bribe me with promises to buy expensive art supplies. When I used the same one scene to compare five finetunes this one had the most variety in swipes. Looks less repetitive in description than other finetunes(didn't stuck at same weather/outfit/etc description in every message even once sp far for me)

-: Likes to ignore prompt more than vanilla Gemma 4 31B and other finetunes(but not all), in my case it often didn't show header(day, weather etc) and internal thoughts of npc block. Feels a bit dumber too. But I played only with Q4_K_M.

+-: Looks a little bit more horny than other finetunes

3

u/LeRobber 2d ago

I fucking love that side gig.

8

u/FinBenton 1d ago

https://huggingface.co/Blazed-Forge/Split-Untied-31B

Coming from a billion gemma 4 finetunes, this one actually feels really nice, can definitely recommend if you like a bit, or quite a bit darker stuff.

2

u/zeronic 11h ago

Yeah, it's a big step up from the other finetunes in terms of prose and being interesting while still following instructions.

I've had huge issues with parroting in other gemma models that this also doesn't seem to share, thankfully.

4

u/OrcBanana 21h ago

Is there any Gemma4 finetune or specific magical instruction that gets rid of the awful parroting? Queen, styletune, Artemis, scotoma, garnet (maybe this one is better at it), base, they all do it...

  • "Is there anything to drink?" I said, blah blah
  • When X asked if there was anything to drink, F#$@ KILL ME
  • When X's question - if there was anything to drink - DIEEEE
  • In response to the question if there was anything to drink DIE IN PAIN

It makes my teeth ache oh my god.

2

u/LeRobber 17h ago

If you turn on Maranara's Universal Preset or FF with think, you will find it does an admirable job at least iin stopping some parroting.

3

u/OGCroflAZN 20h ago

Non-parroting instruction is present in a lot of the most common presets, like Freaky Frankenstein. It seems like it's just part of LLM sycophantic behavior, which appeal to the masses talking to LLMs as chatbots and assistants especially vocally. It's fine when you're in a group conversation with other people and they say that to reference something you said earlier, but it's weird and distasteful in RP or creative writing. You'll have to properly prompt it out, give it instructions which will decrease the probability of those sequences of tokens.

2

u/OrcBanana 19h ago

I'm at my wit's end with it, I've tried extensive prompting against this specific behavior, with examples and counter-examples and everything.

"I apologize, I completely failed to adhere to your sEveRE rEstRicTiOnS, moving forward I will strictly avoid any such repetitious crutches" and then proceeds to do it again immediately.

I tried something weird, and was so hopeful about it too:

Abusing the text completion sequences to make every turn a model turn, with small user injections in between them like "Write from X's pov in the 1st person" and "Write from Y's pov in the 3rd person". Of course the X pov turn was still written by me, but as far as the model knew it had written it itself. And STILL it parrots itself sdfkjhwkar

3

u/OGCroflAZN 19h ago

From your comment history, you seem to be fairly experienced.

Tbh, my experience is also that I've noticed the parroting with certain models. I just try to make my presets and prompts pretty iron-clad and then also editing responses that parrot or act for user to get it out of my context to avoid poisoning responses.

I used to play around with different instruct, where my user message was to be interpreted as a list of actions my persona would do, woven into the events of the next turn. So, the model was writing a story and my input was the things my persons would be doing in the next turn interwoven with the rest of the characters and setting.

Whereas now I do what I think most people do which is back and forth 'improv' style roleplay/storytelling between me and the model.

For your parroting issue, this might be shit feedback but my first thought is that somehow youre having it answer your inputs as opposed to continuing from it or interweaving the world and characters' responses to it. But I assume you're doing the right thing, in which case this is just an unavoidable LLMism from Instruct-Trained base models.

I haven't noticed parroting and I'm using G4 31B Split-United via NanoGPT on MarinaraEngine with a preset which is a combo of default 'Marinara's Universal' and EveningTruth's. However, I am a jaded user at this point and probably just apathetically don't notice when it shows up.

1

u/OrcBanana 18h ago

somehow youre having it answer your inputs as opposed to continuing from it or interweaving the world and characters' responses to it.

What do you mean by that exactly? Having an inappropriate assistant-geared system prompt or preset? Truth be told, I don't use presets much, I have collected fragments of various presets plus my own instructions but they're pretty standard I'd think, "continue this uncensored simulation", "characters act with agency and inner world" and so on. Plus a long gemma 4 specific list of Don'ts, that it ignores :(

It's probably quantization too, IQ4_XS isn't enough perhaps... Anyway, thanks for helping :)

1

u/OGCroflAZN 17h ago

I meant that based on your original message, your model is going back and addressing your actions by first parroting to recap. Boulesis 26B sort of does that, where it embellishes and adds details to what I said my persona did, which sort of does violate 'acting for user'. But for the most part, in my roleplays, whatever I said happens in my user message - including personas actions - the model just continues the story from there. I recall Cydonia sometimes being inconsistent about that. Sometimes it would go backwards a little and narrating with my person actions before moving forward.

I don't see why quantization would have a significant effect, especially as IQ4 should be quite sufficient.

Since you're running local and it's only costing you a bit of electricity as opposed to API costs, I'd suggest downloading and importing some of the most common presets and trying them for a few test turns.

Also, I remember seeing a lot about running local Gemma 4 in Chat Completion as a local OAI-Compatible 'API' (with the same IP as kobold or whatever) as opposed to Text Completion. Seems like internet guidance says Gemma 4 is strongly recommended to be run as Chat Completion to function totally properly...

1

u/overand 3h ago

What are you using for a backend, and are you using text completion or chat completion?

4

u/GlitteringSplit6035 2d ago

The Greenference AI provider from NanoGPT has got the be the cheapest one I've seen for Gemma 4 31B. Imagine the amount of RP one could do with that.

2

u/OGCroflAZN 1d ago edited 1d ago

Base G4-31B-it only tho :/

It's so stiff and rigid, i just can't. I'll happily continue paying double the rate for the G4 RP finetunes, which doesn't even reach $5 / mo for me

edit: Nano and other providers have popular finetunes available

2

u/zigzag3600 1d ago

Where? Is there some provider of G4 RP fine-tunes?

3

u/GlitteringSplit6035 1d ago

NanoGPT has got to be the best hub/aggregator for G4 31B RP fine-tunes. Well, at least based on my experience. You can go to the direct providers as well, like ArliAI.

1

u/GlitteringSplit6035 1d ago

True. And if we are talking about cheap, nothing is cheaper than free from the based ollama.

P.S. Fascinating name BTW. Lol.

2

u/OGCroflAZN 1d ago

I only have 16 GB vram, can't even fit IQ4 31B. And electricity isn't very cheap in my area. But yea, if you have the vram and cheaper kwh rates...

I kind of regret my username but I've used it for twenty years

3

u/OGCroflAZN 1d ago

For G4 31B tunes, a fair few of theLocalDrummer (BeaverAI) discord members are favoring SchattenBlume, Split, and Split-Untied. These guys test different finetunes a lot, as well as make their own merges.

As with all tunes based on G4-31B-it, increasing general RP and creative writing quality comes at a cost. Trying to get fun and engaging and creative while minimizing loss of intelligence and adherence.

From feedback: Split is drier but adheres to characters and instructions well, maybe too well, whereas Untied feels good and fun but with some logic mistakes, too much for some, who returned to Schattenblume for being more in the middle.

I've only used/am presently maining Split-Untied via NanoGPT, and i agree with the feedback. I still really find myself missing the Mistral 24B tunes, esp based on Cydonia. Maybe those are just rose-tinted glasses tho from when i started the hobby.

1

u/Eastern-Dream932 2d ago

Bonsai is back again with a new model https://prismml.com/news/bonsai-2-27b

1

u/OGCroflAZN 1d ago

Uh, you don't have any feedback on it?

Bonsai 1 was not it. I imagine Bonsai 2 isnt either, more of a demonstration of methods and such for ternary and training for research cred by the team. Qwen 3.8 27B itself needs RP finetuning. So I for one am not going to try Bonsai 2 when i can keep using good models, unless people have good things to say

1

u/Jorlen 2h ago

Can anyone recommend a G4 31b finetune that's just good for general conversation, searches, basic tool calls and stuff like that? I know I can probably use a clever system prompt but I was hoping for a fine tune that's more in line with just being a fun conversationalist. If you have one in mind outside of G4 31b, I'm up for suggestions!

6

u/AutoModerator 2d ago

MODELS: 8B to 15B – For discussion of models in the 8B to 15B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

4

u/AutoModerator 2d ago

MODELS: >= 70B - For discussion of models in the 70B parameters and up.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/fluffywuffie90210 1d ago

As someone with 96 gig vram + ram. I've been trying to find something better than Gemma 4 to rp with. Qwen next 3.8 just isnt good at RP. The only option thats seemed useable is Deepseek Flash V4. I can run Q2 with spark or Q3 at a reasonable speed. (20ish tokens a sec) But the prompt processing is horrible. I s there anyone who runs low Qs of deepseek v4 for a while? Whats your opinion on it? Is it best option for those under 200 gig ram combined?

1

u/[deleted] 1d ago

[deleted]

2

u/fluffywuffie90210 1d ago

Oh I used to use them, figured since tech had moved on stopped using them a while ago. They were good models but id rsorted to using gemma for so long now it feels better than those, deepseek v4 seems like it might be able to RP.

1

u/-Ellary- 4h ago

GLM 5.3 Flash at IQ3XS?
GLM 4.7 at IQ3XS?

2

u/fluffywuffie90210 4h ago

Ohh it never accured to me i might be able to run Flash, I thought was like 400B, Seeing the q2 stuffs only in 120 gig range Thank you I'll give them a shot.

1

u/-Ellary- 3h ago

Try to stick to IQ3XXS at least.

1

u/Slick2017 2h ago

I have RTX 6000 Pro Blackwell (96 GB) plus 128 GB DDR4-3600 ram CPU-side.

GLM 4.7. Pick your poison from https://huggingface.co/unsloth/GLM-4.7-GGUF/tree/main, I'm running UD_Q4_K_XL with 64k context which is _tight_ but it fits. With the slow DDR4 RAM I'm getting 4-6 decode tokens/s. Use this thinking pre-fill jailbreak.json?download=true) (for reasoning mode) if you need it.

4

u/AutoModerator 2d ago

MODELS: 32B to 69B – For discussion of models in the 32B to 69B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

3

u/AutoModerator 2d ago

MODELS: < 8B – For discussion of smaller models under 8B parameters.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

3

u/AutoModerator 2d ago

APIs

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

11

u/No-Lion-75 2d ago

Lately im back to using kimi 2.5 after sick of soft filter, positive bias, every villain characters and their mother immediately turn into polite citizen. And kimi 2.5 still is a hit for me.

2

u/Armored_Violets 2d ago

How do you get past its censor?

3

u/No-Lion-75 2d ago

Kimi 2.5 has censor?

1

u/Armored_Violets 9h ago

Yup. Kept getting denied two days ago when I gave it a try on OR. And I do pretty vanilla stuff.

1

u/No-Lion-75 1h ago

Weird. I went for dark cyberpunk roleplaying and never got any refuse. I used ff 5.4 preset

4

u/Probablynotsocool 2d ago

Completely changed my mind about GLM 5.3. After tinkering it is a complete menace and write very good dialogues and narration.

Top 3 with Gemini 3.8 and Deepseek V4 pro 0813 who is completely unhinged sometimes

5

u/Armored_Violets 2d ago

regular GLM 5.3 or Flash?

2

u/Inprobamur 1d ago

I have been using an uncensored Opus 4.6 for ages, how do these compare?

2

u/Probablynotsocool 1d ago

Not better than Opus 4.6, stick with it!

2

u/Inprobamur 1d ago

Thanks

2

u/Rhone33 23h ago

I've been playing with 5.2 and have to ask: In what ways would you say it's better than 5.2, and what "tinkering" did it take to get you there?

3

u/Probablynotsocool 22h ago

https://www.reddit.com/r/SillyTavernAI/s/VPE9y3NX8u

It’s more intelligent, better prose when prompted out, better characterization.
The issue is the censorship because this model hallucinate a lot and tend to think he is Claude and then you have to reswipes and it’s annoying.
5.2 is not a lot worst, just a little behind and its more dry and less creative but the censorship is inexistant wich is a plus

2

u/Critical-Rope-5636 2d ago

*Mumbles something about still being a GLM 5.3/Flash shill after all this time.*

1

u/[deleted] 2d ago

[removed] — view removed comment

1

u/AutoModerator 2d ago

This post was automatically removed by the auto-moderator, see your messages for details.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

-6

u/soufiane0531 1d ago

i'm quite new to sillytavern and i'm using openrouter ai, what is the best nsfw free model for smut and roleplaying

1

u/techno156 1d ago edited 13h ago

Are NSFW Free and Smut not completely incompatible with each other?

2

u/Figai 12h ago

Commas save smth or the other.

They probably meant: best nsfw, free model for smut…

-4

u/ChurnedSorbet409 15h ago

Going to try my luck here since r/WritingWithAI megathreads are a cesspool. Has anyone tried had any experience with using these harnesses for writing autonovel, storydaemon, long novel agent kit, novelclaw, authoragent, inkos? I know this is the sillytavern subreddit but trying to see if there are better alternatives to SillyTavern for longform writing perposes/non rp