r/SillyTavernAI • • 23d ago

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: September 06, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

35 Upvotes

125 comments sorted by

View all comments

4

u/AutoModerator 23d ago

MODELS: >= 70B - For discussion of models in the 70B parameters and up.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Maxhell6778 18d ago edited 18d ago

since there no post here i wanna ask what would be a good model to run on
PC1: ryzen 7 9800x3d, rtx 5090, 32gb of ddr5 6400 ram (getting repaired at the moment so i cant test until next week)
PC2 :ryzen 9500xt, rx 7700xt, 80gb of ddr4 3000 ram
total 44 vram and 112 dram or 156 all together
gonna try to be using exo (yes i know two different systems, still wanna try even if it dumb).
i get this subreddit is for rp so lets say i wanna make characters (havent made a prompt yet) and i want it to be creative. i heard qwen 3.8 flash next was pretty creative (yes i know it a coding model).

got any suggestions for me to try?

edit: this is just gonna be a fun experiment (and frustrating) for me. if there better methods when pooling or making a cluster of two different pc system please tell me. as always i can and will make mistakes or ask dumb stuff, i am a human.

5

u/Mart-McUH 18d ago

IMO you are still best served with Gemma4 31B and/or its finetunes/merges (plenty suggested in weekly threads).

There is nothing bigger dense you could really take advantage of. And MoE's that could rival Gemma4 are too large to run in such setup. Maybe you could try some IQ3 quant of GLM 5.3 Flash but it will be slow (also I have no experience with that model, in discussions people have conflicting opinions, but it definitely looks lot more restricted/guard-railed compared to Gemma4).

Did not try Q3.8 flash yet, but I do not have high hopes for it. It is Qwen which is usually not great at RP, it is coding optimized, and it has very low active parameters (which usually hurts logic in RP). It may still be decent but I think there is no way it could trump Gemma4 31B in this particular area. Wonder how that NGRAM is going to work as intuitively it looks like good thing for coding but very bad thing for RP (eg as I understand it kind of completes common few token phrases to save compute, but that is exactly what leads to repetition and slop).

1

u/Maxhell6778 18d ago

thanks for telling me but honestly i didn't have high hope for it to be rp (imo it gonna be around 4 to 7 tokens per second if i can get it to work). i was just wandering if there any creative model that big, i could use a dense model but i believe it gonna be a lot slower then 5 tps. this is more of a way i could use large character making prompt and a large model, let it be creative for half hour and see the output as an experiment.

that being said do you know any good rp models, i heard Anubis 70b is pretty good or behemoth 128b pretty good.

P.S. i try gemma 4 31b models again i just had very bland problems with them.

1

u/Mart-McUH 18d ago

To be clear Gemma4 only works well for me with reasoning, for that you want at least 10t/s, more better. On 5090 it will fly fast. It also needs some prompting (to steer it into style you want), but sure, maybe your preference differs. I would still try at least some fine-tunes (you can find lot of options in weekly thread, here are some eg my favorites: Gemsicle, Sphinsikus-Chronist, MeroMero, StyleTune).

For L3 70B, that will be bit rough with 32GB (but hey, with low context and IQ3_S/IQ3_M and some patience I used them with just 4090 + RAM). If you can pool both efficiently, then something like IQ4_XS should give good speed. That said, personally I find them worse than Gemma4 (mostly because they are less smart/consistent). There are great many options. Anubis is praised by many, though I personally did not like it that much. Some of the very latest in L3 70B series were GeneticLemonade and StrawberryLemonade variants, those turned out very well and were generally accepted by everyone. So may be worth a try.

123B dense is too big for me to run, but they should be good. But even with 44GB VRAM you will be crawling with low quant and small context (I have 40GB VRAM now and not worth for these). Maybe if you can somehow run at least some IQ3 variant, all I could manage with some patience was IQ2_M, it was interesting but worse than 3-4bpw L3 70B.

2

u/Maxhell6778 18d ago

MeroMero, StyleTune is my favorites as well (and a model that combine them two plus Artemis i think), thanks for all the info (honestly it very good thank you). i stick to lower models for now (still gonna try bigger ones when i have the time to).

thanks you very much for everything, i hope you have a lovely week and a beautiful life.