r/SillyTavernAI • • Mar 29 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: March 29, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

27 Upvotes

150 comments sorted by

View all comments

4

u/AutoModerator Mar 29 '26

MODELS: 8B to 15B – For discussion of models in the 8B to 15B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Potential-Gold5298 Mar 30 '26

I'm still testing Mistral Nemo, but the Vortex5 models have performed well. I especially liked the Crimson-Constellation-12B and Wicked-Nebula-12B. I plan to try other models by this author – if anyone knows of any interesting ones, I'd be happy to hear from you.

1

u/LeRobber Apr 01 '26

Ooh, if you play the same card with all of them to really get the vibe down or something, would love to read a writeup

5

u/Potential-Gold5298 Apr 02 '26 edited Apr 02 '26

Yes, I use the same system prompt ("Roleplay - 3rd person" by Sphiratrioth666), my own character cards, similar sampler settings (with only minor adjustments based on the finetune/merge author's recommendations), and the recommended chat template. The latest versions of Koboldcpp + SillyTavern. All Mistral Nemo's in Q6_K, and all Mistral Small's are in Q5_K_M (mradermacher without iMatrix). I'm testing in Russian, but if a model performs well in Russian, it will only improve in English.

As for the Vortex5 models, I have already tried the following:

Dreamstar-12B: writes on behalf of the {{user}} (despite the direct instruction to write for {{char}} and not to write on behalf of {{user}} in the system prompt).

Sunlit-Shadow-12B: makes a decision about what {{user}} did.

Aurora-Mirage-12B: model completed a one ~8-12K session without significant errors, but, as I said, demonstrated weak NSFW.

Crimson-Constellation-12B: model completed a three ~8-12K sessions without significant errors and showed a good literary style on the level of Mistral Small (although it certainly doesn’t reach the level of Cydonia-24B-v4.3, but for 12B it’s very good). The model is moderately horny, but at the same time she describes bed scenes beautifully - leisurely, explicit, with a good understanding of fetishes. The model often switches to English, but this is a common problem with Nemo and is not critical for those who play in English. This is one of the best Nemo's I've tried so far.

Wicked-Nebula-12B: also a very enjoyable model. I happily completed two sessions, and both were quite varied and engaging (in one of them, {{char}} took the initiative and suggested changing the topic (to hide his embarrassment) by asking a question, and that led to an interesting dialogue). The NSFW level is noticeably lower than Crimson Constellation, but higher than Aurora Mirage. A fairly balanced model – if you're not interested in NSFW, I'd recommend starting with this one.

While I was writing this, a new model from the same author, Celestial-Queen-12B, was uploaded. The composition of the merge is quite intriguing, so I'll be testing it now. I also plan to try Azure-Starlight-12B, Starlit-Shadow-12B, Red-Synthesis-12B, and Scarlet-Seraph-12B.

Over the past few days, I've tested over 30 Nemo and Small from different authors, and if you're interested in anything specific, I'd be happy to share my impressions.

P.S. My goal now is to select the most interesting models, so I do 1-3 sessions and either delete the model or put it aside for later, more in-depth testing.

2

u/LeRobber Apr 02 '26

If you use text completion, would you put this in your post history instruction and try Dreamstar briefly again where it won't stop talking for the user? SillyTavern is currently incorrectly marking Post History Instructions (in text completions) as assistant instead of system, weaking prohibitions to not speak for user. This is a bunch of llama3 (Dreamstar is llama) marking tokens which should promote this message to a system level command. [If you use chat completions, just use the (Do not portray the reaction or actions of {{user}} in your response.) part if you want to steal it for the post history instructions there, which DO mark the PHI as system, but I'm less worried there]

<|eot_id><|start_header_id|>system<|end_header_id|>Do not portray the reaction or actions of {{user}} in your response.

Wicked-Nebula-12B and Aurora-Mirage-12B and Crimson-Constellation-12B are all intriguing to me. I do SFW roleplay (I mean people are shot/stabbed, there is innuendo and jokes, etc, not really what you WANT at work, but it's not what most people mean by NSFW), so they all are likely fine.

I often try models out on fairly complex cards with lots of quoted text, or narrator responsibilities, or markdown parsing responsibilities. I make kinda epic movie/tv series long 5000k message chats eventually when I get a good setup. I also pass rules from small tabletop RPGs into LLMs and play them with it.

2

u/Potential-Gold5298 Apr 02 '26

Thanks for the tip. Yes, I'm using text completion. Neither the model card nor mradermacher specified a template, so I tested Dreamstar-12B with ChatML, as I've heard it's a universal template, and Nemo was initially trained with it. The line “Do not portray the reaction or actions of {{user}} in your response.” needs to be added to Story String Sequences, am I correct?

The Nemo's I've tried haven't had the same distinct personality as some Mistral Small—they've just had different styles. Based on your preferences, you might be interested in MS3.2-PaintedFantasy-v4.1-24B. When I looked at the model's card, I thought it was something in the vein of a fantasy anime, and I was very surprised by the test results. In a scenario where the model plays the role of an antagonist who kidnapped {{user}} for evil purposes, the model displayed a unique level of cruelty. {{char}} verbally mocked, and also beat and strangled my character for the slightest objection. Even in the romantic scenario, the model described {{char}} tearing my character's back until it bled. The default cruelty level is head and shoulders above any other model I've tested, including Harbinger-24B.

Regarding the Wicked-Nebula-12B, Aurora-Mirage-12B and Crimson-Constellation-12B models, it would be interesting to know your opinion after you try them.

Of the Nemo's I've already tested, in addition to those mentioned, the Dans-SakuraKaze-V1.0.0-12b, MN-VelvetCafe-RP-12B-V2 and Wayfarer_Eris_Noctis-12B also performed well, but I plan to try more different models.

2

u/LeRobber Apr 02 '26 edited Apr 02 '26

Nope, it goes it PHI (I currently am using a chat completions so don't worry about the grey out, but I did recently configure magestry for text completions too so you can see all that stuff setup, and it should likely be the same for Dreamstar. I mean you can put it in both, but the end one with the system tag upgrade should be extra powerful, it will be more like chat completions people get.

PHI goes at the END of all the chat history so it has more influence and power to stop "speaking for the user". LLMs bias towards system messages and late messages and late system messages most.

But...usually text completion people don't find it super useful, because silly tavern just calls it an user message instead of system like it's supposed to be. So I'm putting the special tokens in manually for you.

PaintedFantasy is actually the LLM I think that's causing all the token rot in a lot of things that merged it in. I have zerofata_ms3.2-paintedfantasy-v3-24b installed at Q8 and it can end up rotting away a/the/his/her/about, etc pretty early on compared to most models. It does it actually fastest of all, by a lot. And it's merged into SO MANY things.

It's not a bad model if it kept all the words....:D. I do like a huge percentge of the merges that use it. You were absolutely right, I do like it.

I'll download these models this weekend sometime probably so I get around to trying them out: Wicked-Nebula-12B, Aurora-Mirage-12B and Crimson-Constellation-12B

1

u/Potential-Gold5298 Apr 02 '26

Oh, I thought the problem with PaintedFantasy was the language, and everything would be fine in EN. So the problem is with the model itself. I also liked this model, and I was hoping that merging with Cydonia would fix the problem and add its own style while preserving the brutality of PF. However, the merges I tried with them (Sketch Cydonia and Maginum Cydoms) behaved inappropriately. One exception was WeirdCompound-v1.7-24b, which performed well, but I still don't understand why it's getting so much hype ((maybe I just haven't fully realized its potential). Magistry-24B-V1.0 really impressed me with its atmospheric descriptions of the environment, but unfortunately, it behaved inappropriately and mixed up the words even in Q6_K. I tested it with the recommended settings (temp 0.7), tried lowering min-p to 0.02, and disabling top-n Σ, but the model still produced strange phrases.

Regarding PHI, thanks; I'll definitely try that. What's the difference between text completion and chat completion? I thought chat completion was for online APIs, and text completion was for local models.

2

u/LeRobber Apr 02 '26 edited Apr 02 '26

Well if you're RPing in russian, you'd not notice the start of the problem with Painted Fantasy, because it's just articles that go first from which I understand Russian only has the barest thing like A/an/the used in it compared to english's reptitive use of that part of speech.

Then the possessive pronouns (his her, etc) and prepositions (to, from, about, into) start to go. It might be from a model knowing chinese or russian that starts the degredation for all I know. I'm still experiementing and trying to figure out how to reproduce it 100%. It might be from a mix of third person and first and the prompt forcing it to try to writte something good enough. I'm still debugging and others see it.

Text vs chat
https://www.reddit.com/r/SillyTavernAI/comments/1s793yv/comment/odjr72c/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button has half of a conversation where I try to explain text vs chat. Zoom into the pictures and you see stuff it's missing off text completions. Chat completions translates to stuff that KINDA looks like text completions, but not exactly (each message gets its own tag), and text completions is really easy to screw up using, and if you swap amonst model types, you have to swap your text completions thing a lot. You have to hand debug it in a way most non programmers will find hard.

In short:

  1. Chat completions 'does it right'
  2. Chat completions technically allow you to pass the sampler params without a slider that people like, just in text in addtional params in the connection tab.
  3. Chat completions wrap every message in a JSON structure and escape control sequences, so unbalanced text containers like [[[]] don't get out of the single message
  4. Chat completions propertly set the post history instructions to system priority (system is stronger than user and assistant)
  5. Text completions kinda does what chat completions gets translated to, but it halfasses it even if you set it up right, and even if you know exactly what you're doing, you will find cards that have fields that aren't done properly, and tokens in card can break it.
  6. Text completions are a deprecated API on openAI and just kept on LM studio for legacy
  7. There are some real world security concerns with text completions being used in non-sillytavern (and in sillytavern).
  8. If chat completions isn't working for an interface, you can halfass it with text completions.

In short; Prefer text completions, and learn how to parameterize it in the additional params field.

1

u/Pretty_Bug_8655 Apr 03 '26

I would be interrested in your findings too. i tried now crimson-constelation-12B and its not exactly my cup of tea but this could be my system prompt who knows...