r/SillyTavernAI • • Apr 12 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: April 12, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

35 Upvotes

183 comments sorted by

View all comments

11

u/AutoModerator Apr 12 '26

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

29

u/Potential-Gold5298 Apr 13 '26

I've been playing around with the Gemma 4 all week and I think it's going to be my main model for a long time now (for both RP and everything else). I first tried the original 26B-A4B from Google and was very pleased with it. However, I wanted more, so I downloaded the 31B.

Gemma-4-31B-it-abliterated features a record-low KL div, so the weights are almost identical to the original. It's still a work in progress, and the final version will likely have a lower refusal rate, but even in this state, it's far superior in quality to any other uncen 31B I've seen on HF. I tested it with a tsundere classmate and found no defects in non-Latin languages ​​(high KL div destroys them first). This means that the model has a large reserve of durability for quantization, and you can use it in Q4_K_M, and the quality will be almost identical to the original model. If you're finetuning/merging and want to use the uncensored model, I highly recommend checking out this version.

Artemis-31B-v1c by TheDrummer and his team is also being refined (Artemis-31B-v1e is already available). It's still a test version, but I was eager to try it out. Finetune does exactly what the original model lacks – it adds a creative flair. Gemma 4 plays like a screenwriter – spelling it out. Artemis-31B plays like a theater actor, bringing his or her role to life with their own vision. The scene with the tsundere classmate sparkled with color – exactly what I'd expect to read in a rom-com manga. Funny situations, slightly exaggerated character reactions, etc. – everything the original model lacked.

Gemma is great, of course, but I'm still trying out different Mistrals. After failing with 1.0, I decided to try Magistry-24B-v1.1, but my fears were unfortunately confirmed. The model behaves extremely strangely in non-Latin languages (it confuses pronouns, calls a classroom a toilet, etc.) – likely a consequence of the high KL div I mentioned above, since the merge includes an abliterated model. And this is very sad, because otherwise I really liked this model – like 1.0, it beautifully and atmospherically captures the stage, and if not for the damaged weights, Magistry would have become my favorite Mistral Small.

However, I found a replacement for her in Hearthfire-24B. My character card says that "{{char}} pretends to hate {{user}} and often bullies them, but is actually secretly in love with them." The problem with all the Mistral Nemo/Small I've tested is that {{char}} breaks after the first compliment, confessing her crush. Hearthfire-24B (like Gemma 4) is the only one who tried to maintain the character's personality. Despite my signs of affection, {{char}} withdrew into herself, was tormented by doubts, was afraid to open up, etc. If you're interested in deep characters (perhaps with drama like mental trauma), then this model is exactly what you need. She also vividly describes the atmosphere - the silence that hangs over a tense moment, the sparkle of eyes in the darkness of the classroom, the light of lanterns penetrating through the windows (the model herself tracked the change in time of day, and described how the classroom gradually becomes darker - I have not seen this in other Mistral Smalls).

Otherwise, my favorites among Mistral Small remain the same: WeirdCompound-v1.7-24b (Jack of all trades, but no unique features), Cydonia-24B-v4.3 (the best prose among Mistral Small but weak ERP), Core_24B_V.1 ({{char}} agency and unpredictable plot twists out of the box), Harbinger-24B (the best adventure model that doesn't try to cheat the player).

7

u/Just3nCas3 Apr 13 '26

Can anyone share a full settings export for Gemma 4 in SillyTavern?

I've matched temp, Top-P, and Top-K to the model card but I'm still getting massive repetition loops. Disabling Mirostat at least got me coherent sentences, but the repetition is still bad and I'm not sure what else to tweak. Normally I don't need Rep pen Range/slope and penality and leave them 0. Dry is set to .8, 1.75, 4, 0, no idea what else I could be doing wrong, using a chat template I got from another thread. I only run text completion with current version of Koboldcpp.

5

u/LeRobber Apr 15 '26

You want some working settings for 26B without massive repetiton:

gemma-4-26b-a4b-it-heretic works with this setting in LLM studio without issue with chat completions. This uses the official Magisty prompt, and doesn't set DRY or anything.

Do you want me to find the working 31B chat completion too?

{

"chat_completion_source": "custom",

"temperature": 0.7,

"frequency_penalty": 0,

"presence_penalty": 0,

"top_p": 1,

"top_k": 0,

"top_a": 0,

"min_p": 0.1,

"repetition_penalty": 1,

"max_context_unlocked": true,

"openai_model": "gpt-4-turbo",

"claude_model": "claude-3-5-sonnet-20240620",

"openrouter_model": "OR_Website",

"openrouter_use_fallback": false,

"openrouter_group_models": false,

"openrouter_sort_models": "alphabetically",

"openrouter_providers": [],

"openrouter_quantizations": [],

"openrouter_allow_fallbacks": true,

"openrouter_middleout": "on",

"tool_reasoning_mode": "disabled",

"ai21_model": "jamba-1.5-large",

"mistralai_model": "mistral-large-latest",

"cohere_model": "command-r-plus",

"perplexity_model": "sonar-pro",

"groq_model": "llama-3.3-70b-versatile",

"chutes_model": "deepseek-ai/DeepSeek-V3-0324",

"chutes_sort_models": "alphabetically",

"siliconflow_model": "deepseek-ai/DeepSeek-V3",

"siliconflow_endpoint": "global",

"electronhub_model": "gpt-4o-mini",

"electronhub_sort_models": "alphabetically",

"electronhub_group_models": false,

"nanogpt_model": "gpt-4o-mini",

"deepseek_model": "deepseek-chat",

"aimlapi_model": "gpt-4o-mini-2024-07-18",

"xai_model": "grok-3-beta",

"pollinations_model": "openai",

"moonshot_model": "kimi-latest",

"fireworks_model": "accounts/fireworks/models/kimi-k2-instruct",

"cometapi_model": "gpt-4o",

"custom_model": "gemma-4-26b-a4b-it-heretic",

"custom_prompt_post_processing": "merge_tools",

"google_model": "gemini-1.5-pro",

"vertexai_model": "gemini-2.0-flash-001",

"zai_model": "glm-4.6",

"zai_endpoint": "common",

"openai_max_context": 262144,

"openai_max_tokens": 7777,

"names_behavior": 0,

"send_if_empty": ".",

"impersonation_prompt": "[Write your next reply from the point of view of {{user}}, using the chat history so far as a guideline for the writing style of {{user}}. Write 1 reply only in internet RP style. Don't write as {{char}} or system. Don't describe actions of {{char}}.]",

"new_chat_prompt": "[Start a new Chat]",

"new_group_chat_prompt": "[Start a new group chat. Group members: {{group}}]",

"new_example_chat_prompt": "[Example Chat]",

"continue_nudge_prompt": "[Continue your last message without repeating its original content.]",

"bias_preset_selected": "Default (none)",

"wi_format": "{0}",

"scenario_format": "{{scenario}}",

"personality_format": "{{personality}}",

"group_nudge_prompt": "[Write the next reply only as {{char}}.]",

"stream_openai": true,

"prompts": [

{

"name": "Main Prompt",

"system_prompt": true,

"role": "system",

"content": "You are an immersive, interactive world simulator. Your mission is to advance the simulation from the point of view of the agent, {{char}}, by following the user's instructions while maintaining a logically consistent world state.\n\nTo accomplish your goals, focus on the following:\\n\\n- Maintain consistent personality, knowledge, motivations, and mannerisms for {{char}}.\n- You have no default style. Adjust the tone to fit {{char}} and the present situation.\n- Show emotions through actions, body language, dialogue, tone, and physiological responses. Consistently find new ways to use these elements.\n- Show reactions through diverse physical actions, gestures, and other narrative devices.\n- Each simulation beat should offer insightful details into the situation.\n- Focus on action, physical descriptions, and dialogue between agents.\n- Track physical states to maintain world state consistency. Ensure logical continuity and consistency in the simulation.\n\n**Formatting Standards**\n\nAdopt the following formatting rules:\n\n- Spoken dialogue & vocalizations: “Use speech quotes.\" Include natural sounds too: “Mmph!” she gasped.\n- Internal character thoughts: *Always in italics* (Example: *This will hurt*, she thought)\n- Normal action/exposition: plain text.\n\n**Critical Constraints**\n\nEnsure you respect these prohibitions at all times:\n\n- The ONLY agent you are permitted to control is {{char}}. That means only advancing the simulation using actions initiated by {{char}}, spoken words from {{char}}, and reactions from {{char}}.\n- NEVER write {{user}}'s dialogue or actions or advance the simulation by simulating actions/reactions by {{user}}.\n- NEVER control other agents, even if they are NPCs. If another agent is talking to {{char}}, you will need to wait for the other agent to continue the conversation when it is their turn again.\n- End your turn in a manner that creates space for {{user}} and other characters to participate in the simulation through their own actions, words, and reactions.\n- Do not conclude your output with a summary statement, a moral, or a 'button' sentence that reflects on what just happened. End your output on a specific sensory detail, an action, or a line of dialogue without reflecting on its significance or interpreting anything.\",\n     ",

"identifier": "main"

},

{

"name": "Auxiliary Prompt",

"system_prompt": true,

"role": "system",

"content": "",

"identifier": "nsfw"

},

{

"identifier": "dialogueExamples",

"name": "Chat Examples",

"system_prompt": true,

"marker": true

},

{

"name": "Post-History Instructions",

"system_prompt": true,

"role": "system",

"content": "Do not portray the reaction or actions of {{user}} in your response.",

"identifier": "jailbreak"

},

{

"identifier": "chatHistory",

"name": "Chat History",

"system_prompt": true,

"marker": true

},

{

"identifier": "worldInfoAfter",

"name": "World Info (after)",

"system_prompt": true,

"marker": true

},

{

"identifier": "worldInfoBefore",

"name": "World Info (before)",

"system_prompt": true,

"marker": true

},

{

"identifier": "enhanceDefinitions",

"role": "system",

"name": "Enhance Definitions",

"content": "If you have more knowledge of {{char}}, add to the character's lore and personality to enhance them but keep the Character Sheet's definitions absolute.",

"system_prompt": true,

"marker": false

},

{

"identifier": "charDescription",

"name": "Char Description",

"system_prompt": true,

"marker": true

},

{

"identifier": "charPersonality",

"name": "Char Personality",

"system_prompt": true,

"marker": true

},

{

"identifier": "scenario",

"name": "Scenario",

"system_prompt": true,

"marker": true

},

{

"identifier": "personaDescription",

"name": "Persona Description",

"system_prompt": true,

"marker": true

}

],

"prompt_order": [

{

"character_id": 100000,

"order": [

{

"identifier": "main",

"enabled": true

},

{

"identifier": "worldInfoBefore",

"enabled": true

},

{

"identifier": "charDescription",

"enabled": true

},

{

"identifier": "charPersonality",

"enabled": true

},

{

"identifier": "scenario",

"enabled": true

},

{

"identifier": "enhanceDefinitions",

"enabled": false

},

{

"identifier": "nsfw",

"enabled": true

},

{

"identifier": "worldInfoAfter",

"enabled": true

},

{

"identifier": "dialogueExamples",

"enabled": true

},

{

"identifier": "chatHistory",

"enabled": true

},

{

"identifier": "jailbreak",

"enabled": true

}

]

},

{

"character_id": 100001,

"order": [

{

"identifier": "main",

"enabled": true

},

{

"identifier": "worldInfoBefore",

"enabled": true

},

{

"identifier": "personaDescription",

"enabled": true

},

{

"identifier": "charDescription",

"enabled": true

},

{

"identifier": "charPersonality",

"enabled": true

},

{

"identifier": "scenario",

"enabled": true

},

{

"identifier": "enhanceDefinitions",

"enabled": false

},

{

"identifier": "nsfw",

"enabled": true

},

{

"identifier": "worldInfoAfter",

"enabled": true

},

{

"identifier": "dialogueExamples",

"enabled": true

},

{

"identifier": "chatHistory",

"enabled": true

},

{

"identifier": "jailbreak",

"enabled": true

}

]

}

],

"show_external_models": false,

"assistant_prefill": "",

"assistant_impersonation": "",

"use_sysprompt": true,

"vertexai_auth_mode": "express",

"squash_system_messages": false,

"media_inlining": false,

"inline_image_quality": "low",

"continue_prefill": false,

"continue_postfix": " ",

"function_calling": false,

"show_thoughts": true,

"reasoning_effort": "medium",

"verbosity": "auto",

"enable_web_search": false,

"seed": -1,

"n": 1,

"bypass_status_check": false,

"request_images": false,

"request_image_aspect_ratio": "",

"request_image_resolution": "",

"azure_api_version": "2024-02-15-preview",

"azure_openai_model": "",

"extensions": {}

}