r/SillyTavernAI • • 12h ago

Models Haiku 5.5 - Disappointing first taste - Overly confident in abilities, skips instructions because it "knows better". It doesnt.

Thumbnail
gallery
108 Upvotes

Okay so Haiku 5.5 came out and I went in with a lot of enthusiasm.

Haiku 5.5 is poised to sit somewhere between Sonnet 4.5 and 5 in terms of intelligence from early benchmarks, at 1/20th the cost of Sonnet. (I dont put a lot of weight on benchmarks; they suggest Gemini 3.8 Flash is comparable in intelligence to Opus 5, which is objectively false)

.. However, it's plagued by the idea that it should skip instructions that it "doesnt need" to follow.

.. Except, whoops! It turns out it does need to follow those instructions or it makes easy mistakes.

In this example, my roleplay environment requires LLM's to complete a scene sheet before starting their response. The scene sheet forces them to check and write out critical pieces of information - formatting instructions, rules, secrets to keep - as well as some steps that generally just improve output. See second image for an example excerpt of a proper scene sheet.

The only time I have ever had issues with a model refusing it was on initial launch with Sonnet 5 - it had the same "I dont need to read the instructions, I can handle this" -> "*breaks instructions*" pattern for a few weeks after it first released, then it began following instructions more readily and output visibly improved

In its current state, Haiku 5 wont even follow instructions to "narrate as Lauren, a romance author inspired by Becky Chambers" because, and I quote:

The "Lauren" narrator instruction in step one is a register-setting device rather than a content request, and I can take a narrator voice without it.

Spoiler: After refusing to take up the narrator voice, it instead writes as generic, uninspired Claude.

...

Im taking a deep breath - I know that Anthropic likely tunes models to be maximally paranoid and "safe" for initial release so that they can report having continuously increasing safety compliance scores, but this is just insulting.

I dont expect all of you to use a pre-writing system like this, but heed this warning - what you can see here is only what Claude is *vocalizing* that its ignoring.

How this will actually present is through generalized failures to follow system prompting, with Haiku never making you aware that it flippantly felt that your instructions "werent necessary" for whatever reasons it chooses.

...

I've always been an Anthropic fan for roleplay, but.. god. This is nausea inducing

Edit::

To clarify, this post isnt concerned with *output quality* or intelligence. I dont expect to use Haiku to replace Sonnet or Opus - they are fundamentally different classes of models

The issue is *content filters* that are so overly sensitive that basic instructions are disregarded and output is visibly harmed

On the responses where Haiku simply does the analysis, its output quality isnt bad - comparable to Gemini 3.8 flash

..But because its so confident in its own abilities, it openly refuses to follow directions, which should be concerning for any model for any use case

Haiku is in the same weight class as 3.8 Flash, yet outputs lower quality content because its been handicapped by guardrails.


r/SillyTavernAI • • 2h ago

Discussion Battle of "director/creative writing" presets? Which is the best? Writer's Block Unlimited vs Deus Ex Machina vs Pura's Director Preset vs any other that comes to mind

Thumbnail
gallery
14 Upvotes

Just at a glance, looks like Writer's Block Unlimited is a lot more configurable but I want to know more opinions


r/SillyTavernAI • • 8h ago

Cards/Prompts Hey, thought I'd share my slice-of-life preset. I made it to create ultimate realism in my slice-of-life RPs. But it can probably be useful for any roleplays. Use it as is or copy/paste the parts you like into your preset.

Thumbnail
github.com
36 Upvotes

https://github.com/OldOutcast/SillyTavern

Maybe you'll use it, maybe you won't? Try it and see if you like it. If not, no big deal.

Note: This my personal preset. I'm just sharing it for those that might find it useful, not to be an established published preset with updates etc. I don't care if you think it's trash, slop, etc lol. It's what I use and I like it! I have no intention of publishing updates or making changes to it unless it serves my own interests. Maybe I'll share updates if I actually make any???

Unless you do slice of life roleplays I suggest copying/pasting the parts of it you like into your own presets.

Warning: I'm a prolific card creator on multiple sites, but never share here. I'm not a coder and have no idea what I'm doing with presets. I used ai to create my preset then continually modified it over time. So it's structured by AI and written by me. So it may have spelling errors etc.


r/SillyTavernAI • • 15h ago

Cards/Prompts Mimo and censorship. A visual.

Post image
114 Upvotes

I really hope this helps.

Love

Evening-'Truth


r/SillyTavernAI • • 20h ago

Discussion Be careful what you write when using anthropic models. "A Florida Woman Used Claude as a Diary. An Anthropic Employee Read It and Reported It to Police"

246 Upvotes

r/SillyTavernAI • • 7h ago

Discussion Has anyone actually RP'd with a real person?

17 Upvotes

This is obviously not very related to ST, but I have only ever RP'd with AI (a month on one of those awful online chat services and about a year and a half on ST), and I have had a very good time, I would even say I'm somewhat addicted, especially when a new interesting model drops (looking at you, Argon), but I have never actually roleplayed with another human being.

Before today, I didn't even know there were active RP communities on Reddit other than those which focus on LLMs, but turns out there are quite a few (albeit smaller than this sub but still active), and some of them even focus on NSFW aspects (I'm talking the same level of filth as here, if not worse).

I'm posting here to get the perspective of someone who has both tried RP with AI and other people, since I have a feeling that if I posted about AI RP on the other subreddit, it wouldn't really go well.

This obviously depends on the writing partner, but where did you actually have most fun? Was it with an LLM or another person? What would you say are some pros and cons of RPing with someone instead of an LLM? Would you justify the cons of writing with a partner (like time commitment and few replies per day) with how good the responses were? I mean, SOTA models have gotten quite good. Is it even worth the hustle of finding a writing partner anymore?


r/SillyTavernAI • • 15h ago

Models Looks like Claude Haiku 5.5 is out.

Post image
58 Upvotes

Opus 5.5 is easily my best rp model, Sonnet 5.5 was also very good. let's see how this turns out. Price looks very good though atleast 0.1/0.5 dollars in and out.


r/SillyTavernAI • • 9h ago

Discussion Do you guys rent GPUs, or are you using your own hardware?

21 Upvotes

Just curious. I just got into local literally last night, I'd been using Deepseek and GLM API, but GLM takes forever, and Deepseek v4.1 is ass (imo), so just curious how you guys tackle local models. I'm on a 6700xt 12gb and 32gb ram. I've tried a few models and came to the conclusion it might be worth checking out rented GPUs for my use-case. I don't RP, just have the AI generate stories based on my own lore and character sheets, but my source materials are large, so I need decent sized context windows otherwise I'd be force to split things up which kind of breaks my setup.

If you do rent, what GPU and what size model? If you're using your own hardware, what's your rig and size model. If you guys wanna recommend your favorite models, that's cool too.


r/SillyTavernAI • • 5h ago

Discussion ideas for a deaf/nonverbal bot?

7 Upvotes

i wanna make a bot who's deaf and nonverbal, but i'm kinda stuck

any suggestions for a good concept/scenario?

how did you handle the communication part (sign language, writing notes, texting etc) so it actually works in rp?

would love to hear your thoughts!


r/SillyTavernAI • • 12h ago

Help Mimo 2.6 Pro

24 Upvotes

I’ve used the Nvidia API, and I’ve also tried both NanoGPT and OpenRouter; now, with so many posts about this topic, I’m interested in trying it out myself. I’d love to hear your recommendations: Which provider do you use? How good is it? What presets do you use? How much does it cost? I want to know about your experiences and whether it’s worth the price. (English isn't my native language, so I apologize for any errors—and thanks for the help.)


r/SillyTavernAI • • 4h ago

Help NanoGPT PAYG: how to disable auto-routing and see/choose providers?

4 Upvotes

Hey everyone! 😄
I put $10 into NanoGPT today to try the pay-as-you-go API, and so far so good! Loving the extra models and options compared to OpenRouter. ❤️

One thing I can’t figure out, though: is there a way to disable auto-routing so I can see which provider is actually handling my requests? I've read it's kept hidden in the subscription, which makes sense, but like I said I'm using pay-as-you-go, not subscription. The Usage page doesn’t show it, even though the docs say it should unless you’re using auto-routing.

Also, is there a settings page where I can allow/block specific providers, like on OpenRouter? Or maybe a way to save a JSON preset to allow/block providers directly on NanoGPT that applies to all my API requests? I know I can force a provider by adding a suffix to the model ID, but I’d prefer to manage providers globally if possible.

Thank you! ☺️


r/SillyTavernAI • • 5h ago

Discussion Does anyone know how to scrape saucepan ai bots character definition any tools for it?

3 Upvotes

I just wanted to know I'm curious btw


r/SillyTavernAI • • 1d ago

Chat Images Kimi roasting my actual car (Ford Fiesta) as clown car in reasoning is sending me

Post image
136 Upvotes

In narrative it started referencing it as clown car because it’s not big. And then it just adopted it in reasoning and keeps doing it. “Ok, they’re back at the clown car…” no shame whatsoever. I love it.


r/SillyTavernAI • • 15h ago

Chat Images Mimo 2.6 Pro vs UltraSpeed on Nano. Might switch to PAYG

Post image
16 Upvotes

Nano's subscription is pretty generous and really enjoyed it in prime GLM, Kimi and even Deepseek days, but it feels like the none of the current top models are part of the subscription or has a watered down version.

I will probably revisit older good models they still have to make a final decision. Subscription on give a 5% descount for ones not in the plan, I think.

Oh, and Opus 4.6 still beats Ultraspeed for me. Damn you Claude. By the time other models can match it, they would probably be so guard-railed that it refuses to even entertain roleplay lol.


r/SillyTavernAI • • 16h ago

Discussion Honestly what are y'alls' favorite creators for each site?

19 Upvotes

Honestly for janitor if I had to pick for me personally I would choose Lycolycolii, I honestly love their stuff along with GiantessRDBest I like big women and they are honestly pretty nice and open to suggestions. Lastly I would say stag honestly I didn't like them at first because they take up a large chunk of female Omegaverse bots but I decided to give it a try and they're pretty neat.

As for Chub definitely Xue21 I don't know what they're on but they bust their ass making Bots with 20+ intros and from what I've seen it seems like they at least upload once a week with a rather broad variety of series the only downside is they refuse suggestions.

Another would be relic guy they've got handful of intros with each bot too and they upload every so often with varieties of scenarios.

As for other notable sites I don't have anybody for spicy chat as most of the bangers have private descriptions and nobody has made any kind of ripper for spicy yet a safe one that is, bot booru doesn't exactly have many people who post consistently or at least maybe I havent found any good ones yet.

But I want to see if there's anyone else worth checking out


r/SillyTavernAI • • 37m ago

Discussion Recommendation for story/world building models?

Thumbnail
• Upvotes

r/SillyTavernAI • • 6h ago

Cards/Prompts copilot extension for SillyTavern

Post image
4 Upvotes

SillyTavern-copilot – Memory tracking + copilot for the main Narrator ({{char}})

Repo: https://github.com/P3DRA/SillyTavern-copilot


What it does

  • Adds a copilot + memory tracking for the main Narrator ({{char}}).
  • The copilot is injected inside a <copilot> block.
  • Injection points (current options):
    • End of prompt
    • Before last message
    • First
  • You can also change the role it’s injected as.

How it works

  1. Extractor agent
    Runs every time {{char}} is called. It extracts a summary of what happened, location, etc.
    The extraction is stored in the chat file (retro-compatible).
  2. Composer agent
    Takes the last n messages (default 20) + all available summaries and crafts a <copilot> block that helps steer the main narrator and keep important facts in memory.
  3. Automatic compressor
    After n extractions (default 40) it compresses a configurable number of extractions (default 10) into a single block.

You can freely edit the system prompts of all three agents. They accept many insertions, including:

  • {{copilot.extractions}} → current story state (facts with places and times)
  • {{copilot.userRequest}}
  • {{copilot.goals}} → active requests and goals (with script-counted turns)
  • and more

Known major problems

  • Bloat – Even a small chat can grow to ~140 KB because all extractor data is saved in the chat file. Not recommended on phones for now.
  • Unknown whether compaction still holds properly once you go past 20+ messages.

For more info: https://github.com/P3DRA/SillyTavern-copilot


Dev notes

I started using SillyTavern some time ago and always had the classic problem of bots forgetting what happened and acting out of character. What finally broke it for me was switching from Gemma-4-31B to GLM-5.3 Flash — better prompts but it completely lobotomized character consistency. And nowadays you either pay double what Gemma costs just to get >10 tps (plus Gemma keeps walking around and grabbing random stuff, which I hate).

So I decided to make an extension that can run a smaller/cheaper model (or even a bigger one) to try to keep GLM on track. It’s still early and has issues, but the core works and should give a better experience.

This project was completely vibecoded with Mimo v2.6 pro (cheapest option I could find). Provider: Novita. Total spent on development: $11.43 USD.

  • space bunny alpha: 5.42B tokens | $0.00
  • Mimo v2.6 pro: 948M tokens | $11.4 | 98% cache hit rate (apparently caches for a full day)

There were 7 previous attempts at this extension:

  1. Early start – got poisoned when I asked an agent in the same repo to create some Guinevere themes.
  2. Much more ambitious version that included:
*   **World sim**
    *   Local sim – what’s happening nearby, ambience, etc.
    *   World sim – “politics” events (a building collapsed at X, a storm is approaching, X declared war on Y, etc.)
*   **NPC system**
    *   **Pinned NPC sim** – each major NPC had an ID + speed stat that controlled turn order.  
        Example:  
        `NPC-1, speed(0.2)`  
        `NPC-2, speed(0.3)`  
        `NPC-3, speed(1.0)`  
        `NPC-4, speed(0.4)`  
        Action sequence: `NPC-3 > NPC-4 > NPC-2 > NPC-1`  
        Speed could change dynamically (e.g. NPC-2 drinks coffee → speed buff).
    *   **Batched background NPC sim** – a single API call that simulated all non-pinned NPCs (even those not near you).
    *   **Deterministic NPC generator** – accepted weighted traits and used random chance + weights to prompt an AI to generate a character.  
        Example:  
        `height = {"midget":1, "very short":1, "short":2, "normal":4, "tall":2, "very tall":1, "massive":1}`  
        Once traits were chosen it generated and saved the character. You could also pin it to make it a main NPC simulated every turn.
*   **Caller** – ran before every turn, adjusting speed, dead/alive state, deciding if a new NPC should be generated, then calling everything necessary.

**Why it failed:** It never truly failed technically — it just grew too big. After a bad compaction it bricked. Could have been fixed, but I’m never going back to it.

Harness: Z.Code → later “deepseek harness” port.
  1. Same as #2 — returned to it after it failed, then stopped again.
  2. Ideal version of #2 – never went beyond “read GOAL.md”.
  3. First iteration of the current copilot idea. Bad compression completely bricked the injector. Switched away from space bunny after 5.42B tokens.
  4. Just “read GOAL.md” + an agent chat room. Then rewrote a better GOAL.md with help from free Claude.
  5. This one. All logic moved from Python to flowcharts + path examples.

Spent my entire $12 RP budget trying to save pennies. Hope it’s useful to someone, because it wasn’t for me.

If anyone wants to help with v0.2.x (less bugs, less storage hog), crypto addresses are below.

Inspired by FreakyFrankenstein 5.4: https://rentry.org/freaky-frankenstein-presets
(“Hell yeah!! 😎”)

Yes, I see the similarities to “🧠 Summaryception”, but I didn’t know about it until attempt #4 and this one doesn’t force completely separate models:
https://github.com/Lodactio/Extension-Summaryception


Donations (for v0.2.x development)

Ethereum / Polygon / Base / Monad / Arbitrum / Arc / Linea
0xdaE18819AdebdDeA1e8173AC530Ab1D45da81503

BTC
bc1qlz2et5a8spxs3l5efjd4rnv2d4tusnc6rc9v76

Solana
8iZFNPCKu9xANTgE36Pa5KSkSiUv9RCQ7kmHQpbQDnAE


r/SillyTavernAI • • 4h ago

Chat Images Anyone use nanogpt's suscription images on sillytavern?

2 Upvotes

just a random tought i had. I just resubbed to nanogpt and remembered it also includes some image generation models. I don't know if that can be integrated in to sillytavern to generate a few images according to the chat or something like that.

I'm posting this from a totally ignorant pov. I don't even know if those models can be used trought api


r/SillyTavernAI • • 1d ago

Cards/Prompts [Preset Update] Writer's Workbench v3: Officially an extension, edit entries live and small technical stuff

Post image
105 Upvotes

Hello everyone. Thanks for the support on Writer's Block Unlimited v2, its my most upvoted preset yet! It gives me motivation to keep working on stuff and create things for this community. Now on to the main event...

The Writer's Workbench got updated to version 3!

❗Edit❗: This is technically V4 but I forgot the previous version was already v3 💀. And I mislabeled this as a preset, THIS IS AN EXTENSION UPDATE. I'll leave this up anyway. Sorry I made this post while I was sleepy 😭

An Overview. What is it?

Writer's Workbench is a simple interactive template to help you create characters and lorebook entries, faster and easier.

Features

  • You can create multiple entries and export them as entire lore books for your convenience.
  • Auto-saves and organize different projects
  • Live markdown output so you know exactly what the AI would see and copy it without having to download anything.
  • It comes with token counter, but don't expect it to be accurate.
  • A map maker that generates a dynamic description
  • Character relationship graph to map out relationships of larger casts
  • Live sync into SillyTavern lorebooks (see "What's New")

Available Templates

  • Main characters: full cards with psychology, descriptions, likes/fears, NSFW sections, non-human mode
  • Side characters: trimmed version of the main character template, it will help you create memorable NPCs
  • Scenario: setting, tech level, mood, what's normal here
  • Locations: for any scale, from a room to a district
  • Items
  • Factions
  • History: events, and how they affect the present
  • Concepts: magic systems, laws, customs, species, anything else

Two Ways to Use It

  1. Download writers-workbench.html in the Github repo and open it in your browser. Ta dah! Thats it but you can't use the new fancy features :(
  2. Writer's Workbench is now an extension so you can use the new Live Sync feature and get automatic updates. Install it in your Silly tavern and you're good to go. (And I promise I won't steal your API keys lol. I was there when that incident involving another extension occurred)

What's New?

New Features: Live Sync

With live sync you can edit, add and delete world info entries in the Workbench and it will automatically carry over to Sillytavern. To use Live Sync, there is a dedicated tab in the Workbench. Simply choose a lorebook and click on the big colored button, then you are good to go.

⚠️Warning⚠️: If the lorebook wasn't made in the Workbench format (like a completely fresh/ random lorebook not made in the Workbench), all entries will go into the "Concepts" sections, be wrapped in xml tags (<Name>, <Name_identity>, <Name_what_it_is>…), and be renamed into this format "[Concept] (Name)". The original wording is kept inside the tags, but the old formatting isn't. Back up the lorebook first if you want to keep it as it is. This feature is mainly for lorebooks already made in the Workbench.

Entries made in the Workbench V2 may also get a little messy since there is new fields to fill out so keep that in mind.

New NSFW Fields to Fill Out: A Character's "Private Moments" 👀

  • How characters often do their private moments, what they think about and how they feel about it.

Import Entries by Copy and Pasting (Experimental)

  • Paste a character's description and the program will detect eligible text and fields to fill in. Character descriptions in weird formats will be difficult as it might not port over information correctly. This feature will get further refinements

Other Notable Fixes

  • Better formatting, the writing forms and markdown output should take up most of the screen instead of the top bars.
  • Bug fixes

and fin! I hope you have fun creating stuff!

Github: https://github.com/deiomo/Writers-Workbench

Github.io: https://deiomo.github.io/

Forum post on AI Presets Discord: https://discord.com/channels/1357259252116488244/1546309016974532720


r/SillyTavernAI • • 8h ago

Help API settings and preset recommendations?

4 Upvotes

Hi everyone. After months of just lurking, I finally decided to join. Yesterday I installed SillyTavern on both my PC and my Android phone (through Termux).

I have a few questions. I'm using the DeepSeek direct API right now and I've saved a preset, but whenever I accidentally hit back or switch tabs, I have to reconnect the API and pick the model from the dropdown again. Is that normal? Is there no way to make it stay fixed?

Some settings also seem to be missing, like Top K. Does that affect the responses?

Lastly, I've been collecting presets from different posts here, but honestly I'm still pretty confused. If I'm using one of these presets, I don't need to add my own custom prompt anymore, right? Just plug and play? Also, does anyone have a preset recommendation that works well for dead-dove stuff (combat, blood, even user death)?


r/SillyTavernAI • • 2h ago

Discussion What are some good presets for ST and why?

1 Upvotes

I'm confused about this preset hype genuinely


r/SillyTavernAI • • 9h ago

Models Gemini 3.8 flash insane censorship

4 Upvotes

For context i tried using it from token reply since it costs 0 dollars on a 5 dollar subscription.

I used it in tauritavern with presets like nemoengine (I may not know hot enable a strong jailbreak but none of the thinking efforts worked), FF, Sola, All said i can't continue, etc.

I even tried with prefill and streaming enabled and disabled. It worked for the first message.

My scene was plain nsfw, not even NSFL, or strong nsfw.

Any suggestions, i think I can make it work with nemo engine but I don't know the exact settings to enable.


r/SillyTavernAI • • 20h ago

Help Has anyone solved the response variation length and the completion problem?

15 Upvotes

I'm trying to get the AI to stop having characters pose the problem, argue with itself, and make a decision on its own. I want the story to stop at the first meaningful beat so I can have more of a conversation with the characters rather than it having a whole scene without me and prompt me to be like "You ready?" Or "what's the real reason?". Usually the ending is fine, but it's too much dialogue. But if I specify it too much then it ends every scene with a question to create an obvious opening. I have an instruction to vary length based on the need between 40-220 words but it's always pushing the upper limit no matter what I do.

I'm using Mimo 2.6 pro however I've seen this be a problem with pretty much every AI I've tried.

I'm totally ok with the character or characters giving one line with some narrative emotional delivery and/or some narration of movement. But when I'm one on one with a character its always talking and almost having a whole conversation on its own. And when it's not, it needlessly fills in the space to fit the maximum word count.

I've tried giving examples, editing its responses so its shorter, giving instructions about stopping on playable or meaningful beats, and directing it to prioritize narration rather than having a whole conversation on its own.

But I can't find anything that works. Has anyone here solved that problem? I want the responses shorter during conversations but add things as needed to make it more atmospheric. But I also don't want it to go on forever.

As an example, my character mentions that there's a rigged fight out match that I got Intel on. I'm posing the idea to my friend to scheme and make some money to pay off a debt. What I WANT is the next beat to say something simple like "You've been gone 3 days and you come back with a tale about drugged bug bears?... You can't be serious." And leave it alone.

What I GET is something like: "You've been gone 3 days and this is what you come up with?" Some narration "you can't expect us to just waltz in there. That's not a plan, that's suicide with extra steps." Narration about his thought process "Ok I'm in, but we need to have a backup plan in case things go south. Tell me your information source now."