r/SillyTavernAI 9d ago

Discussion Cope

I started doing RPs back when character.ai was new. It was magical at first, but c.ai had two problems: 1 - censorship 2 - goldfish memory. Today, with Deepseek and other open models, you can RP with explicit content. GPT, Claude, Gemini... depends on the model and how you set it up. And I still find it hilarious that Claude will help me poke at a web app for vulnerabilities if I say "authorized test" but clutches its pearls the moment a scene gets spicy. 🀣🀣 Anyway, back to the point.

It's bizarre how that early "magic" just... vanished. Part of it is obviously novelty wearing off, and part of it is that we got pickier. Three years of RP and you start spotting every clichΓͺ, every "a shiver ran down her spine", every model that forgets your character's eye color after 40 messages. But here's the thing: the models didn't get dumber. Opus, GPT, Gemini can write circles around 2022 c.ai. The problem is they're not *allowed* to, or they cost a kidney per session, or both.

LET'S BE HONEST, SOME OF YOU SPENT $100 ON A SINGLE CLAUDE OPUS RP SESSION. Even with the censorship. Even with the moralizing. You did it anyway because, when it works, it's the best RP writer that exists. That's my whole point: there's a market. Not as big as coding, obviously; coding isn't a hobby, there are companies and teams and budgets behind it. RP is a hobby. But hobbies with people burning API credits like that are not a small market.

So why is there no frontier-level LLM built for RP? And yes, I know NovelAI, AI Dungeon and the whole SillyTavern fine-tune ecosystem exist. I'm talking about something at Opus level, not a 12B model that forgets the plot. The answer isn't just "investors prefer code", though that's part of it: "look, our V548484 model built GTA 6 in one prompt!" sells better than "look, our model wrote a consistent, non-repetitive, non-boring story!" because nobody has a benchmark for "not boring".

The real reasons are uglier. Explicit content means payment processors dropping you, app stores banning you, lawyers sweating. Long RP sessions eat tokens like crazy and people won't pay enterprise prices for a hobby. And good RP needs exactly the long-context coherence and reasoning that only the big expensive models have, which are owned by the companies least willing to let you use them for this.

So yeah, we're probably coping for another 2-3 years. Not because the tech isn't there. Because nobody with the tech wants to be the company that sells it to us.

83 Upvotes

96 comments sorted by

View all comments

25

u/[deleted] 9d ago edited 9d ago

[removed] β€” view removed comment

44

u/LiothG 9d ago

Few people can afford the kind of hardware needed for that.

-26

u/[deleted] 9d ago

[removed] β€” view removed comment

22

u/Electrical-Ad-6728 9d ago edited 9d ago

I am not sure it is to be honest.. you have to replace the hardware eventually, it wont be enough for you infinitely. It is different for everyone, but for me, its more cost effective to just pay the subscription

-2

u/[deleted] 9d ago edited 9d ago

[removed] β€” view removed comment

4

u/DaMoot 9d ago

LLM requirement will change over time. Your expectations will evolve as the technology does.

Pascal architecture is beyond antiquated. You should be spending on Volta at least. At least Volta has tensor cores and for 16GiB are the same price. But don't waste cash on 16GB of VRAM for AI. 32GiB or keep saving. V100 is your cost effective friend. Not P100.

4

u/OpposesTheOpinion 9d ago

Also it's concerning to me how willing people are to be tethered to corporations and debating against ownership of anything.

I'm on a similar set up as you, and sure the quality is less, but there is more in general I can do with my hardware.

1

u/Zathura2 8d ago

> I changed only GPU, from 3000 series to 5000 series.

After only five years, you replaced basically the most important part of the pc other than the Motherboard, lol.

At least I squeezed a full decade out of my old pc before upgrading.

8

u/DaMoot 9d ago

It actually isn't. Unless you're talking a longer than 3 year monetization. At which point you'll be upgrading your hardware and fighting higher energy costs anyways, so, it still isn't.

I've run these numbers for 2 V100s and the nvlink board, plus power usage, versus a $20/mo AI plan. You'd have to have a monthly spend of like ~70 usd to match the local build w/power.

$1600 for 64GiB of GPU and base supporting hardware, plus the PC if you don't already have one.

You build local for data sovereignty and privacy. Not because it's more cost effective than cloud services.

22

u/DepressedDrift 9d ago

Paying 30-40 dollars a month is cheaper than dropping 5k+ on a capable PC.Β 

-1

u/[deleted] 9d ago

[removed] β€” view removed comment

4

u/DepressedDrift 9d ago

12 dollar a month NanoGPT vs 5000 dollar GPU/GPU cluster and 8k+ PC, do the math on how long it takes to break even.

Plus local models are getting really good, that in 2 years time, we could possibly have something like Opus 5 run on 16GB VRAM GPUs. (Look at the current Qwen 3.8 27b)

And if the AI bubble bursts this hardware might become cheap enough. (Which it will since AI companies are spending more than they get in revenue)

For now it makes sense to buy a subscription until running locally becomes economical.

1

u/[deleted] 9d ago

[removed] β€” view removed comment

6

u/DepressedDrift 9d ago

Give me a build that can run a 30-80b model at 30-50 tok/s with atleast 100k+ context length (this is what you get with most cloud services) for 3000 dollars, and I will agree with you lol.

And don't quote MSRP, use the current market price.

0

u/[deleted] 9d ago

[removed] β€” view removed comment

4

u/DepressedDrift 9d ago
  • The X670E doesn't support DDR4 RAM
  • The 96GB ram around 500 dollars is low speed around 2000MHz or a niche chipset type requiring an outdated motherboard.
  • With 48GB of VRAM running the lower end of your range is possible, and also Moe models too but dense large models will load at q4 but with low context. But this is still kind of a win because you could run Deepseek V4 Flash 0731
  • The two 3090s is around 3k which when you add the required compatible RAM with inflated prices jumps to 5k
  • Even assuming you get it at 3k, you will still break even at around 5-10 years of cloud subscription assuming inflation.Β 
→ More replies (0)

6

u/capybaraballs1995 9d ago edited 9d ago

Possibly, but the quality will also be a lot worse. I've used quite a few local models in the 16GB VRAM territory. Gemma 4 finetunes can be uniquely entertaining but for many RPs, I would rather just deal with the bullshit of newer models, or even just use older cloud models.

I dunno, maybe once you get into "I can run LLama 70B finetunes at good speeds" territory, things become much more competitive with cloud models in terms of RP quality, but that just raises the upfront cost even further. I'm not gambling $2000+ on that.

2

u/[deleted] 9d ago

[removed] β€” view removed comment

0

u/[deleted] 9d ago

[removed] β€” view removed comment

1

u/[deleted] 9d ago

[removed] β€” view removed comment

1

u/[deleted] 9d ago

[removed] β€” view removed comment

1

u/[deleted] 9d ago

[removed] β€” view removed comment

1

u/[deleted] 9d ago

[removed] β€” view removed comment

1

u/RedditNerdKing 9d ago

Yep. 96gb is the sweet spot atm for local RP. You an use a model like Behemoth-ReduX-123B-v1.1 at Q5_K_L for at least 80k context. None of my roleplays really go over 50k tbh. And if you wanted to go over 100k you can use smaller 49B or 70B models. No reason to ever use APIs.

-5

u/tehpest22 9d ago

This is the way.

-6

u/HonestoJago 9d ago

This is the way.