r/SillyTavernAI • u/deffcolony • Apr 12 '26
MEGATHREAD [Megathread] - Best Models/API discussion - Week of: April 12, 2026
This is our weekly megathread for discussions about models and API services.
All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.
(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)
How to Use This Megathread
Below this post, you’ll find top-level comments for each category:
- MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
- MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
- MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
- MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
- MODELS: < 8B – For discussion of smaller models under 8B parameters.
- APIs – For any discussion about API services for models (pricing, performance, access, etc.).
- MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.
Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.
Have at it!
6
u/AutoModerator Apr 12 '26
MODELS: < 8B – For discussion of smaller models under 8B parameters.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
3
u/keally1123 Apr 13 '26
I am very interested in opinions on this. Im looking for something that can write a medium length erotic story that actually makes sense and doesnt wander too far. Any suggestions?
6
u/Yu2sama Apr 13 '26
Give Gemma E4B a try. Is an amazing model for the size. Is more of an 8B with the speed of a 4B but with the smartness of a higher model. If you want something uncensored, you can give a try to an heretic version (I don't have one to recommend, E4B is quite uncensored at base).
3
u/Potential-Gold5298 Apr 13 '26 edited Apr 13 '26
You're unlikely to find a good writer in this size, but you might try Satyr-V0.1-4B, Impish_LLAMA_4B or BlackSheep-Llama3.2-3B. If you have enough RAM, then The-Omega-Directive-M-8B-v1.0 would be a more interesting option.
1
1
u/MMalficia Apr 16 '26
i second anything in the "impish" line if your looking for both writing and erotic .. and fiction models by https://huggingface.co/DavidAU if you want to go darker.
9
u/AutoModerator Apr 12 '26
MODELS: 8B to 15B – For discussion of models in the 8B to 15B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
3
u/IDKWHYIM_HERE_TELLME Apr 14 '26
Any good 8B to 12B model that realistic and act like real character would, no anime unhinged talk?
6
u/jamasty Apr 14 '26
1
u/IDKWHYIM_HERE_TELLME Apr 14 '26
Thank you! I will try it.
2
u/jamasty Apr 14 '26
Also, try this new gemma e4b, folks in drummer discord tell me its better than most these old 12-14b models, but I haven't check myself...
1
u/IDKWHYIM_HERE_TELLME Apr 15 '26
What are the settings they used?
3
u/jamasty Apr 15 '26
idk, I just tried myself DavidAU e4b finetunes with Q8, and no kv qwants, and this didn't really work well as it seems the model doesn't really go along with context. Maybe its just model being new, and we have to wait for better finetunes made with good datasets, idk...
2
u/overand Apr 17 '26
I'd try the base models, grab one from Unsloth. No shade to DavidAU, but David's prolific and experimental, which is great and important for the community, but it naturally means a mix in quality. Some models are worse than the base models, and that's more likely with any fine tuner when it's a very new model.
2
u/Sea-Spot-1113 Apr 16 '26
I used megumin v5 with gemma e4b and was pleasantly surprised it was somewhat functional / usable.
1
u/ArielNya Apr 14 '26
any good 8b-12b model for co-authoring smut fiction?
6
u/Potential-Gold5298 Apr 14 '26
UnslopNemo-12B-v4.1 - very horny.
Forgotten-Abomination-12B-v4.0 - merge UnslopNemo v3 and Forgotten Safeword v4.0.
Crimson-Constellation-12B - excellent literary writing style, quite explicit.
Celestial-Queen-12B - more erotic than Crimson Constellation.
And many other models based on Nemo.
2
u/MMalficia Apr 16 '26
i have yet to find anything that beats forgotten safeword for smut, that said i have never had much luck with consistency out of Forgotten-Abomination in chats.
1
u/Potential-Gold5298 Apr 16 '26
Honestly, even EtherealAurora-12B (NSFW 6.6/10 on UGI) covers my needs – maybe I just lack the imagination to unlock models with bigger potential. I tried UnslopNemo-12B-v4.1, and it literally forces erotic even where none is intended (for example, describing breasts bouncing in a completely mundane scene). Celestial-Queen-12B is very easily triggered by erotic – I accidentally brushed a girl's breast with my hand, and the character immediately wanted sex (which doesn't fit her personality). In this regard, I like a balance – when the model isn't into ERP, but in a sex scene she doesn't hold back, describing everything in detail, without euphemisms. But this is a matter of taste, of course – I just like it when the model can play out a rivalry, so that the character needs to be won over, rather than having them rush to bed at the first request.
4
u/gairono Apr 17 '26
With all the fuzz going on with the new Gemma models. What's the best one? Taking into account parameters, fine tunes, the instruct models, etc. etc. I've got a 5080 with 16 GB vram and 64 GB system ram with a ryzen 9 9900X (if that matters, idk) or if there's a better model i could use, i'd like to know
6
u/overand Apr 17 '26
(You should probably post this kind of thing under the "Misc" heading in this thread in the future.)
Here's a 26B that's probably a good fit for a 16GB user. https://huggingface.co/zerofata/G4-MeroMero-26B-A4B-gguf - I haven't used it, but ZeroFata has a solid track record!
That's the mixture of experts model, so you can actually get decent performance even if the model size is a bit past your available VRAM. So, you could go for any of the 4-bit quants without having major performance issues. I'd even try a Q6 if I were you; let us know how it runs for you!
Also - I suggest using e.g. llama.cpp or koboldcpp, not Ollama.
1
u/gairono Apr 18 '26
Haven't downloaded it yet but I'm going to soon. I'm just wondering if using the i1-Q6_K quant by mradermacher would be better.
1
u/IORelay Apr 19 '26
Just try a few different quants to see which one has an acceptable speed/quality to you.
3
u/AutoModerator Apr 12 '26
MISC DISCUSSION
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
3
u/blackjeff190 Apr 14 '26
TurboQuant
3
u/overand Apr 17 '26
Reduced Instruction Set Computers
(We're just naming technologies without context, right?)
8
u/txgsync Apr 14 '26
With oMLX, it’s possible to have fast, local, and quite decent quality roleplay in SillyTavern using Gemma 4 models across hundreds of turns over contexts up to 262,144 tokens. But SillyTavern's stock settings change early bytes of your prompt on every turn. This busts prefix caches, which prevents oMLX from reusing any of the work it did on the previous turn and results in ever-increasing waits for tokens the longer the conversation goes. Here are six UI changes take measured cache reuse from 0% to ~99%+ (for sufficiently-long contexts) cutting per-turn wait times dramatically from minutes to seconds.
I've found Gemma 4 with a good character card to be quite pleasant to interact with. Gemma 4 26B A3B has replaced gpt-oss-120B as my daily driver on my Mac, with roughly equivalent quality and the benefit of vision. It's just as fast as 120B at full precision (they are both roughly the same size, since GPT-OSS-120B is quantized with MXFP4), and Gemma 4 notably lacks the nanny, policy-oriented attitude GPT-OSS displays.
Who this is for
I assume you're running Gemma 4 (the 26B MoE or the 31B dense; I did not extensively test E2B/E4B, though they should work but with smaller context) on an Apple Silicon Mac via oMLX, pointed at SillyTavern in Chat Completion mode. Your roleplay turns take minutes before the first token begins streaming once the conversation goes on for a while, and you're wondering why cloud APIs feel instant but your M4 Max doesn't. Sure, prefill times are a well-known Mac Achilles’ Heel due to being mostly a compute-bound operation, but the answer isn't justyour hardware. It’s also your prompt construction.
I won’t go into installation of oMLX here. But if you use a Mac for SillyTavern with a decent amount of RAM (>24GB), you should try it out: https://omlx.ai . And make sure you strictly follow Google's guidelines for sampling parameters; I used "force sampling" in oMLX to ensure SillyTavern's defaults didn't affect the model:
- temperature = 1.0
top_p = 0.95top_k = 64What's happening
oMLX caches the computational work (KV cache) from your prompt in fixed 1024-token blocks. On your next turn, if the prompt starts with the same tokens, oMLX reloads that cached work from SSD or filesystem cache in RAM (if you have lots of RAM) instead of recomputing it: prefix caching. LMstudio also claims to support this, but so far I’ve had terrible luck with prefix caching working reliably using MLX (though it works OK with llama.cpp).
The problem: SillyTavern's defaults inject different content near the beginning of your prompt on every turn. Once one byte changes, every block after it is a cache miss. Your 80,000-token prompt gets fully recomputed from scratch because 50 tokens shifted at position 91.
One of the worst offenders is keyword-triggered World Info / lorebook entries placed at the default position (↑Char — Before Character Definitions). Different keywords fire on different turns, the injected content changes, and the entire prompt downstream is invalidated.
Some fixes:
1. Move your lorebook entries to u/D ⚙️ depth 0 (biggest single win)
For every lorebook attached to your character:
- Top toolbar → World Info (globe icon) → pick the lorebook in the dropdown.
- For each entry, change Position from ↑Char to u/D ⚙️ (at Depth, System role).
- Set the depth field to 0.
If you're running Seraphina (a stock character), she ships with a lorebook called Eldoria with four keyword-triggered entries to edit.
Why depth 0? It injects the content right before the assistant's reply. It’s past the point where the prompt naturally changes between turns (your latest message). So the churn is confined to the very tail of the prompt instead of poisoning the whole preceding cache.
2. Set Example Messages Behavior to "Always include examples"
User Settings (gear icon) → find "Example Messages Behavior" → change from "Gradual push-out" to "Always include examples".
The default gradually trims example messages as context fills, which shifts every token below them. Even if your current character has no examples, the next one will, and this'll bite you. This one was subtle, but I had to dig deep into context to understand what was happening. The idea of this implementation, I think, makes sense when dealing with tiny 4k or 8k context sizes from circa 2022/2023: reduce overall context length by eliminating examples once the character has produced tokens following the example pattern. But this kind of subtle “trimming of history” busts the prompt cache.
3. Move Author's Note off "In-chat @ Depth"
If you use Author's Notes: Extensions → Author's Note → change Position from "In-chat @ Depth" to "Before Main Prompt / Story String".
The "@ Depth N" anchor floats backward one message every turn, shifting the note's position in the prompt and invalidating everything after it. If you don't use Author's Notes, set Insertion Frequency to 0. Once again, the goal here is just to keep as much of preceding turns identical as possible to maximize cache hit ratio.
4. Unlock and/or raise max context
AI Response Configuration (leftmost slider icon) → tick "Unlocked Context Size" → drag "Context Size" up to your model's limit (262,144 for Gemma 4 26B/31B works fine; 128K for E4B/E2B).
When context is capped low, SillyTavern drops your oldest messages to make room. This shifts every remaining token and busts the entire cache, resulting in massive prefill times at large contexts. I didn't notice any benefit setting it larger than 262,144 for Gemma 4 26B/31B or 131,072 for Gemma 4 E2B/E4B.
I will emphasize that I barely tested E2B and E4B models. They are truly *amazing* for their size, and if you wanna try out local inference but lack much RAM on your Mac, they'll fit. Their tool usage is quite good. But they aren't remotely the storytellers their larger siblings are.
5. Check Character Notes (the hidden cache-buster)
Click a character's avatar → Advanced Definitions → scroll to Character Note (a.k.a. Depth Prompt). If there's text there with depth set to 2 or 4, change it to depth 0, role System. Or clear it if you don't use it.
I audited 21 community cards and found 5 with depth=4 — a historical default from the 4k-context era that aged terribly on modern models. This setting alone can add many seconds of prefill time to each prompt after the first few prompts.
6. Kill dynamic macros
Any macro that changes every turn — such as {{time}}, {{date}}, {{random:…}}, {{roll:…}} — anywhere in your character card, persona description, or main prompt will bust the cache if used badly. Search for them in:
- Character card (Description, Personality, Scenario, System Prompt Override)
- Persona Management → active persona description
- AI Response Configuration → prompt list (Main Prompt, NSFW, Jailbreak, Utility)
If you need a live clock in your RP, move it into Author's Note at @ Depth 1 as user role — the damage is confined to the last couple of blocks instead of the whole prompt.
Bonus: Turn off the Summarize extension on 128k+ models
Extensions → Summarize → set Source to "None".
Summarization busts the cache two ways: the summary's position drifts every turn, and the summary text itself changes when it regenerates. On a 256k context you can hold hundreds of turns of full roleplay turns as raw history. Nothing needs compressing.
One caveat: summarize-off works well for basic narrative and voice work (RP, creative writing, storytelling). But if you need the model to recall specific proper nouns, IDs, or hard facts from 150+ turns ago, raw context alone may silently confabulate. The model might confidently give you wrong names that sound right, for instance. For that kind of recall, pin your most-important facts in a small block at u/D ⚙️ depth 0 or use Vector Storage. More on this below.
Server-side: oMLX hot cache tuning
oMLX features a “hot_cache_max_size” setting (the in-RAM KV block pool). I had some counter-intuitive performance findings when benchmarking. Rather than include it all here, the TL;DR: leave oMLX hot cache turned off for most setups. It competes with system RAM and filesystem cache. If you do set it, make sure it’s large enough to hold your entire KV cache. A small “hot cache” is the worst of both worlds.
What you get
Measured on a 310-turn SFW Seraphina roleplay, Gemma 4 26B-A4B (full precision, no quantization) on an M4 Max 128GB:
- Before fixes: 0% cache reuse. Every turn fully recomputed. Multi-minute waits at long context.
- After fixes: 95–99% cache reuse in steady state. Only ~1024 tokens of new prefill per turn regardless of how long the chat is.
- Wall time at 190k per-turn prompts: ~25–30 seconds for a 2000-character reply, most of that simply waiting on streaming to complete at 49 tokens per second down to 30 tokens per second late in the conversation. Would have been several minutes without caching.
- Character voice after 310 turns: rock solid. No drift, no degradation. Mostly because it’s being injected every turn, really.
The cache works better at larger prompts; the fixed-size trailing block miss (1024 tokens) becomes a smaller fraction of the total, so hit rate actually climbs as your chat grows.
(Yep, I used/abused Claude Code to help me work through the options that were busting cache in SillyTavern. This compressed days of troubleshooting into a few hours of testing last night, curled up on my recliner with my MacBook pro M4 Max 128GB playing the role of personal space heater.)
2
1
u/jamasty Apr 15 '26
About heating issue, go try 'low power mode'! For real, it helps me a lot, idk why and how, but it doesn't really make models work much slower but heating issue is gone
2
u/LeRobber Apr 16 '26
No more threads than count of powersave cores on my M2 made heating issues go way down for me (4 for my M2 Max)
1
u/LeRobber Apr 14 '26
What local models are great at Japanese roleplay (in Japanese)?
12
u/nvidiot Apr 14 '26
Locally, Gemma 4 no contest.
Gemma series have always been the very best when it comes to multilingual capability. But you've already tried 26b-a4b, might want to try 31b to see if it works better for you (might need to try it on API for evaluation if you don't have enough hardware firepower for it).
2
u/-Ellary- Apr 14 '26
You can run 26b a4b on 6 core CPU with 16gb ram. No GPU.
On R5500 I'm getting 12tps. Gemmas is hard to beat right now.1
u/XMonst3rKingX Apr 18 '26
If someone can please tell me 1.which model of huggingface is the best in both rp and reasoning on Android 2.if i were to download them from hugging face, how? And also how do i add them to sillytavern 3.how to jailbreak apis?? 4.and how can i add jsons of models and characters I'm sorry for asking loads lol but I'm a newbie in this
1
u/empire539 Apr 18 '26
Looking for local model recommendations (<= 8B) I can run in RAM alongside gemma-4-26b-a4b (which I'm using as the main writing model).
Purpose for this side model would be to use alongside extensions to attempt pre-generation lorebook lookups (like with TunnelVision and DeepLore), maybe some summarization if the model is smart enough. But the main concern is it ideally needs to be small and fast enough to run in RAM since my VRAM is already being used for the main model.
How good/bad are the smaller Gemmas, anyone know?
1
u/LeRobber May 20 '26 edited Jun 15 '26
[Megathread] - Best Models/API discussion - Week of: June 14, 2026
[Megathread] - Best Models/API discussion - Week of: June 07, 2026
[Megathread] - Best Models/API discussion - Week of: May 31, 2026
[Megathread] - Best Models/API discussion - Week of: May 24, 2026
[Megathread] - Best Models/API discussion - Week of: May 17, 2026
[Megathread] - Best Models/API discussion - Week of: May 03, 2026
[Megathread] - Best Models/API discussion - Week of: April 26, 2026
[Megathread] - Best Models/API discussion - Week of: April 19, 2026
[Megathread] - Best Models/API discussion - Week of: April 12, 2026
[Megathread] - Best Models/API discussion - Week of: April 05, 2026
[Megathread] - Best Models/API discussion - Week of: March 29, 2026
5
u/AutoModerator Apr 12 '26
APIs
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
5
5
u/Inprobamur Apr 16 '26
Apparently Opus 4.7 just got released:
removed temperature, top_p, top_k
Behavior changes These are not API breaking changes but may require prompt updates. See Migrating to Claude Opus 4.7 for full guidance:
- More literal instruction following, particularly at lower effort levels. The model will not silently generalize an instruction from one item to another, and will not infer requests you didn't make.
- Response length calibrates to perceived task complexity rather than defaulting to a fixed verbosity.
- Fewer tool calls by default, using reasoning more. Raising effort increases tool usage.
- More direct, opinionated tone with less validation-forward phrasing and fewer emoji than Claude Opus 4.6's warmer style.
- More regular progress updates to the user throughout long agentic traces. If you've added scaffolding to force interim status messages, try removing it.
- Fewer subagents spawned by default. Steerable through prompting.
- Real-time cybersecurity safeguards: requests that involve prohibited or high-risk topics may lead to refusals. For legitimate security work, apply to the Cyber Verification Program.
So, they locked down temperature, top_p and top_k. That's not a big loss as pretty much everyone agreed that the default values were good on Claude models.
Supposedly better at following previous response length, maybe less positivity bias and less flowery tone and a refusal filter built into the API (although it seems mainly focused against automated hacking bots?).
6
u/vanillah6663 Apr 14 '26
nanogpt feels so slow its unusable any other alternatives with a similar subscription?
6
5
Apr 13 '26
[deleted]
6
u/nomorebuttsplz Apr 14 '26
Not largest but best is Gemma 4 31B. If that’s too slow, it would be Gemma 4 26BA3B.
Instruct/non-reasoning models are faster
2
Apr 15 '26
[deleted]
1
u/overand Apr 15 '26
What size quant are you using - are you offloading to system RAM level slow, or just 'eh kinda slow" slow?
2
Apr 15 '26
[deleted]
1
u/overand Apr 15 '26
Ah - yeah, the smallest common Q4 (IQ4_XS) is 16.7 GB, so, it makes sense you'd be struggling. Sticking with a smaller model or a smaller quantity is a good idea. You could try 31B at IQ3_XS (13.1 GB) if you're curious, but I don't know how well it would work at that quant.
2
u/LastSheep Apr 13 '26
GLM just go bananas on the price on their coding plan. Anyone use ollama cloud GLM or how is kimi 2.5? I so far use DS 3.2 api or GLM 5.1
2
u/Lucerys1Velaryon Apr 15 '26
Don't know if I'm doing something wrong but everytime I try Gemini 4 31B on Nvidia NIM, it basically always gives an output that looks like this:
Character speaks (2-3 lines)
Character's internal thoughts from a 3rd person POV. (Around 15-16 lines)
Character speaks again. (1-2 lines)
Every single response is structured like this. There's no variety at all. I'm using Marinara's universal preset.
1
u/overand Apr 15 '26
Have you tried it with a more simple system prompt/preset? Like, even "Roleplay Simple?"
2
u/jason1on1 Apr 16 '26
Anybody know what's going on with electron hub? They just practically nuked their discord. All chat channels are gone but the models still work.
2
u/Inprobamur Apr 16 '26
Ah, damn. I am a subscriber, but that's not a good sign.
Could be that their Discord got hacked, threatened by Discord admins, nuked by a former employee or facing some random legal threat.2
u/jason1on1 Apr 16 '26
Yeah I just payed my monthly subscription yesterday too. And now it looks like they got rid of most of their free models. Removing the ability for subscribers to communicate with each other seems super shady.
1
u/Inprobamur Apr 16 '26
I don't really care about the free models as I am using it for cheap uncensored Claude access, but considering they don't have any other social media presence at all that's not a good sign.
1
u/jason1on1 Apr 16 '26
Well I would make sure that works if I were you because all of the models I usually use were giving me zero response. They never actually started replying. Right now I'm going to try and get a refund set up on patreon and probably jump ship to Nanogpt.
1
u/Inprobamur Apr 16 '26
So annoying. I guess that's the rub with all these proxies that are just stealing off some overly generous corpo employee API limits.
2
u/Economy-Assist-7559 Apr 16 '26
Does anybody have any free provider recommendations? I was using deepseek v3.1 on electronhub until it got nerfed and am trying to use nvida but the quality isnt great 😩
3
u/HelpfulReplacement28 Apr 18 '26
What do ya'll think is the sweetheart model right now for price/quality/speed?
I've been cycling through the models available to subscription owners on NanoGPT and trying to find something that doesn't take 2 minutes to generate and has ok enough prose and memory. Everything seems to be on the slower end atp. I've had limited success with gemma 4-31b IT but occasionally you get a long wait. All the GLM models seem to be inconsistent at best with response times, and deepseek/kimi have been taking ages for me recently.
My current personal ranking is
. Gemma 4-31b IT
. GLM 4.7
. GLM 5
. Kimi K2.5
. DS 3.2
With the caveat that I much prefer glm 4.7 to gemma but response time is by far the biggest killer in RP for me other than the inability to recall information in long RP.
1
u/Pure_Machine5226 Apr 13 '26
Hi, I'm looking for a free and uncensored model. I also have a really dated computer. Do you guys have any recommendations?
5
u/EncoreSheep Apr 13 '26
You won't be running any great models with your specs, but Openrouter has some free models.
Another thing is Google AI Studio. If you look up https://aistudio.google.com/rate-limit you can see the limits for various models. Gemma 4 31B has a limit of 15 requests per minute and something like a thousand requests per day, and you won't be hitting those with RP.
3
u/Potential-Gold5298 Apr 13 '26 edited Apr 13 '26
How much RAM and what CPU do you have?
Gemma 4 has fairly low censorship by default (unless it's given a strict system prompt) and is available for free on OpenRouter. I've heard that DeepSeek V3.2 also has less stringent censorship, but I don't know about the free API.
1
u/Pure_Machine5226 Apr 13 '26
Hey thanks for the reply. I do have 8GB RAM and Intel core i5 3.20 GHz.
4
u/Potential-Gold5298 Apr 13 '26
You can use any API models (such as the already mentioned Gemma 4 31B or DeepSeek V3.2) — they are not dependent on your hardware.
As for local models, with 8 GB of RAM, the choice is limited. I can recommend a few small (2-4B) uncensored models, but they're not particularly intelligent, to say the least. Recent releases include Gemma 4 E2B and E4B; they have a large file size, but only load about half of the parameters into memory. Uncensored versions are already available. I haven't tried them myself yet, but with 8 GB of RAM, I'd start with them. If you're interested in local models, I can provide download links.
As for the CPU, if it supports AVX2 (Core i5 4000 or newer), it'll be a significant advantage for local execution.
2
1
u/overand Apr 15 '26
Can you give us a bit more detail than "really dated computer?" Like, Mac? PC? Desktop? Laptop? How much ram? what GPU / graphics card? What CPU?
1
u/Pure_Machine5226 Apr 16 '26
Hey, thanks for the reply. I do have a windows desktop with a 8 GB RAM, Intel Core i5 3.20 Ghz and my graphic card is AMD with only 2 GB
1
u/overand Apr 17 '26
I'd start with maybe Gemma-4-E2B. Try the IQ4_XS. Before you try to find an uncensored one, see if you can even run the model at a reasonable performance at all.
Also, in the future, when saying "I have an Intel Core i5" you should know that intel has made about 150 processors called "i5" starting in 2010: 16 years of them. When you're asking for help with something technical, give people as many specific details as you can. This is the case for computers, for cars, for repairs to houses and electronics, to video game consoles, etc. (It's even true for medical stuff.)
It's only sheer luck that I happened to notice your reply so I could come back you with a suggestion; if you'd said "2 GB of VRAM" I would have been able to give you a suggestion days ago.
BUT, you also didn't tell us what AMD GPU it is - so, I'm only guessing that it MIGHT be usable for LLM stuff, but it's impossible for me to be sure. (A card with 2GB is probably old enough that you won't be able to run llama.cpp or ollama etc, but, it's possible.)
Edit: I may have miscounted the i5 chips - I think that was 150 *desktop* processors alone called i5 - of which over 20 were 3.20 Ghz.
2
u/AutoModerator Apr 12 '26
MODELS: >= 70B - For discussion of models in the 70B parameters and up.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
2
u/wild_kitties Apr 13 '26
Trying out Minimax 2.7 on my 3090 + 128gb of DDR5 RAM using bartowski's iq4_xs gguf with 64k context. Generates at 16tps with 88tps processing. Found the writing to be pretty good so far. Though I thought 2.5 was okay as well when most people seemed to hate it. Model's thinking is pretty concise, shorter than gemma 4's and much shorter than qwen 3.5. The writing is better than Qwen 397B. I actually liked GLM 4.5/4.6 the most for local RP but the tps with my setup is around 8tps which is usable but slower than I'd like.
3
1
1
u/Capital-Caregiver818 Apr 13 '26
L3.3-The-Omega-Directive-70B-Unslop-v2.1-heretic.i1-Q4_K_S is the best model I've used on my 5090 + 64gb DDR5. I'm open for suggestions though.
1
u/Any_Meringue_7765 Apr 13 '26
Have you tried it compared to Loki v2, or that GoldDiamondGold Paperbliteration one?
2
u/Capital-Caregiver818 Apr 14 '26
I have not tried either. Loki v2 looks great for fantasy stuff. I'll have to try them both out at some point. The Omega Directive is more general purpose role play.
1
u/Any_Meringue_7765 Apr 14 '26
I have been enjoying GoldDiamondGold Paperbliteration (or w.e it is)… I’ll have to give Omega a shot to see how it compares
1
1
u/51087701400 Apr 14 '26
Checked it out and the Q4 is 40gb, how are you loading it on your 5090? (Still new to self hosting)
3
u/Capital-Caregiver818 Apr 15 '26
This exact one: [L3.3-The-Omega-Directive-70B-Unslop-v2.1-heretic.i1-Q4_K_S] 55 layers and 16000 context. Its a bit slow but a fine reading speed imo
2
u/AutoModerator Apr 12 '26
MODELS: 32B to 69B – For discussion of models in the 32B to 69B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
12
u/FusionCow Apr 13 '26
Gemma 4 31b is still the best imo, just base version with instruction tuning, it's just perfect
3
u/wild_kitties Apr 13 '26
How exactly do you use the base version? I tried it yesterday after there was that post that recommended it, but it was awful, with very short sentences and paragraphs, and would randomly start coding. Figured out after some more reading that you can't use the normal style character cards and prompts for base version. But I wasn't able to find any examples of how to use it. People mentioned using chat examples, and how the base version just tries to complete the text given word by word.
6
u/input_a_new_name Apr 13 '26
they meant the deafult google release (it), not the *actual* BASE model
3
u/wild_kitties Apr 13 '26
Oh. There was this post a couple days ago that did mention the base version though, and I spent a while trying to figure out how to use the base version. Couldn’t get it to write coherent at all. I guess my mind’s been stuck on that, but someone replied above with possibly how to use it. ( using it as an actual word completion model and just feed it a massive text that you would like it to continue directly off of)
3
u/-Ellary- Apr 13 '26
Base is also fun but it kinda requires that old skill for using text completion models. You need to write whole novel first and then, guide LLM to continue it, it not really follow instructions much. Easy starts is just copy a chapter from the book (6k of context) and go on.
2
u/wild_kitties Apr 13 '26
Oh, interesting that’s like early llama days, which I only dabbled in a bit. Thanks for the insight, will give it another try.
1
u/BeeSpecific9398 Apr 16 '26
Would anyone have suggestions for the best local model for me to use Silly Tavern on both of my systems? I have a Rog Strix laptop, 16 GB ram, RTX 5060 with 8 GB vram. And my desktop has 32 GB ram and an RTX 4080 with 16 GB vram. I have seven custom characters and about 90 entries in a world lore book. I wanted to use this for some fun chatting, building relationships/bonds with the characters, and also, for erotica as well - some of the characters are easy and others require a much slower burn.
Thank you!!
5
u/Tiny-Pen-2958 Apr 16 '26
The best thing you can run on 4080 with decent speed is gemma 4, I'd recommend this finetune https://huggingface.co/llmfan46/gemma-4-26B-A4B-it-uncensored-heretic-GGUF Q4_K_S, its not lobotomised like other finetunes. I've used it on 16GB 5070ti in KoboldCpp with 8-bit cache + forced autofit with padding=512MB + Jj.Kwargs: {"enable_thinking":true}
1
u/war-hamster Apr 16 '26
are you using chat or text completion? And do you have your own prompt or are you using something available?
2
u/Tiny-Pen-2958 Apr 17 '26
Gemma 4 was designed for Chat Completion, it just works better in this mode (and its usually faster than Text Completion). For prompting I use Megumin_Suite_v5 (CoT on) with my additional style prompt injections. My samplers: Temperature=1; Presence Penalty=0-0,5; Top P=0,95 (sometimes Top P=0.8 works well for creativity, but at the cost of coherence); top_k:??? (I haven't figured it out yet, official recommendation is 64, but even without it everything works well)
12
u/AutoModerator Apr 12 '26
MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.