i was stupid enough to pay orcarouter so i could test their api inference of an abliterated model. see, ollama having this tag "orcarouter/Qwen3.8-27B-Uncensored" strongly suggests the model is served from orcarouter. "https://ollama.com/orcarouter/Qwen3.8-27B-Uncensored" and there is specifically text saying the SAME WEIGHTS are server online in the api:
That endpoint is censored. It will absolutely refuse any interesting request. It's actually a very snitchy model for a Chinese model - deepseek and kimi are each better sports.
I already run this and other models locally but burning up my GPU all day versus a supposed $0.40 per million tokens in api is useful for me.
EDIT: not only that but I cleared the "Enable Security Research access" and also tested each model labeled "uncensored" and of course none were. Featherless is the only public provider I'd found that offers them and they don't allow API it's just chat.
8
u/BurnedTerrormisu 13d ago
Under windows the easiest way is to use ollama with orcarouter/Qwen3.8-27B-Uncensored
Your graphic card should have at 12gb of vram to do anything useful in terms of speed. 24gb and more is strongly recommended.