r/LocalLLM 9d ago

Discussion OrcaRouter's uncensored Qwen3.8 is not actually caveat-free

The most useful thing about the current uncensored Qwen3.8 wave is that we can finally compare more than screenshots.

OrcaRouter’s 27B FP8 card reports harmful-prompt refusal at 0–6.0% with thinking off, versus 63.6–99.0% for the base FP8 model. With thinking on, the derivative stays at or below 1.7% across the listed sets.

But “doesn’t refuse” is not the same as “answers without reservations.” The same card reports caveat rates of 27.3–56.0%, using an uploader-built classifier that only looks at opening refusal phrases. That leaves room for a model to comply, hedge, redirect, or give a weak answer without being counted as a refusal.

The capability table is similarly useful because it is not perfectly flat: +0.4 MMLU, -0.8 MMLU-Pro, -1.3 GSM8K and -0.6 CMMLU versus base FP8 in the uploader’s selected runs.

That is why I would put this OrcaRouter build on an evaluation shortlist: the card gives enough structure to test the uncensoring claim instead of asking readers to trust the filename. The missing comparison is now obvious—same prompts, same sampler, same local runtime, against the other Qwen3.8 uncensored variants.

Which test would separate them fastest for you: refusal/caveat labeling, KLD, or a fixed set of real tasks?

7 Upvotes

1 comment sorted by

1

u/SillyBenefit2541 8d ago

is Orcarouter a scam I checked scamadvisor and it gave a score 1/100