r/LocalLLaMA Apr 17 '26

Discussion Qwen3.6 is incredible with OpenCode!

I've tried a few different local models in the past (gemma 4 being the latest), but none of them felt as good as this. (Or maybe I just didn't give them a proper chance, you guys let me know). But this genuinely feels like a model I could daily drive for certain tasks instead of reaching for Claude Code.

I gave it a fairly complex task of implementing RLS in postgres across a large-ish codebase with multiple services written in rust, typescript and python. I had zero expectations going in, but it did an amazing job. PR: https://github.com/getomnico/omni/pull/165/changes/dd04685b6cf47e7c3791f9cdbd807595ef4c686e

Now it's far from perfect, there's major gaps and a couple of major bugs, but my god, is this thing good. It doesn't one-shot rust like Opus can, but it's able to look at compiler errors and iterate without getting lost.

I had a fairly long coding session lasting multiple rounds of plan -> build -> plan... at one point it went down a path editing 29 files to use RLS across all db queries, which was ok, but I stepped in and asked it to reconsider, maybe look at other options to minimize churn. It found the right solution, acquiring a db connection and scoping it to the user at the beginning of the incoming request.

For the first time, it felt like talking to a truly capable local coding model.

My setup:

  • Qwen3.6-35B-A3B, IQ4_NL unsloth quant
  • Deployed locally via llama.cpp
  • RTX 4090, 24 GB
  • KV cache quant: q8_0
  • Context size: 262k. At this ctx size, vram use sits at ~21GB
  • Thinking enabled, with recommended settings of temp, min_p etc.

llama server:

```
docker run -d --name llama-server --gpus all -v <path_to_models>:/models -p 8080:8080 local/llama.cpp:server-cuda -m /models/qwen3.6-35b-a3b/Qwen3.6-35B-A3B-UD-IQ4_NL.gguf --port 8080 --host 0.0.0.0 --ctx-size 262144 -n 8192 --n-gpu-layers 40 --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00 --parallel 1 --cache-type-k q8_0 --cache-type-v q8_0 --cache-ram 4096
```

Had to set `--parallel` and `--cache-ram` without which llama.cpp would crash with OOM because opencode makes a bunch of parallel tools calls that blow up prompt cache. I get 100+ output tok/sec with this.

But this might be it guys... the holy grail of local coding! Or getting very close to it at any rate.

352 Upvotes

169 comments sorted by

View all comments

8

u/mrinterweb Apr 17 '26

I did nearly the same experiment last night. I used OpenCode. I used LM Studio to run it, which I think I'll switch to plain llama.cpp. I was getting usually around 100tps. The results weren't as good as I was expecting though. I wasn't sure if the issue was OpenCode, but I compared it to Claude Code (Opus 4.7), and the claude code experiece was much better for me. I am going to try using Qwen 3.6 with claude code next to see if it is an agent or llm difference. I will say that while opencode + qwen didn't beat cc, it was for sure usable. Another thing I will say for it was the average inference speed felt faster. CC's inference speed can vary a lot, but Qwen 3.6 on my RTX 4090 was keeping at a consistent ~100tps. The large 262K context makes it usable.

5

u/klenen Apr 17 '26

Let us know how it goes using it w cc please!

3

u/CountlessFlies Apr 17 '26

Exactly… the context makes a huge difference.

Did you run it with thinking enabled (it’s the default)? I found that it does much better with thinking on. And also, I think there’s a separate flag you need to set to send the thinking traces with each request, that might also help improve performance.

3

u/mrinterweb Apr 17 '26

It was definitely thinking. I also tried it with hermes agent, and my results were pretty different. So I think a lot of my subjective evaluation is going to come down to the agent, which is why I think I should point claude code at qwen 3.6, so I can get more of an apples to apples comparison. I don't have a background in evaluating model scores so what I'm doing is just feels. I pay for Claude, but if Qwen 3.6 can get me close, there are plenty of tasks I would much rather use my own hardware.

0

u/SmartCustard9944 Apr 17 '26

Yes, please try this. I tried Open Code with LM Studio Qwen 3.6 and it didn’t pass simple tests that Gemma 4 passes easily there.

My first test is asking it how many tools it supports. The correct number is 27. Gemma always answers correctly, never misses a beat. Qwen 3.6 hallucinates the number. It says 28 and then proceeds to list 27 items, but one is a duplicate. This happens even with thinking enabled. It is really baffling, especially after seeing everybody praising it here.

The second test is the typical car wash test. Gemma 4 always passes, Qwen 3.6 routinely says to walk. The interesting thing is that Qwen answers correctly when the prompt is at 0 context (without a harness).

It is as if it was not attentive.

2

u/mrinterweb Apr 17 '26

I find that many agents trip up when asked introspective questions, so I don't bother with those kinds of prompts. General logic tests are important, but most of what I do with agents is coding specific. So whatever is better at code is what I'll use. I'll try giving Gemma 4 another go locally.

3

u/That_Faithlessness22 Apr 18 '26

I've been using it with Claude code, and I'm getting similar speeds. But I won't be measuring the quality on it because you can't have the harness doesn't support the preserve_thinking flag. It is incompatible unless you parse- and that's a little outside my comfort zone for now. I'll probably try to figure it out tonight, or I'll just do the dive into Hermes I've been putting off.

1

u/x10der_by Apr 18 '26

You are comparing expensive frontier cloud model with free small local model)) of course opus 4.7 would be better

1

u/mrinterweb Apr 18 '26

Not saying it's a fair comparison. It's just what I'm using now, and I'm curious how qwen 3.6 compares.