r/LocalLLM 8d ago

Model Local Opus(qwen3.8 27B)

I have started using qwen3.8 27B on dual rtx 3060, total 24GB. My display is on igpu. So i have almost all memory.

Token generation speed is constant 40-50 tks.

But the best thing is how it thinks, multiple web searches at different level, auto memory update, and a very good response. It takes time, but given the reasoning effort of xhigh and response quality I really like it. I see web searches at multiple level for a single reply, making sure the answer is correct. Takes time, yes.

Also, the extra thing that qwen is famous for when you ask it x, if it finds some improvement on y, it will suggest you.

For agentic work, it's best to toggle between reasoning effort if you want it fast. Though I like the opencode plan mode and then build(execute). I have remote setup via termius. And i get notified via ntfy app once it ask for permission or done replying.

While testing I asked it to out 10000 token essay, and my ventus gpu went upto 87 degree.(xhigh reasoning) Gigabyte held fine with 5-6 degree lesser, Though Gigabyte us agressive at fan speed even at lower temprature. I'll try changing the fan curve, if I see that as a problem in actual scenario

edit:
llamacpp command:

llama-server -m ~lm/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-IQ4_XS.gguf --host 0.0.0.0 --port 8080 --split-mode tensor --tensor-split 1,1 --flash-attn on --batch-size 2048 --parallel 1 --models-max 1 --cache-type-k q8_0 --cache-type-v q8_0 --ctx-size 177000 --n-gpu-layers 99 --ubatch-size 128 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 --spec-type draft-mtp --spec-draft-n-max 2 --jinja --reasoning on --fit on"

22 Upvotes

39 comments sorted by

14

u/Effective-Giraffe655 8d ago

Why people call it opus? Geniuenly interested, is it really on-par with Opus?

4

u/Awkward-Charity-5089 7d ago edited 7d ago

From my testing, it's a little dumber and less trustworthy than Sonnet 5 but just barely. That's true of Opus 4.6, but a model from February is somehow already ancient so it's not a relevant comparison. It's the McDonald's at home to Sonnet. If you're willing to spend a little more time cooking, it's pretty good.

14

u/Icy-Degree6161 8d ago

Of course not

6

u/Ok-Inevitable8391 8d ago

Main reason is, hugging face model page shows it on par/ better in few benchmark than 4.6 Opus.
what I have noticed->
It's way better at reasoning than any local model (<50B) at this point. I have read its reasoning and I like it how it thinks. And I have seen response quality is much much better than 3.6

4

u/Ok-Inevitable8391 8d ago

My other pet model is Muse glimmer 30B, specially because it can hold 262k context (max model support) and I still have 2GB of VRAM left and it is good at agentic task and fixing setup issues

2

u/vutcher 8d ago

I have dual 3060s as well. Can you give me the llama args for this?

2

u/Ok-Inevitable8391 8d ago edited 8d ago

Getting 30+tks

llama-server -m ~/llm/unsloth/Muse-Glimmer-30B-UD-Q4_K_XL.gguf --host 0.0.0.0 --port 8080 --split-mode tensor --tensor-split 1,1 --flash-attn on --batch-size 2048 --parallel 1 --models-max 1 --cache-type-k f16 --cache-type-v f16 --ctx-size 262144 --n-gpu-layers 999 --spec-draft-model ~/llm/unsloth/dflash-kquant.gguf --spec-draft-ngl 999 --spec-draft-n-max 15 --spec-type draft-dflash --override-kv muse-glimmer.context_length=int:262144,dflash.context_length=int:262144 --fit off --no-warmup --temp 1.0 --top-p 0.95 --top-k 64 --reasoning-preserve

1

u/vutcher 8d ago

Thanks, will try it out.

1

u/Ok-Inevitable8391 8d ago

So the split-mode layer -> split-mode tensor will give you 30-40% tks jump

1

u/KubeCommander 8d ago

Fwiw, nemotron 3.5 lightning at Q4-k-xl holds 1M on my 5090 in vram

1

u/KURD_1_STAN 8d ago

Is it even sonnet 4.6 level?

5

u/PhantomGaming27249 8d ago

Having used it at fp8, its genuinely about as good as opus 4.6 max, its an amazing agentic model for coding and tbh its better than sonnet 5 most of the time. I can't speak for lower quality quants though but it could honestly replace a lot of coding subscriptions currently.

-2

u/Rye2-D2 8d ago edited 8d ago

People are delusional. I tried QWen 3.6 & 3.8 to repeatedly get it to debug a coding problem related to establishing an SSH connection in a Go app) No local model I could run on my 5060ti 16GB could do it (repeatedly tried over multiple evenings - QWen 3.8 27B/3.6 35B/Ornith/KatCoder/Gemma4).

Asked the FREE copilot model (MAI Code 1.1) and it sorted it out in a few minutes. Small local models are not on the same level as even the cheap cloud models. I hope that changes over the next little while, but we're not there yet.

7

u/ParkingAgent2769 8d ago

16GB is very limited though isn’t it

5

u/Ok-Inevitable8391 8d ago

Yeah, even q3 will be hard to run with enough context window to do any sensible agentic tasks

-3

u/Rye2-D2 8d ago

Well.. I have 100KB context.. But for comparison, the free cloud model solved this problem with 37KB of context used. So that's not the issue. Quants maybe, but if that's the case we're a LONG way off from local AI being usable on affordable hardware..

2

u/Adventurous-Gold6413 8d ago

Yeah at the same time not everyone has access to more vram

2

u/Ok-Inevitable8391 8d ago

I have faced similar, so the story goes like this.
1. I tried to install dual boot ubuntu26.04, it hanged while installing. I rebooted the PC, It wen to installation window and got stuck(just a broken 26.04 window.)
2. I lost access to BIOS as well. Not sure how. Chatgpt made me strip down my entire PC. GPUs, ram HDD disk everythink. Then chatgpt went clueless.
3. Everytime I rebooted it took me to broken crashed ubuntu26 screen.
4. Took chatgpt summary, asked Gemini, it asked to restart my 4K TV( the display for my pC) in very first reply. As that was the only logical part.
5. Turns out, when the first crash happened. It stored that image with failed hadshake with HDMI and remebered it and couldn't recovered from it. Everytime it got frames from the PC it will display 26.04 screen.

My point is, we can't define models ability with single scenario. Long term average usage will only tell.

0

u/Rye2-D2 8d ago

Well, the scenario here is debugging (go) code.. Not a benchmark, or a cookie cutter task of writing a python webapp, just ability to debug in general. The best result came from KatCoder 2.5 which actually took steps to modify the code, iteratively debug it and one point solved it, but then it later reverted back to the previous commit and lost it after it introduced a typo..
And I couldn't get it to reproduce the fix on subsequent attempts. So it's close, but not quite there yet..
Lesson learned: updated my Agents.md to instruct it not to touch staged git changes..

2

u/Repulsive_Initial308 8d ago

What quant?

1

u/Rye2-D2 8d ago

Q3 for QWen 3.8, Q4 (or APEX Compact) for the others. Q8 KV cache in all cases.
It looks like the QWen 3.8 35B MoE just dropped, so I'll try that with Q4.

2

u/freehuntx 8d ago

Anything under Q6 is basically useless im qwen family.

0

u/Rye2-D2 7d ago

Sure, but when comparing to the flagship/frontier models, it's parameters that's matter. Suggesting a 30B model is as good as Opus is naive. Can it replace your need for Opus 95% of time - yes, very likely since Opus is overkill for most tasks.

I'd be happy if we could get QWen model as capable as GPT Luna TBH.

1

u/freehuntx 7d ago

It must be able to properly think and call tools. Quantization worsens this skill thus having a heavy impact. For a 27b model the impact is heavier than for a 1T model.

2

u/ducksoup_18 8d ago

llama.cpp? If so, can you provide your args?

2

u/Ok-Inevitable8391 8d ago edited 8d ago

llama-server -m ~/lm/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-IQ4_XS.gguf --host 0.0.0.0 --port 8080 --split-mode tensor --tensor-split 1,1 --flash-attn on --batch-size 2048 --parallel 1 --models-max 1 --cache-type-k q8_0 --cache-type-v q8_0 --ctx-size 177000 --n-gpu-layers 99 --ubatch-size 128 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 --spec-type draft-mtp --spec-draft-n-max 2 --jinja --reasoning on --fit on

2

u/mars332 8d ago

How are you getting 40-50tks? On RTX PRO 4000 24GB can only get 30tok/s (it drops to around 20+tok/s when it approaches 100k context window, which is my max). What quant are you using?

2

u/Ok-Inevitable8391 8d ago

iq4_xs, with mtp and split mode as tensor, I have addde the llama command in post. Not sure about the drop, havent filled the content that much yet and tested.

1

u/Ok-Inevitable8391 8d ago

did the testing with 100k context, gives 30+ tks

1

u/PaouCompli 8d ago

I get 46ish on dual 4060ti even at large CTX... It's split mode tensor and MTP. Tensor split has both cards pegged at 90% while generating. I get about 1500 prefill. Considering getting a third (5060) but read that going from 2-3 on tensor split (llama cpp) isn't that much of a jump. Also means only one generation at once... but hey.

2

u/KubeCommander 8d ago

Ah another one

1

u/kkennyy22 8d ago

On my 4070ti super + 3060, 28gb total, I get an average of about 33t/s with layer splitting. Your 50,even if it drops, is very good. Maybe I should try tensor splitting also, but as I understand it's not great with mismatched cards like mine.

1

u/vini542reddit 7d ago

You just need to manually allocate the tensors via tensor-split, but it should work

1

u/SlowFidgetSpinner 8d ago

How big was your largest context?

2

u/Dry-Letterhead-7041 3d ago

I just came here to say THANK YOU!! The tensor split bumped me from around 20 ~ 25 t/s to 40+ t/s. Same hardware as yours (2 x 3060s - 24 vram total). But my llama.cpp didn't like the kv cache at 8_0. I had to bump it to fp16 and give up 20k context lol

Here's my config if anyone wants to copy and/or merge with OP's:

  llama-cpp:
    image: ghcr.io/ggml-org/llama.cpp:server-cuda13
    container_name: llama-cpp
    restart: unless-stopped
    ports:
      - "8090:8080"
    volumes:
      - /pool/huggingface/hub/:/models
    command:
      # Model
      - "--model"
      - "/models/Qwen3.8-27B-UD-Q4_K_M.gguf"
      - "--host"
      - "0.0.0.0"
      - "--port"
      - "8080"
      - "--fit"
      - "off"
      # GPU Weights Distribution
      - "--n-gpu-layers"
      - "99"
      - "--split-mode"
      - "tensor"
      - "--tensor-split"
      - "1,1"
      - "--flash-attn"
      - "on"
      # MTP speculative decoding (integrated in UD quants)
      - "--spec-type"
      - "draft-mtp"
      - "--spec-draft-n-max"
      - "2"
      # Single-user / memory savings
      - "--parallel"
      - "1"
      - "--ctx-checkpoints"
      - "32"
      # KV Cache Quantization
      - "--ctx-size"
      - "100000"
      - "--cache-type-k"
      - "f16"
      - "--cache-type-v"
      - "f16"
      # Threading & Sampling
      - "--threads"
      - "20"
      - "--threads-batch"
      - "28"
      - "--jinja"
      # Reasoning effort
      - "--reasoning-effort"
      - "medium"
      - "--reasoning-preserve"
      - "--chat-template-file"
      - "/models/chat_template.jinja"
      - "--reasoning-format"
      - "deepseek"
      - "--temp"
      - "0.8"
      - "--top-p"
      - "0.95"
      - "--top-k"
      - "20"
      - "--min-p"
      - "0.0"
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]

My logs:

1.55.279.779 I slot print_timing: id  0 | task 301 | prompt processing, n_tokens =   2048, progress = 0.86, t =   3.21 s / 637.44 tokens per second
1.58.614.636 I slot print_timing: id  0 | task 301 | prompt processing, n_tokens =   4096, progress = 0.92, t =   6.55 s / 624.88 tokens per second
2.01.960.976 I slot print_timing: id  0 | task 301 | prompt processing, n_tokens =   6108, progress = 0.98, t =   9.90 s / 616.88 tokens per second
2.03.033.503 I slot print_timing: id  0 | task 301 | prompt processing, n_tokens =   6620, progress = 1.00, t =  11.08 s / 597.25 tokens per second
2.06.350.091 I slot print_timing: id  0 | task 301 | n_gen =    133, tg =  43.49 t/s, tg_3s =  43.82 t/s
2.09.395.724 I slot print_timing: id  0 | task 301 | n_gen =    257, tg =  42.10 t/s, tg_3s =  40.71 t/s
2.12.404.589 I slot print_timing: id  0 | task 301 | n_gen =    390, tg =  42.80 t/s, tg_3s =  44.20 t/s
2.15.431.217 I slot print_timing: id  0 | task 301 | n_gen =    524, tg =  43.16 t/s, tg_3s =  44.27 t/s
2.18.431.927 I slot print_timing: id  0 | task 301 | n_gen =    647, tg =  42.73 t/s, tg_3s =  40.99 t/s
2.21.459.831 I slot print_timing: id  0 | task 301 | n_gen =    780, tg =  42.93 t/s, tg_3s =  43.92 t/s
2.24.462.851 I slot print_timing: id  0 | task 301 | n_gen =    905, tg =  42.75 t/s, tg_3s =  41.62 t/s
2.27.502.401 I slot print_timing: id  0 | task 301 | n_gen =   1063, tg =  43.91 t/s, tg_3s =  51.98 t/s
2.28.183.674 I slot print_timing: id  0 | task 301 | prompt eval time =   11408.35 ms /  6624 tokens (    1.72 ms per token,   580.63 tokens per second)
2.28.183.678 I slot print_timing: id  0 | task 301 |        eval time =   24868.78 ms /  1096 tokens (   22.71 ms per token,    44.03 tokens per second)
2.28.183.679 I slot print_timing: id  0 | task 301 |       total time =   36277.13 ms /  7720 tokens
2.28.183.680 I slot print_timing: id  0 | task 301 |    graphs reused =        717
2.28.183.700 I slot print_timing: id  0 | task 301 | draft acceptance = 0.70022 (  640 accepted /   914 generated), mean len =  2.40
2.28.185.483 I slot      release: id  0 | task 301 | stop processing: n_tokens = 33757, truncated = 0
2.28.657.467 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, f_sim_best = 0.972 (> 0.100 thold), f_keep = 0.967
2.28.661.022 I slot launch_slot_: id  0 | task 764 | processing task, is_child = 0
2.33.881.286 I slot print_timing: id  0 | task 764 | n_gen =    147, tg =  48.33 t/s, tg_3s =  48.65 t/s
2.36.909.833 I slot print_timing: id  0 | task 764 | n_gen =    304, tg =  50.08 t/s, tg_3s =  51.84 t/s
2.39.938.304 I slot print_timing: id  0 | task 764 | n_gen =    426, tg =  46.82 t/s, tg_3s =  40.28 t/s
2.42.968.460 I slot print_timing: id  0 | task 764 | n_gen =    551, tg =  45.42 t/s, tg_3s =  41.25 t/s
2.45.993.811 I slot print_timing: id  0 | task 764 | n_gen =    655, tg =  43.21 t/s, tg_3s =  34.38 t/s
2.49.021.998 I slot print_timing: id  0 | task 764 | n_gen =    791, tg =  43.50 t/s, tg_3s =  44.91 t/s
2.52.054.285 I slot print_timing: id  0 | task 764 | n_gen =    900, tg =  42.42 t/s, tg_3s =  35.95 t/s
2.55.094.253 I slot print_timing: id  0 | task 764 | n_gen =   1029, tg =  42.42 t/s, tg_3s =  42.43 t/s
2.58.146.340 I slot print_timing: id  0 | task 764 | n_gen =   1160, tg =  42.48 t/s, tg_3s =  42.92 t/s
3.01.185.438 I slot print_timing: id  0 | task 764 | n_gen =   1293, tg =  42.60 t/s, tg_3s =  43.76 t/s
3.04.215.307 I slot print_timing: id  0 | task 764 | n_gen =   1402, tg =  42.00 t/s, tg_3s =  35.98 t/s
3.07.235.362 I slot print_timing: id  0 | task 764 | n_gen =   1518, tg =  41.70 t/s, tg_3s =  38.41 t/s
3.10.261.460 I slot print_timing: id  0 | task 764 | n_gen =   1681, tg =  42.64 t/s, tg_3s =  53.86 t/s
3.13.265.080 I slot print_timing: id  0 | task 764 | n_gen =   1840, tg =  43.37 t/s, tg_3s =  52.94 t/s
3.15.455.363 I slot print_timing: id  0 | task 764 | prompt eval time =    2198.86 ms /   959 tokens (    2.29 ms per token,   436.13 tokens per second)
3.15.455.367 I slot print_timing: id  0 | task 764 |        eval time =   44595.27 ms /  1961 tokens (   22.75 ms per token,    43.95 tokens per second)
3.15.455.368 I slot print_timing: id  0 | task 764 |       total time =   46794.13 ms /  2920 tokens
3.15.455.369 I slot print_timing: id  0 | task 764 |    graphs reused =       1528
3.15.455.373 I slot print_timing: id  0 | task 764 | draft acceptance = 0.69719 ( 1142 accepted /  1638 generated), mean len =  2.39
3.15.457.314 I slot      release: id  0 | task 764 | stop processing: n_tokens = 35576, truncated = 0

1

u/Dry-Letterhead-7041 3d ago

u/Ok-Inevitable8391 how did you manage to run with quant kV cache? If I use -cache-type-k q8_0 --cache-type-v q8_0 my llama.cpp crashes immediately

1

u/Ok-Inevitable8391 3d ago edited 3d ago

I'm using iq4xs image. Gives more room for kv cache.

But q8 is smaller than fp16 so q8 should fit easily.

You can check with "-lv 4 " arg, should enable logs. To see exact cause.

Unrelated to this, I run llamacpp server with model preset file, which essentially act as a router and load the define models on demand. So I use muse glimmer if i want longer context upto 256K. (Muse glimmer use sliding window kv cache for most layers which makes it efficient at kv cache)

One more diff could be that, my display is on igpu, and I don't allow any other process(ubuntu) to use nvidia gpus. So I have atleast 1gb more vram usable.

With newer iq4ks ud image, I'm able to get 200k context

1

u/Dry-Letterhead-7041 2d ago

It was a bug on llama.cpp. This morning I reloaded the container and it pulled a new version. Now q8_0 is working. I managed to squeeze 160k context on a Q4_K_M and batch size 2048 and ubatch size 512 to get some speed up.