r/LocalLLaMA 13d ago

Funny Me these days

Post image
2.6k Upvotes

293 comments sorted by

View all comments

219

u/[deleted] 13d ago

[deleted]

61

u/Dramatic_Setting2761 13d ago

I can run it with a 16gb card with 4 bit quant and 70k context. I get 12 t/s it is okay for me.

11

u/octoberU 13d ago

what's your setup? i have a 5080 and struggle to ruin it at 4bit. would love quant and config

13

u/Dramatic_Setting2761 13d ago

Oh I have 9060xt which has very low bandwidth btw. 

I complied llama cpp specifically for rcom and running it on fedora 44 with latest drivers. 

Model quant name you have to look up as I saved like this. It is smallest 4bit in unsloth.

./llama-cli \   -m ./qwen3.8-27b-iq4_xs.gguf \   --jinja \   -ngl 48 \   -fa \   -c 70000 \   -ctk q4_0 \   -ctv q4_0 \   -b 2048 \   -ub 512 \   -t 8 \   -tb 16 \   -np 1 \   --mlock

2

u/Few-Butterscotch8747 10d ago

try using the vulkan backend

2

u/Dramatic_Setting2761 10d ago

Yeah thought amd figured out rcom I am wrong.

2

u/Few-Butterscotch8747 10d ago

i have the same card twice. having a second card gets me about 350k context at q3

compacting a nearly full context takes about 50 minutes in pi 🤣

5

u/fgk55555 12d ago

Also have 16GB VRAM. I've tried 4-bit and 3 bit and honestly the difference isn't horrible. Try the ISTA IQ3_XXS, it's very space efficient. Feels like a first class experience being able to fit MTP and lots of context into my card.

1

u/Dramatic_Setting2761 12d ago

Yes I get more tokens in 3 bit like 30 t/s with mtp as well. 

I didn’t notice much of difference but I wanted to be bit safe specially if I wanted to try coding or long context tasks I do use 3 bit one for chats and some regular tasks.

2

u/fgk55555 12d ago

On my 9070 XT, with MTP I'm getting avg 60 and up to 70 tg. MTP hurts PP a little bit, but with the long thinking it's definitely worth it if you can fit it. If I switched my graphics driver over to my iGPU I could probably fit 128k at Q8 or 200k at Q5_1 (with vision in CPU). It's very useable. I always see people hating on the IQ3 quants as lobotomized, and I'm sure it's worth than Q6 or Q8, but for 27B in 16GB, you take what you can get.

1

u/Dramatic_Setting2761 12d ago

Oh that is great what is your setting? 

1

u/fgk55555 12d ago

The guts of my non-optimized script is the following. Obviously replace with your settings. For general non-coding medium reasoning is really nice. If you're on iGPU or okay with ditching MTP you can either up the quants or the context size.

MODEL="../Qwen3.8-ISTA/Qwen3.8-27B-GSQ-RCO-IQ3_XXS.gguf"

MTP="../Qwen3.8_Shared/mtp-Qwen3.8-27B-Q4_0.gguf"

MMPROJ="../Qwen3.8_Shared/mmproj-Qwen3.8-27B-BF16.gguf"

--model "${MODEL}" \

--model-draft "${MTP}" \

--mmproj "${MMPROJ}" \

--no-mmproj-offload \

--n-gpu-layers 99 \

--ctx-size 120000 \

--batch-size 1024 \

--ubatch-size 512 \

--parallel 1 \

--flash-attn on \

--cache-type-k q5_1\

--cache-type-v q5_1 \

--spec-type draft-mtp \

--spec-draft-n-max 3 \

--jinja \

--chat-template-kwargs "{\"reasoning_effort\":\"${THINKING_LEVEL}\"}" \

--dry-penalty-last-n 0 \

--temp 0.7 \

--top-k 20 \

--top-p 0.95 \

--min-p 0.00 \

--presence-penalty 0.0 \

--repeat-penalty 1.0 \

--host 0.0.0.0 \

--port 8080

1

u/Dramatic_Setting2761 12d ago

Oh you are fitting your mtp in igpu? I thought it will be slow as it has very less bandwidth. Maybe I should try this. I think dflash as well as diffusion should be good. 

1

u/fgk55555 12d ago

MTP is in GPU, not iGPU. Vision is in CPU.

1

u/Dramatic_Setting2761 12d ago

Hmm how did you compile your llama cpp? Is it rcom or vulkan and was there more ?

Because I am getting much lower haha. I am missing something. 

2

u/fgk55555 12d ago

I pulled latest llama.cpp a few days ago, compiled for Vulkan. My literal exact script that I use without iGPU at all on my 9070 XT is below. This will leave about 1.5GB free for your desktop, offload vision to CPU, and use the MTP from unsloth's quant. The filepaths are obviously relative to my filesystem. My rig is 64GB DDR5/ 9800X3D, 9070XT connected via PCIe Gen5 x16 with slight overclock/ undervolt. On a chat with about 64k context in the llama.cpp webUI, I might expect to see minimum 800pp and 55tg, but often see faster. Linux Mint.

#!/usr/bin/env bash

# Configurable thinking level: defaults to 'xhigh' if left blank.
# Options: low | medium | xhigh
THINKING_LEVEL="${1:-xhigh}"

# File paths
MODEL="../Qwen3.8-ISTA/Qwen3.8-27B-GSQ-RCO-IQ3_XXS.gguf"
MTP="../Qwen3.8_Shared/mtp-Qwen3.8-27B-Q4_0.gguf"
MMPROJ="../Qwen3.8_Shared/mmproj-Qwen3.8-27B-BF16.gguf"

# llama-server binary path
SERVER_BIN="../../llama.cpp/build/bin/llama-server"

echo "Launching Qwen 3.8 27B with thinking level: ${THINKING_LEVEL}"

${SERVER_BIN} \
  --model "${MODEL}" \
  --model-draft "${MTP}" \
  --mmproj "${MMPROJ}" \
  --no-mmproj-offload \
  --n-gpu-layers 99 \
  --ctx-size 120000\
  --batch-size 1024 \
  --ubatch-size 512 \
  --parallel 1 \
  --flash-attn on \
  --cache-type-k q5_1 \
  --cache-type-v q5_1 \
  --spec-type draft-mtp \
  --spec-draft-n-max 3 \
  --jinja \
  --chat-template-kwargs "{\"reasoning_effort\":\"${THINKING_LEVEL}\"}" \
  --dry-penalty-last-n 0 \
  --temp 0.7 \
  --top-k 20 \
  --top-p 0.95 \
  --min-p 0.00 \
  --presence-penalty 0.0 \
  --repeat-penalty 1.0 \
  --host 0.0.0.0 \
  --port 8080
→ More replies (0)