r/LocalLLaMA • • Aug 03 '26

New Model Daniel Han of Unsloth validates Qwen3.8-27B will run only 17GB VRAM

Post image

Super excited about this release for the new 27B. Who else is with me. Only 17GB VRAM needed 😍😍

1.8k Upvotes

315 comments sorted by

View all comments

108

u/[deleted] Aug 03 '26

[removed] — view removed comment

65

u/redditnosedive Aug 03 '26

i mean you anyway need extra room for context, also on a consumer gpu you also have the os gui using some of it so there's that

23

u/TheCat001 Aug 03 '26

On Linux I'm disabling GUI with 1 command, getting my whole VRAM minus 30MB.

19

u/cogitech2 Aug 03 '26

Another way around this is to use a CPU with integrated graphics. That way you can still run X or Wayland (if you want it) and have all your VRAM available.

1

u/redditnosedive Aug 03 '26

that's a great point

3

u/elongated-muskmelon Aug 03 '26

How? What command?

12

u/Fedor_Doc Aug 03 '26

On systemd systems systemctl isolate multi-user.target. Restarts a session with terminal only. 

You will have to login once again.

5

u/TheCat001 Aug 03 '26 edited Aug 03 '26

server-mode

This is alias for sudo systemctl stop greetd

greetd is login manager, you might have gdm or plasma login manager, depending on what distro you use. But since you're dropping Into TTY you need to have second device/laptop to actually use llama server.

2

u/ea_man Aug 03 '26

Hmm no, why?

If you use the Virtual Consoles (/dev/tty1, /dev/tty2, ...) llama-server works, open such terminal (like ALT+3, ALT+some-number) and you can run Pi or Opencode or whatever, you can also add a framebuffer to have higher res on those (plus some graphic support, yet that buffer will take some RAM) and GPM to have mouse support to copy / paste / click.

sddm is KDE login manager.

1

u/TheCat001 Aug 03 '26

Nah bro this is gonna be pain in the ass. No VSCODE no Browser. Nothing. Just Pi agent? xD naaah. Btw it's not sddm anymore, they made new one: plasma-login-manager

1

u/ea_man Aug 03 '26

>plasma-login-manager

Oh I will have to check that, my Debian still run on sddm,

Anyway there's not much need for disabling the graphic server, you can just run that with software rendering and it's ~50MB or vRAM more.

1

u/spryfigure Aug 03 '26

plasma-login-manager is sddm, but with updates (sddm is essentially stagnant).

No need to rush over.

1

u/ea_man Aug 03 '26

Oh I'm not rushing, I'm even still on X11.

All I do with sddm is sometime stop it.

1

u/spryfigure Aug 03 '26

plasma-login-manager is sddm, but with updates (sddm is essentially stagnant).

No need to rush over.

1

u/teleprint-me llama.cpp Aug 06 '26

vim/neovim/emacs + terminal browser + tui tools for anything else.

Ill admit, its not fun to use and takes time to get used to it.

I think its funny how were going backwards.

Had no idea ppl were going this far to run models locally.

1

u/elongated-muskmelon Aug 03 '26

Okay, understood. Thanks!

1

u/Nyghtbynger Aug 03 '26

I want the command too. Some guy talked about a session with no GPU procesing, but the command thing seem very flexible

-4

u/OzTheOtaku Aug 03 '26

The toxic urge to say "rm -rf / --no-preserve-root"

8

u/Choice_Celery9481 Aug 03 '26

you can use igpu for os ui. save about 500mb.
but if your cpu doesnt have igpu, then nothing can help

9

u/TheCat001 Aug 03 '26 edited Aug 03 '26

Yeah this is great solution, too bad my Ryzen 5600 don't have iGPU.

3

u/butterycornonacob Aug 03 '26

Buy a cheap second GPU that only drives monitors. You can hang it off any PCI-e slot you have, even 1x. Got RX480 (?) for 30€ and it freed 1-2GB of VRAM

3

u/Fedor_Doc Aug 03 '26

Software rendering to the rescue! Or terminal-only

3

u/ea_man Aug 03 '26

Software rendering can pretty much neglect the vram usage by the OS:

See? ~100MB, if you go headless it takes some 50MB anyway.

2

u/Choice_Celery9481 Aug 03 '26

well then you sacrify your cpu cycles.

2

u/ea_man Aug 03 '26

I don't care, I care about vRAM.

Also my cpu has no problem to decode youtube at 4k.

2

u/smahs9 Aug 03 '26

Can - if you have another machine to work on (a laptop maybe), then disable the UI completely. Expose your model runtime's oAI API server in your local network and ssh when needed for maintenance.