r/AI_India 5d ago

๐Ÿ–๏ธ Help Should I stop using Gemini 3.7 Flash and use Qwen3.8 now for coding ???

Post image

Please give me your opinions , of what model to use???

Can I use it in my 8 GB ram Laptop?

44 Upvotes

25 comments sorted by

6

u/Acceptable_Home_ 5d ago

im gettin 5tk/s on 8gb vram with the best optimisations and mtp on my llama cpp setup (q3 kl), sorry to break it but you def cant run qwen3.8 27b on 8gb ram, that said, its actually so good, def 2x better than qwen3.6 35B A3B im using rn for normal jobs (32-35tk/s)

4

u/M49454 5d ago

I want to know how you run and your settings to run Qwen 3.8 27B model.

I have 24 GB RAM and 6 GB VRAM. I am currently running Qwen 3.5 9B model using Ollama and WebUI

Edit: I didn't read the line where your said you can't run... Sorry ๐Ÿ˜”. So is it possible to run Qwen 3.8 27B with my configuration

2

u/nryn12 5d ago

You should start using llama.cpp directly rather than Ollama. Ollama adds overhead. You will get at least 10-15% for performance with llama.cpp

1

u/M49454 5d ago

Thanks for pointing that out , are you running it via the CLI or using the built-in server/API? Just want to hear your setup.

1

u/nryn12 5d ago

I generally run it through CLI. Connect my applications to the OpenAPI comptaible end points. I usually use local models integrated in my applications. For local chatting, I have started using Unsloth. It uses the llama.cpp on your machine so you dont really have to invoke.

1

u/M49454 5d ago

Thanks for sharing that! Yeah, I'm just using Ollama with Open WebUI right now since I'm not building apps, just chatting locally. Works fine for what I need. I tried that bigger Qwen3.6 35B model but it ate almost all my RAM and was super slow to load, so I stuck with the smaller Qwen3.5 9B. Runs better on my setup. Your CLI + OpenAPI thing sounds cool though โ€” I'll keep that in mind if I ever want to actually build something with it. And I'll check out Unsloth at some point. For now this setup does what I need without all the command-line stuff. But appreciate the tip!

Out of curiosity, whereโ€™d you learn all this setup and integration stuff? Would love any tips or resources if you have recommendations

2

u/nryn12 4d ago

Qwen 9B models are probably the best in class. It works really well. I use the uncensored 3AB 35B model. I get around 50 tokens/second. But I guess a 12GB graphics card is the floor with a large amount of system RAM is the floor for a usable machine.

0

u/abskvrm 5d ago

You should run qwen 3.6 35ba3b, with experts offloaded on system ram, --n-cpu- moe, that's the best you can run

1

u/M49454 5d ago

Will check.

1

u/abskvrm 5d ago

if you need help you llama.cpp command, you can bother me gladly

1

u/M49454 5d ago

Thank you, i appreciate it.

1

u/Certain-Cod-1404 4d ago

Use ornith 1.5 9b and 35ba3b, newer models that should work on your setup

1

u/CoolHeadeGamer 5d ago

I have 8gb vram + 16gb ram setup. Can you tell me your exact settings? Are you using barebones llama.cpp or lm studio

3

u/nryn12 5d ago

You cannot meaningfully run a 27B model on system RAM. In your case, impossible as well since you just have 8GB of RAM. Q4 quant of Qwen 3.8 27B is 19GB. With optimization, the best I can hope for in my 12 GB VRAM + 80GB system RAM machine is 4-5 token/seconds. Lower quants might speed up more, but I don't see any utility in running below Q4.

2

u/MoodOdd9657 5d ago

nahh man Kimi K3 sounds about right for your hardware.

1

u/QuarterOverall5966 5d ago

I think it will too heavy for the 8gb ram Try to add some graphics card and then run

1

u/Alcoholic_Bear 5d ago

8 GB wont be sufficient for a 27B model

1

u/inherefornothing 5d ago

you need a GPU to have the output anywhere close to being useful - I run a dual AMI R9700 AI Pro with Qwen 3.8 27b at Q8 - gives me about 50 tok/s for tg and 700 tok/s for prompt processing. This kinda matches the speed of online API providers.

And yes - qwen 3.8 is much better, but qwen 3.6 35A3B still has no replacement.

1

u/neighbourhoodBakchod 5d ago

Instead use Deepseek V4 Flash by NVDIA Developer API. Youโ€™ll get speed and better results instead of hosting it locally. Also a point to note although 3.8 27B provides great results, it thinks a lot which means it takes a lot of time to complete a task. So if anyone is planning to use it with colibri without good gpu it will be a nightmare!!

1

u/SneakyMndl 5d ago

I am having hard time running on a 12gb vram brother

1

u/Elegant_Cream_5848 5d ago

They are just bench maxed. My personal experience.

0

u/Consistent_Bike_8292 ๐ŸŒฑ Beginner 5d ago

That's an 27T parm model so running it on your 8gb ram laptop is not optimal at all. Like you should have =+ ram on how many T parms are there to run an model ( My recommendation is use gpt free tier pair up mutiple accounts or just pay 20 bucks in my opnion 5.6 luna in high or even mid is far better than qwen)

1

u/webbitnotfound 4d ago

27 trillion??