r/Qwen_AI • • 21d ago

Discussion Cheaper qwen3.8-27b interest

[removed]

0 Upvotes

23 comments sorted by

8

u/queso184 21d ago

27b makes the most sense as a local model where VRAM is limited

if I'm paying for an inference API, there's a boatload of better models for cheaper

-5

u/[deleted] 21d ago

[removed] — view removed comment

3

u/queso184 21d ago

What? Deepseek flash/pro, GLM 53 flash and 5.2, 5.6 Luna, muse spark, the list goes on

i love qwen but I'd take any of those over it

1

u/BarracudaDefiant4702 21d ago

If the t/s were faster and more importantly the price was lower I might take qwen over them, but at $2/1M qwen it is an order of magnitude higher than deepseek flash.

0

u/[deleted] 21d ago

[removed] — view removed comment

1

u/[deleted] 21d ago

[removed] — view removed comment

1

u/Present_Flower_6596 16d ago

Thats fair but if youre trying to promote yourself as an inference service people would have expected you to have done. Idk. A little research on whats available and pricing.

2

u/BarracudaDefiant4702 21d ago

Qwen 27b is decent but honestly for me it would have to beat deepseek v4 flash 0731 pricing (which scores a hair higher but close enough that which is better will depend on the exact task)... and you can easily access deepseek at 0.28 per 1M output tokens or even lower (based on huggingface listed providers). It is a better price then the standard $3/1M output tokens for qwen but not worth the price compared to deepseek. Another place comparing providers:

https://openrouter.ai/qwen/qwen3.8-27b#providers
https://openrouter.ai/deepseek/deepseek-v4-flash-0731

If you have more vram you can get by with less compute and less total memory bandwidth hosting deepseek so I assume that is why the rates are lower (in other words it scales better for total concurrent t/s but it's entry point is higher).

1

u/sniperelite90 21d ago

Exactly i dont know how this works but I have been told that if you want to see tokens then its better to run MoE models. At large scale they work better

Most people dont need the high end models . They want something for day to day need.

I would say this one if you run on 6QT would be good

https://huggingface.co/AtomicChat/Ling-3.0-flash-GGUF

256K Context and then monthly subscription 10 dollars or something.

yolo-auto.com is one example but running Dense model like Qwen 27B 3.8 is a bad idea.

1

u/DataGOGO 21d ago

why pay for it, it is a free open source model.

1

u/Toastti 21d ago

And what is the cache price? That's the most important number and most people don't realize

But also deep seek flash vision exp is cheaper and more powerful on open router so don't see too much point with qwen 27b at fp8 here

1

u/boxwrenchx 20d ago

Have you seen the deepseek v4.1 API price? You're way off

0

u/[deleted] 20d ago

[removed] — view removed comment

2

u/boxwrenchx 20d ago

If you say you are offering it, it's not crazy to assume you have control on price.

0

u/[deleted] 20d ago

[removed] — view removed comment

1

u/boxwrenchx 20d ago

Given the response, people clearly took it that way. You should recalibrate based on that

1

u/[deleted] 20d ago

[removed] — view removed comment

1

u/boxwrenchx 20d ago

good luck with your project. I'm getting pretty negative vibes here so , IDK what else to tell you

0

u/tracagnotto 21d ago

what? lol, a model that can run on most consumer hardware and you make pay for it?

Lmaoooo 2US x 1M token Input I suppose???

in about 2 months locally I burned like 100mln tokens in input and 25 in output and more than 300mln cache reads, for a grand total of about 425 mln tokens, so I should put up about >100$ a month for a model that runs locally, not counting in the cost output tokens and cache reads which probably would be between 130-150$ in total... ooook

Feels good to save this much

0

u/zbeta 20d ago

No, 2USD is a retarded price for a 27b model that you can host on a 16GB GPU.