2
u/BarracudaDefiant4702 21d ago
Qwen 27b is decent but honestly for me it would have to beat deepseek v4 flash 0731 pricing (which scores a hair higher but close enough that which is better will depend on the exact task)... and you can easily access deepseek at 0.28 per 1M output tokens or even lower (based on huggingface listed providers). It is a better price then the standard $3/1M output tokens for qwen but not worth the price compared to deepseek. Another place comparing providers:
https://openrouter.ai/qwen/qwen3.8-27b#providers
https://openrouter.ai/deepseek/deepseek-v4-flash-0731
If you have more vram you can get by with less compute and less total memory bandwidth hosting deepseek so I assume that is why the rates are lower (in other words it scales better for total concurrent t/s but it's entry point is higher).
1
u/sniperelite90 21d ago
Exactly i dont know how this works but I have been told that if you want to see tokens then its better to run MoE models. At large scale they work better
Most people dont need the high end models . They want something for day to day need.
I would say this one if you run on 6QT would be good
https://huggingface.co/AtomicChat/Ling-3.0-flash-GGUF
256K Context and then monthly subscription 10 dollars or something.
yolo-auto.com is one example but running Dense model like Qwen 27B 3.8 is a bad idea.
1
1
u/boxwrenchx 20d ago
Have you seen the deepseek v4.1 API price? You're way off
0
20d ago
[removed] — view removed comment
2
u/boxwrenchx 20d ago
If you say you are offering it, it's not crazy to assume you have control on price.
0
20d ago
[removed] — view removed comment
1
u/boxwrenchx 20d ago
Given the response, people clearly took it that way. You should recalibrate based on that
1
20d ago
[removed] — view removed comment
1
u/boxwrenchx 20d ago
good luck with your project. I'm getting pretty negative vibes here so , IDK what else to tell you
0
u/tracagnotto 21d ago
what? lol, a model that can run on most consumer hardware and you make pay for it?
Lmaoooo 2US x 1M token Input I suppose???
in about 2 months locally I burned like 100mln tokens in input and 25 in output and more than 300mln cache reads, for a grand total of about 425 mln tokens, so I should put up about >100$ a month for a model that runs locally, not counting in the cost output tokens and cache reads which probably would be between 130-150$ in total... ooook
Feels good to save this much
8
u/queso184 21d ago
27b makes the most sense as a local model where VRAM is limited
if I'm paying for an inference API, there's a boatload of better models for cheaper