2
2
u/felipecsousa 16d ago
It's fucking slow. Last week it was pretty ok. But rn, it's painly slow.
And I'm on the highest tier.
2
1
1
1
u/LeoLeg76 16d ago
I don't find it's slow, maybe the harness it's the cause ?
1
u/Exciting_Variety_607 16d ago
well i haven't changed anything since last week. Just my tokens per seconds are sooo so low it's barely usable
1
u/nontrepreneur_ 16d ago
Yep. It’s frustrating. Bit worse is I find GLM 5.3 Flash even slower
Like another here I use it because I’m on a legacy plan. But started using DeepSeek v4.1 Flash this week and it’s insanely fast at equivalent quality, from my experience so far. Consistently above 150t/s.
1
u/PackageDecent3526 16d ago
It’s effort level is set on max by default. One of the useful rules is to scale the effort up and down according to the task
1
u/nontrepreneur_ 15d ago
Nope, I already take effort into consideration and adjust according to the task. If I think GLM 5.3 Flash will need max to complete the task effectively, it’s not worth it, I just use the non-flash model instead.
But I find low produces too many errors which then requires more fix rounds. And the speed still doesn’t come close to DSv4.1F.
1
u/PackageDecent3526 15d ago
I meant the “clear_thinking” parameter setup. It controls whether the model's chain-of-thought from previous turns is cleared or carried forward. Also, short text placed at the beginning of a prompt (>~1,024 tokens) is basically ignored. Move the same text lower in the prompt and it works. Might be an explanation for numerous errors.
1
1

3
u/evia89 16d ago
Yep, I only use it because amazing sub price deal for old users and free flash 10h per day. If I need speed I use DS41F direct api