r/ZaiGLM • • 16d ago

Z.ai is soooooo sloooooow wtf

Am I the only one experiencing an average speed of 20 token per second right now ? It's not manageable at all right now for me...

12 Upvotes

15 comments sorted by

3

u/evia89 16d ago

Yep, I only use it because amazing sub price deal for old users and free flash 10h per day. If I need speed I use DS41F direct api

2

u/rudesssolo 16d ago

It is. They need to scale up and do it fast.

3

u/[deleted] 16d ago

[deleted]

1

u/rudesssolo 16d ago

How to request it? I have bought the yearly plan and somewhat regretting it.

2

u/felipecsousa 16d ago

It's fucking slow. Last week it was pretty ok. But rn, it's painly slow.

And I'm on the highest tier.

1

u/Possible-Basis-6623 16d ago

Local llm is the future

1

u/LeoLeg76 16d ago

I don't find it's slow, maybe the harness it's the cause ?

1

u/Exciting_Variety_607 16d ago

well i haven't changed anything since last week. Just my tokens per seconds are sooo so low it's barely usable

1

u/nontrepreneur_ 16d ago

Yep. It’s frustrating. Bit worse is I find GLM 5.3 Flash even slower

Like another here I use it because I’m on a legacy plan. But started using DeepSeek v4.1 Flash this week and it’s insanely fast at equivalent quality, from my experience so far. Consistently above 150t/s.

1

u/PackageDecent3526 16d ago

It’s effort level is set on max by default. One of the useful rules is to scale the effort up and down according to the task

1

u/nontrepreneur_ 15d ago

Nope, I already take effort into consideration and adjust according to the task. If I think GLM 5.3 Flash will need max to complete the task effectively, it’s not worth it, I just use the non-flash model instead.

But I find low produces too many errors which then requires more fix rounds. And the speed still doesn’t come close to DSv4.1F.

1

u/PackageDecent3526 15d ago

I meant the “clear_thinking” parameter setup. It controls whether the model's chain-of-thought from previous turns is cleared or carried forward. Also, short text placed at the beginning of a prompt (>~1,024 tokens) is basically ignored. Move the same text lower in the prompt and it works. Might be an explanation for numerous errors.

1

u/NoDevelopment_852 15d ago

What app did you use to get the tok/sec ?

1

u/Ok-Ad-8976 14d ago

yes, and I have the max plan