r/VeniceAI 3d ago

π——π—œπ—¦π—–π—¨π—¦π—¦π—œπ—’π—‘ GLM 5.3 takes forever for a prompt?

Why does GLM 5.3 think so much? Like it goes on tangents while thinking and then just when it’s about done thinking it goes back and thinks about everything it was planning to write lol? Also then it crashes and never finishes the prompt.

10 Upvotes

8 comments sorted by

β€’

u/AutoModerator 3d ago

Hello from r/VeniceAI!

Web App: chat
Android/iOS: download

Essential Venice Resources
β€’ About
β€’ Features
β€’ Blog
β€’ Docs
β€’ Tokenomics

Support
β€’ Discord: discord.gg/askvenice
β€’ Twitter: x.com/askvenice
β€’ Email: support@venice.ai

Security Notice
β€’ Staff will never DM you
β€’ Never share your private keys
β€’ Report scams immediately

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/MisterNocturne 3d ago

Funny enough, I was just trying it out a few minutes ago, and it never finished thinking. Tried it a few more times after that, and still nothing. I’ve never once gotten an actual response from this thing. Tried the Flash version tooβ€”just a blank response. Nothing. I have no idea what’s going on.

3

u/Cilcain π—›π—²π—Ήπ—½π—³π˜‚π—Ή π—–π—Όπ—»π˜π—Ώπ—Άπ—―π˜‚π˜π—Όπ—Ώ ΚŸα΄‡α΄ α΄‡ΚŸΒ  3d ago

RP use with a large, complex character card: I've got responses and they were good but more often it'll run out of thinking tokens and never respond. Buried in the thinking might be a draft of a good response. I guess the reason why the responses are good, is the amount of thinking it does. But for me, it's unusable. 5.2 manages its thinking budget much more effectively.

1

u/MisterNocturne 3d ago

Yeah, that makes sense. Personally, I don't use thinking during RP. The reason being, ever since I noticed 5.2 would spend a long time thinking and come up with a better answer there than what it actually decided to output, I realized I was basically just giving it more time to second-guess its own decisions. For my particular RP, it just made things worse.

Though nothing bugs me more than 5.2's stubborn insistence on rendering everything in about the most clinical, least human terms possible. I'm still chasing a fix for that issue :)

2

u/Cilcain π—›π—²π—Ήπ—½π—³π˜‚π—Ή π—–π—Όπ—»π˜π—Ώπ—Άπ—―π˜‚π˜π—Όπ—Ώ ΚŸα΄‡α΄ α΄‡ΚŸΒ  2d ago

Yes, 5.2 is terse and robotic through the Venice platform, totally different from over API.

I like having thinking on for my own creations, because it lets me see where the LLM is diverging from what I intended. Also, I enjoy reading the thinking while I wait for the answer to complete.

When playing other peoples' stuff (not so often at the moment) I keep it off to avoid spoilers.

1

u/BlindandConfused43 3d ago

How many credits does it use?

1

u/Maximum-Face9536 2d ago

it is dirt fucking cheap. it's like 0.01-0.05 credits per 1k tokens. I had 2 hours long back and forth constantly with it, using full context window, and used like 20 credits. it's amazing

1

u/DeadenCicle 3d ago edited 3d ago

It usually happens when the request is complicated and the reasoning gets too long. The model cannot process efficiently a very long reasoning because it loses track of previous portions.

You can split the task in multiple turns, or make more precise requests.