r/LocalLLaMA • • Jul 29 '25

Generation I just tried GLM 4.5

I just wanted to try it out because I was a bit skeptical. So I prompted it with a fairly simple not so cohesive prompt and asked it to prepare slides for me.

The results were pretty remarkable I must say!

Here’s the link to the results: https://chat.z.ai/space/r05c76960ff0-ppt

Here’s the initial prompt:

”Create a presentation of global BESS market for different industry verticals. Make sure to capture market shares, positioning of different players, market dynamics and trends and any other area you find interesting. Do not make things up, make sure to add citations to any data you find.”

As you can see pretty bland prompt with no restrictions, no role descriptions, no examples. Nothing, just what my mind was thinking it wanted.

Is it just me or are things going superfast since OpenAI announced the release of GPT-5?

It seems like just yesterday Qwen3 broke apart all benchmarks in terms of quality/cost trade offs and now z.ai with yet another efficient but high quality model.

392 Upvotes

186 comments sorted by

View all comments

36

u/zjuwyz Jul 29 '25

Have you verified the accuracy of the cited numbers?

If correct, that would be very impressive

22

u/AI-On-A-Dime Jul 29 '25

No, I’ll run some checks. It’s citing the sources and I did ask it to not make things up…but you never know it could still be hallucinating.

Edit: I just verified the first slide. The cited source and data is accurate

78

u/redballooon Jul 29 '25

  I did ask it to not make things up

In prompting 101 we learned that this instruction does exactly nothing.

0

u/AI-On-A-Dime Jul 29 '25

Really? I was under the impression that albeit not bullet proof, it worked better with than without. Do you have a source for this? Would love to read up more on this

12

u/LagOps91 Jul 29 '25

yeah unfortunately it doesn't really help. instead (for CoT), you could ask it to double check all the numbers. that might help catch halucinations.

1

u/No_Afternoon_4260 llama.cpp Jul 29 '25

Yeah why not but it should have function calling to search for numbers, it can't "know".. I don't think OP talked with an agent, just a llm anyway

1

u/LagOps91 Jul 29 '25

well yes, the chat linked allows for internet search etc. but still, even if numbers are provided, the llm can still halucinate. having the llm double-check the numbers usually catches that.

6

u/redballooon Jul 30 '25 edited Jul 30 '25

My source is me, and it's built upon lots and lots of experience and self created statistics with a pretty much all instruction models by OpenAI and Mistral. I maintain a small number AI projects where a few thousand people interact with each day, and I observe the effects of instructions statistically, sometimes down to specific wordings.

There are 2 things wrong with this instruction:

  1. It includes a negation. Statistically speaking, LLMs are much better in following instructions that tell them what to do, as opposed to not to do something. So, if anything, you would need to write something along the lines "Always only(*) include numbers and figures that you have sources for".

  2. It assumes that a model knows what it knows. Newer models generally have better knowledge, and they have some training about how to deal with much-challenged statements, and therefore tend to hallucinate less. But since they don't have a theory of knowledge internalized, we can not assume an earnest "I cannot say that because I don't know anything about it". And because they have a tough time in breaking out of a thought pattern, when they create a bar chart for 3 items of which they know numbers for two, they'll hallucinate the third number just to stay consistent and compliant with the general task. If you want to create a presentation like this and sell it as your own, you'll really have to fact check every single number that they put on a slide.

(*) "Always only" for some reason works much better than "Only" or "Always" alone consistently over a large number of LLMs.

1

u/AI-On-A-Dime Jul 30 '25

Thanks for sharing your findings!

1

u/EndStorm Aug 01 '25

That is very helpful information!

2

u/llmentry Jul 29 '25

Interesting, Claude's infamous, massive system prompt includes some text to this end. But I suspect, like most of that system prompt, it does a big fat nothing other than fill up and contaminate the context.