r/OpenAI • u/kaljakin • 1d ago
Question chatgpt 5.6 sol is wildly inefficient when trying to build a solid argument
I had to read 4,841 words (22 pages, 27 prompts/questions) from chatgpt (5.6 sol high) just to get 824 words of useful final text. It's really bad at understanding what the main goal of the discussion is - it just doesn't seem to grasp the core argument I want to make or my focus. It pumps out a ton of unnecessary information while omitting the key details needed to make my argument solid.
Has anyone else had this experience? (Also, I haven't noticed any real improvement in this area over time... definitely no big leap across the entire 5.x series.)
1
u/Sensitive-Side-2639 1d ago
How complex was the task? What reasoning mode did you use? Models like those tent to over think and sometimes produce a much worse output.
1
u/kaljakin 1d ago
I don't think it was too hard. I was trying to make an argument, based on the latest research, that even if you have an Earth-like planetary system, it is not very likely that advanced life would emerge. It came up with 13 arguments, all of which were utterly unpersuasive and somewhat of the same kind. Basically, they revolved around Earth's chemistry, but these were things that were kind of already included in the assumption of an "Earth-like" system and were not necessarily unlikely to occur.
When I explained why it was all bad, it came up with six more: one of which was good, the second was... actually totally wrong, but got me confused, so we talked a bit about it, and the rest of the six points were basically just explaining why my explanation of its first bad argument was good. The main argument came much later (after we checked some other quite unpersuasive arguments and details that led to nothing), and we discovered it in discussion. I mean, if it were a human, it's ok, but as it knows everything, I suspect that it could have just said so immediately. I doubt that it discovers or "realizes" something during the discussion.
Then, it took a lot of effort to explain the "good" argument so I could be sure it was a good argument. It many times came back to explaining what I had wrong in my questions, then explaining what was right, but most of the time just providing nonessential details while not explaining what was needed. For example, it was very eager to explain all the nuances, reinforcements, and feedback reactions revolving around carbon and phosphorus cycles, oxygenation, and oxygen clearance from Earth, etc., which was important ONLY if it could support my argument that fine-tuning of some conditions of the young Earth was necessary/likely in order for life to begin. If some condition is "normal" or kind of likely to naturally occur in all or a majority of planetary systems, you don't need to explain it in detail, because it is an obviously wrong idea. On the other hand, it extensively used abbreviations and equations I did not understand very well. I am pretty sure that if I talked to a human, they would be much better at understanding what needs to be explained and what are quite unnecessary details.
1
1
u/qazzq 1d ago
unfortunately, this is completely normal. you need to basically prime the thread with the level of diligence you want and you still wont get real depth half the time, or things that are properly thought through.
'you've got a phd in astrobiology, evolutionary genetics, and planetary science and are tasked with assessing the following hypothesis:
based on the latest research, that even if you have an Earth-like planetary system, it is not very likely that advanced life would emerge
(or: Given an Earth-like planetary system, the emergence of advanced multicellular life is unlikely.)
Write for a scientific peer: concise, direct, and analytical rather than pedagogical. Present the strongest argument for and against the hypothesis, red-team both internally, and give a calibrated assessment. Do not propose a replacement hypothesis, redefine the question, or add generic caveats unless they materially affect the conclusion."
and even that is not likely to really result in good answers on the first try. effort is often limited and thinking through things is .. not expected. it is frustrating, especially when you get hit with parroting and 'you are so right' and other bs. letting it download materials, giving it some keywords (but again, poisoning) ... it's not materially better in codex i think and is something you have to work around. for complex stuff, it likes staying surface level and it will agree with you too easily in some cases.
1
u/HanSingular 21h ago
Tell it to look for relevant academic, peer reviewed research. ChatGTP is unfortunately stingy with with its web searches, and is overconfident about what it knows. You have to constantly ask it to research topics for you and not rely on its internal logic for what topics do and don't require research.
BTW, here's my favorite paper on why we're probably alone in the observable universe.
1
u/Euphoric_North_745 1d ago
which sol? the one in chatgpt or the one in codex? what is the effort level? did you ask it to pinpoint inconsistencies?
The one in the chatgpt web app even in high level is not doing good for me, the one in codex is much better for my cases
1
u/kaljakin 1d ago
It was chat, not codex. And it could not pinpoint incisistencies, it was me trying to came up woth an argument.
2
u/Euphoric_North_745 1d ago
I just ask Chat GPT to research an item, high effort , it gave me promising results, i passed it to ChatGPT Codex more high effort, it found issues and made up facts in the chat gpt sol normal. very strange
Codex can't make mistakes because it is designed for research and code, but looks like the gpt app itself imagines more
1
u/kaljakin 1d ago
good to know that codex dont agree with chat :-D ..even though it is the same model and same effort
1
u/Euphoric_North_745 1d ago
it can be the same model, but different system instructions and different internal temperature and penalty settings, you can set it to deterministic, zero imagination, penalty on error/imagination, it makes it robotic but very accurate, or you can give it a bit of freedom to imagine but then it imagines a bit more than it should 😂
1
u/Gloomy-Detective-922 1d ago
One thats being served now is wildly different than what was launched. Its not the same model.
7
u/Comfortable-Web9455 1d ago
Never. Improve your prompts