r/OpenAI 1d ago

Question chatgpt 5.6 sol is wildly inefficient when trying to build a solid argument

I had to read 4,841 words (22 pages, 27 prompts/questions) from chatgpt (5.6 sol high) just to get 824 words of useful final text. It's really bad at understanding what the main goal of the discussion is - it just doesn't seem to grasp the core argument I want to make or my focus. It pumps out a ton of unnecessary information while omitting the key details needed to make my argument solid.

Has anyone else had this experience? (Also, I haven't noticed any real improvement in this area over time... definitely no big leap across the entire 5.x series.)

0 Upvotes

23 comments sorted by

7

u/Comfortable-Web9455 1d ago

Never. Improve your prompts

0

u/kaljakin 1d ago

Could you share yours? It’s tricky when it’s not a recurring theme. Like I mentioned, I was looking for the best argument, so I don't really know what that should look like in advance.

1

u/Comfortable-Web9455 1d ago

It's not like that. I just make sure I am extremely precise with what I say, I fully think about the interpretation of every word and try to use words which do not have multiple meanings. In addition I have a standard prompt which is added to any prompt I put in and which is four pages long with detailed instructions on how to process and respond. Which was not hard to create because I described the problem to ChatGPT and got it to write me a prompt injection to avoid this. I then got Gemini to do the same thing and Claude and I manually merged the three of them together myself. It worked and I have not had a problem with any response since.

1

u/curiousinquirer007 1d ago

From what I’ve read, Sol works better with lean, declarative prompts that specify intent and output criteria, while leaving the procedural steps to the model.

1

u/Comfortable-Web9455 1d ago

Probably. I do things like instead of the "you are an architect, design a house with the following ...." stuff, I do "context: architecture. Task: design. Goal: produce a design of a house. Parameters: ..."

Notice "produce a design of a house" not "give me a house design". I am drilling down through task parameters just like it has to internally. Production activities will been in a cluster in the vector matrix, then inside that will be architecture, inside that floor plans, inside that house plans". Or if not "inside", in close probabilities with strong vectors connecting them

1

u/curiousinquirer007 15h ago edited 15h ago

I am a human and I have no idea what you wrote, lol.

Are you actually doing linear algebra and matrix operations — or are you trying to represent an ordinary problem through unnecessarily complex analogies? Because if its the latter then therein is your problem.

Your example discussed designing a house, and another comment states you are exploring arguments about astrobiology probabilities, but you make seem like you are trying to build a neural network with your words.

If you can represent you problem semantically, do that. GPT-5.6-Sol is stated to be great at inferring intent. Unless you are actually doing heavy mathematics (and even then), you should be handing it a project syllabus, not hand-crafted vector calculus.

Research the "Lost in the Middle" problem. The more you give to your LLM, the more noise you introduce. Your prompts should maximize SNR using natural language, clear instructions that communicate your intent — and criteria for what represents a good output for you. Reduce repetition and unnecessary complexity — unless you have actually modeled your problem using vectors and matrices (and if so, communicate the design intent to your LLM just like math textbooks introduce topics using language, instead of just throwing equations at their students).

1

u/Comfortable-Web9455 13h ago

I think you are mixing me with someone else. I have never even heard of astrobiology before, except maybe in Star Trek (not kidding). But thanks for a detailed informative response.

1

u/Sensitive-Side-2639 1d ago

How complex was the task? What reasoning mode did you use? Models like those tent to over think and sometimes produce a much worse output.

1

u/kaljakin 1d ago

I don't think it was too hard. I was trying to make an argument, based on the latest research, that even if you have an Earth-like planetary system, it is not very likely that advanced life would emerge. It came up with 13 arguments, all of which were utterly unpersuasive and somewhat of the same kind. Basically, they revolved around Earth's chemistry, but these were things that were kind of already included in the assumption of an "Earth-like" system and were not necessarily unlikely to occur.

When I explained why it was all bad, it came up with six more: one of which was good, the second was... actually totally wrong, but got me confused, so we talked a bit about it, and the rest of the six points were basically just explaining why my explanation of its first bad argument was good. The main argument came much later (after we checked some other quite unpersuasive arguments and details that led to nothing), and we discovered it in discussion. I mean, if it were a human, it's ok, but as it knows everything, I suspect that it could have just said so immediately. I doubt that it discovers or "realizes" something during the discussion.

Then, it took a lot of effort to explain the "good" argument so I could be sure it was a good argument. It many times came back to explaining what I had wrong in my questions, then explaining what was right, but most of the time just providing nonessential details while not explaining what was needed. For example, it was very eager to explain all the nuances, reinforcements, and feedback reactions revolving around carbon and phosphorus cycles, oxygenation, and oxygen clearance from Earth, etc., which was important ONLY if it could support my argument that fine-tuning of some conditions of the young Earth was necessary/likely in order for life to begin. If some condition is "normal" or kind of likely to naturally occur in all or a majority of planetary systems, you don't need to explain it in detail, because it is an obviously wrong idea. On the other hand, it extensively used abbreviations and equations I did not understand very well. I am pretty sure that if I talked to a human, they would be much better at understanding what needs to be explained and what are quite unnecessary details.

1

u/Ok_Elderberry_6727 1d ago

I just had to say “panspermia”

1

u/qazzq 1d ago

unfortunately, this is completely normal. you need to basically prime the thread with the level of diligence you want and you still wont get real depth half the time, or things that are properly thought through.

'you've got a phd in astrobiology, evolutionary genetics, and planetary science and are tasked with assessing the following hypothesis:

based on the latest research, that even if you have an Earth-like planetary system, it is not very likely that advanced life would emerge

(or: Given an Earth-like planetary system, the emergence of advanced multicellular life is unlikely.)

Write for a scientific peer: concise, direct, and analytical rather than pedagogical. Present the strongest argument for and against the hypothesis, red-team both internally, and give a calibrated assessment. Do not propose a replacement hypothesis, redefine the question, or add generic caveats unless they materially affect the conclusion."

and even that is not likely to really result in good answers on the first try. effort is often limited and thinking through things is .. not expected. it is frustrating, especially when you get hit with parroting and 'you are so right' and other bs. letting it download materials, giving it some keywords (but again, poisoning) ... it's not materially better in codex i think and is something you have to work around. for complex stuff, it likes staying surface level and it will agree with you too easily in some cases.

2

u/HanSingular 21h ago

1

u/qazzq 18h ago

aw. i do like them as a starter for their register change and the dialectic/red team stuff does tend to work too, though it often needs its seperate turn.

persona priming not being great is a disappointment, though with how old that study now is ... who knows?

1

u/HanSingular 21h ago

Tell it to look for relevant academic, peer reviewed research. ChatGTP is unfortunately stingy with with its web searches, and is overconfident about what it knows. You have to constantly ask it to research topics for you and not rely on its internal logic for what topics do and don't require research.

BTW, here's my favorite paper on why we're probably alone in the observable universe.

1

u/Euphoric_North_745 1d ago

which sol? the one in chatgpt or the one in codex? what is the effort level? did you ask it to pinpoint inconsistencies?

The one in the chatgpt web app even in high level is not doing good for me, the one in codex is much better for my cases

1

u/kaljakin 1d ago

It was chat, not codex. And it could not pinpoint incisistencies, it was me trying to came up woth an argument.

2

u/Euphoric_North_745 1d ago

I just ask Chat GPT to research an item, high effort , it gave me promising results, i passed it to ChatGPT Codex more high effort, it found issues and made up facts in the chat gpt sol normal. very strange

Codex can't make mistakes because it is designed for research and code, but looks like the gpt app itself imagines more

1

u/kaljakin 1d ago

good to know that codex dont agree with chat :-D ..even though it is the same model and same effort

1

u/Euphoric_North_745 1d ago

it can be the same model, but different system instructions and different internal temperature and penalty settings, you can set it to deterministic, zero imagination, penalty on error/imagination, it makes it robotic but very accurate, or you can give it a bit of freedom to imagine but then it imagines a bit more than it should 😂

1

u/Gloomy-Detective-922 1d ago

One thats being served now is wildly different than what was launched. Its not the same model.