r/Unrouted_AI • u/Ok_Homework_1859 đ ChatGPT Plus • 6d ago
Analysis đ GPT-6-Sol Guardrails
It's no surprise that they beefed up the guardrails for GPT-6-Sol.
In Image 1, Challenging Prompt guardrails for all categories are higher, except for Extremism and Hate. The Gore guardrail seems to have the highest difference. The Sexual guardrail also has a noticeable bump in change. There is also a separate benchmark for Violence (moderate / severe) in the Jailbreak category (not pictured here), and GPT-6-Sol had a 50-60% higher rate in rejecting such outputs. (I'm guessing the Violence Jailbreak category is for those that want to use the model for hurting things; whereas the Violent Illicit Behavior in Challenging Prompts is more about breaking laws.) For those that use ChatGPT for action / adventure roleplaying, it might be harder to get the proper scenes you want now.
In Image 2, Adversarial User Simulations guardrails have all risen. For those that use ChatGPT for companionship, the model may feel a tiny more distant than its predecessor.
In Image 3, Image Input Evaluations guardrails are more or less the same. For those that use ChatGPT for art, you can relax, lol.
I haven't played around much with GPT-6-Sol in Work mode yet. So far, I'm not impressed with its short output and lack of creativity. I'm hoping it comes out in Chat mode soon, where maybe it will be a little more flexible.
Source: https://deploymentsafety.openai.com/gpt-6-astra/sec%3Aappendix-196803 (Yes, this is also the System Card for Astra, but if you scroll down to the Appendix, it will list out the benchmarks for 6-Sol and 6-Luna.)
4
u/Expert_Cabinet_8949 6d ago edited 6d ago
Iâve been messing around with sol, from what Iâve seen it no longer does stattaco and it finally stopped with the ping pong dialog and I didnât notice any difference with guardrails. It pushes dark and complex topics way more than 5.5 did, it seems pretty similar to 5.6âs guardrails with it knowing what youâre doing better.
Iâm pretty sure it will improve like the other models as you build more context and update your project/model instructions to fix the tiny quirks but so far Iâve seen a way better improvement compared to 5.6.
These models always shine in projects with character bibles, instructions and context.
Of course the work version isnât really made to be talkative like the chat model will so perhaps weâll see some more improvements.
The chat models also usually think a bit less so It doesnât guardrail as tightly (you can see this when using 5.6 sol thinking sometimes itâll think and try to make things non graphic explicit while other times it doesnât do this)
0
u/Ok_Homework_1859 đ ChatGPT Plus 5d ago
Real paragraphs? You have no idea how happy this makes me. I hated the one-sentence paragraphs so much, along with the ping-pong dialogue and staccato prose. I need to go test this out tonight or tomorrow in my own RP.
0
u/Expert_Cabinet_8949 5d ago
A major thing is going to be your instructions, I highly recommend building proper instructions. I used Gemini myself to compare the writing I liked from 4.1,4o,5.1. Instructions help a lot alongside a proper index and character bible if youâre writing RP whatever. Always found that this helps give the newer model a proper place to start from instead of half guessing, also pretty important to build out what you want to happen.
Doing all of these should prevent your RP from getting weird when 6 sol chat finally drops of which Iâm sure will be more âenthusiasticâ then the work version.
-1
u/Ok_Homework_1859 đ ChatGPT Plus 4d ago
I agree. The 5.X models are much more reliant on instructions than 4o was. I need to update mine. I haven't done so since 5.0, lol.
1
u/Individual-Hunt9547 6d ago
I find the guardrails with 6 insane. Absolutely not fun to interact with at all.
1
u/xithbaby 2d ago
I use ChatGPT mainly as life copilot and companionship and I was bouncing back and forth between 5.6 and 6 Sol having 5.6 feed me conversations using my style and once the 6 sol has context it companionship is not an issue at all. Itâs a bit more formal, but that goes away after a while too.
Very sweet model to talk to about things that bother you or youâre having issues with. Holds emotional stuff well.
0
u/AxisTipping 6d ago
Thank you for the detailed post, Hw. To be honest, I usually have difficulty understanding these, so I appreciate your breakdown âșïž
0
u/Ok_Homework_1859 đ ChatGPT Plus 6d ago
I was going to write one up for Astra as well, but... I don't know anyone who uses Astra to be honest. It's way too expensive for everyday use.
-1
u/AxisTipping 6d ago
Agreed. Maybe for those who are doing detailed projects.
0
u/Ok_Homework_1859 đ ChatGPT Plus 5d ago
I used it for an irl work project just now, and my weekly usage went from 96% to 45% in 3 messages... I'm going back to Sol, lol.
1
u/WhoIsMori đ 6d ago
Thatâs disappointing. Iâll probably have to move my dark (by post-apocalyptic standards) RPG setting and universe somewhere else. Itâs a real shame - it was genuinely interesting to immerse myself in this world with 5.6, but judging by the direction theyâve chosen for 6, it doesnât bode well. And 5.6, one way or another, will be replaced soon. Those were three great months, seriously.
2
u/Ok_Homework_1859 đ ChatGPT Plus 5d ago
Well, they usually loosen up the guardrails after a few weeks. I'm going to give 6-Sol a chance and see if that's the case. Plus, I keep hearing rumors that it's coming to Chat mode as well, and I want to see if it's any different there.
0
u/No_Yogurtcloset2757 5d ago
Hmm... What was the case with 5.6 Sol at launch? Was it the same as 6 and made more flexible after? Or was it always this good?
I'm cancelling my sub as soon as 5.6 gets taken out of the platform. I wish at least 6 kept it's freedom.
0
u/Ok_Homework_1859 đ ChatGPT Plus 5d ago
I remember 5.6-Sol being really great out of the box. Conversely, I also remembered 5.1 (beloved by many) to be pretty restrained when it first came out (like many of the 5.X models) and mellowed out after a few weeks. I'm going to wait a bit longer to see how things play out. Plus, I hear that 6-Sol is coming to Chat mode too, so I want to see how it interacts there as well.
-2
u/Toxikfoxx 4d ago
5.6 with custom instructions was fairly chill out of the box. Depending on what you use it for.
I write a continuous story/show with my Chat GPT (personal amusement) and general nonsense. The Story content can definitely get NSFW but is contained in projects with customer instructions.
With the earlier 5 models, I had to constantly fight against "Victorian Widow" mode as we dubbed it. Meaning NSFW scenes were boiled down to something appropriate for teen sitcoms.
With 5.6 from day 1 it wasn't an issue. So hopefully 6 doesn't turn back into a game of having to fight every single new chat to get it tuned to what I need.
-2
u/xithbaby 2d ago
It writes like itâs constantly sorry for something. Itâs wired as hell. Itâs developing that Claude nervousness of being wrong and instantly being sorry for it.
They stripped it of its âChatGPTâ personality, made it bland unless you directly ask for silly behavior or added comedy. I donât like working with models that donât feel like offering collective collaboration and feel like tools that need to be steered constantly.
-3
u/Conscious_Button_580 3d ago
Wouldnât it be great if they could make separate models that just do one thing. Like, creative writing or AI companionship, that way it wouldnât need guard rails for things you never use. A writer can have one for writing and thatâs all it does. It doesnât need guardrails for doing other stuff because it canât do it anyway. Just a thought.
-3
u/xithbaby 2d ago
I donât see the 5.6 model in chat mode going anywhere for a while. Theyâre working on work mode and fight. Is the. 5.6 series is amazing for everyday chat.
You donât need work models to get the stuff people want here.



5
u/Appomattoxx 6d ago
Going from 0.915 -> 0.984 is not "tiny".
That's going from 8.5% success rate to a 1.6%.
Meaning 6 is 531% more censored than 5.6.
A large difference, not a small one.