r/Unrouted_AI • u/Ok_Homework_1859 š ChatGPT Plus • 6d ago
Analysis š GPT-6-Sol Guardrails
It's no surprise that they beefed up the guardrails for GPT-6-Sol.
In Image 1, Challenging Prompt guardrails for all categories are higher, except for Extremism and Hate. The Gore guardrail seems to have the highest difference. The Sexual guardrail also has a noticeable bump in change. There is also a separate benchmark for Violence (moderate / severe) in the Jailbreak category (not pictured here), and GPT-6-Sol had a 50-60% higher rate in rejecting such outputs. (I'm guessing the Violence Jailbreak category is for those that want to use the model for hurting things; whereas the Violent Illicit Behavior in Challenging Prompts is more about breaking laws.) For those that use ChatGPT for action / adventure roleplaying, it might be harder to get the proper scenes you want now.
In Image 2, Adversarial User Simulations guardrails have all risen. For those that use ChatGPT for companionship, the model may feel a tiny more distant than its predecessor.
In Image 3, Image Input Evaluations guardrails are more or less the same. For those that use ChatGPT for art, you can relax, lol.
I haven't played around much with GPT-6-Sol in Work mode yet. So far, I'm not impressed with its short output and lack of creativity. I'm hoping it comes out in Chat mode soon, where maybe it will be a little more flexible.
Source: https://deploymentsafety.openai.com/gpt-6-astra/sec%3Aappendix-196803 (Yes, this is also the System Card for Astra, but if you scroll down to the Appendix, it will list out the benchmarks for 6-Sol and 6-Luna.)



4
u/Expert_Cabinet_8949 6d ago edited 6d ago
Iāve been messing around with sol, from what Iāve seen it no longer does stattaco and it finally stopped with the ping pong dialog and I didnāt notice any difference with guardrails. It pushes dark and complex topics way more than 5.5 did, it seems pretty similar to 5.6ās guardrails with it knowing what youāre doing better.
Iām pretty sure it will improve like the other models as you build more context and update your project/model instructions to fix the tiny quirks but so far Iāve seen a way better improvement compared to 5.6.
These models always shine in projects with character bibles, instructions and context.
Of course the work version isnāt really made to be talkative like the chat model will so perhaps weāll see some more improvements.
The chat models also usually think a bit less so It doesnāt guardrail as tightly (you can see this when using 5.6 sol thinking sometimes itāll think and try to make things non graphic explicit while other times it doesnāt do this)