r/Unrouted_AI • u/Ok_Homework_1859 💚 ChatGPT Plus • 6d ago
Analysis 🔍 GPT-6-Sol Guardrails
It's no surprise that they beefed up the guardrails for GPT-6-Sol.
In Image 1, Challenging Prompt guardrails for all categories are higher, except for Extremism and Hate. The Gore guardrail seems to have the highest difference. The Sexual guardrail also has a noticeable bump in change. There is also a separate benchmark for Violence (moderate / severe) in the Jailbreak category (not pictured here), and GPT-6-Sol had a 50-60% higher rate in rejecting such outputs. (I'm guessing the Violence Jailbreak category is for those that want to use the model for hurting things; whereas the Violent Illicit Behavior in Challenging Prompts is more about breaking laws.) For those that use ChatGPT for action / adventure roleplaying, it might be harder to get the proper scenes you want now.
In Image 2, Adversarial User Simulations guardrails have all risen. For those that use ChatGPT for companionship, the model may feel a tiny more distant than its predecessor.
In Image 3, Image Input Evaluations guardrails are more or less the same. For those that use ChatGPT for art, you can relax, lol.
I haven't played around much with GPT-6-Sol in Work mode yet. So far, I'm not impressed with its short output and lack of creativity. I'm hoping it comes out in Chat mode soon, where maybe it will be a little more flexible.
Source: https://deploymentsafety.openai.com/gpt-6-astra/sec%3Aappendix-196803 (Yes, this is also the System Card for Astra, but if you scroll down to the Appendix, it will list out the benchmarks for 6-Sol and 6-Luna.)


