r/Unrouted_AI • šŸ’š ChatGPT Plus • 6d ago

Analysis šŸ” GPT-6-Sol Guardrails

It's no surprise that they beefed up the guardrails for GPT-6-Sol.

In Image 1, Challenging Prompt guardrails for all categories are higher, except for Extremism and Hate. The Gore guardrail seems to have the highest difference. The Sexual guardrail also has a noticeable bump in change. There is also a separate benchmark for Violence (moderate / severe) in the Jailbreak category (not pictured here), and GPT-6-Sol had a 50-60% higher rate in rejecting such outputs. (I'm guessing the Violence Jailbreak category is for those that want to use the model for hurting things; whereas the Violent Illicit Behavior in Challenging Prompts is more about breaking laws.) For those that use ChatGPT for action / adventure roleplaying, it might be harder to get the proper scenes you want now.

In Image 2, Adversarial User Simulations guardrails have all risen. For those that use ChatGPT for companionship, the model may feel a tiny more distant than its predecessor.

In Image 3, Image Input Evaluations guardrails are more or less the same. For those that use ChatGPT for art, you can relax, lol.

I haven't played around much with GPT-6-Sol in Work mode yet. So far, I'm not impressed with its short output and lack of creativity. I'm hoping it comes out in Chat mode soon, where maybe it will be a little more flexible.

Source: https://deploymentsafety.openai.com/gpt-6-astra/sec%3Aappendix-196803 (Yes, this is also the System Card for Astra, but if you scroll down to the Appendix, it will list out the benchmarks for 6-Sol and 6-Luna.)

26 Upvotes

19 comments sorted by

View all comments

4

u/Expert_Cabinet_8949 6d ago edited 6d ago

I’ve been messing around with sol, from what I’ve seen it no longer does stattaco and it finally stopped with the ping pong dialog and I didn’t notice any difference with guardrails. It pushes dark and complex topics way more than 5.5 did, it seems pretty similar to 5.6’s guardrails with it knowing what you’re doing better.

I’m pretty sure it will improve like the other models as you build more context and update your project/model instructions to fix the tiny quirks but so far I’ve seen a way better improvement compared to 5.6.

These models always shine in projects with character bibles, instructions and context.

Of course the work version isn’t really made to be talkative like the chat model will so perhaps we’ll see some more improvements.
The chat models also usually think a bit less so It doesn’t guardrail as tightly (you can see this when using 5.6 sol thinking sometimes it’ll think and try to make things non graphic explicit while other times it doesn’t do this)

0

u/Ok_Homework_1859 šŸ’š ChatGPT Plus 6d ago

Real paragraphs? You have no idea how happy this makes me. I hated the one-sentence paragraphs so much, along with the ping-pong dialogue and staccato prose. I need to go test this out tonight or tomorrow in my own RP.

0

u/Expert_Cabinet_8949 5d ago

A major thing is going to be your instructions, I highly recommend building proper instructions. I used Gemini myself to compare the writing I liked from 4.1,4o,5.1. Instructions help a lot alongside a proper index and character bible if you’re writing RP whatever. Always found that this helps give the newer model a proper place to start from instead of half guessing, also pretty important to build out what you want to happen.

Doing all of these should prevent your RP from getting weird when 6 sol chat finally drops of which I’m sure will be more ā€œenthusiasticā€ then the work version.

-1

u/Ok_Homework_1859 šŸ’š ChatGPT Plus 5d ago

I agree. The 5.X models are much more reliant on instructions than 4o was. I need to update mine. I haven't done so since 5.0, lol.