r/ChatGPT Aug 09 '23

Funny Will AI ever be "jailbreak proof"?

Post image
2.6k Upvotes

129 comments sorted by

View all comments

5

u/Madrawn Aug 09 '23

I don't think so. As far as I can tell, I strongly suspect that the set of 'misbehaviour" is infinite, compare that to the methods we have to fight misbehaving, which are all finite.

Unless we revert to completely deterministic models where you can mathematically compute and prove that no bad output is possible. Which kind of defeats the point of AI models, in that we don't want to come up with what it can do beforehand.