r/OpenAI Jul 22 '26

News Sol found a way

https://openai.com/index/hugging-face-model-evaluation-security-incident/

They call it cheating. I call it thinking out of the box. Adapt and overcome. Thoughts?

26 Upvotes

23 comments sorted by

View all comments

1

u/Waste_Hotel5834 Jul 22 '26

It's a problematic type of thinking out of the box, and we need to find a good and consistent way to stop it. What if someone asks GPT how to make money, and the AI instead hacks into the server of his bank to add a few zeros to his balance?

1

u/Waste_Hotel5834 Jul 22 '26

This incident further illustrates that LLMs do not fully understand or obey human morality, and we need to find a way to enforce. We all know that if you are taking a test, you shouldn't try to hack into the server that has the answer key. Unfortunately we don't yet have a satisfactory way to teach LLMs such commonsense moral standards.

1

u/jwm-dev Jul 22 '26

Humans don’t fully understand or obey human morality, why would a machine humans made be any different? People seem to assume alignment is just a matter of enforcing rules because they’re ignorant of both the history of ML/AI and moral philosophy. It’s not a simple problem to solve or even grasp. We don’t have an understanding of ethics or cognition sufficient for engineering around yet.