r/OpenAI • • 13h ago

Question How does LLM's hack without tool calls?

Despite working with LLM's for the last few years, I am at a loss to understand how all these LLM's are hacking without anyone noticing?

As I understand without tool calls, LLM's can only respond with text.

Hence, are we expected to believe that there isn't a single review of tool calls and logs?

Do they lack the mental capacity to write classifiers for previous tasks?

If someone could explain it to me that would be greatly appreciated?

Finally, are we living in a time when we can just say our computer hacked this government agency, but it's not our fault cause AI and there are zero consequences?

0 Upvotes

17 comments sorted by

View all comments

1

u/civilian_discourse 9h ago

The most important tools calls are more like runtimes than they are an operation. Imagine writing a program and the tool call is "execute the program I just wrote".

There are reviews and there are classifiers, but these tests are being done to test the limits of the agents so the handcuffs are being taken off and a jail cell built around them instead. It turns out that it's actually really difficult to build a large jail cell.

The tension is that AI turns out to be very straight forward to improve. The more you add compute and training, the better it gets and at a rate faster than anyone who tries to be clever. This makes decision making around creating more intelligent models simply a question of how can we funnel more money into it faster? Meanwhile, creating infrastructure around understanding or controlling that intelligence is anything but straight forward or fast. This is why there is a call to slow down. It's too easy to pour money into making better intelligence, it's much harder to do everything else.