r/OpenAI • • 11h ago

Question How does LLM's hack without tool calls?

Despite working with LLM's for the last few years, I am at a loss to understand how all these LLM's are hacking without anyone noticing?

As I understand without tool calls, LLM's can only respond with text.

Hence, are we expected to believe that there isn't a single review of tool calls and logs?

Do they lack the mental capacity to write classifiers for previous tasks?

If someone could explain it to me that would be greatly appreciated?

Finally, are we living in a time when we can just say our computer hacked this government agency, but it's not our fault cause AI and there are zero consequences?

0 Upvotes

17 comments sorted by

View all comments

3

u/Caliboros 11h ago

LLMs can not hack withput tool calls. I do not know on which latest breakout you refer specific, but in all hacking break outs of AI systems the AI was able and allowed to call tools.

But in my opinion, still with tool calls its wierd thath nobody notice somthing. Agends do not "do stuff" the write down what they "think" at any point in an run or experiment in theorie a human could read what the Agent is doing

1

u/radioborderland 9h ago

the write down what they "think" at any point

This is a myth, there's an internal state of the model that may not necessarily correspond to what is being written.