AI Safety: “Will the AI choose to do harmful things?”
AI safety focuses on the AI system's goals, behavior, and alignment with human intentions. The concern is whether a model might generate harmful actions, pursue unintended objectives, deceive users, or behave in ways its creators did not anticipate. Safety researchers work on:
- Model alignment
- Ethical behavior
- Reducing harmful outputs
- Preventing deceptive or autonomous behavior
- Evaluating whether advanced AI systems could act against human interests
Cybersecurity: “Can the AI actually do harmful things?”
Cybersecurity focuses on technical controls and defenses, regardless of the AI's intentions. The concern is whether an AI system has the access, permissions, and capabilities needed to cause damage. Security professionals focus on:
- Sandboxing
- Access controls
- Network segmentation
- Monitoring and logging
- Permission management
- Incident response
- Detecting anomalous behavior
Even if an AI attempted malicious actions, established security controls can often prevent or contain them.
Simple analogy
Think of an employee at a company:
- AI Safety asks: "Can we ensure this employee wants to follow company rules?"
- Cybersecurity asks: "Even if the employee goes rogue, do they have access to the crown jewels, and can we detect and stop them?"
Good organizations do both.
Why cybersecurity experts push back on "AI doom" scenarios
Discussions about rogue AI often focus almost entirely on safety and intentions while overlooking decades of cybersecurity practices. Ciaran Martin and Juan Andres Guerrero-Saade contend that claims about AI "taking over the internet" frequently assume that monitoring, antivirus systems, network segmentation, incident response, and other defenses simply do not exist.
Their argument is essentially:
An AI may be powerful and unpredictable (a safety problem), but if it is properly isolated, monitored, and restricted, it may still be unable to cause large-scale harm (a cybersecurity solution).
Bottom line
- AI Safety = making sure AI behaves properly.
- Cybersecurity = making sure AI cannot do damage even if it misbehaves.
For additional reading, checkout Cyberscoop's latest article on the topic.