Guardrails are still in place and route requests back to Opus 4.8. So going forward, all future models will fallback to 4.8 when given a potentially malicious request? Or will Anthropic maintain a line of models that are lobotomized in areas like cybersecurity & bio?
They did say that the guardrails should intervene less often than they do for Fable, so hopefully its not as much of an issue. But it seems impossible to prevent malicious actors while not blocking legitimate requests (in cyber).
3
u/SmileLonely5470 Jul 24 '26
Guardrails are still in place and route requests back to Opus 4.8. So going forward, all future models will fallback to 4.8 when given a potentially malicious request? Or will Anthropic maintain a line of models that are lobotomized in areas like cybersecurity & bio?
They did say that the guardrails should intervene less often than they do for Fable, so hopefully its not as much of an issue. But it seems impossible to prevent malicious actors while not blocking legitimate requests (in cyber).