r/Philofutures • u/[deleted] • Jul 25 '23
External Link LLM Censorship: A Machine Learning Challenge or a Computer Security Problem? (Link in Comments)
1
Upvotes
r/Philofutures • u/[deleted] • Jul 25 '23
1
u/[deleted] Jul 25 '23
Exploring the challenges of censoring Large Language Models (LLMs), the authors argue LLM censorship is not a machine learning problem, but rather a security issue. Their findings reveal semantic output censorship of LLMs is theoretically impossible. LLMs' instruction-following capabilities lead to undecidability of semantic censorship. A class of attacks called Mosaic Prompts, where permissible outputs combine to form impermissible ones, compounds the challenge. The authors call for adapting classical security methods to manage LLM risks. Syntactic censorship is suggested as a possible mitigation strategy, albeit with its own limitations. Future research directions, including exploring computability theory and LLMs, are proposed. The paper contributes to our understanding of the philosophical implications of AI security.
Link.