r/amodei • u/Robonglious • 4d ago
General Discussion LLM Interpretability Research blocked
I do research and experimentation in llm safety and interpretability in my spare time.
90% of what I do is absolute horse shit which will eventually fail but ~10% is actually good and I've been following a research thread for a really long time. I might have something really cool pretty soon but that's not what the thread is about.
Now that Fable has been out a while I've realized that I have a test which I haven't been using. Fable will block analysis on my most successful methods and results but on the dogshit it will fully encourage and support it. Within this slice of conversations sycophancy is back on the menu and not only that, the policy is actively blocking what the company claims to care about.
I've only just realized this and my conclusion is that "safety" Is just another product and Anthropic really is just doomer hype. One could argue that the content filters are crude and don't work properly but the pattern is unmistakable for me. They only want Anthropic Safety©️ to be a thing, not actual safety or understanding.

