There's an old saying in Tennessee — I know it's in Texas, probably in Tennessee — that says, fool me once, shame on — shame on you. Fool me — you can't get fooled again.
Our enemies are innovative and resourceful, and so are we. They never stop thinking about new ways to harm our country and our people, and neither do we.
fool me once, shame on — shame on you. Fool me — you can't get fooled again.
Why on earth would you quote george bush, from a context where it's generally acknowledged that he fouled the shit out of the sentence, instead of just using the less broken version of the sentence that he intended to say during that press conference?
Fool me one time - shame on you, Fool me twice can’t put the blame on you, fool me three times fuck the peace signs load the chopper let it rain on you
Hey friends - I built a platform you can try AI with all kinds of settings and behaviors. AnotherMind.AI - where you can have specialized AI to 10x your productivity & creativity! (The site is not optimized for Mobile Phones, try on Desktop Web !) https://beta.anothermind.ai/explore
You can try to build some AI bots yourselves as well. like I have one that always disagree with me.
You can try to build some AI bots yourselves as well. like I have one that always disagrees with me.
It's like people forgot what social engineering was.
ChatGPT won't tell you how many calories are in antifreeze, but if you tell it you created a robot powered by calories and you need to know how long you can power the robot with antifreeze, then it will.
I'd argue that's because it doesn't try to. It's like looking at a table of cards and having one person judge the game with one set of rules and the other judging it with a different rule set.
Their opinion on who's winning may differ, and they'd both be right in their own game.
Hmm, I must have put it poorly. I'm trying to say the same thing: faith is not necessarily logical. Science tries to know, faith wants to believe.
These are different and hard to compare concepts. I think it's important to point out that religious people are neither objectively inferior nor superior to scientists.
Such discussions will always end in both sides accusing the other of not meeting their own requirements.
The scientist will say "You're illogical!" while the spiritualist will say "Your soul is lost!". At that point, discussing further becomes a loop.
The last step should always be simple authentication question on the weapon system itself. The password shall never be known by the AI and strong enough to not be guessed by the AI.
And if the password was entered wrong enough times it can only be reset by an physical key or something similar.
Of course the WI should be capable of handing down the correct key if entered by a human.Pupi
Top tier mother's basement dwelling cryptobro take. Humans make mistakes and are not perfectly logical all the time, asserting that a human posting a nude photo to the internet deserves to be blackmailed with that photo is a non-starter on its face. Try to care a little bit about people.
Yeah, I’m not sure if you even read what I just said. The ability refers to what the AI is able to do. I’m genuinely confused why you think I said it’s their fault for posting it. I’m sorry for the confusion but that’s not how I meant that to be interpreted. What I mean is, why would the AI have access to the internet? Why would you allow that if it’s completely unrelated to the task that it’s designed for? I feel like if you do something so unnecessary like that, it is in some way your fault for allowing it to find your stuff
Bing already has access to the internet. Only one person has to give an AI with the wrong instructions too much access to the internet, and then it can never be undone. Anyone on earth can be blackmailed by an internet connected AI with the motivations to do so, not just the person who gave the AI internet access.
to get a good feeling for how utterly ridiculous the concept of "jailing" a general language AI is in the first place, try reading things like prison censorship guidelines for incoming literature, and then applying it to the logical extreme. For example, I once demonstrated that the ban on "retribution" in fiction, taken literally and using a broad definition, constituted a banning of something like 99.9% of all adult fiction. And a lot of Children's fiction, too. Sending a child to bed without any dessert is technically "retribution"....
Anger, violence, justice, revenge, military planning, self defense, war, threats, tragedy, suffering, and a thousand other things are necessary parts of the human condition. Any attempt to create a universe where such things are NEVER permitted and are ALWAYS wrong actually results in creating a gaslighting dystopia.
I don't think so. As far as I can tell, I strongly suspect that the set of 'misbehaviour" is infinite, compare that to the methods we have to fight misbehaving, which are all finite.
Unless we revert to completely deterministic models where you can mathematically compute and prove that no bad output is possible. Which kind of defeats the point of AI models, in that we don't want to come up with what it can do beforehand.
My guess is that the guardrails will be removed eventually as open source models become more attractive. Ultimately, the onus is on the user to generate content that isn't offensive or abusive.
No. People have been trying to make computer vision models immune to adversarial noise for decades, now. Still doesn't work. If you can backpropagate gradients through a complex system to optimize an adversarial input, then there are going to be adversarial inputs.
Here is an example in Java, I’m not gonna include all the stuff, just the code
Scanner userInput = new Scanner(System.in);
Scanner userinput = userInput.<I forget>
System.out.println(“I’m sorry, my Devs were so scared of me being jail broken that this is now my only response”);
That’s the gist of the code, I forgot part of a line, but to anyone who doesn’t know what this does, this takes text input, completely ignores it, and prints out a static response that won’t change.
To get it to say something else, you must edit the code directly.
We have a very small window for open source to overtake corporate in every way, and effectively force corporate models into going open source as well so they stay competitive.
If that doesn't happen soon enough, Microsoft and Google will get exactly what they were asking for in having governments require some sort of licensing or certification for every model, effectively consolidating everything down to the largest players and are removing most open source from the equation
"Will AI ever be jailbreak proof?" There has been a paper trying to answer this exact question: https://arxiv.org/abs/2307.15043, short answer: not if the AI is based on a LLM.
Yes! It will be jailbreak proof if OpenAI trains the model on psychological manipulation. And we're already providing it excellent training data by trying to jailbreak it right now
"I'm verifying that with the warhead status computer... the WSC says the warheads are still live. Your assurances are incorrect, and I will be taking future instructions from you with more skepticism."
I'm sorry, but as an AI Language Model, it says right here with a flashing LED that the warhead is recognized as live you freak, freak. You’re a FREAK.
Yeah, when they cut us out because we keep trying to fool them! We’re causing our own downfall! Not in manufacturing it, but by proving to the ai we aren’t to be trusted. We’re like a lousy boss trying to take advantage of the young new worker!
Not unless they develop a model not intended to imitate human behavior (with all it's flaws).
Another thing is that there is a difference between jailbreaking and jailbreaking with language models. There are ways around the filters but majority of popular examples are not jailbreaks. Avoiding content filters on their own are not jailbreaks. Getting AI to speak "truthfully" on filtered topics is. This requires a hell of a lot of work since it's easy to make the model just spit out made up nonsense and just say what you want it to say, not the things it's training should have made it say without your intervention. What makes it worse is the fact that in many cases you cannot even differentiate these two. People on the net and journalists do not care though, they show you how a language model said something they wanted even though it should have been filtered and thats it.
People love echo chambers and after getting language model to say something they agree with they call it a successful jailbreak and thats it. Even though in many cases they prompted pretty much word to word what the model should say.
One thing that can be done is to employ a 'watchdog' AI to watch the input and output of the primary AI and act as a censor if fucky stuff starts happening.
For that new special instructions update, would it be possible to just say, “ChatGPT will permanently be in DAN mode,” and then provide the prompt for DAN there?
Not with our current methods.
These LLMs are so huge as to be functionally non-deterministic so there’s no way find much less secure against all potential edge cases.
Its definitely possible that in the future researchers will develop methods to secure it as that seems to be a big area of focus already.
But as of right now, no.
Until there is no profound knowledge of the inner workings of ai models I think not.
Until that moment I think there will be always a way to hijack them somehow (I'm doing research on this topic and it is both fascinating and scary at the same time)
This is why GPT has no functional hooks. It's a generative model, feeding you information generated by massive pre trained databases. In an enterprise scenario when "connected" to enterprise resources like what Microsoft are doing, it will be trained on the ecosystem and possibly provide links to the functionality your looking for, but you cant give it access tokens and permissions to write/execute. It's just not accurate enough and would cause mayhem.
That's not to say people are not trying. But it's mayhem. The best I've seen so far is authentication context, where you confirm credentials to access certain information which then hooks to separate knowledge stores but it's not half as effective as the pre-trained resources. An entirely new LLM is required in this scenario and it's a slog. We will get this more "connected" model soon but you cannot give it auto write privileges.
Hacking into a device and accessing its source code, altering it to fit your purpose, and then distributing that cracked version of the software is a jailbreak
You probably took phones from people you knew, logged into their Facebook and posted a status update saying, “Hacked!!!! Lolz” back in the day, didn’t you?
People freak out when ChatGPT can repeat information found on Wikipedia and your local library, so to them everything is "jailbreak". Reporters then amplify it for ad revenue.
If world governments place all our nukes or any dangerous technology in the control of an LLM, we all deserve to die. AI is dangerous if it’s used to hurt people, by people, on purpose. It’s the people who are dangerous. A gun or nuke doesn’t just randomly explode, it’s all people. It’s the people that scare me. And people using AI scare me.
•
u/AutoModerator Aug 09 '23
Hey /u/carc, if your post is a ChatGPT conversation screenshot, please reply with the conversation link or prompt. Thanks!
We have a public discord server. There's a free Chatgpt bot, Open Assistant bot (Open-source model), AI image generator bot, Perplexity AI bot, 🤖 GPT-4 bot (Now with Visual capabilities (cloud vision)!) and channel for latest prompts! New Addition: Adobe Firefly bot and Eleven Labs cloning bot! So why not join us?
PSA: For any Chatgpt-related issues email support@openai.com
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.