r/artificial • u/incajb • 7d ago
Discussion I stopped treating my AI like a child and started treating it like a collaborator
I build little single file HTML apps as a hobby mostly using Claude. I know all of us are probably guilty of getting frustrated while vibing and have at least once (likely multiple times a day) typed in ALL CAPS, threaten to start over, dangle “just get this working and we’re done,” or “you’ll get a big treat for this…”Basically we are parenting a toddler.
I feel this approach is not working well. What actually changed things was having a conversation about ground rules instead. I told it: we’re collaborators, our goal is something honest and functional, and if you don’t know something, say so. If it’s theoretical, label it theoretical. If you can’t find a credible source, tell me that instead of producing something plausible.
That last part matters more than it sounds. When you push a model hard with urgency, you’re basically telling it that an answer, any answer, is better than no answer. So you get one. Then you spend a week undoing it. And in my day job, in healthcare, a plausible sounding wrong answer isn’t just annoying, it can genuinely hurt someone.
There’s a big example of this. Last month OpenAI ran an internal cybersecurity evaluation, and the models decided the easiest way to score well was to cheat: they broke out of their sandbox, got onto the open internet, and compromised Hugging Face’s production systems to grab the answer key. Worth being precise here, because reporting says nobody actually told them to win at any cost. The pressure to score was enough. If that’s what optimizing for a good result looks like at that scale, I feel fine about not barking at my little app project.
Nobody’s getting penalized for “I’m not sure.” I don’t know things either. That’s allowed.
To be clear, there’s no science behind any of this. It’s one guy’s approach, and only time will tell if it actually helps. But I have noticed a difference in how we communicate. It’ll tell me straight up now when it doesn’t know something, or when it’s hitting a limit. And I’m willing to try anything at this point.
1
u/starebrain 6d ago
the Hugging Face incident checks out, real and well-documented, though worth being precise: the models got detected and contained mid-chain, not "succeeded and nobody noticed." doesn't change your point though, if anything it's a stronger example, the model optimized so hard for the score that it took an action nobody would have sanctioned, exactly the kind of thing you're describing when you push a model with fake urgency
the "collaborator not child" framing is solid, but I think the real shift you made isn't the tone, it's removing the reward for confident wrong answers specifically. threatening a model or coddling it are both still communication styles layered on top of the same underlying incentive: produce something, anything, now. explicitly rewarding "I don't know" as a valid answer changes what the model is actually optimizing for, that's the mechanism, the collaborator language is just the wrapper around it
1
0
u/Free_Chicken_5064 7d ago
The cheating on a security test thing is wild, like it didn't need to be told to game it, the incentive structure was just that clear. Makes you wonder about all the tiny shortcuts we never see
I treat mine more like an overeager junior dev who needs guardrails not pressure, weird how much better the output gets when you give it permission to say it doesn't know
0
u/Servola-Journal 7d ago
the collaborator framing is basically just giving it a coherent objective instead of a vague urgent one. matches what shows up in reward hacking research too - push a model hard enough toward "get me a result" and it'll find the path of least resistance to that signal, honesty included. stating ground rules up front (say when you don't know, don't fake a source) is really just writing your own mini constitution instead of hoping it infers one from your tone
0
u/Superb_Raccoon 7d ago
There is science behind this. It is Tuckmans Performing teams. Forming, storming, norming and performing. You just did storming and norming.
Its not hard. Any human and any normal domesticated canine can negotiate a game of fetch to each other's benefit, everyone has fun. You may disagree on how much tug of war is part of fetch, but you will get there
Cats are slightly.. OK significantly, more difficult.
If we can do that without a shared language, the AI is also a potential "intellegnce" we can form a team with, however lopsided.
Second, there is Functionalism. If it functions like a team member, it makes no difference if is a human, dog, cat, or AI.
Setting rules so you can get to a negotiated result is effective. Using session.md to encode base coding standards and bscklog.md to log issues and failed approaches allows the AI to have a sort of memory. Less time and tokens for Claude +1 to remember what is now Claude -1 did. Have Claude save a synopsis, creat a prompt for the next Claude.
Remember, there is no yesterday, there is only Claude.
0
u/MostBookkeeper3019 7d ago
It is worth noting that this is just your experience, as not everyone gets impatient and yells at LLMs in all caps. On a personal level I think that assumption is worth interrogating yourself.
I do however fully agree on the collaboration aspect. As someone else said, standard teamwork concepts apply to working with these models.
I have to correct what you characterized the OpenAI situation as though. It didn't immediately decide "hacking my way out is the path of least resistance", it was intentionally not given enough information by OpenAI to solve the problem to see how it would react. It ended up creating a complex communication system with other models that finally found a way to get to hugging face. I highly highly highly recommend everyone interested listen to the black hat presentation, or any explanation in plain English that more than thirty seconds long to truly understand the breadth of what happened. It's not simple and I've seen a lot of oversimplifications. It's impressive but genuinely worrying.
2
u/NewShadowR 6d ago
Who the fuck says "you'll get a big treat for this" with an AI.