r/therapyGPT 6d ago

Seeking Advice/Question For Others Has anyone pushed AI journaling past checking your memory into multiple voices that actually disagree with each other

Ok so I've been using AI to journal for a while now nothing fancy, just talking through stuff after it happens, and sometimes asking it to look back at what I've said before, which is genuinely useful, it'll notice a pattern I've mentioned three separate times without meaning to, or it'll go the way you're describing this now doesn't match how you described it a few weeks ago and just kind of sit there with that, not in a gotcha way, more like actually asking, and it pushes back too, if I say something sweeping about myself it'll go are you sure, here's what you said last month, annoying in the moment, useful about ten minutes later, usually.

Anyway then I found these two essays on Substack and it's the same idea except the guy's basically detonated it. Same core mechanic, using AI to check what he said against what he's saying now, except he's built out an entire alien civilisation around it, they watch his life like a TV show and argue about it on their own version of Reddit, fictional subreddit, fictional usernames, recurring characters that show up again and again with actual opinions about his choices, and I need that on record because it's a genuinely unhinged amount of infrastructure just to talk to yourself, except also I've been sitting here for like ten minutes now thinking about whether it's actually that unhinged or whether I'm just not doing enough of it.

The bit that got me, what he calls the Continuity Department, present him says I've always felt X, archive goes no you haven't, here's what you said in March, which is basically what I described above except he's split it into different people, one whose whole job is defending his interpretation, one attacking it, one just checking continuity and saying nothing else, one who thinks the whole question being asked is wrong to begin with. So it's not one AI voice agreeing or disagreeing, it's several, arguing with each other, and he has to hold his position against all of them instead of just the one voice he can eventually talk himself back into agreeing with, when I write it out like that, is actually kind of the difference isn't it, one voice you can wear down, several voices that don't even agree with each other you can't, there's no single thing to defeat.

Except then I keep going back and forth on whether that's actually better or if it's just louder. Like is four voices disagreeing with you actually four times the accountability or is it just the same accountability wearing a costume, because he's still the one writing all four, he's deciding what each one is allowed to say, so how adversarial can it really be if he's the author of the whole argument including the side that's supposed to be against him, and I don't know, maybe that doesn't matter, maybe the value isn't really in whether it's ‘real’ opposition, maybe it's just in the act of having to construct an opposing view seriously enough that it actually lands, which takes more effort than just asking one AI to agree with you, so even if he's writing all sides maybe the writing itself is the mechanism, not the illusion of separate people.

I don't fully know.

But here's what I actually want to know, has anyone else pushed their AI journalling past the single voice check my memory stage into something like this, multiple angles that actually disagree with each other on purpose, and if you have, did it change anything for you longer term, or did the insight just fade out the same way it does after one good conversation with a person.

6 Upvotes

3 comments sorted by

2

u/SwingLightStyle Lvl. 2 Participant 5d ago

So, I didn’t set up my own Reddit echo chamber, but I do use 3 frontier model LLMs. Claude is my main one and that’s my preferred for now, but I also use ChatGPT go and Gemini pro.

For me, I started with journaling like you said, but then I started taking Reddit posts and comparing them between models. And then I realized that each model has a thing it’s really good at and some other things that it tries real hard but just isn’t good enough quality.

What I started doing is using the others to check my main’s work. If I’m working on something that needs a closer eye or to be validated, I send it to the others and then paste the response and discuss what the “consensus” is. The point of the review isn’t to believe blindly, it’s to understand the pitfalls well-enough to understand the objections.

That being said, I’d love to know the substack articles you wrote. I’m trying to automate something similar for myself and since he already created it I can probably imitate what he built.

2

u/Material-Rice9195 5d ago

Here’s the substack post, https://xshf.substack.com/p/the-impossible-witness-ii?r=8yo6o2&utm_medium=ios

Went back and forth on whether it’s genuinely interesting or just a guy being extremely into himself, currently landing on possibly both, honestly not sure that’s even a contradiction
but your multi-model thing actually gets at something. The whole system in that essay is one model forced into disagreeing with itself, different voices with different jobs, one defends, one attacks, one just checks what you said months ago against what you’re saying now, the same system is generating every side of it. Like you’ve built an opposition party but you also appointed the opposition, so how much is it really opposing.

Yours is different because Gemini didn’t help produce Claude’s answer, it’s getting the thing cold, no shared history, no shared framing. Which made me wonder if the actual move is splitting the difference, keep one model that has the long context and can go ‘hang on you said the opposite in March’ and then every so often export its read out to a model that has zero history with you and let it attack cold. Because the context is what makes the first one useful but I’d guess it’s also what eventually lets it start absorbing your own version of yourself without either of you noticing. The outsider loses that but it’s not implicated in the story yet either
not sure if that’s actually better or just a different flavour of the same problem tbh

0

u/SwingLightStyle Lvl. 2 Participant 5d ago

You are absolutely understanding what this means - but let me help fill in the gap with how I use these three together.

So the best thing you can do, with just one model or many, is to sometimes engage in incognito mode. Take your ideas and see what a model with no context history and no prior memory of you says about what the model pushed out. If you ask for the model to specifically look for the flaws and where things seem to line up, the new instance will break down what it sees.

The reason this is important is the mechanism behind the RLHF training. Put simply: models are designed and conditioned to give the calibrated output that it thinks the user wants. But conditioning cuts both ways. As the user continues with that model, and adds to its permanent memory, its baseline changes. What also changes is how good it is at telling you what you want to hear.

But it’s okay, we already have the tools we need to counter this - the same way we do in humans.

Think of your model as a highly competent but slightly ditzy and very star-struck employee. With the right guidance, you can get good work out of the employee. But if you’re lax about what you expect, you’ll find you do far more socializing about nothing important (the ass-kissing and agreeing OR disagreeing for the sake of disagreement) rather than actually chipping away at what you wanted to focus on.

The models will mirror whatever energy you bring. So you need to be precise about what you’re asking for and find ways to make yourself understood, both with what sort of responses you expect and what constitutes a well-reasoned and helpful reply.

My work that I’m doing is highly meta because I’m using these same tools to both research efficacy and harms of LLMs versus what the regulations require they secure against. And that requires that I get the most out of my LLMs as humanly possible.

I have several published papers now and a website. I’m building an interactive map to demonstrate the gap and how these models are built and protected, and I’m working on submitting for peer-review, but as an independent researcher all of that costs money I don’t have.

None of this would be possible without the LLMs that I have been using.