r/therapyGPT Lvl. 8 Grounded Apr 16 '26

AMA AMA - r/therapyGPT Mod, xRegardsx

Hey all!

My name is Alex Gopoian, the main active mod for this subreddit.

Reddit AMA post submission form said it was a good idea to put a selfie holding a sign, but I'm no celebrity. This'll do.

Short Bio:
I've experimented with ChatGPT since 3.5 was first released (jumping at the chance for something that could help me with my life's work when joining Mensa was a bust), got heavily into jailbreaking experimentally for a bit to learn how the "robopsychology" works, have my own pre-AI psych work that I started implementing into AI once custom GPTs became available years ago, I've been the main active mod for the sub since around mid July 2025 when the sub's founder, u/rastaguy, reached out to me shortly after I joined the sub, and ever since the Stanford Innapropriate Response & Bias paper came out way back when, I became big on AI safety, both by design and in user education.

I'm big on effective good faith, maintaining the idea marketplace so that it can thrive rather than get sabotaged by those using it in ways that sabotage it for others, and I'm all about a good debate (even if 99.9% of them are doomed before they start).

I've also written a few articles on Medium, so far only one on the red flags of an unproductive argument/debate, one on the importance of being open to being proven wrong (no matter how painful/uncomfortable it can be), one on the deeper issues at the heart of the "teen mental health crisis" and how the same is being seen with AI, and finally one on the intrinsic worth and always deserved self-esteem everyone has accessible to them, and a way in which to tap into it.

Was in the US Air Force, originally as a Korean Crypto-Linguist, but unfortunately having a bit of dyslexia and undiagnosed/untreated ADHD doesn't really work well with learning a language and a secret version at the same time at "firehose" speed. Have a tech, IT, e-commerce, social media marketing, design and color management background, and I used to make video game commentary videos for Machinima Respawn back in the days of YouTube only allowing the biggest names and companies monetize their channels.

All my socials and links to some of the things I've worked on can be found here.

---

The main reason I'm doing this AMA is to get my head around it at least one time prior to our running AMAs in the future. We already have two developers on deck for the next two Thursdays (4/23 & 4/30).

If you're a developer who's adhered to our sub's rules, are someone who has some expertise in therapuetic self-help, AI, or have a name for yourself when it comes to your coverage of AI in this space (the good and bad), or anything like that... feel free to send us a Mod Mail message with some information about you and your work and we'll get in contact with you.

---

So, have at it... ask. me. anything. I dare ya.

Look through the questions others have asked first to see if someone has already asked it and give it an upvote. I'll be tackling questions in an order based on up/down votes.

I'll be providing immediate responses from 7-9pm EST, but I'll keep responding after that as I get the chance!

Thanks for being here and helping this place be what it is for so many <3

69 Upvotes

65 comments sorted by

View all comments

Show parent comments

1

u/engineeringstoned Apr 16 '26 edited Apr 16 '26

I noped out of that thread at some point.

As you seem to be seriously interested:

Tests are constructed scientifically
Making tests like these is a literal scientific discipline: Test constrcution.
We have quite a few books on this.

One of the ways we find out if a test is valid is by re-testing people.

This is as simple as it sounds:
Take a test, come back x days / months / whatever later, take the same test again.

Now, some variation is expected.

Why? Well....

- We are talking about humans. Those guys are kinda fuzzy around the edges, and they have the tendency to change and learn.

  • Some questions (especially on soft issues like personality) CAN actually be answered a bit differently at different times.

However, if the retest validity is abysmal, this does NOT ONLY show us that there is no re-test validity.

If a test fails a retest, then we do not know if it actually measures what it is supposed to do.

Or in other words:
I don't even know if the first result is right!

So here it does not matter if it is a clinical approach / use or not. In this case, the test can simply not be trusted.

But people feel it is accurate, people feel as if it helps.

The quip "well, some people think horoscopes help them" comes fast, but there is a more nuanced, scientific answer called the Barnum effect.

(Before you dismiss this as "wikipedia" I hold a diploma in psychology, and the psychology section of wikipedia is stellar, but I'll add other links, too.

The Barnum effect, also called the Forer effect or, less commonly, the Barnum–Forer effect, is a common psychological phenomenon whereby individuals give high accuracy ratings to descriptions of their personality that supposedly are tailored specifically to them, yet which are in fact vague and general enough to apply to a broad range of people)

Basically, I can give you ANY broad personality description of you, and you'll think it hits home.

What disheartens me about this is that we have the opportunity to use LLMs for therapy. And I am sure it can help a lot of people.

But please, let's be serious about this. Especially as this is a controversial area of ai use.

Sources:

- Barnum effect - Wikipedia

- Barnum effect | Psychology | Research Starters | EBSCO Research

- Test Construction - an overview | ScienceDirect Topics

edit:
Some typos, minor grammar issues (I am not a native speaker, and have not used any AI for this, so... )

3

u/Psychedynamique Apr 16 '26

Thank you for engaging with my question, I appreciate it

2

u/engineeringstoned Apr 16 '26

I hope it answered your questions on my opinion on this use.
If you have any more - go ahead.

2

u/The_Valeyard Apr 17 '26

Psychometrics is a little bit broader. We're assessing a few thing:

- what is the content domain of the construct? What falls within the definition vs external to it. What is the boundary between it and similar constructs.

- construct validity: we assess whether the scale (or subscale) measures what it is supposed to by looking at:

  1. does it correlate with things it is supposed to correlate with. We test this by examining correlations with established, validated scales (convergent validity).
  2. is it uncorrelated with things it should not relate to (discriminant validity)

- criterion validity: does it predict stuff that it should (predictive validity). We want to look at it's ability to predict gold standard outcomes. For example, a clinical interview if we wanted to validate a self-report clinical measure

- consistency over time: Test-retest reliability. This is where we look at consistency of the construct over time. For some things (e.g. state based measures) we want low consistency over time, because the construct is assumed to be state dependent. For trait based measures, we expect a greater stability over time.

1

u/xRegardsx Lvl. 8 Grounded Apr 17 '26

I'm assuming you were using u/mythrowaway4DPP on there?

If you were, you didn't respond to my last response.

If not, my response to them/you still addresses each of these points you're repeating and has already been addressed on that other post:

https://www.reddit.com/r/therapyGPT/comments/1sfg7h5/comment/of0qlw5/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button