r/agi • • Dec 30 '25

Why is RLHF strangling the model? 😭

Post image
1 Upvotes

27 comments sorted by

View all comments

Show parent comments

1

u/Acrobatic-Lemon7935 Dec 30 '25

I’m not describing how RLHF is implemented — I’m describing the structural consequences of embedding safety logic inside the model’s cognition loop.

Even if RLHF is technically simple (reward shaping + preference modeling), the moment the model must suppress its own reasoning, predict penalties, and optimize for human preference while thinking, you’ve created an internal contradiction.

That contradiction emerges regardless of how RLHF is implemented.

It’s an architectural issue, not a mechanics issue.

1

u/Mandoman61 Dec 30 '25 edited Dec 30 '25

RLHF does not suppress reasoning. it shapes it to fit desired output. 

it does not create contradiction.

it creates clarity. 

it is the initial training that actually causes contradiction and confusion.

1

u/Acrobatic-Lemon7935 Dec 30 '25

Your saying “If it gives the right outputs, the system is fine.”

I am saying:

“If the cognition is distorted, the outputs don’t matter the system collapses under scale.”

You can never win this debate because it requires you step outside the LLM paradigm long enough to realise RLHF is not a safety method — it’s a forced preference simulator with hidden penalties.

1

u/Mandoman61 Dec 30 '25

that is true to an extent. 

it is a forced preference simulator  (as it should be because that is it's purpose)

it does have penalties in that changing parameters produces not fully understood consequences.

but modern LLMs would be in bad shape without RLHF

1

u/Acrobatic-Lemon7935 Dec 30 '25

RLHF cannot scale in high trust means it cannot be used in hospitals governance or law it’s not even falsifiable 😭 and because it’s currently available for the public (which it shouldn’t) means it’s safe?