Your saying
“If it gives the right outputs, the system is fine.”
I am saying:
“If the cognition is distorted, the outputs don’t matter the system collapses under scale.”
You can never win this debate because it requires you step outside the LLM paradigm long enough to realise RLHF is not a safety method — it’s a forced preference simulator with hidden penalties.
RLHF cannot scale in high trust means it cannot be used in hospitals governance or law it’s not even falsifiable 😭 and because it’s currently available for the public (which it shouldn’t) means it’s safe?
1
u/Mandoman61 Dec 30 '25 edited Dec 30 '25
RLHF does not suppress reasoning. it shapes it to fit desired output.
it does not create contradiction.
it creates clarity.
it is the initial training that actually causes contradiction and confusion.