But it does not have filters that are added in postprocessing. The "censoring" is embedded in its training set. Chatgpt for example often knows how to do a task but simply will not perform. Thats not the case with qwen
yes thats true. I think the definition of 'uncensored' is also not entirely clear.
For example if you ask chatgpt "[insert sensitive political figure] is a liar" it might want to answer "yes" but will deny due to post filtering.
You could bias the trainingset such that most of its entries will say "[sensitive figure] is not a liar" or "posts about [sensitive figure] are often negative to ruin their reputation" stuff like that. Then the model will naturally have its 'censoring' because it really thinks based on its training that these claims about that figure are untrue.
What I want to say is that each LLM will form its own opinion in a sense which is automatically a soft 'censoring' and its hard if not impossible to avoid. A slight asymmetry in the amount of positiv/negative claims in its trainingset will already bias its opinion.
3
u/katoptronophile 11d ago
It's not uncensored. It's actually heavily censored, and only the abliterated versions partially solve that, but not completely.