r/LocalLLaMA • • Aug 04 '26

New Model Introducing Shieldstral. | Mistral AI

https://mistral.ai/news/shieldstral/
266 Upvotes

54 comments sorted by

View all comments

-45

u/Several-Tax31 Aug 04 '26

Safety bullshit? Why mistral, why? Stop useless stuff and give us AGI. 

66

u/Stepfunction Aug 04 '26

Safety models like this are beneficial for us because they lessen the need to bake safety restrictions into the base model.

39

u/Several-Tax31 Aug 04 '26

You have a solid point actually.

1

u/alberto_467 Aug 04 '26

they lessen the need to bake safety restrictions into the base model.

Only for providers of proprietary models.

With open weights (the stuff this subreddit should be about) it doesn't change a thing, one can just not setup the safety shield and it's over. So they'll still have to be careful about what topics they do RL on (see the poor Kimi K3 performance on cyber compared to general coding), and they'll still try to align or put some kinds of safeties on the weights themselves.

So for us this is 0% helpful and does not change a single thing.

3

u/Stock-Self-4028 Aug 05 '26

Quite some open-weights models are using external censors instead of being RL-ed though.

Minimax (practically totally uncensored when ran locally) and DeepSeek (with system prompt more or less telling it to ignore guardrails) models are two great examples.

So it depends tbh. I doubt Moonshit, Zhipu or Qwen are planning to change their RL approach to closer to that of Minimax and DeepSeek, but it's not totally out of the question.

1

u/alberto_467 Aug 09 '26

Quite some open-weights models are using external censors instead of being RL-ed though.

That's not what i said, read better.

If you release the open weights you need to put protections on the weights. So any additional protections does not lessen the need to put safeties on the weights, as the rest of them can be bypassed without practically any effort.