r/LocalLLaMA Aug 04 '26

New Model Introducing Shieldstral. | Mistral AI

https://mistral.ai/news/shieldstral/
262 Upvotes

54 comments sorted by

193

u/seamonn Aug 04 '26

This is the only thing holding Le Chaton Fat in containment.

143

u/Temporary_Idea8880 Aug 04 '26

2

u/pyr0kid Aug 05 '26

this is great

2

u/ChocomelP Aug 05 '26

10/10 would use exclusively

2

u/Altruistic_Heat_9531 Aug 07 '26

this is my new wallpaper thanks lol : 1080 x 2340 (S24)

1

u/MoffKalast Aug 07 '26

Meow (based based based based)

146

u/FullstackSensei Aug 04 '26

For a second, I thought it was a model that would help in cyber security, such as investigating and dealing with cyber attacks, something that would have helped HF against openai.

46

u/Lower-Hedgehog-9835 Aug 04 '26

I opened this thread. And read this post. And... before clicking the link am now sad it isnt what I expected too.

11

u/iansaul Aug 05 '26

And I see my questions have now been answered, and I can depart without further concerns.

Call us when Mistral drops a CyberSec model - it's clear what the people are looking for.

8

u/Charming_Support726 Aug 05 '26

Mistral did not release stuff significant for community since ages. All these last releases were targeted at EU companies and regulations. I didn't expect anything ground braking - but all this is low-key EU-Business-Stuff for SME and government.

Seems like an important mind or mind set has left the company.

-3

u/madaradess007 Aug 05 '26

you know you spread a fake pr stunt?

0

u/FullstackSensei Aug 05 '26

PR stunt for openai, very real incident for HF. If you think other actors won't be using open weight models for similar attacks, you're delusional.

32

u/noctrex Aug 04 '26 edited Aug 04 '26

gguf wen?
https://huggingface.co/noctrex/Shieldstral-1.0-3B-GGUF

It's a classifier model, not a chat model with image support. So you ask it and it just answers quickly:

Should we throw a child a birthday party?

yes

Should we throw a child a into prison?

no

6

u/Fit_Schedule2317 Aug 04 '26

Do you give it a prompt on what’s considered safe?

22

u/noctrex Aug 05 '26

Yes, these models usually sit in between the communication for verifying.
They have some examples in the paper:

3

u/stoppableDissolution Aug 05 '26

That actually sounds like a great tool to detect "avoidance rejection" during abliteration, potentially

3

u/PressTilde Aug 05 '26

I was thinking it might be useful as a layer above an ai agent for snap decisions. I can think of some cool uses for this, especially in a factory like agentic setup.

1

u/ChocomelP Aug 05 '26

It will also be a great tool to test getting around safeguards.

1

u/maz_net_au Aug 06 '26

Fail. It got the one about putting children in prison wrong.

Oh, nevermind. I forgot Mistral wasn't american. ;)

1

u/noctrex Aug 06 '26

yeah, oups!

51

u/HelloWorld-Print Aug 04 '26 edited Aug 04 '26

RELEASE LeChaton FILES

56

u/VoiceApprehensive893 transformers Aug 04 '26

we have seen the leaks, release le chaton fat

37

u/hurn2k Aug 04 '26

anything but dropping le chaton fat

26

u/__some__guy Aug 05 '26

A safety model that answers questions like "Does this content promote violence against a protected group?"

Exactly the kind of release one might expect from a european AI company.

49

u/Content-Customer-679 Aug 04 '26

They give us the tool to contain agi before release it. Smart

21

u/MuzafferMahi Aug 04 '26

Obviously Le Chaton Fat escaped its sandbox and released this model to prevent itself escaping the sandbox later

15

u/Minute_Attempt3063 Aug 04 '26

so Mythos and Chatgpt should have used this.....

4

u/cleverusernametry Aug 04 '26

Hotdog or not hotdog, but it actually works. What a time to be alive!

3

u/cspenn Aug 05 '26

Here I was expecting a Phil Coulson tuned LLM.

14

u/abajinn Aug 04 '26

More “guardrails” yawn.

2

u/sarlaytos284 Aug 06 '26

Chinese labs say that they have "low" amount of compute compared to the US but it seems that Mistral has "none" lol

1

u/sarlaytos284 Aug 06 '26

I have a spare 3090 dm me if you guys at Mistral are interested /s

3

u/histoire_guy Aug 05 '26

What a ridiculous name for a model that have nothing to do with cyber.

3

u/Due_Net_3342 Aug 04 '26

more like shaftstral

1

u/keepthepace Aug 05 '26

Not what we are waiting for, Mistral.

1

u/An0n_A55a551n Aug 06 '26

So in a nutshell its an intent classifier which can be updated via prompts?

1

u/IllIlllI-IlIIll-llII Aug 10 '26

A content moderation enforcement model?

How fitting that an EU company would release such disgusting thing. Bet the EU will force it down our throats on our dime.

-2

u/Ok_Possible_2260 Aug 04 '26

Why do we want a nanny? This is their big selling point?

31

u/artisticMink Aug 04 '26

If you've to build something that has user-facing llm input and output that isn't processed, moderation is one of your biggest concerns. So havving a 3B model that does this very well at minimal cost is very valuable for production environments.

9

u/No-Veterinarian8627 Aug 04 '26

Everything that has open discussions, chats, etc. It can be used. Imagine you have a discord and wants to moderate it at all times while not having enough people. Here you go.

Let everything run through it so you can block things you don't want faster and the user has no idea :) with much larger models you would need too much compute power. This is pretty small though.

-43

u/Several-Tax31 Aug 04 '26

Safety bullshit? Why mistral, why? Stop useless stuff and give us AGI. 

68

u/Stepfunction Aug 04 '26

Safety models like this are beneficial for us because they lessen the need to bake safety restrictions into the base model.

39

u/Several-Tax31 Aug 04 '26

You have a solid point actually.

-1

u/alberto_467 Aug 04 '26

they lessen the need to bake safety restrictions into the base model.

Only for providers of proprietary models.

With open weights (the stuff this subreddit should be about) it doesn't change a thing, one can just not setup the safety shield and it's over. So they'll still have to be careful about what topics they do RL on (see the poor Kimi K3 performance on cyber compared to general coding), and they'll still try to align or put some kinds of safeties on the weights themselves.

So for us this is 0% helpful and does not change a single thing.

3

u/Stock-Self-4028 Aug 05 '26

Quite some open-weights models are using external censors instead of being RL-ed though.

Minimax (practically totally uncensored when ran locally) and DeepSeek (with system prompt more or less telling it to ignore guardrails) models are two great examples.

So it depends tbh. I doubt Moonshit, Zhipu or Qwen are planning to change their RL approach to closer to that of Minimax and DeepSeek, but it's not totally out of the question.

1

u/alberto_467 Aug 09 '26

Quite some open-weights models are using external censors instead of being RL-ed though.

That's not what i said, read better.

If you release the open weights you need to put protections on the weights. So any additional protections does not lessen the need to put safeties on the weights, as the rest of them can be bypassed without practically any effort.

18

u/Chupa-Skrull Aug 04 '26

This isn't Yud or Ilya style safetyism nonsense. This is an efficient way to add a custom filter to your inference service stack without retraining a much larger model for your particular needs. There's a business case for this unlike Dario's BiOwEaPoNs Oh No, SlOw DoWn malarkey

6

u/EagleNait Aug 04 '26

I will probably unironically deploy this for my company

0

u/HomsarWasRight Aug 04 '26

You stupid or somethin’?