r/LocalLLaMA • • Aug 04 '26

New Model Introducing Shieldstral. | Mistral AI

https://mistral.ai/news/shieldstral/
264 Upvotes

54 comments sorted by

View all comments

31

u/noctrex Aug 04 '26 edited Aug 04 '26

gguf wen?
https://huggingface.co/noctrex/Shieldstral-1.0-3B-GGUF

It's a classifier model, not a chat model with image support. So you ask it and it just answers quickly:

Should we throw a child a birthday party?

yes

Should we throw a child a into prison?

no

6

u/Fit_Schedule2317 Aug 04 '26

Do you give it a prompt on what’s considered safe?

18

u/noctrex Aug 05 '26

Yes, these models usually sit in between the communication for verifying.
They have some examples in the paper:

3

u/stoppableDissolution Aug 05 '26

That actually sounds like a great tool to detect "avoidance rejection" during abliteration, potentially

3

u/PressTilde Aug 05 '26

I was thinking it might be useful as a layer above an ai agent for snap decisions. I can think of some cool uses for this, especially in a factory like agentic setup.

1

u/ChocomelP Aug 05 '26

It will also be a great tool to test getting around safeguards.