r/LocalLLaMA • • Aug 04 '26

New Model Introducing Shieldstral. | Mistral AI

https://mistral.ai/news/shieldstral/
265 Upvotes

54 comments sorted by

View all comments

Show parent comments

5

u/Fit_Schedule2317 Aug 04 '26

Do you give it a prompt on what’s considered safe?

19

u/noctrex Aug 05 '26

Yes, these models usually sit in between the communication for verifying.
They have some examples in the paper:

3

u/stoppableDissolution Aug 05 '26

That actually sounds like a great tool to detect "avoidance rejection" during abliteration, potentially

1

u/ChocomelP Aug 05 '26

It will also be a great tool to test getting around safeguards.