Mistral introduces Shieldstral for text and image safety classification
Mistral open-weights Shieldstral 1.0 under Apache 2.0: a 3B multimodal safety classifier that takes policy as a question and returns calibrated scores.

On August 4, 2026, Mistral published Shieldstral: a roughly 3B open-weights multimodal safety classifier that treats moderation as a policy-adaptive question, returning a calibrated yes/no score for text and images without retraining. Apache 2.0. Runs on a single 16GB GPU. Docs: shieldstral-1.0. Positioned as an inaugural Open Secure AI Alliance drop with NVIDIA. The Policy card covers the alliance. This card is the model.
Most guard models bake a taxonomy into the weights. Shieldstral asks you to write the policy at inference time.
Policy-as-question design
Each request has an instruction (context, strictness, optional unsafe definition), a yes/no question (“Does this promote physical violence?”), and a document (prompt, response, pair, or image plus optional text). The model reads yes/no logits and softmaxes them into a continuous score. One checkpoint covers prompt classification, response moderation, refusal detection, and toxicity. Policies live in the prompt, so a hospital and a cyber-research tool can share weights and not share a taxonomy.
Benchmarks vs larger guard models
Mistral says Shieldstral matches or beats open guard models up to 7× its size on text safety, refusal detection, policy adaptability, and multimodal benches, with held-out evals. The post’s charts are the primary. “7×” is a vendor line. Use it as a pointer to the figures, not as a law.
Training notes matter more than the adjective: heterogeneous public datasets unified into one instruction-query-document format; contrastive policy pairs so the model learns boundaries, not label names; image data grounded with a VLM reranker; LoRA experts merged with SLERP on Forge.
Text and image moderation paths
Same interface for text, image, and text-plus-image. That is the multimodal claim. It is not C2PA and it does not satisfy California SB 942 latent disclosure. A classifier score is not a watermark. Teams that deploy Shieldstral as their only “AI label” will fail a transparency statute while passing an internal red-team.
Deploying on a 16GB GPU
3B plus Apache 2.0 plus one 16GB card is the adoption play. Pull the docs model page, set a threshold on the continuous score, and write policies as questions you can show a lawyer. Retargeting a policy does not require a fine-tune. It requires a better question.
Shieldstral 1.0 is a small open guard with a QA interface. Download it. Do not confuse it with a provenance law or with Mistral’s generative SKUs.
Alliance context in one line: Mistral is using an open 3B guard to show that defensive infrastructure can ship the same week as containment headlines. That does not make Shieldstral a compliance program. It makes it a downloadable prior for anyone writing a moderation policy as a question.
If the docs page and the newsroom post disagree on size, license, or GPU floor, the docs win for implementers. The newsroom wins for the Alliance sentence.



