Mistral introduces Shieldstral for text and image safety classification
Mistral open-weights Shieldstral 1.0 under Apache 2.0: a 3B multimodal safety classifier that takes policy as a question and returns calibrated scores.

On August 4, 2026, Mistral published Shieldstral: a roughly 3B open-weights multimodal safety classifier that treats moderation as a policy-adaptive question, returning a calibrated yes/no score for text and images without retraining. Apache 2.0. Runs on a single 16GB GPU. Docs: shieldstral-1.0. Positioned as an inaugural Open Secure AI Alliance drop with NVIDIA. The Policy card covers the alliance. This card is the model.
Most guard models bake a taxonomy into the weights. Shieldstral asks you to write the policy at inference time.
Policy-as-question design
Each request has an instruction (context, strictness, optional unsafe definition), a yes/no question (“Does this promote physical violence?”), and a document (prompt, response, pair, or image plus optional text). The model reads yes/no logits and softmaxes them into a continuous score. One checkpoint covers prompt classification, response moderation, refusal detection, and toxicity. Policies live in the prompt, so a hospital and a cyber-research tool can share weights and not share a taxonomy.
Benchmarks vs larger guard models
Mistral says Shieldstral matches or beats open guard models up to 7ร its size on text safety, refusal detection, policy adaptability, and multimodal benches, with held-out evals. The post’s charts are the primary. “7ร” is a vendor line. Use it as a pointer to the figures, not as a law.
Training notes matter more than the adjective: heterogeneous public datasets unified into one instruction-query-document format; contrastive policy pairs so the model learns boundaries, not label names; image data grounded with a VLM reranker; LoRA experts merged with SLERP on Forge.
Text and image moderation paths
Same interface for text, image, and text-plus-image. That is the multimodal claim. It is not C2PA and it does not satisfy California SB 942 latent disclosure. A classifier score is not a watermark. Teams that deploy Shieldstral as their only “AI label” will fail a transparency statute while passing an internal red-team.
Deploying on a 16GB GPU
3B plus Apache 2.0 plus one 16GB card is the adoption play. Pull the docs model page, set a threshold on the continuous score, and write policies as questions you can show a lawyer. Retargeting a policy does not require a fine-tune. It requires a better question.
Shieldstral 1.0 is a small open guard with a QA interface. Download it. Do not confuse it with a provenance law or with Mistral’s generative SKUs.
Alliance context in one line: Mistral is using an open 3B guard to show that defensive infrastructure can ship the same week as containment headlines. That does not make Shieldstral a compliance program. It makes it a downloadable prior for anyone writing a moderation policy as a question.
If the docs page and the newsroom post disagree on size, license, or GPU floor, the docs win for implementers. The newsroom wins for the Alliance sentence.



