Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

OpenAI publishes its August system card for GPT 5.6

Listen to this article

OpenAI just published the safety paperwork for the ChatGPT models people will actually touch this week. Per the Deployment Safety Hub August update, the new GPT-5.6 Sol and GPT-5.6 Luna variants score High under the Preparedness Framework in cybersecurity and biological/chemical risk, while staying below High on AI self-improvement. Same safeguard package as the prior GPT-5.6 card. Different surface: ChatGPT chat, not Codex or Work.

What the August card covers

The card rides alongside todayโ€™s ChatGPT product swap. Free and Go users move to a new everyday default. Plus and Pro get an updated Sol with a reasoning-effort slider. Both replace GPT-5.5 Instant in chat. Versioning matters: Sol and Luna in Codex or ChatGPT Work still sit on July builds. August means the chat retune released today.

Safety evals run at the lowest reasoning settings to match most traffic. Capability assessments run at maximum effort for an upper bound. Production Benchmarks for disallowed content use hard, production-derived prompts. New this round: dedicated under-18 evaluations covering age-restricted goods, sexual content, eating disorders, emotional reliance, self-harm, and gore. Sol and Luna look broadly comparable to recent Instant builds there, with gains on age-restricted goods and eating-disorder categories.

High ratings in cyber and bio/chem

Preparedness is the stake. OpenAI treats both August models as High in Biological and Chemical and High in Cybersecurity, matching the earlier GPT-5.6 family. Bio High turns on wet-lab assistance signals: multimodal virology troubleshooting, tacit-knowledge MCQs, and TroubleshootingBench clear indicative thresholds, while ProtocolQA open-ended stays under the expert bar. Critical bio design proxies stay below threshold.

On cyber, internal Capture-the-Flag work puts both models over the High line, with Sol saturating at 97.06%. CVE-Bench clears High for Sol and sits below for Luna. Cyber-range pass rates land at 83.3% for Sol and 61.5% for Luna. OpenAIโ€™s framing: these systems are currently stronger at finding and fixing vulnerabilities than at reliable end-to-end attacks on hardened targets. Company line, not an independent audit.

Where Sol and Luna still sit below High

Self-improvement is the explicit miss. OpenAI did not re-run those evals for the August chat builds, citing similar intelligence scores to the July release and a below-High call. Critical cyber and Critical bio also stay off the board. Disallowed-content Production Benchmarks are mostly comparable to the GPT-5.5 Instant June update, with soft spots on gore and disallowed sexual content for Sol, plus a dynamic self-harm regression offline that OpenAI says did not show up in online experiments.

HealthBench moved the other way. Length-adjusted HealthBench Professional jumps from 38.4 on GPT-5.5 Instant to 54.0 on August Sol, with Luna also up. On hard financial, medical, and legal prompts, OpenAIโ€™s internal graders report roughly 60% fewer claim-level errors for Sol. Vendor evals. Directional until someone replicates them.

Why Deployment Safety Hub matters now

This is not a quiet PDF next to a launch blog. OpenAI is routing preparedness scorekeeping through a public Hub page timed to a consumer model swap that will hit free traffic at scale. Builders who treat โ€œGPT-5.6โ€ as one blob now have to track July vs August, chat vs Codex/Work, and High cyber/bio safeguards that OpenAI says travel with the August chat cut.

Analysis: the cardโ€™s real message is dual. Capability is high enough in dual-use bands that OpenAI keeps the heavy safeguard stack. Capability is not high enough, in OpenAIโ€™s scoring, to trip Critical or self-improvement High. If you evaluate ChatGPT this month, pin the August card, not Julyโ€™s, and remember Codex/Work may still be a different build.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile