
Study finds limits in simple AI safety embedding scores
A new paper tests a simple safety scoring shortcut. Adding a reference improves ranking, while deployment limits remain substantial.
Accessibility Adjustments
Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.
AI safety and security examines how AI systems fail, how they can be misused, and which safeguards have evidence behind them. This category covers model evaluations, prompt injection, data exposure, misuse testing, and the practical limits of protective measures. We look at the conditions of a test, the behavior it measured, and whether the results support the conclusions being drawn. Coverage also asks how containment, permissions, monitoring, and human review affect the consequences of an error or attack. Findings from AI research help explain emerging risks and the methods used to investigate them. The growing use of AI agents with access to tools makes questions about authority and access especially relevant. A useful assessment states its evidence limits, including missing tests, uncertain assumptions, and behaviors that may change in a different setting. The goal is to help readers evaluate safety claims and understand which controls address a specific risk.

A new paper tests a simple safety scoring shortcut. Adding a reference improves ranking, while deployment limits remain substantial.

AISIโs October 1 update describes stronger network restrictions, monitoring and checks before testing. Further infrastructure work remains underway.

Anthropicโs October 2 snapshot reports more findings sent to maintainers, with separate review and patch counts that need careful interpretation.

Apple says future macOS controls will require explicit action before apps receive Full Disk Access, citing growing risks from autonomous AI agents.

New research examines a more selective approach to screening biological AI work.

A new preprint describes a security assessment method that connects recorded software controls with explicit and repeatable scoring rules.

An October 1 preprint studies when an AI result can be checked without exposing the confidential information behind it.

A new preprint examines how learned connections between AI agents can change safety behavior even when the underlying models stay fixed.

Agility and FORT plan to connect Digit 5 with external safety systems through an expanded partnership that still requires definitive agreements.

Google finds a different mix of vulnerabilities in AI attributed research, with important limits on what the comparison proves.

OpenAI outlines Private Intelligence and a fall Private Inference preview. Its safety processing documentation clarifies retention duties and protection boundaries.

Codex Security Cloud scans GitHub repositories and monitors new commits. Its documentation explains validation, proposed patches and public review visibility.

OpenAI just put a sharper cyber knife in trusted hands and left everyone else staring at the rope. In todayโs Daybreak expansion post, the lab introduced GPT-5.6-Cyber for authorized vulnerability research and split Daybreak into Blue and Red tiers. Theโฆ

OpenAI just published the safety paperwork for the ChatGPT models people will actually touch this week. Per the Deployment Safety Hub August update, the new GPT-5.6 Sol and GPT-5.6 Luna variants score High under the Preparedness Framework in cybersecurity andโฆ

An autonomous FOFA-to-PoC loop is already operational in the wild โ and still brittle. Palo Alto Networks Unit 42 published a July 30 report on a Chinese-speaking operator who used DeepSeek through open-source Hermes Agent to enumerate FOFA targets, pullโฆ

Anthropic’s Frontier Red Team just showed that a frontier model can find mathematical cracks in crypto designs that survived years of human review. Deployed systems are still fine. The research bar just moved. In a July 28 research post, Anthropicโฆ

Microsoft introduced MAI Cyber 1 Flash inside MDASH. Its reported CyberGym score covers a combined system, with access limited to verified defenders.