
GitHub upgrades AI secret detection and outlines new checks
Existing alert scans get a new model while additional checks have separate rollout and billing plans.
Accessibility Adjustments
Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.
AI safety and security examines how AI systems fail, how they can be misused, and which safeguards have evidence behind them. This category covers model evaluations, prompt injection, data exposure, misuse testing, and the practical limits of protective measures. We look at the conditions of a test, the behavior it measured, and whether the results support the conclusions being drawn. Coverage also asks how containment, permissions, monitoring, and human review affect the consequences of an error or attack. Findings from AI research help explain emerging risks and the methods used to investigate them. The growing use of AI agents with access to tools makes questions about authority and access especially relevant. A useful assessment states its evidence limits, including missing tests, uncertain assumptions, and behaviors that may change in a different setting. The goal is to help readers evaluate safety claims and understand which controls address a specific risk.

Existing alert scans get a new model while additional checks have separate rollout and billing plans.

Strands Box combines local isolation with rules based on earlier actions. Its macOS preview leaves direct file grants outside policy history.

A new benchmark separates safe task completion from efficient resource use.

A preprint tests what electricity measurements can establish about computation.

A new preprint measures the cost and detection tradeoff in agent monitoring.

Claude Code 2.1.292 repairs plugin permission checks and uses base branch CLAUDE.md rules when a pull request edits them.

The edition tests agents against cybersecurity job roles.

A September check of three synthetic images found uneven use of embedded provenance, with important limits on what the results establish.

The evaluator tests security boundaries before measuring a modelโs cyber capabilities.

Claude Code 2.1.290 fixes permission enforcement and task recovery with a compaction limit.

The offering combines Rubrik controls with LTM assessments, deployment and ongoing governance.

The collaboration links network protections with software fixes. Zscalerโs shield remains in early access.

API customers can enable textGrain for selected models. OpenAI also plans an EU rollout for ChatGPT and Codex in the coming weeks.

Experimental API credits leave providers able to see prompts

C1 details credential vending and destination controls while its linked AppHub code still requires production validation

Memory checks reduce attacks in a controlled agent study

BlackFog introduces prompt checks and optional auditing as ADX Vision expands into AI agent workflows

Meta updates its framework with training controls and open weight considerations

VeriSpec checks written AI rules and leaves final judgments to reviewers

Claude Code 2.1.289 repairs policy checks and extends teammate events

Puzzle study separates useful warnings from weak model signals

RSA plans a November release for agent discovery and access controls, with broader governance scheduled for 2027

Microsoftโs annual security report separates controlled AI evaluations from observed intrusions and emphasizes identity controls.

A new paper tests a simple safety scoring shortcut. Adding a reference improves ranking, while deployment limits remain substantial.