AI Security
In-depth guides and analysis on ai security from the Safeguard engineering team.
786 articles
Daybreak vs. Mythos: 2026 Is the Year the Frontier Labs Entered Defensive Security
OpenAI's Daybreak and Anthropic's Mythos both bet that frontier models can find and fix vulnerabilities at scale. The discovery race is real — but the bottleneck, the cost curve, and the winning strategy all point the same direction: be model-agnostic.
Patch the Planet: What AI-Generated Fixes Actually Mean for Open-Source Maintainers
OpenAI's Patch the Planet, co-founded with Trail of Bits, wants to move widely-used open-source projects from findings to fixes. The ambition is right — but it shifts the bottleneck to maintainer review, patch provenance, and the trust of machine-authored code.
OpenAI's Daybreak: An Honest Assessment of Codex Security, GPT-5.5-Cyber, and the Find-Validate-Patch Loop
Daybreak is the most complete attempt yet to turn a frontier model into a vulnerability-finding-and-fixing system. We break down what it gets right, where the verification and economics still bite, and how it fits alongside a purpose-built engine.
What Is a Claude Code Skill, and How Do You Secure One?
A Claude Code skill is a folder of Markdown instructions and scripts that an AI agent loads on demand. Because it can carry executable code, it deserves the same review as any dependency.
Securing AI coding assistants (Claude Code, Copilot, etc.)
AI coding assistants like Claude Code and Copilot introduce new supply chain risks. Here's what's actually going wrong and how to secure your pipeline.
Prompt Injection Examples: Attacks Seen in the Wild
From hidden text in resumes to poisoned web pages that hijack AI browsing agents, prompt injection has moved from research demos to real incidents. Here are the patterns and what actually blunts them.
GPT-5.5-Cyber and Trusted Access: The Dual-Use Governance Questions Defenders Should Be Asking
OpenAI's Daybreak ships a permissive, offensive-capable model behind a tiered Trusted Access program and a wave of government partnerships. Here's what model-risk, procurement, and security-policy teams should demand before they rely on it.
Agentic AI Security: Why Architecture Beats Model Size in Vulnerability Discovery
The CyberGym leaderboard shows the lead in AI vulnerability discovery moving to multi-agent orchestration, not raw model scale. Here is what that means for security teams betting on agentic AI.
What is AI Code Remediation?
AI code remediation turns vulnerability findings into ready-to-merge patches. Here's how it works, where Veracode's approach falls short, and how Safeguard closes the gap.
SPDX 3.0 AI Profile: Building an AIBOM in Practice
SPDX 3.0 was published in March 2025 with a dedicated AI profile and a Dataset profile. We walk through how to produce a defensible AIBOM in SPDX format alongside or in place of CycloneDX.
What is Vibe Coding (and its security risk)?
Vibe coding lets AI write your app while you skip the review. Veracode found 45% of AI-generated code is vulnerable. Here's the risk, and how Safeguard closes the gap.
OpenRouter API Security: Using the Unified LLM Gateway Safely
The OpenRouter API routes your prompts through one endpoint to many model providers. Convenient, but it changes where your data goes and where your keys live.
Self-healing security runs on Safeguard.
Your first fix PR is minutes away.
No sales call required, even your agent can complete the purchase over MCP.