AI Security
In-depth guides and analysis on ai security from the Safeguard engineering team.
100 articles
Authorization Belongs in the Tool, Not in the Prompt
Your support chatbot can now issue refunds and look up orders, because someone connected it to real tools. Every one of those actions sits behind a customer-facing text box, protected however carefully the prompt was worded.
A Poisoned Memory Outlives the Conversation That Created It
Prompt injection in a single turn affects one response, then the next request starts clean. A persistent memory feature breaks that boundary by design, so a single successful injection becomes durable, recalled and trusted in every future session.
Writing an MCP Server That Holds No Credentials
Give a model a tool that writes and you have added a route into whatever sits behind it. The design that keeps it a new shape rather than a new privilege, and the four things that were not obvious.
The AI Code Percentage On Your Dashboard Is a Floor, Not a Measurement
Commit-level attribution answers 'lines added by commits an assistant co-authored'. That is a different sentence from 'lines an AI wrote', and the gap between them is where governance metrics go wrong.
Your Git History Already Knows Which AI Wrote Your Code
Coding assistants sign their own work in the commit trailer block. That makes 'how much of this was AI-written' a parsing problem, not a heuristic one — as long as your tooling reads the commit body, which most of it does not.
Indirect Prompt Injection Stopped Being a Demo. Google Is Measuring It at Web Scale.
Malicious injected instructions are now tracked across billions of crawled pages a month, every AI browser tested at Black Hat proved vulnerable, and one framework bug turned a prompt into RCE.
ChatGPT Atlas and the Permanent Browser-Agent Injection Problem
OpenAI shipped ChatGPT Atlas in October 2025 and admitted by December that prompt injection in AI browsers may never be fully solved. Defenders need a posture, not a patch.
You Cannot Defend an MCP Server You Do Not Know You Are Running
Tool poisoning is the most impactful client-side MCP vulnerability, and the defensive research is solid. All of it assumes you know which MCP servers you connect to. Almost nobody does.
Every AI Coding Tool Has the Same Vulnerability, and It Isn't a Bug
Sandbox escapes in Claude Code, critical CVEs in Cursor, a 10.0 in Gemini CLI, prompt injection in Copilot. Different vendors, one shared cause: the agent must hold elevated access to be useful.
2,130 AI-Related CVEs and Counting: The Surge Is Structural, Not a Blip
AI-related CVEs rose 34.6% year over year and more than 200% since 2023. The interesting question is what kind of vulnerabilities they are — because most of them are not model flaws at all.
What agentic AI security means and why traditional AppSec...
Traditional AppSec was built for static code, not decision-making agents. Here's what agentic AI security actually covers—and why autonomous agents need a new defense model.
Identity and access management for non-human AI agents
AI agents now hold production credentials the way employees do, except most are never offboarded. Here's how AI agent identity and access management closes that gap.
Self-healing security runs on Safeguard.
Your first fix PR is minutes away.
No sales call required, even your agent can complete the purchase over MCP.