AI Security
In-depth guides and analysis on ai security from the Safeguard engineering team.
779 articles
AI Agent Memory: Security Risks
Persistent memory makes AI agents more useful and more dangerous. A security engineer's walkthrough of how agent memory gets poisoned, exfiltrated, and weaponised, with concrete 2025 examples.
Gemini 2.5 Pro and the Late Safety Report
Google released Gemini 2.5 Pro Experimental on March 25, 2025 without a contemporaneous safety report. The UK reaction set a precedent.
Vector DB Security Considerations
Vector stores hold derivatives of your most sensitive text. We cover the access, isolation, and integrity controls production deployments of Pinecone and Weaviate need.
How Snyk Agent Fix's agentic retry loop self-corrects fai...
A technical look at how Snyk's Agent Fix uses a bounded, feedback-driven retry loop to validate and self-correct AI-generated vulnerability fixes before they reach a pull request.
GenAI Code Review Tools: A 2025 Field Test
We field-tested five GenAI code review tools against 240 seeded security defects to see which catch real issues and which hallucinate findings.
Supply Chain Attacks Targeting AI/ML Pipelines
AI and ML pipelines introduce unique supply chain risks -- from poisoned training data to compromised model registries. Here is what attackers are targeting and how to defend.
Open-Weight Model Sandboxing Patterns
Running an open-weight model inside an enterprise perimeter seems safer than calling a hosted API. It is, and it isn't. The sandboxing patterns that actually produce the safety properties.
GPT-5 Launch: Reading the System Card for Supply-Chain Risk
GPT-5 shipped August 13, 2025 under OpenAI's Preparedness Framework v2. Here's what the system card tells security teams about deployment risk.
Local LLM Deployment: Enterprise Risks
Running LLMs on local hardware eliminates some risks and introduces others. A clear-eyed look at the enterprise risk profile of on-premise and on-device model deployments.
The Prompt Injection Problem Hiding Inside Everyday Code ...
AI coding assistants read untrusted files as instructions, not data. Here's how prompt injection sneaks malicious code into your commits — and how to catch it before it ships.
Model Training Data and the Propagation of Insecure Codin...
LLM coding assistants inherit insecure patterns from their training data — from SQLi-prone snippets to hallucinated packages attackers exploit. Here's how the risk propagates.
Securing AI Agents: MCP Protocol Risks and Mitigations
The Model Context Protocol is transforming how AI agents interact with tools, but it introduces new attack surfaces. Here is what security teams need to understand.
Self-healing security runs on Safeguard.
Your first fix PR is minutes away.
No sales call required, even your agent can complete the purchase over MCP.