AI Security
In-depth guides and analysis on ai security from the Safeguard engineering team.
100 articles
Understanding membership inference attacks against traine...
Membership inference attacks let adversaries confirm if your data trained a model, exposing privacy leakage in ML models and training data inference risks.
What shadow AI is and how to discover unsanctioned AI use...
Shadow AI risk is spreading faster than governance can keep up. Here's what unsanctioned AI use looks like inside real enterprises and how to discover it before data leaks.
Glossary of AI Trust, Risk, and Security Management (AI T...
A glossary of AI trust risk security management concepts: the Gartner AI TRiSM framework, its four pillars, AI risk taxonomy, and adversarial threats.
How to build an AI-specific incident response playbook
A step-by-step guide to building an AI incident response plan — covering scoping, escalation, detection, containment, and post-incident review for LLM and agent failures.
CoSAI Releases Model Signing and Incident Response Frameworks
The Coalition for Secure AI published two operational frameworks in November 2025: Signing ML Artifacts and AI Incident Response. We unpack what each contains and how to adopt them.
Prompt Injection Detection in Retrieval Systems
Indirect prompt injection arrives through your retrieval corpus, not your chat box. We cover the detection strategies that survive when attackers write your RAG content.
The Hugging Face Breach: What Changes When an AI Agent Runs the Intrusion
On 16 July 2026 Hugging Face disclosed that a malicious dataset gave an attacker code execution inside its data-processing pipeline, escalating to node-level access and internal cluster credentials over a single weekend — driven by an autonomous agent framework executing thousands of actions. Here's the anatomy, and what it means for anyone who treats a model registry as a trusted input.
vLLM CVE-2025-62164: Tensor Deserialization RCE
vLLM 0.10.2-0.11.0 deserialized user-supplied PyTorch tensors via torch.load() in the Completions API. Memory corruption, potential RCE.
Claude Opus 4.5 System Card: Defender Takeaways
Anthropic released Claude Opus 4.5 on November 24, 2025 with the most detailed safety section of any system card to date. We pull out what enterprise defenders should change.
AI Model Watermarking and Provenance
Watermarking and provenance are the two most confused terms in AI security. A practical breakdown of what each actually does, where the 2025 techniques break, and what to ship in the meantime.
vLLM CVE-2025-66448: Auto-Map RCE via Model Configs
A critical RCE in vLLM allows malicious model configs to bypass trust_remote_code=False. We analyze the bug, the patch, and what every vLLM operator should do.
Gemini 3 Pro and the Frontier Safety Framework Report
Google released Gemini 3 Pro on November 18, 2025 with the most thorough Frontier Safety Framework evaluation yet. We unpack what was disclosed and how it changes downstream defender posture.
Self-healing security runs on Safeguard.
Your first fix PR is minutes away.
No sales call required, even your agent can complete the purchase over MCP.