Safeguard
Topic

AI Security

In-depth guides and analysis on ai security from the Safeguard engineering team.

100 articles

AI Security

Understanding membership inference attacks against traine...

Membership inference attacks let adversaries confirm if your data trained a model, exposing privacy leakage in ML models and training data inference risks.

Jul 30, 20267 min read
AI Security

What shadow AI is and how to discover unsanctioned AI use...

Shadow AI risk is spreading faster than governance can keep up. Here's what unsanctioned AI use looks like inside real enterprises and how to discover it before data leaks.

Jul 30, 20267 min read
AI Security

Glossary of AI Trust, Risk, and Security Management (AI T...

A glossary of AI trust risk security management concepts: the Gartner AI TRiSM framework, its four pillars, AI risk taxonomy, and adversarial threats.

Jul 30, 20268 min read
AI Security

How to build an AI-specific incident response playbook

A step-by-step guide to building an AI incident response plan — covering scoping, escalation, detection, containment, and post-incident review for LLM and agent failures.

Jul 30, 20268 min read
AI Security

CoSAI Releases Model Signing and Incident Response Frameworks

The Coalition for Secure AI published two operational frameworks in November 2025: Signing ML Artifacts and AI Incident Response. We unpack what each contains and how to adopt them.

Jul 30, 20267 min read
AI Security

Prompt Injection Detection in Retrieval Systems

Indirect prompt injection arrives through your retrieval corpus, not your chat box. We cover the detection strategies that survive when attackers write your RAG content.

Jul 28, 20265 min read
AI Security

The Hugging Face Breach: What Changes When an AI Agent Runs the Intrusion

On 16 July 2026 Hugging Face disclosed that a malicious dataset gave an attacker code execution inside its data-processing pipeline, escalating to node-level access and internal cluster credentials over a single weekend — driven by an autonomous agent framework executing thousands of actions. Here's the anatomy, and what it means for anyone who treats a model registry as a trusted input.

Jul 28, 20266 min read
AI Security

vLLM CVE-2025-62164: Tensor Deserialization RCE

vLLM 0.10.2-0.11.0 deserialized user-supplied PyTorch tensors via torch.load() in the Completions API. Memory corruption, potential RCE.

Jul 26, 20265 min read
AI Security

Claude Opus 4.5 System Card: Defender Takeaways

Anthropic released Claude Opus 4.5 on November 24, 2025 with the most detailed safety section of any system card to date. We pull out what enterprise defenders should change.

Jul 26, 20267 min read
AI Security

AI Model Watermarking and Provenance

Watermarking and provenance are the two most confused terms in AI security. A practical breakdown of what each actually does, where the 2025 techniques break, and what to ship in the meantime.

Jul 25, 20267 min read
AI Security

vLLM CVE-2025-66448: Auto-Map RCE via Model Configs

A critical RCE in vLLM allows malicious model configs to bypass trust_remote_code=False. We analyze the bug, the patch, and what every vLLM operator should do.

Jul 24, 20267 min read
AI Security

Gemini 3 Pro and the Frontier Safety Framework Report

Google released Gemini 3 Pro on November 18, 2025 with the most thorough Frontier Safety Framework evaluation yet. We unpack what was disclosed and how it changes downstream defender posture.

Jul 23, 20267 min read

Self-healing security runs on Safeguard.

Your first fix PR is minutes away.

No sales call required, even your agent can complete the purchase over MCP.