llm-security
Safeguard articles tagged "llm-security" — guides, analysis, and best practices for software supply chain and application security.
120 articles
What is Training Data Poisoning
Training data poisoning corrupts an ML model's training data to plant hidden backdoors. Learn how it works, real incidents, and how to detect it.
What is a Model Inversion Attack
Model inversion attacks reconstruct sensitive training data from a model's outputs. Learn how they work, real cases, and how to defend your ML APIs.
What is Insecure Output Handling in LLMs
Insecure output handling lets LLM-generated text execute code, alter queries, or render unsanitized HTML — a real, exploitable OWASP LLM05:2025 risk.
What is Sensitive Information Disclosure in LLMs
LLM sensitive information disclosure leaks training data, prompts, and secrets through model outputs. Real incidents, causes, and defenses explained.
What is RAG (Retrieval-Augmented Generation) Security
RAG pipelines blend retrieved data with model instructions, creating prompt injection, poisoning, and embedding-leak risks traditional AppSec tools miss.
RAG Pipeline Security Controls in 2026
Retrieval-augmented generation pipelines have become a primary breach vector for LLM products. The controls that contain the risk without breaking the use case.
The OWASP Top 10 for Large Language Model Applications: A Field Guide
A working breakdown of the OWASP Top 10 for Large Language Model Applications — what each risk actually looks like in production and how teams are testing for it.
The Real Security Concerns About AI in 2026
The most grounded concerns about AI are not sci-fi scenarios; they are prompt injection, data leakage, supply chain risk in models, and opaque decisions. Here is how each one actually shows up.
Prompt Injection Defense Architectures in 2026
Prompt injection remains the LLM01 entry on the OWASP LLM Top 10 for a reason. A pragmatic look at the defense architectures that hold up in production this year.
LLM Output Filtering as a Security Control
Output filters are the last line before the user and the tool call. We cover when they work, when they fail, and how to measure them honestly in production.
Artificial Intelligence and Cyber Security: Risks and Benefits
A balanced look at what artificial intelligence and cyber security actually means in practice today, the concrete benefits teams are seeing and the new risks AI introduces into the same systems.
AI Agent Tool-Scope Enforcement Patterns
Agents get tool lists, not tool boundaries. We walk through scoping patterns that actually hold when Claude 4 or GPT-5 picks the wrong function at runtime.
Self-healing security runs on Safeguard.
Your first fix PR is minutes away.
No sales call required, even your agent can complete the purchase over MCP.