Safeguard
Model family · Capabilities

Every task. Which model handles it.

The comprehensive task-to-model mapping for the lineup. Every capability the family ships, mapped to the model that actually performs it, with a per-cell verdict: full support, partial support, not supported, or dispatched to a sibling via the auto-router.

◈ the matrix, scanned · 36 capabilities × 4 models
Lineup recap

Three families, one corpus.

Lion

Inline · sub-100 ms · on-device · the commit-time gut check.

Read the deep dive

Eagle

Wide-angle triage · ranks and clusters candidate paths across the repo.

Read the deep dive

Griffin

Deep reasoning · the hypothesis engine.

Read the deep dive
Master capability matrix

Capability by model.

36 tasks across 7 bands. Marks reflect actual capability rather than badge-collection upgrades.

Full supportPartialNot supportedroutingVia auto-routing
Capability
Zero
z20
Lion
15B
Eagle
70B
Griffin
750B
Detection & triage
Inline sink detection (deserialisation, SSRF-able URL builders, unsafe SQL)
Secret pattern detection (Gitleaks-class)
Sanitiser-quality scoring
Cross-scanner finding dedup
Taint path enumeration (single-package)
Taint path enumeration (cross-package)
Path ranking + clustering
Confidence scoring per candidate path
Reasoning & hypothesis
Exploit-class hypothesis (CWE category mapping)
Exploit-trigger input synthesis
Cross-package taint chain reasoning (≤ 4 hops)
Cross-package taint chain reasoning (≤ 12 hops)
Cross-package taint chain reasoning (> 12 hops)
Multi-finding correlation in a single reasoning pass
Adversarial disproof pass (refute own hypothesis)
Structured reasoning trace output
Remediation
Single-finding fix suggestion
Auto-fix PR with diff
Auto-fix PR with cited reasoning trace
Sanitiser-aware patch synthesis
Multi-service auto-fix campaign
Upstream coordinated-disclosure patch + draft
Eval & gates
PR-time gate decision
Pre-merge policy evaluation
SARIF / CycloneDX / SPDX emit
Eval-harness scoring of candidate patches
Context & scale
Repo-wide reasoning (1k–5k packages)
Portfolio-wide reasoning (multi-repo)
Deployment shape
On-device inference (no network egress)
Shared cloud
Dedicated cluster
VPC-isolated
Air-gapped / sovereign
AI & MCP governance
Sensitive-data egress scanning
Prompt audit-log signing
MCP tool-call inspectionrouting
"Via routing" means the model does not perform the capability itself; the auto-router dispatches to a sibling when that capability is requested.
Per-model quick reference

One card per model.

Lion 15B

Inline gut check at the keystroke.

Best at
  • Sub-100 ms inline sink + sanitiser detection
  • Secret patterns and obvious unsafe primitives
  • On-device, fully offline operation

In the pipeline: Sits in the IDE, CLI, and pre-commit hook.

Deep dive

Eagle 70B

Wide-angle triage across the repo.

Best at
  • Ranking and clustering candidate taint paths
  • Cross-scanner deduplication and confidence scoring
  • Batched full-repo sweeps after CI

In the pipeline: Runs after CI; feeds the auto-router queue.

Deep dive

Griffin 750B

Deep reasoning, the hypothesis engine.

Best at
  • Multi-hop cross-package exploit hypothesis
  • Cited auto-fix PRs with full trace
  • Coordinated-disclosure draft synthesis

In the pipeline: Candidates from Eagle's queue; the survivors-of-disproof tier.

Deep dive
How the auto-router decides

Triage decides what Griffin proves.

Lion checks the commit, Eagle ranks what it finds across the repo, and Griffin reasons about the candidates worth proving.

01
Lion runs at the commit

Inline on the developer machine, sub-100 ms, no network egress. Catches the obvious unsafe primitives before they ever reach a build server.

02
Eagle sweeps the repo

Post-CI batched scan across every package. Ranks and clusters candidate paths and assigns each one a confidence score on the dataflow head.

03
Griffin proves or refutes

Deep reasoning posits an exploit chain and a CWE class, then a second pass tries to refute it. Survivors land in your queue with the reasoning trace.

Below 0.4, Eagle's verdict ships as-is and no Griffin pass is requested.

Honest about scope

What's NOT on this matrix.

The lineup is weighted for cybersecurity. It deliberately doesn't try to do the things below — a different model class would be the right tool.

Out of scope by design

  • Zero is on the matrix and marked — throughout: it runs no model, so none of these inference capabilities apply to it. What it does instead is turn a question into a structured query, deterministically and with no tokens.
  • General-purpose code generation outside a security context.
  • Image, audio, or video generation of any kind.
  • Customer-support or open-domain chat.
  • Translation, summarisation of marketing copy, or content rewriting.
  • Autocomplete for unrelated business logic.
  • Code review for style, lint, or developer-experience concerns.
  • Synthesising training data for unrelated downstream models.
  • Hardware fault diagnosis or non-software risk analysis.

The corpus is curated to defenders, taint graphs, CVE bodies, and patch diffs. We'd rather be excellent at a small set of security tasks than mediocre at a large set of unrelated ones.

Self-healing security runs on Safeguard.

Your first fix PR is minutes away.

No sales call required, even your agent can complete the purchase over MCP.