Safeguard
Model family · Griffin

Griffin. The hypothesis engine.

Griffin is the heavyweight reasoning model, weighted purely on a cybersecurity corpus. It hypothesises exploit chains, cites the call-graph path, attempts a disproof against the project's sanitiser config, and writes the patch.

✦ the watch · the beam sweeps, the hypothesis lands
Specs

Griffin at a glance.

Parameters, context window, price per 1M tokens and where it runs, as the model catalog lists them.

Parameters
750B
Context window
1.1M
Input per 1M tokens
$0.0001
Output per 1M tokens
$0.0001
RunsPublic cloudPrivate cloudOn-premAir-gapped245 countries

Early access: Griffin is not callable through the API yet. This request works once Safeguard opens it and switches the gateway on for your workspace.

See Griffin in the model catalog
Architecture

The internals that earn the verdict.

01

Security-augmented tokeniser with ~28k extra tokens covering CWE / CVE IDs, taint operators, package coordinates, and attack-pattern shorthand.

02

Sliding-window plus landmark attention for long-context call-graph reasoning.

03

Structured reasoning trace: hypothesise the exploit, cite the path, propose a disproof, propose a patch.

04

Security-domain RLHF using preference data labelled by senior offensive-security engineers, not generic annotation vendors.

Eval highlights

Measured against known ground truth.

81%
Exploit-hypothesis accuracy
98%
Adversarial prompt resistance
0.6%
Hallucination rate on security Q&A
94%
Top-5 candidate path retention vs CVE ground truth
Reasoning trace

The trace is the finding.

Every Griffin call emits a four-stage trace. Reviewers see the chain, not a single label, and can reject at any stage.

griffin · finding #4129structured trace
[01] HYPOTHESIS
class: CWE-502 (unsafe deserialization)
entry: HTTP POST /api/import-config
gadget: pkg:maven/com.fasterxml.jackson.core/jackson-databind@2.9.10
[02] CITED PATH
handler.parseRequest() -> service.importConfig()
-> codec.decode(bytes) -> ObjectMapper.readValue(InputStream, Object.class)
6 hops, 3 package boundaries, 1 sanitiser bypassed (allow-list mismatch).
[03] DISPROOF ATTEMPT
- polymorphic typing disabled? no (DefaultTyping.NON_FINAL active)
- allow-list enforced? partial; missing on nested key 'plugins'
- sandbox or seccomp profile? none on this code path
refutation failed; finding stands.
[04] PROPOSED PATCH
- replace ObjectMapper.readValue with constrained reader
using ALLOWED_TYPES allow-list
- bump jackson-databind to >= 2.15.2 (advisory-aligned)
- add SecurityManager-equivalent unit test covering nested 'plugins'.
Triage first

Eagle ranks, Griffin reasons.

A triage score from Eagle decides which candidates reach Griffin, so its reasoning budget goes to the chains that need proving.

01 · Triage score

Eagle assigns a complexity score from the call graph: depth, sanitiser ambiguity, cross-package edges, sink severity.

02 · Reasoning pass

Griffin runs the hypothesise / cite / disprove / patch trace. The trace ships with the finding so reviewers can audit how it was reached.

Development history

How Griffin got to where it is.

Three years, one corpus discipline. Each release earned its slot against the eval set, not a roadmap deadline.

2023

Internal prototype, "Aegis-0".

Started as a research prototype to test whether a transformer trained narrowly on CVE descriptions, exploit write-ups, and patched diffs would outperform a general-purpose model on three tasks: CWE classification, taint-path hypothesis generation, and patch suggestion. Initial parameter count under 1B. Eval against an internal held-out set of 4,200 disclosed CVEs showed a 28-point F1 lift over an off-the-shelf code model — enough to justify scaling.

Q2 2024

First production-grade model.

Scaled the architecture, introduced the security-augmented tokeniser (~28k additional tokens for CWE/CVE IDs, taint operators, package coordinates). Trained on a curated corpus that grew from 1.2M to 11M security-domain documents. Adversarial prompt-injection eval rate dropped from 42% to 6%.

Q4 2024

The structured trace contract.

Introduced the structured reasoning trace as a first-class output (HYPOTHESIS / CITED PATH / DISPROOF / PROPOSED PATCH). Eval methodology shifted from "did the model find the bug" to "did the model find the bug and refute its own hypothesis under sanitiser-aware constraints." This is what later became the disproof pass.

Q2 2025

Long-context attention.

Long-context attention added: sliding-window plus landmark, taking usable context from 32k to 128k. Distillation experiments started in parallel: early Lion prototypes derived from Griffin's reasoning traces.

Now

Current research direction.

Three concurrent research tracks: (1) on-device distillation of larger reasoning traces into Lion-class students, with the goal of pushing more reasoning depth into the IDE without breaking the sub-100ms latency budget; (2) adversarial training against prompt-injection attacks observed in real MCP-server traffic; (3) longer-horizon agentic workflows for coordinated disclosure, where Griffin proposes upstream patches, runs them through the maintainer's project test suite, and drafts the disclosure thread.

Release pipeline

How a Griffin release actually ships.

Every Griffin release moves through the same six gates. Each one can block the release on its own.

01 · Curation pass

Corpus is filtered against the security-only criteria (no general web crawl, no LLM-generated text, no customer code), deduplicated, and the held-out eval set is rotated.

02 · Pretraining + security RLHF

Base pretraining on the curated corpus, followed by RLHF where the preference data is labelled by senior offensive-security engineers, not crowdworkers.

03 · Adversarial red team

Internal red team runs prompt-injection, jailbreak, and refusal-rate suites against every checkpoint. A checkpoint that regresses the adversarial scores does not ship, regardless of capability gains.

04 · Eval gate + cited-trace audit

Quantitative eval against the held-out CVE / taint-path / patch suites, plus a manual audit of 300 reasoning traces by the engineering team. Trace-quality regressions block release the same way capability regressions do.

05 · Staged rollout

A release ships to the shared-cloud tier first (smallest blast radius), then dedicated cluster and VPC-isolated, then sovereign. Each tier has a 14-day soak period.

06 · Post-release telemetry

Anonymised, aggregated usage and trace-quality metrics feed back into the next curation pass. Customer code never enters the loop.

Self-healing security runs on Safeguard.

Your first fix PR is minutes away.

No sales call required, even your agent can complete the purchase over MCP.