Safeguard
Company · SSL

Safeguard Superintelligence Lab. Superintelligence, built for defenders.

SSL is Safeguard's research lab for AI that understands software more deeply than any human team: systems that find the flaw before it is exploited, fix it with proof, and stay safely under the control of the people who run them. Humanity comes first: if this work ever points toward human mass extinction, we shut everything down immediately.

Join SSL
◈ the drill — every capability earns its release through evaluation, red-teaming and a written safety case
Humanity First
Human survival outranks everything
Security-first
Every system aimed at defence
Open
Methods and evaluations published
Gated
Capability released only under the RSP
Humanity First

Human survival outranks everything we build.

SSL is Humanity First. The lab exists to protect people, and no capability, deadline or commercial goal outranks human survival. If our work ever points toward human mass extinction, we shut everything down immediately.

  • If any SSL system, experiment or result shows credible signs of contributing to human mass extinction or catastrophic, irreversible harm to humanity, all related training, deployment and research stops at once.
  • The shutdown is immediate and unconditional. It does not wait for a product decision, a customer commitment or a funding conversation.
  • Any member of the lab can trigger the stop. Restarting requires an independent review and a written safety case, never the stop being quietly lifted.
  • We would rather lose a capability, a product or the lab itself than build something that endangers humanity.
Mission

Why a superintelligence lab.

The Safeguard Superintelligence Lab exists to build AI systems that reason about software better than any human team can, and to make sure that capability works for the people defending software rather than the people attacking it.

Software now ships faster than anyone can read it. Attackers already use AI to find weaknesses at machine speed. Defence has to reach the same speed and go past it: systems that understand an entire codebase, its dependencies and its build, find the flaw before it is exploited, and fix it without breaking anything.

SSL is where Safeguard does the long-horizon work toward that goal. What we learn feeds the Griffin model family and the platform, and what we publish is written so that other defenders can use it too.

Scale

From one country to the edge of the universe.

Zoom out from the United States to Earth, the Solar System, the Milky Way, Andromeda and past the observable universe. Every ring is drawn to its true size. This is the future a superintelligence could help humanity reach, and why it has to be built with people first.

UNITED STATESview ≈ 5,870 km wideeach ring drawn to true scale
01 / 12

United States

~4,500 km coast to coast

Light crosses it in 15 milliseconds

Where Safeguard started. A continent of software, networks and people depending on code nobody has fully read.

The Kardashev scale of civilisations.

Sagan's form: K = (log₁₀ P − 6) / 10
Humanity ≈ 0.73
I
II
III

Humanity uses roughly 19 terawatts today, which puts us at about 0.73: not yet Type I. Every step up the scale is more power, and more that can go wrong. Humanity First means the step is only worth taking if people survive it.

Research

What SSL works on.

Autonomous security reasoning

Models that hold a whole codebase, its dependency graph and its build in mind at once, and reason across them the way a senior security engineer would, only exhaustively.

Zero-day discovery at scale

Finding vulnerability classes before they have a CVE: reachability, data-flow and exploitability reasoning that goes deeper than pattern matching or signature lookup.

Verified self-repair

Fixes that come with evidence: a patch, the test that proves the flaw is gone, and the check that proves nothing else broke. A fix nobody can verify is not a fix.

Alignment and control

Keeping highly capable systems inside the bounds their operators set: structured traces, refusal boundaries, tool-use limits, and oversight that still works as capability grows.

Adversarial robustness

Prompt injection, poisoned training data, tampered weights and hostile tool output. A security model is itself a target, so we attack our own systems before anyone else does.

Evaluation science

Benchmarks that measure what matters to a defender, resist contamination and gaming, and tell us when a system is ready for release and when it is not.

Method

From problem to defenders, through every gate.

  1. 01

    Start from a real defensive problem

    Every project begins with a failure defenders actually face today: a class of bug that slips through, a fix that takes weeks, a supply chain signal nobody reads.

  2. 02

    Build the evaluation first

    Before training anything, we decide how success will be measured and how the measurement could be fooled. The evaluation is reviewed separately from the work it judges.

  3. 03

    Train, then attack it

    Every candidate system goes through internal red-teaming aimed at both its security task and its own safety: can it be turned, misled or made to overreach.

  4. 04

    Write the safety case

    Nothing leaves the lab without a written case for why it is safe to release at its level under the Responsible Scaling Policy, signed off by people who did not build it.

  5. 05

    Publish and ship

    What passes goes into the Griffin models and the platform. Methods, evaluations and negative results are published so the wider security community can check and reuse them.

Principles

How the lab holds itself to account.

SSL works under Safeguard's Responsible Scaling Policy. These are the commitments on top of it.

Humanity first

Human survival and wellbeing come before every capability and every commercial goal. If our work points toward human mass extinction, we shut everything down immediately.

Defence first

SSL builds capability for defenders. We do not release offensive tooling, and we decide what to publish with the risk of misuse in view.

Capability follows safety

A system moves to a higher capability level only when its safeguards have been shown to hold at that level. Scaling waits for the safety case, not the other way round.

Human oversight stays meaningful

Operators must be able to see what a system did and why, and to stop it. We treat oversight that only works on weaker systems as a research problem, not a solved one.

Show the work

Claims about capability come with the evaluation behind them. Where we are uncertain, we say so. Negative results are published alongside the positive ones.

Work with SSL.

Research partnerships, design-partner access to early systems, and roles on the lab. Email research@safeguard.sh.

Self-healing security runs on Safeguard.

Your first fix PR is minutes away.

No sales call required, even your agent can complete the purchase over MCP.