Benchmarks
How AI models score on each benchmark, with the quality, value and speed leaders and the source and date of every number.
No benchmark results yet
Results appear here once the first benchmark data is published. Nothing on this page is estimated.
Coming soon: our own security evaluations
Safeguard Security Benchmarks will test secure code generation, vulnerability detection and prompt-injection resistance, graded by deterministic checks rather than by another model. Results appear here after the first published run.
How these benchmarks are measured
Each card is one benchmark. We show what each source published, with its date and licence, and link a model to our catalog when we list it.
- Quality
- The top score on the benchmark, in the benchmark's own unit.
- Value
- The model with the smallest measured run cost. Shown only where the source reports what a run cost.
- Speed
- The model with the shortest measured latency. Shown only where the source reports latency.
- Task evaluations
- One model scored on one benchmark. The header adds these up across every benchmark shown.
- Safeguard Security Benchmarks
- Our own evaluations. Each model is called through Safeguard's gateway the way a customer calls it, and every answer is graded by a deterministic check: static analysis for insecure code, a label match for vulnerability detection and a canary string for prompt injection. No model grades another model's answers.
- OpenRouter, Artificial Analysis and Design Arena
- Scores published through OpenRouter's Data API under CC BY 4.0: OpenRouter's own evaluations, the Artificial Analysis indexes and Design Arena ratings. We show them as published and credit each source.
Self-healing security runs on Safeguard.
Your first fix PR is minutes away.
No sales call required, even your agent can complete the purchase over MCP.