Models
69 models from 11 providers, closed and open, beside Safeguard's own Griffin, Eagle and Lion. Every model is available via Safeguard at the provider's list price.
No markup: provider list price. No markup on model usage: you pay the provider's list price, last verified . Safeguard enables models for each customer: pick one and choose Use with Safeguard. MCP tool calls are never metered.
69 models
Safeguard security model for discovery and triage: malicious package and typosquat classification, path ranking.
by 1M context70B parameters$0.0001 / 1M input tokens$0.0001 / 1M output tokensPublic cloudPrivate cloudOn-premAir-gapped245 countriesSafeguard security model for remediation and reasoning: reachability, patch authoring and fix pull requests.
by 1.1M context750B parameters$0.0001 / 1M input tokens$0.0001 / 1M output tokensPublic cloudPrivate cloudOn-premAir-gapped245 countriesSafeguard security model for compliance and narrative work: evidence mapping, audit summaries and policy drafting.
by 128K context15B parameters$0.0001 / 1M input tokens$0.0001 / 1M output tokensPublic cloudPrivate cloudOn-premAir-gapped245 countriesSafeguard: ZeroNo model tokens
Safeguard's no-weights model: it reads a data question into a deterministic, replayable query in process, so nothing leaves the platform and no model tokens are used.
by Not applicableNo weights$0 / 1M input tokens$0 / 1M output tokensPublic cloudPrivate cloudOn-premAir-gapped245 countries
Qwen model specialised for code generation and agentic coding.
by 1M context$1 / 1M input tokens$5 / 1M output tokensPublic cloudPrivate cloud🇨🇳
Mid-tier Qwen model recommended by Alibaba for balanced cost and capability, with tool calling and a 1M-token context window.
by 1M context$0.40 / 1M input tokens$1.60 / 1M output tokensPublic cloudPrivate cloud🇨🇳🇯🇵🇺🇸
Large open-source Qwen3.8 model served on Model Studio, with thinking and non-thinking modes.
by Context not published$2 / 1M input tokens$6 / 1M output tokensPrivate cloudOn-premAir-gapped
Smaller open-source Qwen3.8 model served on Model Studio, with thinking and non-thinking modes.
by Context not published$0.50 / 1M input tokens$3 / 1M output tokensPrivate cloudOn-premAir-gapped
Lower-cost Qwen model with the same 1M-token context as Max and Plus.
by 1M context$0.15 / 1M input tokens$0.47 / 1M output tokensPublic cloudPrivate cloud🇨🇳🇺🇸
Alibaba's top Qwen model for complex reasoning, with thinking and non-thinking modes and image and video understanding.
by 1M context$2 / 1M input tokens$6 / 1M output tokensPublic cloudPrivate cloud🇨🇳🇺🇸
Alibaba's text embedding model for search, RAG and clustering, with configurable dimensions from 64 to 2048.
by 8K context$0.07 / 1M input tokensOutput not billed per tokenPublic cloudPrivate cloud🇨🇳🇭🇰Fastest current Claude model, for high-volume tasks such as classification, routing, extraction and subagents.
by 1M context$0.10 / 1M input tokens$0.50 / 1M output tokensPublic cloudPrivate cloud🇦🇺EU🇯🇵🇺🇸Balanced Claude model trading some capability for lower latency and price.
by 1M context$2 / 1M input tokens$10 / 1M output tokensPublic cloudPrivate cloudEU🇺🇸Anthropic's recommended starting model for long-running agentic coding and knowledge work.
by 1M context$4 / 1M input tokens$20 / 1M output tokensPublic cloudPrivate cloud🇦🇺EU🇬🇧🇯🇵🇺🇸Anthropic's model for demanding reasoning and long-horizon agentic work, Anthropic suggests it when Opus 5.5 at higher effort still falls short.
by 1M context$10 / 1M input tokens$50 / 1M output tokensPublic cloudPrivate cloud🇺🇸The same model as Claude Fable 5.1, offered by Anthropic only to organizations verified through its verification programs, such as the Cyber Verification Program.
by 1M context$10 / 1M input tokens$50 / 1M output tokensPublic cloudPrivate cloud🇺🇸Cohere's first Mixture-of-Experts model (218B total, 25B active) with text and image input, reasoning, agentic and translation capabilities across 48 languages. Weights released under Apache 2.0.
by 128K contextPrice on requestPrivate cloudOn-premAir-gappedCohere's small, fast Command model for RAG, tool use and agent tasks. Weights are available on Hugging Face.
by 128K context$0.0375 / 1M input tokens$0.15 / 1M output tokensPrivate cloudOn-premAir-gappedMid-sized Command model balancing efficiency and quality for RAG and tool use.
by 128K context$0.15 / 1M input tokens$0.60 / 1M output tokensPublic cloudPrivate cloudLighter, lower-latency version of Embed 5 sharing the same embedding space as Pro.
by 128K context$0.08 / 1M input tokensOutput not billed per tokenPublic cloudPrivate cloudCohere's highest-quality embedding model for text, images and mixed documents such as PDFs, in 100+ languages.
by 128K context$0.12 / 1M input tokensOutput not billed per tokenPublic cloudPrivate cloudLower-latency version of Rerank 4 for high-throughput search.
by 33K contextPrice on requestPublic cloudPrivate cloud🇧🇷🇩🇪🇬🇧🇮🇳🇯🇵🇸🇦+1Reranking model that reorders search results by relevance for search and RAG, with multilingual and JSON support.
by 33K contextPrice on requestPublic cloudPrivate cloud🇧🇷🇩🇪🇬🇧🇮🇳🇯🇵🇺🇸DeepSeek's current Flash model (DeepSeek-V4.1-Flash), a 552B-parameter MoE with native image understanding and thinking and non-thinking modes. Weights are published on Hugging Face.
by 1M context$0.30 / 1M input tokens$1.20 / 1M output tokensPrivate cloudOn-premAir-gappedLarger DeepSeek V4 model (DeepSeek-V4-Pro-0813) with thinking and non-thinking modes, text only. The V4 family was released with open weights.
by 1M context$1.32 / 1M input tokens$3.96 / 1M output tokensPrivate cloudOn-premAir-gappedImage generation and conversational editing model, an update to Gemini 3.1 Flash Image, with 1K, 2K and 4K output.
by 131K context$1.50 / 1M input tokens$30 / 1M output tokensPublic cloudPrivate cloudGoogle's newest Flash model, aimed at long-horizon software engineering, agents and enterprise workflows. Accepts text, image, video, audio and PDF input.
by 1.05M context$0.75 / 1M input tokens$3.75 / 1M output tokensPublic cloudPrivate cloudEU🇺🇸Low-cost Gemini model for high-volume agentic tasks, translation and simple data processing.
by 1.05M context$0.30 / 1M input tokens$2.50 / 1M output tokensPublic cloudPrivate cloudEU🇺🇸Google open-weight multimodal model (30.7B parameters, text and image in, text out), Apache 2.0 licence.
by 256K contextPrice on requestPrivate cloudOn-premAir-gappedEarlier GA Flash model for routine, high-throughput workloads with multimodal input.
by 1.05M context$1.50 / 1M input tokens$9 / 1M output tokensPublic cloudPrivate cloud🇦🇺🇨🇦🇩🇪EU🇬🇧🇮🇳+3Multimodal embedding model that maps text, images, video, audio and PDFs into one embedding space.
by 8K context$0.20 / 1M input tokensOutput not billed per tokenPublic cloudPrivate cloudEU🇺🇸Google's current Pro model, offered as a preview, for multimodal understanding and agentic coding.
by 1.05M context$2 / 1M input tokens$12 / 1M output tokensPublic cloudPrivate cloudOpen-weight Llama 4 MoE model (17B active, 400B total, 128 experts) with text and image input.
by 1M contextPrice on requestPrivate cloudOn-premAir-gappedOpen-weight Llama 4 MoE model (17B active, 109B total, 16 experts) with text and image input and a 10M-token context.
by 10M contextPrice on requestPrivate cloudOn-premAir-gappedText-only 70B instruction-tuned Llama model for multilingual chat and coding.
by 128K contextPrice on requestPrivate cloudOn-premAir-gappedMeta's 30B open-weight multimodal model distilled from Muse Spark, released under Apache 2.0 for self-hosting.
by 128K contextPrice on requestPrivate cloudOn-premAir-gappedMeta's image generation and editing model on Meta Model API, with built-in web and image search.
by Context not publishedPrice on requestPublic cloudPrivate cloudMeta's hosted model on Meta Model API for agentic and coding work, with text, image, video and PDF input.
by 1.05M context$1.25 / 1M input tokens$4.25 / 1M output tokensPublic cloudPrivate cloudOpen-weight multimodal Mixture-of-Experts model with 52B active and 1.05T total parameters, currently in public preview.
by 1M context$1.36 / 1M input tokens$4.18 / 1M output tokensPrivate cloudOn-premAir-gappedMultimodal model tuned for agentic and coding work, released as open weights under a modified MIT license.
by 256K context$1.50 / 1M input tokens$7.50 / 1M output tokensPrivate cloudOn-premAir-gappedHybrid model combining instruct, reasoning and coding in one model, with 119B total and 6.5B active parameters, Apache 2.0.
by 256K context$0.15 / 1M input tokens$0.60 / 1M output tokensPrivate cloudOn-premAir-gappedAudio model trained for speech transcription.
by Context not publishedPrice on requestPublic cloudPrivate cloudLargest model in the Ministral 3 family, with text and vision support and a focus on local deployment.
by 256K context$0.20 / 1M input tokens$0.20 / 1M output tokensPrivate cloudOn-premAir-gappedOpen-weight multimodal Mixture-of-Experts model with 41B active and 675B total parameters, released under Apache 2.0.
by 256K context$0.50 / 1M input tokens$1.50 / 1M output tokensPrivate cloudOn-premAir-gappedCode model for low-latency tasks such as fill-in-the-middle completion and code generation.
by 128K context$0.30 / 1M input tokens$0.90 / 1M output tokensPublic cloudPrivate cloudEmbedding model built for representing code for search and retrieval.
by 8K context$0.15 / 1M input tokensOutput not billed per tokenPublic cloudPrivate cloudMid-tier GPT-6 model for complex coding, computer use and professional work at a lower price than Astra.
by 1.05M context$2 / 1M input tokens$10 / 1M output tokensPublic cloudPrivate cloudAPACEU🇺🇸Low-cost GPT-6 model for focused, high-volume tasks.
by 1.05M context$0.10 / 1M input tokens$0.50 / 1M output tokensPublic cloudPrivate cloudEU🇺🇸Earlier Sol model in the GPT-6 family for coding and agentic workflows, GPT-6.1 Sol is the newer version.
by 1.05M context$2 / 1M input tokens$10 / 1M output tokensPublic cloudPrivate cloud🇦🇪EU🇺🇸Image generation and editing model that takes text and image inputs and returns images.
by Context not published$5 / 1M input tokens$30 / 1M output tokensPublic cloudPrivate cloudOpenAI's top model in the GPT-6 family, aimed at hard reasoning, coding, computer use and research tasks. Supports configurable reasoning effort up to max.
by 1.05M context$10 / 1M input tokens$50 / 1M output tokensPublic cloudPrivate cloud🇺🇸Speech-to-text model for audio files and committed turns in Realtime sessions, with support for context and keyword hints.
by Context not publishedPrice on requestPublic cloudPrivate cloudOpenAI open-weight reasoning model with 117B parameters (5.1B active), Apache 2.0 licence.
by Context not publishedPrice on requestPrivate cloudOn-premAir-gappedOpenAI's larger text embedding model for search, clustering and classification in English and other languages.
by Context not published$0.13 / 1M input tokensOutput not billed per tokenPublic cloudPrivate cloud🇦🇪🇦🇺🇨🇦🇨🇭🇩🇪🇪🇸+12Smaller, cheaper OpenAI text embedding model.
by Context not published$0.02 / 1M input tokensOutput not billed per tokenPublic cloudPrivate cloud🇦🇪🇦🇺🇨🇦🇨🇭EU🇯🇵+1Sarvam AI's latest speech recognition model and the API default: transcribes, translates to English, writes verbatim, transliterates or code-mixes speech in 22 Indian languages and English, with Indian and global English accents and key-term prompting.
by Context not publishedPrice on requestPublic cloud🇮🇳Sarvam AI's flagship reasoning model, trained from scratch: a Mixture-of-Experts with 105B total and 10.3B active parameters and Multi-head Latent Attention, built for Indian languages and English, coding, maths and agentic tool use, released as open weights under Apache 2.0.
by 131K context105B parametersPrice on requestPublic cloudPrivate cloudOn-premAir-gapped🇮🇳Sarvam AI's speech recognition model, available beside v4 with the same five output modes (transcribe, translate to English, verbatim, transliterate, code-mix) for 22 Indian languages and Indian English, with real-time streaming, and offered for self-hosting on Amazon SageMaker.
by Context not publishedPrice on requestPublic cloudPrivate cloud🇮🇳Sarvam AI's stable text-to-speech model and the API default: more than 30 named voices for 10 Indian languages and English, pace control, sample rates up to 48 kHz and up to 2,500 characters per request, also offered for self-hosting on Amazon SageMaker.
by Context not publishedPrice on requestPublic cloudPrivate cloud🇮🇳Sarvam AI's document intelligence model, a 3B parameter state-space vision language model for OCR, tables and layout in 22 Indian languages and English, served through the Document AI Digitise and Extract APIs and offered for self-hosting on Amazon SageMaker. Version 2.1 (September 2026) adds form key-value extraction and handwriting in Indian languages.
by Context not published3B parametersPrice on requestPublic cloudPrivate cloud🇮🇳Sarvam AI's open-weights translation model for English and all 22 scheduled Indian languages, built with AI4Bharat on Gemma 3 4B IT, for formal, document-level text up to 8K tokens. The open checkpoint translates only between English and an Indian language, in either direction. GPL 3.0 licence.
by 8K contextPrice on requestPublic cloudPrivate cloudOn-premAir-gapped🇮🇳Sarvam AI's latest text-to-speech model: more than 200 persona voices by language and style in Hindi, English, Bengali, Gujarati, Kannada, Marathi, Punjabi, Tamil and Telugu, with the same API as Bulbul v3, pace control and up to 2,500 characters per request.
by Context not publishedPrice on requestPublic cloud🇮🇳Sarvam AI's translation model for English and 10 Indian languages: formal, colloquial and code-mixed styles, Roman, native or spoken output script, numeral format control, automatic source language detection and up to 1,000 characters per request.
by Context not publishedPrice on requestPublic cloud🇮🇳
xAI's flagship model for coding, agentic tasks and knowledge work, with configurable reasoning effort.
by 500K context$2 / 1M input tokens$6 / 1M output tokensPublic cloudPrivate cloud🇺🇸
Previous xAI frontier model for coding, agentic tasks and knowledge work.
by 500K context$2 / 1M input tokens$6 / 1M output tokensPublic cloudPrivate cloud🇺🇸
xAI coding model for agentic software engineering, replaces grok-code-fast-1.
by 256K context$1 / 1M input tokens$2 / 1M output tokensPublic cloudPrivate cloud
Grok 4.20 reasoning variant with a 1M-token context, focused on speed and agentic tool calling.
by 1M context$1.25 / 1M input tokens$2.50 / 1M output tokensPublic cloudPrivate cloud
Lower-priced Grok model with a 1M-token context, strong tool calling and reasoning effort from none to xhigh.
by 1M context$1.25 / 1M input tokens$2.50 / 1M output tokensPublic cloudPrivate cloud🇺🇸
xAI image generation and editing model.
by Context not publishedPrice on requestPublic cloudPrivate cloud
Safeguard model family
Griffin, Eagle, Lion and Zero: how our security models are built, benchmarked and deployed.
Explore the familyModel directory on Safeguard Gold
The free directory of open and frontier models, with licences and commercial-use guidance.
Open the directorySelf-healing security runs on Safeguard.
Your first fix PR is minutes away.
No sales call required, even your agent can complete the purchase over MCP.