DeepSeek
DeepSeek V4.1 Flash
deepseek/deepseek-flash
Request access
Not served yet. Ask for access and we will tell you when it is live.
DeepSeek's current Flash model (DeepSeek-V4.1-Flash), a 552B-parameter MoE with native image understanding and thinking and non-thinking modes. Weights are published on Hugging Face.
Price
- Input
- $0.3 per 1M tokens
- Output
- $1.2 per 1M tokens
- Notes
- Peak-hour list price (cache miss). Off-peak is half: $0.15 in / $0.60 out. Peak hours 01:00-04:00 and 06:00-10:00 UTC Mon-Fri excluding Chinese public holidays. Cache-hit input $0.006 peak.
- Source
- DeepSeek pricing , last verified October 8, 2026
No markup on model usage: you pay the provider's list price.
Model
- Context
- 1M (1,000,000 tokens)
- Max output
- 384,000 tokens
- Modality
- Text and image
- Input
- text, image
- Output
- text
- Weights
- Open weights
- Released
- September 10, 2026
- Tags
- open-weights, reasoning, coding, vision, long-context, low-cost
Data policy
- Prompts go to
- DeepSeek (People's Republic of China)
- Retention
- Privacy policy: personal data stored in the PRC and kept as long as necessary to provide the service, may be used to train models, with a right to opt out. No API-specific retention period or ZDR offering published
- Zero data retention
- Not stated by the provider
- Region
- China (PRC)
- Policy
- DeepSeek data policy
Summarised from the provider's published terms. The provider's own policy is what applies.
Call it with any OpenAI SDK
Available when the model is live
from openai import OpenAI
client = OpenAI(
base_url="https://api.safeguard.sh/v1",
api_key="SAFEGUARD_API_KEY",
)
response = client.chat.completions.create(
model="deepseek/deepseek-flash",
messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)Base URL https://api.safeguard.sh/v1, model deepseek/deepseek-flash.
Self-healing security runs on Safeguard.
Your first fix PR is minutes away.
No sales call required, even your agent can complete the purchase over MCP.