DeepSeek V3 0324
DeepSeekfamily · DeepSeek-V3huggingface.co/deepseek-ai/DeepSeek-V3-0324
DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...
Updated 3 h ago · first seen 11 Sept 2026
model_01M294WWHTCRTKN8AHRV7H4SBP
Overview
Identity
Identity block not returned by the API for this entity.
Openness
Openness not classified yet — no sourced evidence to place this model in the ontology.
Key facts
- Release date
Source:OpenRouter public model & pricing listingT2observed 10 h agomedium
- Status
Source:DeepSeek — site & API docsT2observed 11 h agomediumLLM-extracted
- Version
Source:DeepSeek — site & API docsT2observed 11 h agomediumLLM-extracted
- Knowledge cutoff
Source:OpenRouter public model & pricing listingT2observed 10 h agomedium
- Model card
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 10 h agomedium
- Repository
Source:DeepSeek — site & API docsT2observed 11 h agomediumLLM-extracted
- Openrouter id
Source:OpenRouter public model & pricing listingT2observed 10 h agomedium
Architecture
- Architecture
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 10 h agomedium
- Model type
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 10 h agomedium
- Parameters
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 10 h agomedium
- Tokenizer
Source:OpenRouter public model & pricing listingT2observed 10 h agomedium
- Weights dtype
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 10 h agomedium
- File size
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 10 h agomedium
- Library name
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 10 h agomedium
- Pipeline tag
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 10 h agomedium
- Hugging Face repo
Source:OpenRouter public model & pricing listingT2observed 10 h agomedium
Capabilities
Modalities
- Modalities
- text
- Input
- text
- Output
- text
Capabilities
Tool calling
Yes
OpenRouter public model & pricing listing · T2
Structured output
Yes
OpenRouter public model & pricing listing · T2
Reasoning
Yes
DeepSeek — site & API docs · T2
Vision
Unavailable
Audio
Unavailable
Fine-tuning available
Unavailable
- Context window
Source:OpenRouter public model & pricing listingT2observed 7 h agomedium
- Max output
Source:OpenRouter public model & pricing listingT2observed 10 h agomedium
- Knowledge cutoff
Source:OpenRouter public model & pricing listingT2observed 10 h agomedium
- Tokenizer
Source:OpenRouter public model & pricing listingT2observed 10 h agomedium
Benchmarks12
Compare with another model →Comparable same task and conditions · Partially comparable same task, conditions differ (effort, temperature, judge) · Not comparable different variant or metric
No benchmark results recorded
Providers & Pricing1
All offers in the price terminal →USD per 1M tokens as published by each provider (USD). Rows are append-only: every change is kept in the history below.
Price history
Output price · USD / 1M tokens 1 provider
- DeepSeek API
- DeepSeek APIfirst observed $111 Sept 2026
Input price · USD / 1M tokens 1 provider
- DeepSeek API
- DeepSeek APIfirst observed $0.2511 Sept 2026
Hardware fit37
Assumptions (6)
- Estimated, not measured: weights = parameters × bytes/param × 1.15 runtime overhead.
- bytes/param: 4bit = 0.5, 8bit = 1.0, fp16 = 2.0 (uniform quantization, no per-layer exceptions).
- KV cache approximated at 0.5 GB per 8 192 tokens of context, independent of architecture (GQA/MLA models need less).
- A model 'fits' when the estimate is at most the device memory minus 2 GB reserved for the OS and framework.
- Mixture-of-experts models are estimated on total parameters (all experts must be resident); active parameters are ignored.
- Device memory uses the largest configuration when several are listed (e.g. Apple silicon tiers).
Papers1
- arXiv:2412.19437Active35
Timeline20
Full timeline →DeepSeek V3 0324: context length changed from 128000 to 163840
Context window128K tokens→163.8K tokensopenrouterDeepSeek V3 0324 scores 42% on SWE-bench Verified
swebench_leaderboardDeepSeek V3 0324 scores 15.15% on Terminal-Bench
artificial_analysisDeepSeek V3 0324 scores 13.86% on Terminal-Bench
artificial_analysisDeepSeek V3 0324 scores 0% on Terminal-Bench
artificial_analysisDeepSeek V3 0324 scores 47.08% on τ²-bench
artificial_analysisDeepSeek V3 0324 scores 41.02% on IFBench
artificial_analysisDeepSeek V3 0324 scores 39% on SciCode
artificial_analysisDeepSeek V3 0324 scores 4.74% on Humanity's Last Exam
artificial_analysisDeepSeek V3 0324 scores 65.45% on GPQA
artificial_analysisDeepSeek V3 0324 scores 9.72 on Artificial Analysis Intelligence Index
artificial_analysisDeepSeek V3 0324: context length changed from 163840 to 128000
Context window163.8K tokens→128K tokensartificial_analysisDeepSeek V3 0324 scores 99.6% on Aider polyglot
aider_leaderboardDeepSeek V3 0324 scores 55.1% on Aider polyglot
aider_leaderboardDeepSeek API lists DeepSeek V3 0324 at $0.25 in / $1 out per 1M tokens
openrouterDeepSeek V3 0324: reasoning changed from false to true
ReasoningNo→YesdeepseekDeepSeek V3 0324: license changed from mit to MIT License
Licensemit→MIT LicensedeepseekDeepSeek V3 0324: openness changed from open-weights to open-source
Opennessopen-weights→open-sourcedeepseekDeepSeek V3 0324: status changed from deprecated to announced
Statusdeprecated→announceddeepseek
Change history
Model typemodel_type1
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
Provenance
Attributed facts
39
Source tiers
T239
Freshest observation
3 h ago
Conflicts
None
Source documents 8
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.
Data quality (68/100) measures how well AI Atlas knows this entity — completeness, primary-source ratio, freshness, conflicts — never how good the model is. Methodology →