DeepSeek V4 Flash Vision Exp
DeepSeekfamily · DeepSeekhuggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vis
DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...
Updated 3 h ago · first seen 11 Sept 2026
model_01M294AJ4V6A3PXCEM2FCQ5NQV
Overview
Identity
- Canonical model
- Yesidentity confidence: mediumOne row per real model release. Artifacts (checkpoints, quantisations, conversions) and folded evaluation variants point here.
- Official checkpoints
- official_checkpoints = hf_repo identifiers carried by the model itself; artifacts are separate entities pointing here through canonical_id.
- Artifacts
- None recordedSeparate entities (checkpoint · quantization · conversion · packaging) pointing to this model through canonical_id.
- Provider deployments
- 3
- API aliases
- deepseek/deepseek-v4-flash-vision-expfireworks/deepseek-v4-flash-vision-expIdentifiers under which providers and evaluators refer to this model.
- Folded evaluation variants
- 0Effort / thinking variants (…-high, …-non-reasoning) are result configurations of this model, not separate models. Their old URLs redirect here.
Openness
Open weights— weights downloadable under MIT; commercial use allowed; redistribution allowed; derivatives allowed; 4 dimensions unknown.
Weights downloadable under a permissive or Creative Commons licence allowing commercial use; code or data may be missing.
Weights
Yes
Inference code
—
Training code
—
Training data
—
Dataset
—
Commercial use
Yes
Redistribution
Yes
Derivatives
Yes
Licence: MIT License (permissive · SPDX MIT · stated as “mit”)
dimensions marked null are unknown, not false
Key facts
- Release date
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 6 d agomedium
- Status
Source:DeepSeek — site & API docsT2observed 6 d agomediumLLM-extracted
- Model card
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 6 d agomedium
- Openrouter id
Source:OpenRouter public model & pricing listingT2observed 6 d agomedium
Architecture
- Architecture
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 6 d agomedium
- Model type
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 6 d agomedium
- Parameters
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 6 d agomedium
- Tokenizer
Source:OpenRouter public model & pricing listingT2observed 6 d agomedium
- Weights dtype
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 6 d agomedium
- File size
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 6 d agomedium
- Library name
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 6 d agomedium
- Pipeline tag
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 6 d agomedium
- Hugging Face repo
Source:OpenRouter public model & pricing listingT2observed 6 d agomedium
Capabilities
Modalities
- Modalities
- imagetext
- Input
- imagetext
- Output
- text
Capabilities
Tool calling
Yes
OpenRouter public model & pricing listing · T2
Structured output
Yes
OpenRouter public model & pricing listing · T2
Reasoning
Yes
OpenRouter public model & pricing listing · T2
Vision
Yes
OpenRouter public model & pricing listing · T2
Audio
Unavailable
Fine-tuning available
Unavailable
- Context window
Source:OpenRouter public model & pricing listingT2observed 6 d agomedium
- Max output
Source:OpenRouter public model & pricing listingT2observed 7 h agomedium
- Tokenizer
Source:OpenRouter public model & pricing listingT2observed 6 d agomedium
Benchmarks8
Compare with another model →Comparable same task and conditions · Partially comparable same task, conditions differ (effort, temperature, judge) · Not comparable different variant or metric
Current rows only, grouped by benchmark → canonical metric → comparability group (task configuration). Effort variants folded into this model appear as rows of the same group. 8 current rows in total. “vs leader” compares with the current leader of the benchmark's primary group only; other groups are not directly comparable. Comparability rules →
Providers & Pricing5
All offers in the price terminal →USD per 1M tokens as published by each provider; native units (per-request fees, flex/priority tiers) are kept verbatim. Rows are append-only — every price change is kept in the history below. Cost of a workload →
Price history
Output price · USD / 1M tokens 3 providers
- Fireworks AI
- DeepSeek API
- OpenRouter
- OpenRouter$0.33 → $0.64718 Sept 2026
- OpenRouter$0.66 → $0.3312 Sept 2026
- OpenRouterfirst observed $0.6612 Sept 2026
- DeepSeek API$0.66 → $0.3311 Sept 2026
- DeepSeek APIfirst observed $0.6611 Sept 2026
- Fireworks AIfirst observed $0.6611 Sept 2026
Input price · USD / 1M tokens 3 providers
- Fireworks AI
- DeepSeek API
- OpenRouter
- OpenRouter$0.11 → $0.21618 Sept 2026
- OpenRouter$0.22 → $0.1112 Sept 2026
- OpenRouterfirst observed $0.2212 Sept 2026
- DeepSeek API$0.22 → $0.1111 Sept 2026
- DeepSeek APIfirst observed $0.2211 Sept 2026
- Fireworks AIfirst observed $0.2211 Sept 2026
Hardware fit37
Assumptions (7)
- Estimated, not measured: weights = parameters × bytes/param × 1.15 runtime overhead (or the observed artifact file size when one is recorded).
- bytes/param: 4bit = 0.5, 8bit = 1.0, fp16 = 2.0 (uniform quantization, no per-layer exceptions).
- KV cache: 2 × layers × kv_heads × head_dim × 2 bytes × context × batch when the architecture is known; otherwise 0.5 GB per 8 192 tokens (× batch), independent of architecture (GQA/MLA models need less).
- A model 'fits' when the estimate is at most the device memory minus 2 GB reserved for the OS and framework.
- Mixture-of-experts models are estimated on total parameters (all experts must be resident); active parameters are ignored.
- Device memory uses the largest configuration when several are listed (e.g. Apple silicon tiers).
- Multi-GPU: device memories are summed; interconnect bandwidth, tensor-parallel replication and pipeline bubbles are not modelled.
Versions & Artifacts0
Version history
Max output1 change
11 Sept 2026→18 Sept 2026current
Openness1 change
11 Sept 2026→11 Sept 2026
Context windowfirst observation only
11 Sept 2026current
Licensefirst observation only
11 Sept 2026current
Parametersfirst observation only
11 Sept 2026current
Statusfirst observation only
11 Sept 2026current
Each hop is a claim: click a value for its source, tier and observation time. Nothing is overwritten — a new observation closes the previous claim.
Artifacts 0
No artifact (checkpoint, quantisation, conversion or packaging) points to this model yet.
Timeline6
Full timeline →OpenRouter changed pricing for DeepSeek V4 Flash Vision Exp: $0.22 in / $0.66 out per 1M tokens → $0.2156 in / $0.6468 out per 1M tokens
$0.22 in / $0.66 out→$0.216 in / $0.647 outopenrouterDeepSeek V4 Flash Vision Exp: max output tokens changed from 943718 to 262144
Max output943.7K tokens→262.1K tokensopenrouterOpenRouter lists DeepSeek V4 Flash Vision Exp at $0.11 in / $0.33 out per 1M tokens
openrouterOpenRouter lists DeepSeek V4 Flash Vision Exp at $0.22 in / $0.66 out per 1M tokens
openrouterDeepSeek V4 Flash Vision Exp: release date changed from 2026-08-31 to 2026-08-21
Release date31 Aug 2026→21 Aug 2026openrouterDeepSeek V4 Flash Vision Exp: release date changed from 2026-08-21 to 2026-08-31
Release date21 Aug 2026→31 Aug 2026huggingface
Change history89
Viewing AI Atlas as of 17 Sept 2026 — attributes exactly as the atlas knew them on that day; later corrections are not shown.
Back to today →Attributes as of 17 Sept 2026 46 claims in force
- Family
- DeepSeek-V4
- Release date
- 31 Aug 2026
- Status
- preview
- Openness
- open-weights
- License
- MIT
- Architecture
- DeepseekV4ForCausalLM
- Parameters
- 304.6B
- Context window
- 1.05M tokens
- Max output
- 943.7K tokens
- Modalities
- image, text
- Input modalities
- image, text
- Output modalities
- text
- Tokenizer
- DeepSeek
- Quantization format
- fp8
- File size
- 167.8 GB
- Hugging Face repo
- deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
- Pipeline tag
- image-text-to-text
- Downloads
- 668,317
- Likes
- 918
- Access
- open
- Artifact kind
- quantization
- Commercial use allowed
- Yes
- Derivatives allowed
- Yes
- Gated
- No
- Hf inference providers
- deepinfra, fireworks-ai, novita
- Is quantized
- Yes
- Last modified
- 2026-09-01T09:22:10+00:00
- Library name
- transformers
- License key
- MIT
- License raw
- mit
- Livebench hf link
- https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
- Downloads all time
- 668,317
- Model type
- deepseek_v4
- Openrouter id
- deepseek/deepseek-v4-flash-vision-exp
- Openrouter listed at
- 21 Aug 2026
- Reasoning
- Yes
- Redistribution allowed
- Yes
- Structured output
- Yes
- Supported parameters
- frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p
- Tags
- transformers, safetensors, deepseek_v4, text-generation, image-text-to-text, 8-bit, fp8
- Tool calling
- Yes
- Vision
- Yes
- Weights available
- Yes
- Weights dtype
- BF16, F32, F8_E4M3, I64, I8
Familyfamily1
Release daterelease_date6
Statusstatus1
Opennessopenness3
Licenselicense1
Architecturearchitecture1
Parametersparameter_count1
Context windowcontext_length1
Max outputmax_output_tokens2
Modalitiesmodalities1
Input modalitiesmodalities_input1
Output modalitiesmodalities_output1
Tokenizertokenizer1
Quantization formatquant_format1
File sizefile_size_gb1
Hugging Face repohf_repo1
Pipeline tagpipeline_tag1
Model cardmodel_card_url1
Downloadsmetric.downloads7
Likesmetric.likes24
Accessaccess1
Artifact kindartifact_kind1
Commercial use allowedcommercial_use_allowed1
Derivatives allowedderivatives_allowed1
Descriptiondescription1
Gatedgated1
Hf inference providershf_inference_providers1
Is quantizedis_quantized1
Last modifiedlast_modified1
Library namelibrary_name1
License keylicense_key1
License rawlicense_raw1
Livebench hf linklivebench_hf_link1
Downloads all timemetric.downloads_all_time7
Model typemodel_type1
Openrouter idopenrouter_id1
Openrouter listed atopenrouter_listed_at1
Reasoningreasoning1
Redistribution allowedredistribution_allowed1
Structured outputstructured_output1
Supported parameterssupported_parameters1
Tool callingtool_calling1
Visionvision1
Weights availableweights_available1
Weights dtypeweights_dtype1
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
Provenance
Attributed facts
46
Source tiers
T246
Freshest observation
4 h ago
Conflicts
None
Source documents 7
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.
Data quality (72/100) measures how well AI Atlas knows this entity — completeness, primary-source ratio, freshness, conflicts — never how good the model is. Methodology →