GLM 5.3 Flash
Z.ai (Zhipu AI)huggingface.co/zai-org/GLM-5.3-Flash
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
Updated 4 h ago · first seen 11 Sept 2026
model_01M294AJ50P57GSKVAN937X1AH
Overview
Identity
Identity block not returned by the API for this entity.
Openness
Openness not classified yet — no sourced evidence to place this model in the ontology.
Key facts
- Release date
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 6 h agomedium
- Model card
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 11 h agomedium
- Openrouter id
Source:OpenRouter public model & pricing listingT2observed 11 h agomedium
Architecture
- Architecture
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 10 h agomedium
- Model type
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 10 h agomedium
- Parameters
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 11 h agomedium
- Weights dtype
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 10 h agomedium
- File size
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 10 h agomedium
- Library name
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 10 h agomedium
- Pipeline tag
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 11 h agomedium
- Hugging Face repo
Source:OpenRouter public model & pricing listingT2observed 11 h agomedium
Capabilities
Modalities
- Modalities
- imagetextvideo
- Input
- imagetextvideo
- Output
- text
Capabilities
Tool calling
Yes
OpenRouter public model & pricing listing · T2
Structured output
Yes
OpenRouter public model & pricing listing · T2
Reasoning
Yes
OpenRouter public model & pricing listing · T2
Vision
Yes
OpenRouter public model & pricing listing · T2
Audio
Unavailable
Fine-tuning available
Unavailable
- Context window
Source:OpenRouter public model & pricing listingT2observed 8 h agomedium
- Max output
Source:OpenRouter public model & pricing listingT2observed 11 h agomedium
- Languages
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 10 h agomedium
Benchmarks14
Compare with another model →Comparable same task and conditions · Partially comparable same task, conditions differ (effort, temperature, judge) · Not comparable different variant or metric
No benchmark results recorded
Providers & Pricing4
All offers in the price terminal →USD per 1M tokens as published by each provider (USD). Rows are append-only: every change is kept in the history below.
Price history
Output price · USD / 1M tokens 3 providers
- Fireworks AI
- Z.ai API
- Together AI
- Together AIfirst observed $0.5011 Sept 2026
- Z.ai API$0.50 → $0.2511 Sept 2026
- Z.ai APIfirst observed $0.5011 Sept 2026
- Fireworks AIfirst observed $0.5011 Sept 2026
Input price · USD / 1M tokens 3 providers
- Fireworks AI
- Z.ai API
- Together AI
- Together AIfirst observed $0.1511 Sept 2026
- Z.ai API$0.15 → $0.07511 Sept 2026
- Z.ai APIfirst observed $0.1511 Sept 2026
- Fireworks AIfirst observed $0.1511 Sept 2026
Hardware fit37
Assumptions (6)
- Estimated, not measured: weights = parameters × bytes/param × 1.15 runtime overhead.
- bytes/param: 4bit = 0.5, 8bit = 1.0, fp16 = 2.0 (uniform quantization, no per-layer exceptions).
- KV cache approximated at 0.5 GB per 8 192 tokens of context, independent of architecture (GQA/MLA models need less).
- A model 'fits' when the estimate is at most the device memory minus 2 GB reserved for the OS and framework.
- Mixture-of-experts models are estimated on total parameters (all experts must be resident); active parameters are ignored.
- Device memory uses the largest configuration when several are listed (e.g. Apple silicon tiers).
Papers1
- arXiv:2602.15763Active35
Timeline23
Full timeline →GLM 5.3 Flash: context length changed from 1000000 to 1310720
Context window1M tokens→1.31M tokensopenrouterGLM 5.3 Flash scores 52.821% on LiveBench
livebench_leaderboardGLM 5.3 Flash scores 77.307% on LiveBench
livebench_leaderboardGLM 5.3 Flash scores 76.401% on LiveBench
livebench_leaderboardGLM 5.3 Flash scores 81.245% on LiveBench
livebench_leaderboardGLM 5.3 Flash scores 56.768% on LiveBench
livebench_leaderboardGLM 5.3 Flash scores 78.95% on LiveBench
livebench_leaderboardGLM 5.3 Flash scores 77.644% on LiveBench
livebench_leaderboardGLM 5.3 Flash scores 71.591% on LiveBench
livebench_leaderboardGLM 5.3 Flash scores 84.27% on Terminal-Bench
artificial_analysisGLM 5.3 Flash scores 32.83% on Terminal-Bench
artificial_analysisGLM 5.3 Flash scores 51.62% on SciCode
artificial_analysisGLM 5.3 Flash scores 39.85% on Humanity's Last Exam
artificial_analysisGLM 5.3 Flash scores 91.21% on GPQA
artificial_analysisGLM 5.3 Flash scores 41.91 on Artificial Analysis Intelligence Index
artificial_analysisGLM 5.3 Flash: context length changed from 1310720 to 1000000
Context window1.31M tokens→1M tokensartificial_analysisTogether AI lists GLM 5.3 Flash at $0.15 in / $0.5 out per 1M tokens
together_pricingGLM 5.3 Flash: release date changed from 2026-08-25 to 2026-08-26
Release date25 Aug 2026→26 Aug 2026openrouterGLM 5.3 Flash: release date changed from 2026-08-26 to 2026-08-25
Release date26 Aug 2026→25 Aug 2026huggingfaceZ.ai API lists GLM 5.3 Flash at $0.075 in / $0.25 out per 1M tokens
openrouterZ.ai API lists GLM 5.3 Flash at $0.15 in / $0.5 out per 1M tokens
openrouterFireworks AI lists GLM 5.3 Flash at $0.15 in / $0.5 out per 1M tokens
fireworks_pricing
Change history37
Viewing AI Atlas as of 12 Sept 2026 — attributes exactly as the atlas knew them on that day; later corrections are not shown.
Back to today →Attributes as of 12 Sept 2026 44 claims in force
- Release date
- 25 Aug 2026
- Openness
- open-weights
- License
- MIT
- Architecture
- Glm5NextForConditionalGeneration
- Parameters
- 321.3B
- Context window
- 1.31M tokens
- Max output
- 131.1K tokens
- Modalities
- image, text, video
- Input modalities
- image, text, video
- Output modalities
- text
- Languages
- en, zh
- Quantization format
- fp8
- File size
- 328.4 GB
- Hugging Face repo
- zai-org/GLM-5.3-Flash
- Pipeline tag
- image-text-to-text
- Model card
- Downloads
- 1,173,520
- Likes
- 2,253
- Aa context window
- 1,000,000
- Aa openness
- open-weights
- Commercial use allowed
- Yes
- Derivatives allowed
- Yes
- Gated
- No
- Hf inference providers
- baseten, deepinfra, fireworks-ai, novita, together, zai-org
- Is quantized
- Yes
- Last modified
- 2026-09-07T12:13:47+00:00
- Library name
- transformers
- Livebench hf link
- https://huggingface.co/zai-org/GLM-5.3-Flash
- Aa median output tokens per second
- 97.9
- Downloads all time
- 1,173,520
- Model type
- glm5_next
- Openrouter expiration date
- 31 Dec 2098
- Openrouter id
- z-ai/glm-5.3-flash
- Openrouter listed at
- 26 Aug 2026
- Reasoning
- Yes
- Redistribution allowed
- Yes
- Structured output
- Yes
- Supported parameters
- frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p
- Tags
- transformers, safetensors, glm5_next, image-text-to-text, en, zh, fp8
- Tool calling
- Yes
- Vision
- Yes
- Weights available
- Yes
- Weights dtype
- BF16, F32, F8_E4M3
Release daterelease_date2
Opennessopenness1
Licenselicense1
Architecturearchitecture1
Parametersparameter_count1
Context windowcontext_length1
Max outputmax_output_tokens1
Modalitiesmodalities1
Input modalitiesmodalities_input1
Output modalitiesmodalities_output1
Languageslanguages1
Quantization formatquant_format1
File sizefile_size_gb1
Hugging Face repohf_repo1
Pipeline tagpipeline_tag1
Model cardmodel_card_url1
Downloadsmetric.downloads1
Likesmetric.likes2
Descriptiondescription1
Gatedgated1
Hf inference providershf_inference_providers1
Is quantizedis_quantized1
Last modifiedlast_modified1
Library namelibrary_name1
Downloads all timemetric.downloads_all_time1
Model typemodel_type1
Openrouter expiration dateopenrouter_expiration_date1
Openrouter idopenrouter_id1
Reasoningreasoning1
Structured outputstructured_output1
Supported parameterssupported_parameters1
Tool callingtool_calling1
Visionvision1
Weights dtypeweights_dtype1
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
Provenance
Attributed facts
38
Source tiers
T238
Freshest observation
4 h ago
Conflicts
None
Source documents 8
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.
Data quality (72/100) measures how well AI Atlas knows this entity — completeness, primary-source ratio, freshness, conflicts — never how good the model is. Methodology →