bartowski/Qwen_Qwen3-Next-80B-A3B-Thinking-GGUF
published by bartowskihuggingface.co/bartowski/Qwen_Qwen3-Next-80B-A3B-Thinkin
This is a quantization of Qwen3 Next 80B A3B Thinking, not an independent model. Parameters, benchmarks, prices and lineage are recorded on the canonical model. Open Qwen3 Next 80B A3B Thinking →
Updated 3 h ago · first seen 11 Sept 2026
model_01M294XK19W64JKZJBPGJBS2QM
- File size
- —
- Format
- Downloads
- Published
Artifact facts
- Quantization format
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 12 h agomedium
- Quantization
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 12 h agomedium
- Downloads
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 6 h agomedium
- Likes
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 6 h agomedium
- Hugging Face repo
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 12 h agomedium
- Base model
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 12 h agomedium
- Quantized by
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 12 h agomedium
- Pipeline tag
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 12 h agomedium
- Gated
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 12 h agomedium
- License
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 12 h agomedium
- Release date
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 12 h agomedium
- Last modified
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 12 h agomedium
- Model card
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 12 h agomedium
- Tags
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 12 h agomedium
21 quantisation levels in this repository (IQ1_M, IQ1_S, IQ2_M, IQ2_S, IQ2_XS, IQ2_XXS, IQ3_M, IQ3_XS…).
Hardware fit (this packaging)
Assumptions (7)
- Estimated, not measured: weights = parameters × bytes/param × 1.15 runtime overhead (or the observed artifact file size when one is recorded).
- bytes/param: 4bit = 0.5, 8bit = 1.0, fp16 = 2.0 (uniform quantization, no per-layer exceptions).
- KV cache: 2 × layers × kv_heads × head_dim × 2 bytes × context × batch when the architecture is known; otherwise 0.5 GB per 8 192 tokens (× batch), independent of architecture (GQA/MLA models need less).
- A model 'fits' when the estimate is at most the device memory minus 2 GB reserved for the OS and framework.
- Mixture-of-experts models are estimated on total parameters (all experts must be resident); active parameters are ignored.
- Device memory uses the largest configuration when several are listed (e.g. Apple silicon tiers).
- Multi-GPU: device memories are summed; interconnect bandwidth, tensor-parallel replication and pipeline bubbles are not modelled.
Timeline
bartowski/Qwen_Qwen3-Next-80B-A3B-Thinking-GGUF: parameter count changed from 80000000000 to 79674391296
Parameters80B→79.7Bhuggingfacebartowski/Qwen_Qwen3-Next-80B-A3B-Thinking-GGUF: parameter count changed from 79674391296 to 80000000000
Parameters79.7B→80Bhuggingface
Provenance
Attributed facts
25
Source tiers
T225
Freshest observation
3 h ago
Conflicts
None
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.