EstimatedAll figures are estimates, not measurements.256 GB · 4-bit · 8k tokens624 of 658 models with a known parameter count fit
- Estimated, not measured: weights = parameters × bytes/param × 1.15 runtime overhead.
- bytes/param: 4bit = 0.5, 8bit = 1.0, fp16 = 2.0 (uniform quantization, no per-layer exceptions).
- KV cache approximated at 0.5 GB per 8 192 tokens of context, independent of architecture (GQA/MLA models need less).
- A model 'fits' when the estimate is at most the device memory minus 2 GB reserved for the OS and framework.
- Mixture-of-experts models are estimated on total parameters (all experts must be resident); active parameters are ignored.
- Device memory uses the largest configuration when several are listed (e.g. Apple silicon tiers).
| Model | Params | Quant | Est. memory | Headroom | Fits | Compare |
|---|---|---|---|---|---|---|
| MiniMax-M3-MXFP8Open weightsMiniMaxestimated: 440.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 440.3B | 4bit | 253.7 GB | +0.3 GB | ✓ fits | |
| MiniMax-M3Open weightsMiniMaxestimated: 427.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 427B | 4bit | 246.1 GB | +8.0 GB | ✓ fits | |
| ERNIE-4.5-VL-424B-A47B-Base-PTOpen weightsBaiduestimated: 423.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 423.5B | 4bit | 244.0 GB | +10.0 GB | ✓ fits | |
| ERNIE 4.5 VL 424B A47BOpen weightsBaiduestimated: 423.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 423.5B | 4bit | 244.0 GB | +10.0 GB | ✓ fits | |
| amd/GLM-5.2-MXFP4Open weightsAMDestimated: 412.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 412.3B | 4bit | 237.6 GB | +16.4 GB | ✓ fits | |
| meta-llama/Llama-3.1-405B-FP8RestrictedMeta AIestimated: 405.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 405.9B | 4bit | 233.9 GB | +20.1 GB | ✓ fits | |
| Llama-3.1-405BRestrictedMeta AIestimated: 405.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 405.9B | 4bit | 233.9 GB | +20.1 GB | ✓ fits | |
| ai21labs/AI21-Jamba-Large-1.7-FP8RestrictedAI21 Labsestimated: 398.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 398.6B | 4bit | 229.7 GB | +24.3 GB | ✓ fits | |
| AI21-Jamba-Large-1.7RestrictedAI21 Labsestimated: 398.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 398.6B | 4bit | 229.7 GB | +24.3 GB | ✓ fits | |
| AI21-Jamba-Large-1.5RestrictedAI21 Labsestimated: 398.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 398.6B | 4bit | 229.7 GB | +24.3 GB | ✓ fits | |
| AI21-Jamba-Large-1.6RestrictedAI21 Labsestimated: 398.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 398.6B | 4bit | 229.7 GB | +24.3 GB | ✓ fits | |
| Intel/Qwen3.5-397B-A17B-int4-AutoRoundOpen weightsIntelestimated: 397.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 397B | 4bit | 228.8 GB | +25.2 GB | ✓ fits | |
| amd/GLM-5.3-Quark-MXFP4-AttnFP8Open weightsAMDestimated: 384.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 384.3B | 4bit | 221.5 GB | +32.5 GB | ✓ fits | |
| nvidia/GLM-5.2-NVFP4Open weightsNVIDIAestimated: 381.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 381B | 4bit | 219.6 GB | +34.4 GB | ✓ fits | |
| amd/DeepSeek-R1-MXFP4Open weightsAMDestimated: 370.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 370.9B | 4bit | 213.8 GB | +40.3 GB | ✓ fits | |
| GLM 4.5Open weightsZ.ai (Zhipu AI)estimated: 358.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 358.3B | 4bit | 206.5 GB | +47.5 GB | ✓ fits | |
| GLM 4.7Open weightsZ.ai (Zhipu AI)estimated: 358.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 358.3B | 4bit | 206.5 GB | +47.5 GB | ✓ fits | |
| amd/DeepSeek-R1-0528-MXFP4Open weightsAMDestimated: 355.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 355.7B | 4bit | 205.0 GB | +49.0 GB | ✓ fits | |
| amd/DeepSeek-R1-0528-MXFP4-v2Open weightsAMDestimated: 350.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 350B | 4bit | 201.7 GB | +52.3 GB | ✓ fits | |
| cerebras/DeepSeek-V3.2-REAP-345B-A37BOpen weightsCerebras Systemsestimated: 344.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 344.9B | 4bit | 198.8 GB | +55.2 GB | ✓ fits | |
| GLM 5.3 FlashOpen weightsZ.ai (Zhipu AI)estimated: 321.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 321.3B | 4bit | 185.3 GB | +68.7 GB | ✓ fits | |
| DeepSeek V4 Flash Vision ExpOpen weightsDeepSeekestimated: 304.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 304.6B | 4bit | 175.7 GB | +78.3 GB | ✓ fits | |
| DeepSeek V4 Flash (0731)Open weightsDeepSeekestimated: 304.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 304.2B | 4bit | 175.4 GB | +78.6 GB | ✓ fits | |
| ERNIE-4.5-300B-A47B-PTOpen weightsBaiduestimated: 300.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 300.5B | 4bit | 173.3 GB | +80.7 GB | ✓ fits | |
| ERNIE-4.5-300B-A47B-PaddleOpen weightsBaiduestimated: 300.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 300.5B | 4bit | 173.3 GB | +80.7 GB | ✓ fits | |
| ERNIE-4.5-300B-A47B-Base-PTOpen weightsBaiduestimated: 299.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 299.5B | 4bit | 172.7 GB | +81.3 GB | ✓ fits | |
| tencent/Hy3-FP8Open weightsTencentestimated: 298.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 298.8B | 4bit | 172.3 GB | +81.7 GB | ✓ fits | |
| Hy3Open weightsTencentestimated: 298.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 298.8B | 4bit | 172.3 GB | +81.7 GB | ✓ fits | |
| Hy3 previewOpen weightsTencentestimated: 298.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 298.8B | 4bit | 172.3 GB | +81.7 GB | ✓ fits | |
| DeepSeek-V4-Flash-BaseOpen weightsDeepSeekestimated: 292.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 292B | 4bit | 168.4 GB | +85.6 GB | ✓ fits | |
| DeepSeek V4 Flash 0423Open weightsDeepSeekestimated: 290.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 290.9B | 4bit | 167.8 GB | +86.2 GB | ✓ fits | |
| cerebras/GLM-4.7-REAP-268B-A32BOpen weightsCerebras Systemsestimated: 268.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 268.8B | 4bit | 155.1 GB | +99.0 GB | ✓ fits | |
| cerebras/GLM-4.6-REAP-268B-A32BOpen weightsCerebras Systemsestimated: 268.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 268.8B | 4bit | 155.1 GB | +99.0 GB | ✓ fits | |
| unsloth/Inkling-Small-GGUFOpen weightsUnslothestimated: 263.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 263.7B | 4bit | 152.1 GB | +101.9 GB | ✓ fits | |
| Intern-S1Open weightsInternLM (Shanghai AI Laboratory)estimated: 240.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 240.7B | 4bit | 138.9 GB | +115.1 GB | ✓ fits | |
| amd/MiniMax-M3-MXFP4Open weightsAMDestimated: 231.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 231.9B | 4bit | 133.8 GB | +120.2 GB | ✓ fits | |
| MiniMax M2.5Open weightsMiniMaxestimated: 228.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 228.7B | 4bit | 132.0 GB | +122.0 GB | ✓ fits | |
| MiniMax M2Open weightsMiniMaxestimated: 228.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 228.7B | 4bit | 132.0 GB | +122.0 GB | ✓ fits | |
| MiniMax M2.1Open weightsMiniMaxestimated: 228.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 228.7B | 4bit | 132.0 GB | +122.0 GB | ✓ fits | |
| MiniMax M2.7Open weightsMiniMaxestimated: 228.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 228.7B | 4bit | 132.0 GB | +122.0 GB | ✓ fits | |
| amd/Qwen3.5-397B-A17B-MXFP4Open weightsAMDestimated: 222.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 222.2B | 4bit | 128.3 GB | +125.7 GB | ✓ fits | |
| command-a-plus-05-2026-w4a4Open weightsCohereestimated: 218.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 218.8B | 4bit | 126.3 GB | +127.7 GB | ✓ fits | |
| command-a-plus-05-2026-bf16Open weightsCohereestimated: 218.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 218.8B | 4bit | 126.3 GB | +127.7 GB | ✓ fits | |
| cerebras/GLM-4.7-REAP-218B-A32BOpen weightsCerebras Systemsestimated: 218.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 218.4B | 4bit | 126.1 GB | +127.9 GB | ✓ fits | |
| cerebras/GLM-4.6-REAP-218B-A32BOpen weightsCerebras Systemsestimated: 218.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 218.4B | 4bit | 126.1 GB | +127.9 GB | ✓ fits | |
| north-small-translate-1-0Cohereestimated: 218.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 218B | 4bit | 125.8 GB | +128.2 GB | ✓ fits | |
| amd/Qwen3.5-397B-A17B-MXFP4-AttnFP8-V2Open weightsAMDestimated: 203.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 203.2B | 4bit | 117.3 GB | +136.7 GB | ✓ fits | |
| unsloth/Qwen3.8-Flash-Next-GGUFOpen weightsUnslothestimated: 176.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 176.9B | 4bit | 102.2 GB | +151.8 GB | ✓ fits | |
| cerebras/MiniMax-M2.5-REAP-172B-A10BOpen weightsCerebras Systemsestimated: 172.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 172.5B | 4bit | 99.7 GB | +154.3 GB | ✓ fits | |
| cerebras/MiniMax-M2-REAP-172B-A10BOpen weightsCerebras Systemsestimated: 172.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 172.5B | 4bit | 99.7 GB | +154.3 GB | ✓ fits | |
| DeepSeek-V4-Flash-DSparkOpen weightsDeepSeekestimated: 165.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 165.3B | 4bit | 95.5 GB | +158.5 GB | ✓ fits | |
| cerebras/MiniMax-M2-REAP-162B-A10BOpen weightsCerebras Systemsestimated: 162.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 162B | 4bit | 93.6 GB | +160.4 GB | ✓ fits | |
| Step-3.5-Flash-REAP-149B-A11BOpen weightsCerebras Systemsestimated: 149.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 149.4B | 4bit | 86.4 GB | +167.6 GB | ✓ fits | |
| fireworks-ai/mixtral-8x22b-instruct-ohOpen weightsFireworks AIestimated: 140.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 140.6B | 4bit | 81.4 GB | +172.6 GB | ✓ fits | |
| cerebras/MiniMax-M2.1-REAP-139B-A10BOpen weightsCerebras Systemsestimated: 139.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 139.2B | 4bit | 80.5 GB | +173.5 GB | ✓ fits | |
| cerebras/MiniMax-M2.5-REAP-139B-A10BOpen weightsCerebras Systemsestimated: 139.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 139.2B | 4bit | 80.5 GB | +173.5 GB | ✓ fits | |
| cerebras/MiniMax-M2-REAP-139B-A10BOpen weightsCerebras Systemsestimated: 139.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 139.2B | 4bit | 80.5 GB | +173.5 GB | ✓ fits | |
| Mistral-Medium-3.5-128BOpen weightsMistral AIestimated: 127.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 127.7B | 4bit | 73.9 GB | +180.1 GB | ✓ fits | |
| NVIDIA-Nemotron-3-Super-120B-A12B-BF16Open weightsNVIDIAestimated: 123.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 123.6B | 4bit | 71.6 GB | +182.4 GB | ✓ fits | |
| Step-3.5-Flash-REAP-121B-A11BOpen weightsCerebras Systemsestimated: 121.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 121B | 4bit | 70.1 GB | +183.9 GB | ✓ fits | |
| amd/Qwen3-VL-235B-A22B-Instruct-MXFP4RestrictedAMDestimated: 118.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 118.8B | 4bit | 68.8 GB | +185.2 GB | ✓ fits | |
| gpt-oss-120bOpen weightsOpenAIestimated: 116.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 116.8B | 4bit | 67.7 GB | +186.3 GB | ✓ fits | |
| command-a-vision-07-2025RestrictedCohereestimated: 111.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 111.9B | 4bit | 64.8 GB | +189.2 GB | ✓ fits | |
| c4ai-command-a-03-2025RestrictedCohereestimated: 111.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 111.1B | 4bit | 64.4 GB | +189.6 GB | ✓ fits | |
| zai-org/GLM-4.5-Air-FP8Open weightsZ.ai (Zhipu AI)estimated: 110.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 110.5B | 4bit | 64.0 GB | +190.0 GB | ✓ fits | |
| GLM 4.5 AirOpen weightsZ.ai (Zhipu AI)estimated: 110.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 110.5B | 4bit | 64.0 GB | +190.0 GB | ✓ fits | |
| Llama 4 ScoutRestrictedMeta AIestimated: 108.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 108.6B | 4bit | 63.0 GB | +191.0 GB | ✓ fits | |
| Llama-3.2-90B-Vision-InstructRestrictedMeta AIestimated: 88.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 88.6B | 4bit | 51.4 GB | +202.6 GB | ✓ fits | |
| HunyuanImage-3.0-InstructOpen weightsTencentestimated: 83.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 83B | 4bit | 48.2 GB | +205.8 GB | ✓ fits | |
| cerebras/GLM-4.5-Air-REAP-82B-A12BOpen weightsCerebras Systemsestimated: 81.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 81.9B | 4bit | 47.6 GB | +206.4 GB | ✓ fits | |
| Hunyuan A13B InstructOpen weightsTencentestimated: 80.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 80.4B | 4bit | 46.7 GB | +207.3 GB | ✓ fits | |
| bartowski/Qwen_Qwen3-Next-80B-A3B-Thinking-GGUFOpen weightsbartowskiestimated: 80.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 80B | 4bit | 46.5 GB | +207.5 GB | ✓ fits | |
| Intel/Qwen3.8-Flash-Next-W4A16-AutoRoundOpen weightsIntelestimated: 75.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 75.4B | 4bit | 43.8 GB | +210.2 GB | ✓ fits | |
| UI-TARS-72B-DPOOpen weightsByteDanceestimated: 73.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 73.4B | 4bit | 42.7 GB | +211.3 GB | ✓ fits | |
| Kimi-Dev-72BOpen weightsMoonshot AIestimated: 72.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 72.7B | 4bit | 42.3 GB | +211.7 GB | ✓ fits | |
| amd/Llama-3.3-70B-Instruct-FP8-KVOpen weightsAMDestimated: 70.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 70.5B | 4bit | 41.1 GB | +212.9 GB | ✓ fits | |
| NousResearch/Meta-Llama-3.1-70B-InstructOpen weightsNous Researchestimated: 70.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 70.5B | 4bit | 41.1 GB | +212.9 GB | ✓ fits | |
| Llama-3.1-70B-InstructRestrictedMeta AIestimated: 70.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 70.5B | 4bit | 41.1 GB | +212.9 GB | ✓ fits | |
| NousResearch/Meta-Llama-3-70B-InstructOpen weightsNous Researchestimated: 70.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 70.5B | 4bit | 41.1 GB | +212.9 GB | ✓ fits | |
| Llama 3.3 70B InstructRestrictedMeta AIestimated: 70.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 70.5B | 4bit | 41.1 GB | +212.9 GB | ✓ fits | |
| Meta-Llama-3-70BRestrictedMeta AIestimated: 70.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 70.5B | 4bit | 41.1 GB | +212.9 GB | ✓ fits | |
| fireworks-ai/llama-3-firefunction-v2Open weightsFireworks AIestimated: 70.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 70.5B | 4bit | 41.1 GB | +212.9 GB | ✓ fits | |
| Hermes-4-70BOpen weightsNous Researchestimated: 70.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 70.5B | 4bit | 41.1 GB | +212.9 GB | ✓ fits | |
| nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4Open weightsNVIDIAestimated: 67.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 67.2B | 4bit | 39.2 GB | +214.8 GB | ✓ fits | |
| nvidia/Qwen3.5-122B-A10B-NVFP4Open weightsNVIDIAestimated: 64.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 64.6B | 4bit | 37.6 GB | +216.4 GB | ✓ fits | |
| amd/gpt-oss-120b-w-mxfp4-a-fp8Open weightsAMDestimated: 59.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 59.5B | 4bit | 34.7 GB | +219.3 GB | ✓ fits | |
| ai21labs/AI21-Jamba-Mini-1.7-FP8RestrictedAI21 Labsestimated: 51.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 51.6B | 4bit | 30.2 GB | +223.8 GB | ✓ fits | |
| ai21labs/AI21-Jamba2-Mini-FP8Open weightsAI21 Labsestimated: 51.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 51.6B | 4bit | 30.2 GB | +223.8 GB | ✓ fits | |
| Jamba-v0.1Open weightsAI21 Labsestimated: 51.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 51.6B | 4bit | 30.1 GB | +223.8 GB | ✓ fits | |
| AI21-Jamba-Mini-1.7RestrictedAI21 Labsestimated: 51.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 51.6B | 4bit | 30.1 GB | +223.8 GB | ✓ fits | |
| AI21-Jamba2-MiniOpen weightsAI21 Labsestimated: 51.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 51.6B | 4bit | 30.1 GB | +223.8 GB | ✓ fits | |
| AI21-Jamba-Mini-1.5RestrictedAI21 Labsestimated: 51.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 51.6B | 4bit | 30.1 GB | +223.8 GB | ✓ fits | |
| AI21-Jamba-Mini-1.6RestrictedAI21 Labsestimated: 51.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 51.6B | 4bit | 30.1 GB | +223.8 GB | ✓ fits | |
| Kimi-Linear-48B-A3B-BaseOpen weightsMoonshot AIestimated: 49.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 49.1B | 4bit | 28.8 GB | +225.3 GB | ✓ fits | |
| Kimi-Linear-48B-A3B-InstructOpen weightsMoonshot AIestimated: 49.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 49.1B | 4bit | 28.8 GB | +225.3 GB | ✓ fits | |
| firefunction-v1Open weightsFireworks AIestimated: 46.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 46.7B | 4bit | 27.4 GB | +226.7 GB | ✓ fits | |
| Nous-Hermes-2-Mixtral-8x7B-DPOOpen weightsNous Researchestimated: 46.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 46.7B | 4bit | 27.4 GB | +226.7 GB | ✓ fits | |
| function-calling-v1Open weightsFireworks AIestimated: 46.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 46.7B | 4bit | 27.4 GB | +226.7 GB | ✓ fits | |
| Mixtral-8x7B-Instruct-v0.1Open weightsMistral AIestimated: 46.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 46.7B | 4bit | 27.4 GB | +226.7 GB | ✓ fits | |
| DFN2B-CLIP-ViT-L-14-39BOpen weightsAppleestimated: 39.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 39B | 4bit | 22.9 GB | +231.1 GB | ✓ fits | |
| NousResearch/Hermes-4.3-36B-GGUFOpen weightsNous Researchestimated: 36.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 36.1B | 4bit | 21.3 GB | +232.7 GB | ✓ fits | |
| Seed-OSS-36B-BaseOpen weightsByteDanceestimated: 36.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 36.1B | 4bit | 21.3 GB | +232.7 GB | ✓ fits | |
| Seed-OSS-36B-InstructOpen weightsByteDanceestimated: 36.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 36.1B | 4bit | 21.3 GB | +232.7 GB | ✓ fits | |
| Qwen/Qwen3.6-35B-A3B-FP8Open weightsQwenestimated: 36.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 36B | 4bit | 21.2 GB | +232.8 GB | ✓ fits | |
| Qwen3.6 35B A3BOpen weightsQwenestimated: 36.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 36B | 4bit | 21.2 GB | +232.8 GB | ✓ fits | |
| bartowski/endless-frontier_BigBang-v1-GGUFOpen weightsbartowskiestimated: 35.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 35.5B | 4bit | 20.9 GB | +233.1 GB | ✓ fits | |
| unsloth/Qwen3.6-35B-A3B-MTP-GGUFOpen weightsUnslothestimated: 35.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 35.5B | 4bit | 20.9 GB | +233.1 GB | ✓ fits | |
| cerebras/Kimi-Linear-REAP-35B-A3B-InstructOpen weightsCerebras Systemsestimated: 35.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 35.1B | 4bit | 20.7 GB | +233.3 GB | ✓ fits | |
| perplexity-ai/pplx-computer-qwen-3-6-35b-a3b-nvfp4-20260709Open weightsPerplexity AIestimated: 35.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 35.1B | 4bit | 20.7 GB | +233.3 GB | ✓ fits | |
| bartowski/Qwen_Qwen3.6-35B-A3B-GGUFOpen weightsbartowskiestimated: 35.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 35B | 4bit | 20.6 GB | +233.4 GB | ✓ fits | |
| bartowski/Qwen_Qwen3.5-35B-A3B-GGUFOpen weightsbartowskiestimated: 35.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 35B | 4bit | 20.6 GB | +233.4 GB | ✓ fits | |
| c4ai-command-r-v01RestrictedCohereestimated: 35.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 35B | 4bit | 20.6 GB | +233.4 GB | ✓ fits | |
| bartowski/XYZAILab_XYZ-Aquila-mini-GGUFOpen weightsbartowskiestimated: 34.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 34.7B | 4bit | 20.4 GB | +233.6 GB | ✓ fits | |
| bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUFOpen weightsbartowskiestimated: 34.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 34.7B | 4bit | 20.4 GB | +233.6 GB | ✓ fits | |
| unsloth/Qwen3.6-35B-A3B-GGUFOpen weightsUnslothestimated: 34.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 34.7B | 4bit | 20.4 GB | +233.6 GB | ✓ fits | |
| bartowski/thomsonreuters_Thomson-1.0-Small-GGUFOpen weightsbartowskiestimated: 34.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 34.7B | 4bit | 20.4 GB | +233.6 GB | ✓ fits | |
| perplexity-ai/pplx-computer-qwen-3-6-35b-a3b-mlx-20260709Open weightsPerplexity AIestimated: 34.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 34.7B | 4bit | 20.4 GB | +233.6 GB | ✓ fits | |
| Nous-Hermes-2-Yi-34BOpen weightsNous Researchestimated: 34.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 34.4B | 4bit | 20.3 GB | +233.7 GB | ✓ fits | |
| GKA-primed-HQwen3-32B-ReasonerOpen weightsAmazon Web Servicesestimated: 34.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 34.1B | 4bit | 20.1 GB | +233.9 GB | ✓ fits | |
| MiniMax-H3Open weightsMiniMaxestimated: 33.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 33.1B | 4bit | 19.6 GB | +234.4 GB | ✓ fits | |
| nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8Open weightsNVIDIAestimated: 33.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 33B | 4bit | 19.5 GB | +234.5 GB | ✓ fits | |
| DeepSeek-R1-Distill-Qwen-32BOpen weightsDeepSeekestimated: 32.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 32.8B | 4bit | 19.3 GB | +234.7 GB | ✓ fits | |
| SynLogic-32BOpen weightsMiniMaxestimated: 32.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 32.8B | 4bit | 19.3 GB | +234.7 GB | ✓ fits | |
| SynLogic-Mix-3-32BOpen weightsMiniMaxestimated: 32.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 32.8B | 4bit | 19.3 GB | +234.7 GB | ✓ fits | |
| Qwen3 32BOpen weightsQwenestimated: 32.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 32.8B | 4bit | 19.3 GB | +234.7 GB | ✓ fits | |
| c4ai-command-r-08-2024RestrictedCohereestimated: 32.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 32.3B | 4bit | 19.1 GB | +234.9 GB | ✓ fits | |
| aya-expanse-32bRestrictedCohereestimated: 32.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 32.3B | 4bit | 19.1 GB | +234.9 GB | ✓ fits | |
| OLMo-2-0325-32B-InstructOpen weightsAllen Institute for AIestimated: 32.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 32.2B | 4bit | 19.0 GB | +235.0 GB | ✓ fits | |
| FLUX.2-devRestrictedBlack Forest Labsestimated: 32.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 32.2B | 4bit | 19.0 GB | +235.0 GB | ✓ fits | |
| Nemotron 3 Nano 30B A3BOpen weightsNVIDIAestimated: 31.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 31.6B | 4bit | 18.7 GB | +235.3 GB | ✓ fits | |
| Gemma 4 31BOpen weightsGoogleestimated: 31.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 31.3B | 4bit | 18.5 GB | +235.5 GB | ✓ fits | |
| mlx-community/gemma-4-31b-it-4bitOpen weightsMLX Communityestimated: 31.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 31.3B | 4bit | 18.5 GB | +235.5 GB | ✓ fits | |
| GLM 4.7 FlashOpen weightsZ.ai (Zhipu AI)estimated: 31.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 31.2B | 4bit | 18.4 GB | +235.6 GB | ✓ fits | |
| unsloth/gemma-4-31B-it-qat-GGUFOpen weightsUnslothestimated: 30.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 30.7B | 4bit | 18.1 GB | +235.8 GB | ✓ fits | |
| unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUFOpen weightsUnslothestimated: 30.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 30.5B | 4bit | 18.1 GB | +235.9 GB | ✓ fits | |
| mlx-community/Qwen3-30B-A3B-Instruct-2507-4bitOpen weightsMLX Communityestimated: 30.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 30.5B | 4bit | 18.1 GB | +235.9 GB | ✓ fits | |
| browsesafeOpen weightsPerplexity AIestimated: 30.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 30.5B | 4bit | 18.1 GB | +235.9 GB | ✓ fits | |
| unsloth/Qwen3-30B-A3B-Thinking-2507-GGUFOpen weightsUnslothestimated: 30.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 30.5B | 4bit | 18.1 GB | +235.9 GB | ✓ fits | |
| North Mini Code (free)Open weightsCohereestimated: 30.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 30.5B | 4bit | 18.0 GB | +236.0 GB | ✓ fits | |
| Hy-MT2-30B-A3BOpen weightsTencentestimated: 30.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 30.1B | 4bit | 17.8 GB | +236.2 GB | ✓ fits | |
| north-mini-code-1-0Cohereestimated: 30.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 30B | 4bit | 17.8 GB | +236.3 GB | ✓ fits | |
| ibm-granite/granite-4.2-30b-GGUFOpen weightsIBMestimated: 30.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 30B | 4bit | 17.8 GB | +236.3 GB | ✓ fits | |
| ERNIE-4.5-VL-28B-A3B-ThinkingOpen weightsBaiduestimated: 29.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 29.7B | 4bit | 17.6 GB | +236.4 GB | ✓ fits | |
| ERNIE-4.5-VL-28B-A3B-PTOpen weightsBaiduestimated: 29.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 29.4B | 4bit | 17.4 GB | +236.6 GB | ✓ fits | |
| granite-4.1-30bOpen weightsIBMestimated: 28.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 28.9B | 4bit | 17.1 GB | +236.9 GB | ✓ fits | |
| Qwen/Qwen3.6-27B-FP8Open weightsQwenestimated: 27.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 27.8B | 4bit | 16.5 GB | +237.5 GB | ✓ fits | |
| Qwen/Qwen3.8-27B-FP8Open weightsQwenestimated: 27.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 27.8B | 4bit | 16.5 GB | +237.5 GB | ✓ fits | |
| Qwen3.6 27BOpen weightsQwenestimated: 27.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 27.8B | 4bit | 16.5 GB | +237.5 GB | ✓ fits | |
| Qwen3.8 27BOpen weightsQwenestimated: 27.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 27.8B | 4bit | 16.5 GB | +237.5 GB | ✓ fits | |
| mlx-community/Qwen3.8-27B-8bitOpen weightsMLX Communityestimated: 27.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 27.4B | 4bit | 16.2 GB | +237.8 GB | ✓ fits | |
| mlx-community/Qwen3.8-27B-4bitOpen weightsMLX Communityestimated: 27.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 27.4B | 4bit | 16.2 GB | +237.8 GB | ✓ fits | |
| unsloth/Qwen3.6-27B-MTP-GGUFOpen weightsUnslothestimated: 27.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 27.3B | 4bit | 16.2 GB | +237.8 GB | ✓ fits | |
| perplexity-ai/pplx-computer-qwen-3-8-27b-dflash2-gguf-20260826Open weightsPerplexity AIestimated: 27.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 27.3B | 4bit | 16.2 GB | +237.8 GB | ✓ fits | |
| unsloth/Qwen3.8-27B-GGUFOpen weightsUnslothestimated: 27.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 27.3B | 4bit | 16.2 GB | +237.8 GB | ✓ fits | |
| bartowski/Qwen3.8-27B-GGUFOpen weightsbartowskiestimated: 27.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 27B | 4bit | 16.0 GB | +238.0 GB | ✓ fits | |
| bartowski/Fara1.5-27B-GGUFOpen weightsbartowskiestimated: 27.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 27B | 4bit | 16.0 GB | +238.0 GB | ✓ fits | |
| unsloth/Qwen3.6-27B-GGUFOpen weightsUnslothestimated: 26.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 26.9B | 4bit | 16.0 GB | +238.0 GB | ✓ fits | |
| Gemma 4 26B A4BOpen weightsGoogleestimated: 25.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 25.8B | 4bit | 15.3 GB | +238.7 GB | ✓ fits | |
| unsloth/gemma-4-26B-A4B-it-qat-GGUFOpen weightsUnslothestimated: 25.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 25.2B | 4bit | 15.0 GB | +239.0 GB | ✓ fits | |
| unsloth/gemma-4-26B-A4B-it-GGUFOpen weightsUnslothestimated: 25.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 25.2B | 4bit | 15.0 GB | +239.0 GB | ✓ fits | |
| cerebras/Qwen3-Coder-REAP-25B-A3BOpen weightsCerebras Systemsestimated: 24.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 24.9B | 4bit | 14.8 GB | +239.2 GB | ✓ fits | |
| unsloth/Qwen3.6-35B-A3B-NVFP4Open weightsUnslothestimated: 24.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 24.6B | 4bit | 14.7 GB | +239.3 GB | ✓ fits | |
| Voxtral Small 24B 2507Open weightsMistral AIestimated: 24.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 24.3B | 4bit | 14.4 GB | +239.6 GB | ✓ fits | |
| Devstral-Small-2-24B-Instruct-2512Open weightsMistral AIestimated: 24.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 24B | 4bit | 14.3 GB | +239.7 GB | ✓ fits | |
| Mistral Small 3.2 24BOpen weightsMistral AIestimated: 24.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 24B | 4bit | 14.3 GB | +239.7 GB | ✓ fits | |
| mlx-community/Devstral-Small-2-24B-Instruct-2512-4bitOpen weightsMLX Communityestimated: 24.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 24B | 4bit | 14.3 GB | +239.7 GB | ✓ fits | |
| Mistral Small 3.1 24BOpen weightsMistral AIestimated: 24.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 24B | 4bit | 14.3 GB | +239.7 GB | ✓ fits | |
| cerebras/GLM-4.7-Flash-REAP-23B-A3BOpen weightsCerebras Systemsestimated: 23.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 23B | 4bit | 13.7 GB | +240.3 GB | ✓ fits | |
| ERNIE-4.5-21B-A3B-PTOpen weightsBaiduestimated: 21.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 21.9B | 4bit | 13.1 GB | +240.9 GB | ✓ fits | |
| ERNIE-4.5-21B-A3B-Base-PTOpen weightsBaiduestimated: 21.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 21.8B | 4bit | 13.1 GB | +240.9 GB | ✓ fits | |
| ERNIE-4.5-21B-A3B-ThinkingOpen weightsBaiduestimated: 21.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 21.8B | 4bit | 13.1 GB | +240.9 GB | ✓ fits | |
| gpt-oss-safeguard-20bOpen weightsOpenAIestimated: 21.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 21.5B | 4bit | 12.9 GB | +241.1 GB | ✓ fits | |
| unsloth/Qwen3.6-27B-NVFP4Open weightsUnslothestimated: 21.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 21.2B | 4bit | 12.7 GB | +241.3 GB | ✓ fits | |
| gpt-oss-20bOpen weightsOpenAIestimated: 20.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 20.9B | 4bit | 12.5 GB | +241.5 GB | ✓ fits | |
| mlx-community/gpt-oss-20b-MXFP4-Q8Open weightsMLX Communityestimated: 20.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 20.9B | 4bit | 12.5 GB | +241.5 GB | ✓ fits | |
| nvidia/Gemma-4-31B-IT-NVFP4Open weightsNVIDIAestimated: 20.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 20.9B | 4bit | 12.5 GB | +241.5 GB | ✓ fits | |
| gpt-neox-20bOpen weightsEleutherAIestimated: 20.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 20.7B | 4bit | 12.4 GB | +241.6 GB | ✓ fits | |
| unsloth/MiniMax-H3-GGUFOpen weightsUnslothestimated: 20.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 20.1B | 4bit | 12.1 GB | +241.9 GB | ✓ fits | |
| internlm/internlm2_5-20b-chat-ggufOpen weightsInternLM (Shanghai AI Laboratory)estimated: 20.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 20B | 4bit | 12.0 GB | +242.0 GB | ✓ fits | |
| internlm2-base-20bOpen weightsInternLM (Shanghai AI Laboratory)estimated: 20.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 20B | 4bit | 12.0 GB | +242.0 GB | ✓ fits | |
| internlm2-20bOpen weightsInternLM (Shanghai AI Laboratory)estimated: 20.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 20B | 4bit | 12.0 GB | +242.0 GB | ✓ fits | |
| unsloth/Qwen3.8-27B-NVFP4Open weightsUnslothestimated: 19.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 19.9B | 4bit | 11.9 GB | +242.1 GB | ✓ fits | |
| internlm2-chat-20bOpen weightsInternLM (Shanghai AI Laboratory)estimated: 19.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 19.9B | 4bit | 11.9 GB | +242.1 GB | ✓ fits | |
| Intel/Qwen3.5-122B-A10B-int4-AutoRoundOpen weightsIntelestimated: 19.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 19.7B | 4bit | 11.8 GB | +242.2 GB | ✓ fits | |
| pplx-qwen-3-8-27b-dflash2-20260819Open weightsPerplexity AIestimated: 18.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 18.8B | 4bit | 11.3 GB | +242.7 GB | ✓ fits | |
| pplx-computer-qwen-3-8-27b-dflash2-20260824Open weightsPerplexity AIestimated: 18.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 18.8B | 4bit | 11.3 GB | +242.7 GB | ✓ fits | |
| nvidia/Qwen3.6-35B-A3B-NVFP4Open weightsNVIDIAestimated: 18.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 18.7B | 4bit | 11.2 GB | +242.8 GB | ✓ fits | |
| nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4Open weightsNVIDIAestimated: 18.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 18.3B | 4bit | 11.0 GB | +243.0 GB | ✓ fits | |
| nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4Open weightsNVIDIAestimated: 17.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 17.8B | 4bit | 10.8 GB | +243.3 GB | ✓ fits | |
| Kimi-VL-A3B-ThinkingOpen weightsMoonshot AIestimated: 16.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 16.4B | 4bit | 9.9 GB | +244.1 GB | ✓ fits | |
| Kimi-VL-A3B-InstructOpen weightsMoonshot AIestimated: 16.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 16.4B | 4bit | 9.9 GB | +244.1 GB | ✓ fits | |
| Kimi-VL-A3B-Thinking-2506Open weightsMoonshot AIestimated: 16.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 16.4B | 4bit | 9.9 GB | +244.1 GB | ✓ fits | |
| Moonlight-16B-A3BOpen weightsMoonshot AIestimated: 16.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 16B | 4bit | 9.7 GB | +244.3 GB | ✓ fits | |
| Moonlight-16B-A3B-InstructOpen weightsMoonshot AIestimated: 16.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 16B | 4bit | 9.7 GB | +244.3 GB | ✓ fits | |
| DeepSeek-V2-LiteOpen weightsDeepSeekestimated: 15.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 15.7B | 4bit | 9.5 GB | +244.5 GB | ✓ fits | |
| DeepSeek-Coder-V2-Lite-InstructOpen weightsDeepSeekestimated: 15.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 15.7B | 4bit | 9.5 GB | +244.5 GB | ✓ fits | |
| bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUFOpen weightsbartowskiestimated: 15.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 15.7B | 4bit | 9.5 GB | +244.5 GB | ✓ fits | |
| amd/Qwen3.8-27B-Quark-AWQ-MXFP4Open weightsAMDestimated: 15.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 15.6B | 4bit | 9.5 GB | +244.5 GB | ✓ fits | |
| BAGEL-7B-MoTOpen weightsByteDanceestimated: 14.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 14.7B | 4bit | 8.9 GB | +245.1 GB | ✓ fits | |
| Phi 4Open weightsMicrosoftestimated: 14.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens | 14.7B | 4bit | 8.9 GB | +245.1 GB | ✓ fits |
Headroom = device memory − 2 GB reserve − estimated need. Sorted by the API (fitting models first). Compare the models you shortlist with the + buttons. Devices with these memory sizes: hardware listing.