Skip to content
AI Atlas
Hardware Fit tool

What fits my machine?

Pick a device memory size, a quantization and a context window: the atlas estimates each model's memory need from its parameter count. Every figure here is an estimate, never a measurement.

The custom field wins over the preset when both are filled. Estimates reserve 2 GB for the OS and framework.

EstimatedAll figures are estimates, not measurements.128 GB · 4-bit · 8k tokens575 of 654 models with a known parameter count fit

  • Estimated, not measured: weights = parameters × bytes/param × 1.15 runtime overhead.
  • bytes/param: 4bit = 0.5, 8bit = 1.0, fp16 = 2.0 (uniform quantization, no per-layer exceptions).
  • KV cache approximated at 0.5 GB per 8 192 tokens of context, independent of architecture (GQA/MLA models need less).
  • A model 'fits' when the estimate is at most the device memory minus 2 GB reserved for the OS and framework.
  • Mixture-of-experts models are estimated on total parameters (all experts must be resident); active parameters are ignored.
  • Device memory uses the largest configuration when several are listed (e.g. Apple silicon tiers).

Show only fitting modelsMethod →

Estimated hardware fit
ModelParamsQuantEst. memoryHeadroomFitsCompare
north-small-translate-1-0Cohereestimated: 218.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens218B4bit125.8 GB+0.1 GB✓ fits
amd/Qwen3.5-397B-A17B-MXFP4-AttnFP8-V2Open weightsAMDestimated: 203.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens203.2B4bit117.3 GB+8.7 GB✓ fits
unsloth/Qwen3.8-Flash-Next-GGUFOpen weightsUnslothestimated: 176.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens176.9B4bit102.2 GB+23.8 GB✓ fits
cerebras/MiniMax-M2.5-REAP-172B-A10BOpen weightsCerebras Systemsestimated: 172.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens172.5B4bit99.7 GB+26.3 GB✓ fits
cerebras/MiniMax-M2-REAP-172B-A10BOpen weightsCerebras Systemsestimated: 172.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens172.5B4bit99.7 GB+26.3 GB✓ fits
DeepSeek-V4-Flash-DSparkOpen weightsDeepSeekestimated: 165.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens165.3B4bit95.5 GB+30.5 GB✓ fits
cerebras/MiniMax-M2-REAP-162B-A10BOpen weightsCerebras Systemsestimated: 162.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens162B4bit93.6 GB+32.4 GB✓ fits
Step-3.5-Flash-REAP-149B-A11BOpen weightsCerebras Systemsestimated: 149.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens149.4B4bit86.4 GB+39.6 GB✓ fits
fireworks-ai/mixtral-8x22b-instruct-ohOpen weightsFireworks AIestimated: 140.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens140.6B4bit81.4 GB+44.6 GB✓ fits
cerebras/MiniMax-M2.1-REAP-139B-A10BOpen weightsCerebras Systemsestimated: 139.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens139.2B4bit80.5 GB+45.5 GB✓ fits
cerebras/MiniMax-M2.5-REAP-139B-A10BOpen weightsCerebras Systemsestimated: 139.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens139.2B4bit80.5 GB+45.5 GB✓ fits
cerebras/MiniMax-M2-REAP-139B-A10BOpen weightsCerebras Systemsestimated: 139.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens139.2B4bit80.5 GB+45.5 GB✓ fits
Mistral-Medium-3.5-128BOpen weightsMistral AIestimated: 127.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens127.7B4bit73.9 GB+52.1 GB✓ fits
NVIDIA-Nemotron-3-Super-120B-A12B-BF16Open weightsNVIDIAestimated: 123.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens123.6B4bit71.6 GB+54.4 GB✓ fits
Step-3.5-Flash-REAP-121B-A11BOpen weightsCerebras Systemsestimated: 121.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens121B4bit70.1 GB+55.9 GB✓ fits
amd/Qwen3-VL-235B-A22B-Instruct-MXFP4RestrictedAMDestimated: 118.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens118.8B4bit68.8 GB+57.2 GB✓ fits
gpt-oss-120bOpen weightsOpenAIestimated: 116.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens116.8B4bit67.7 GB+58.3 GB✓ fits
command-a-vision-07-2025RestrictedCohereestimated: 111.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens111.9B4bit64.8 GB+61.2 GB✓ fits
c4ai-command-a-03-2025RestrictedCohereestimated: 111.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens111.1B4bit64.4 GB+61.6 GB✓ fits
zai-org/GLM-4.5-Air-FP8Open weightsZ.ai (Zhipu AI)estimated: 110.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens110.5B4bit64.0 GB+62.0 GB✓ fits
GLM 4.5 AirOpen weightsZ.ai (Zhipu AI)estimated: 110.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens110.5B4bit64.0 GB+62.0 GB✓ fits
Llama 4 ScoutRestrictedMeta AIestimated: 108.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens108.6B4bit63.0 GB+63.0 GB✓ fits
Llama-3.2-90B-Vision-InstructRestrictedMeta AIestimated: 88.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens88.6B4bit51.4 GB+74.6 GB✓ fits
HunyuanImage-3.0-InstructOpen weightsTencentestimated: 83.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens83B4bit48.2 GB+77.8 GB✓ fits
cerebras/GLM-4.5-Air-REAP-82B-A12BOpen weightsCerebras Systemsestimated: 81.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens81.9B4bit47.6 GB+78.4 GB✓ fits
Hunyuan A13B InstructOpen weightsTencentestimated: 80.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens80.4B4bit46.7 GB+79.3 GB✓ fits
bartowski/Qwen_Qwen3-Next-80B-A3B-Thinking-GGUFOpen weightsbartowskiestimated: 80.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens80B4bit46.5 GB+79.5 GB✓ fits
Intel/Qwen3.8-Flash-Next-W4A16-AutoRoundOpen weightsIntelestimated: 75.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens75.4B4bit43.8 GB+82.2 GB✓ fits
UI-TARS-72B-DPOOpen weightsByteDanceestimated: 73.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens73.4B4bit42.7 GB+83.3 GB✓ fits
Kimi-Dev-72BOpen weightsMoonshot AIestimated: 72.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens72.7B4bit42.3 GB+83.7 GB✓ fits
NousResearch/Meta-Llama-3-70B-InstructOpen weightsNous Researchestimated: 70.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens70.5B4bit41.1 GB+84.9 GB✓ fits
fireworks-ai/llama-3-firefunction-v2Open weightsFireworks AIestimated: 70.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens70.5B4bit41.1 GB+84.9 GB✓ fits
Llama 3.3 70B InstructRestrictedMeta AIestimated: 70.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens70.5B4bit41.1 GB+84.9 GB✓ fits
Llama-3.1-70B-InstructRestrictedMeta AIestimated: 70.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens70.5B4bit41.1 GB+84.9 GB✓ fits
amd/Llama-3.3-70B-Instruct-FP8-KVOpen weightsAMDestimated: 70.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens70.5B4bit41.1 GB+84.9 GB✓ fits
Hermes-4-70BOpen weightsNous Researchestimated: 70.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens70.5B4bit41.1 GB+84.9 GB✓ fits
NousResearch/Meta-Llama-3.1-70B-InstructOpen weightsNous Researchestimated: 70.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens70.5B4bit41.1 GB+84.9 GB✓ fits
Meta-Llama-3-70BRestrictedMeta AIestimated: 70.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens70.5B4bit41.1 GB+84.9 GB✓ fits
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4Open weightsNVIDIAestimated: 67.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens67.2B4bit39.2 GB+86.8 GB✓ fits
nvidia/Qwen3.5-122B-A10B-NVFP4Open weightsNVIDIAestimated: 64.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens64.6B4bit37.6 GB+88.4 GB✓ fits
amd/gpt-oss-120b-w-mxfp4-a-fp8Open weightsAMDestimated: 59.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens59.5B4bit34.7 GB+91.3 GB✓ fits
ai21labs/AI21-Jamba-Mini-1.7-FP8RestrictedAI21 Labsestimated: 51.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens51.6B4bit30.2 GB+95.8 GB✓ fits
ai21labs/AI21-Jamba2-Mini-FP8Open weightsAI21 Labsestimated: 51.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens51.6B4bit30.2 GB+95.8 GB✓ fits
AI21-Jamba2-MiniOpen weightsAI21 Labsestimated: 51.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens51.6B4bit30.1 GB+95.8 GB✓ fits
AI21-Jamba-Mini-1.5RestrictedAI21 Labsestimated: 51.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens51.6B4bit30.1 GB+95.8 GB✓ fits
AI21-Jamba-Mini-1.7RestrictedAI21 Labsestimated: 51.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens51.6B4bit30.1 GB+95.8 GB✓ fits
AI21-Jamba-Mini-1.6RestrictedAI21 Labsestimated: 51.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens51.6B4bit30.1 GB+95.8 GB✓ fits
Jamba-v0.1Open weightsAI21 Labsestimated: 51.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens51.6B4bit30.1 GB+95.8 GB✓ fits
Kimi-Linear-48B-A3B-InstructOpen weightsMoonshot AIestimated: 49.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens49.1B4bit28.8 GB+97.3 GB✓ fits
Kimi-Linear-48B-A3B-BaseOpen weightsMoonshot AIestimated: 49.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens49.1B4bit28.8 GB+97.3 GB✓ fits
Nous-Hermes-2-Mixtral-8x7B-DPOOpen weightsNous Researchestimated: 46.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens46.7B4bit27.4 GB+98.7 GB✓ fits
firefunction-v1Open weightsFireworks AIestimated: 46.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens46.7B4bit27.4 GB+98.7 GB✓ fits
function-calling-v1Open weightsFireworks AIestimated: 46.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens46.7B4bit27.4 GB+98.7 GB✓ fits
Mixtral-8x7B-Instruct-v0.1Open weightsMistral AIestimated: 46.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens46.7B4bit27.4 GB+98.7 GB✓ fits
DFN2B-CLIP-ViT-L-14-39BOpen weightsAppleestimated: 39.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens39B4bit22.9 GB+103.1 GB✓ fits
Seed-OSS-36B-BaseOpen weightsByteDanceestimated: 36.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens36.1B4bit21.3 GB+104.7 GB✓ fits
NousResearch/Hermes-4.3-36B-GGUFOpen weightsNous Researchestimated: 36.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens36.1B4bit21.3 GB+104.7 GB✓ fits
Seed-OSS-36B-InstructOpen weightsByteDanceestimated: 36.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens36.1B4bit21.3 GB+104.7 GB✓ fits
Qwen/Qwen3.6-35B-A3B-FP8Open weightsQwenestimated: 36.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens36B4bit21.2 GB+104.8 GB✓ fits
Qwen3.6 35B A3BOpen weightsQwenestimated: 36.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens36B4bit21.2 GB+104.8 GB✓ fits
bartowski/endless-frontier_BigBang-v1-GGUFOpen weightsbartowskiestimated: 35.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens35.5B4bit20.9 GB+105.1 GB✓ fits
unsloth/Qwen3.6-35B-A3B-MTP-GGUFOpen weightsUnslothestimated: 35.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens35.5B4bit20.9 GB+105.1 GB✓ fits
cerebras/Kimi-Linear-REAP-35B-A3B-InstructOpen weightsCerebras Systemsestimated: 35.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens35.1B4bit20.7 GB+105.3 GB✓ fits
perplexity-ai/pplx-computer-qwen-3-6-35b-a3b-nvfp4-20260709Open weightsPerplexity AIestimated: 35.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens35.1B4bit20.7 GB+105.3 GB✓ fits
bartowski/Qwen_Qwen3.5-35B-A3B-GGUFOpen weightsbartowskiestimated: 35.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens35B4bit20.6 GB+105.4 GB✓ fits
bartowski/Qwen_Qwen3.6-35B-A3B-GGUFOpen weightsbartowskiestimated: 35.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens35B4bit20.6 GB+105.4 GB✓ fits
c4ai-command-r-v01RestrictedCohereestimated: 35.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens35B4bit20.6 GB+105.4 GB✓ fits
bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUFOpen weightsbartowskiestimated: 34.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens34.7B4bit20.4 GB+105.6 GB✓ fits
bartowski/thomsonreuters_Thomson-1.0-Small-GGUFOpen weightsbartowskiestimated: 34.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens34.7B4bit20.4 GB+105.6 GB✓ fits
bartowski/XYZAILab_XYZ-Aquila-mini-GGUFOpen weightsbartowskiestimated: 34.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens34.7B4bit20.4 GB+105.6 GB✓ fits
unsloth/Qwen3.6-35B-A3B-GGUFOpen weightsUnslothestimated: 34.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens34.7B4bit20.4 GB+105.6 GB✓ fits
perplexity-ai/pplx-computer-qwen-3-6-35b-a3b-mlx-20260709Open weightsPerplexity AIestimated: 34.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens34.7B4bit20.4 GB+105.6 GB✓ fits
Nous-Hermes-2-Yi-34BOpen weightsNous Researchestimated: 34.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens34.4B4bit20.3 GB+105.7 GB✓ fits
GKA-primed-HQwen3-32B-ReasonerOpen weightsAmazon Web Servicesestimated: 34.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens34.1B4bit20.1 GB+105.9 GB✓ fits
MiniMax-H3Open weightsMiniMaxestimated: 33.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens33.1B4bit19.6 GB+106.5 GB✓ fits
nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8Open weightsNVIDIAestimated: 33.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens33B4bit19.5 GB+106.5 GB✓ fits
SynLogic-32BOpen weightsMiniMaxestimated: 32.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens32.8B4bit19.3 GB+106.7 GB✓ fits
DeepSeek-R1-Distill-Qwen-32BOpen weightsDeepSeekestimated: 32.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens32.8B4bit19.3 GB+106.7 GB✓ fits
SynLogic-Mix-3-32BOpen weightsMiniMaxestimated: 32.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens32.8B4bit19.3 GB+106.7 GB✓ fits
Qwen3 32BOpen weightsQwenestimated: 32.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens32.8B4bit19.3 GB+106.7 GB✓ fits
aya-expanse-32bRestrictedCohereestimated: 32.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens32.3B4bit19.1 GB+106.9 GB✓ fits
c4ai-command-r-08-2024RestrictedCohereestimated: 32.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens32.3B4bit19.1 GB+106.9 GB✓ fits
OLMo-2-0325-32B-InstructOpen weightsAllen Institute for AIestimated: 32.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens32.2B4bit19.0 GB+107.0 GB✓ fits
FLUX.2-devRestrictedBlack Forest Labsestimated: 32.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens32.2B4bit19.0 GB+107.0 GB✓ fits
Nemotron 3 Nano 30B A3BOpen weightsNVIDIAestimated: 31.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens31.6B4bit18.7 GB+107.3 GB✓ fits
mlx-community/gemma-4-31b-it-4bitOpen weightsMLX Communityestimated: 31.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens31.3B4bit18.5 GB+107.5 GB✓ fits
Gemma 4 31BOpen weightsGoogleestimated: 31.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens31.3B4bit18.5 GB+107.5 GB✓ fits
GLM 4.7 FlashOpen weightsZ.ai (Zhipu AI)estimated: 31.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens31.2B4bit18.4 GB+107.5 GB✓ fits
unsloth/gemma-4-31B-it-qat-GGUFOpen weightsUnslothestimated: 30.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens30.7B4bit18.1 GB+107.8 GB✓ fits
browsesafeOpen weightsPerplexity AIestimated: 30.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens30.5B4bit18.1 GB+107.9 GB✓ fits
mlx-community/Qwen3-30B-A3B-Instruct-2507-4bitOpen weightsMLX Communityestimated: 30.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens30.5B4bit18.1 GB+107.9 GB✓ fits
unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUFOpen weightsUnslothestimated: 30.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens30.5B4bit18.1 GB+107.9 GB✓ fits
unsloth/Qwen3-30B-A3B-Thinking-2507-GGUFOpen weightsUnslothestimated: 30.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens30.5B4bit18.1 GB+107.9 GB✓ fits
North Mini Code (free)Open weightsCohereestimated: 30.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens30.5B4bit18.0 GB+108.0 GB✓ fits
Hy-MT2-30B-A3BOpen weightsTencentestimated: 30.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens30.1B4bit17.8 GB+108.2 GB✓ fits
north-mini-code-1-0Cohereestimated: 30.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens30B4bit17.8 GB+108.3 GB✓ fits
ibm-granite/granite-4.2-30b-GGUFOpen weightsIBMestimated: 30.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens30B4bit17.8 GB+108.3 GB✓ fits
ERNIE-4.5-VL-28B-A3B-ThinkingOpen weightsBaiduestimated: 29.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens29.7B4bit17.6 GB+108.4 GB✓ fits
ERNIE-4.5-VL-28B-A3B-PTOpen weightsBaiduestimated: 29.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens29.4B4bit17.4 GB+108.6 GB✓ fits
granite-4.1-30bOpen weightsIBMestimated: 28.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens28.9B4bit17.1 GB+108.9 GB✓ fits
Qwen/Qwen3.6-27B-FP8Open weightsQwenestimated: 27.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens27.8B4bit16.5 GB+109.5 GB✓ fits
Qwen3.6 27BOpen weightsQwenestimated: 27.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens27.8B4bit16.5 GB+109.5 GB✓ fits
Qwen3.8 27BOpen weightsQwenestimated: 27.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens27.8B4bit16.5 GB+109.5 GB✓ fits
Qwen/Qwen3.8-27B-FP8Open weightsQwenestimated: 27.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens27.8B4bit16.5 GB+109.5 GB✓ fits
mlx-community/Qwen3.8-27B-4bitOpen weightsMLX Communityestimated: 27.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens27.4B4bit16.2 GB+109.8 GB✓ fits
mlx-community/Qwen3.8-27B-8bitOpen weightsMLX Communityestimated: 27.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens27.4B4bit16.2 GB+109.8 GB✓ fits
unsloth/Qwen3.8-27B-GGUFOpen weightsUnslothestimated: 27.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens27.3B4bit16.2 GB+109.8 GB✓ fits
unsloth/Qwen3.6-27B-MTP-GGUFOpen weightsUnslothestimated: 27.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens27.3B4bit16.2 GB+109.8 GB✓ fits
perplexity-ai/pplx-computer-qwen-3-8-27b-dflash2-gguf-20260826Open weightsPerplexity AIestimated: 27.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens27.3B4bit16.2 GB+109.8 GB✓ fits
bartowski/Qwen3.8-27B-GGUFOpen weightsbartowskiestimated: 27.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens27B4bit16.0 GB+110.0 GB✓ fits
bartowski/Fara1.5-27B-GGUFOpen weightsbartowskiestimated: 27.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens27B4bit16.0 GB+110.0 GB✓ fits
unsloth/Qwen3.6-27B-GGUFOpen weightsUnslothestimated: 26.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens26.9B4bit16.0 GB+110.0 GB✓ fits
Gemma 4 26B A4BOpen weightsGoogleestimated: 25.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens25.8B4bit15.3 GB+110.7 GB✓ fits
unsloth/gemma-4-26B-A4B-it-GGUFOpen weightsUnslothestimated: 25.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens25.2B4bit15.0 GB+111.0 GB✓ fits
unsloth/gemma-4-26B-A4B-it-qat-GGUFOpen weightsUnslothestimated: 25.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens25.2B4bit15.0 GB+111.0 GB✓ fits
cerebras/Qwen3-Coder-REAP-25B-A3BOpen weightsCerebras Systemsestimated: 24.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens24.9B4bit14.8 GB+111.2 GB✓ fits
unsloth/Qwen3.6-35B-A3B-NVFP4Open weightsUnslothestimated: 24.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens24.6B4bit14.7 GB+111.3 GB✓ fits
Voxtral Small 24B 2507Open weightsMistral AIestimated: 24.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens24.3B4bit14.4 GB+111.5 GB✓ fits
Devstral-Small-2-24B-Instruct-2512Open weightsMistral AIestimated: 24.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens24B4bit14.3 GB+111.7 GB✓ fits
mlx-community/Devstral-Small-2-24B-Instruct-2512-4bitOpen weightsMLX Communityestimated: 24.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens24B4bit14.3 GB+111.7 GB✓ fits
Mistral Small 3.2 24BOpen weightsMistral AIestimated: 24.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens24B4bit14.3 GB+111.7 GB✓ fits
Mistral Small 3.1 24BOpen weightsMistral AIestimated: 24.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens24B4bit14.3 GB+111.7 GB✓ fits
cerebras/GLM-4.7-Flash-REAP-23B-A3BOpen weightsCerebras Systemsestimated: 23.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens23B4bit13.7 GB+112.3 GB✓ fits
ERNIE-4.5-21B-A3B-PTOpen weightsBaiduestimated: 21.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens21.9B4bit13.1 GB+112.9 GB✓ fits
ERNIE-4.5-21B-A3B-Base-PTOpen weightsBaiduestimated: 21.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens21.8B4bit13.1 GB+113.0 GB✓ fits
ERNIE-4.5-21B-A3B-ThinkingOpen weightsBaiduestimated: 21.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens21.8B4bit13.1 GB+113.0 GB✓ fits
gpt-oss-safeguard-20bOpen weightsOpenAIestimated: 21.5B params × 0.5 B × 1.15 + KV cache for 8192 tokens21.5B4bit12.9 GB+113.1 GB✓ fits
unsloth/Qwen3.6-27B-NVFP4Open weightsUnslothestimated: 21.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens21.2B4bit12.7 GB+113.3 GB✓ fits
gpt-oss-20bOpen weightsOpenAIestimated: 20.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens20.9B4bit12.5 GB+113.5 GB✓ fits
mlx-community/gpt-oss-20b-MXFP4-Q8Open weightsMLX Communityestimated: 20.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens20.9B4bit12.5 GB+113.5 GB✓ fits
nvidia/Gemma-4-31B-IT-NVFP4Open weightsNVIDIAestimated: 20.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens20.9B4bit12.5 GB+113.5 GB✓ fits
gpt-neox-20bOpen weightsEleutherAIestimated: 20.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens20.7B4bit12.4 GB+113.6 GB✓ fits
unsloth/MiniMax-H3-GGUFOpen weightsUnslothestimated: 20.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens20.1B4bit12.1 GB+113.9 GB✓ fits
internlm2-20bOpen weightsInternLM (Shanghai AI Laboratory)estimated: 20.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens20B4bit12.0 GB+114.0 GB✓ fits
internlm2-base-20bOpen weightsInternLM (Shanghai AI Laboratory)estimated: 20.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens20B4bit12.0 GB+114.0 GB✓ fits
internlm/internlm2_5-20b-chat-ggufOpen weightsInternLM (Shanghai AI Laboratory)estimated: 20.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens20B4bit12.0 GB+114.0 GB✓ fits
unsloth/Qwen3.8-27B-NVFP4Open weightsUnslothestimated: 19.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens19.9B4bit11.9 GB+114.1 GB✓ fits
internlm2-chat-20bOpen weightsInternLM (Shanghai AI Laboratory)estimated: 19.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens19.9B4bit11.9 GB+114.1 GB✓ fits
Intel/Qwen3.5-122B-A10B-int4-AutoRoundOpen weightsIntelestimated: 19.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens19.7B4bit11.8 GB+114.2 GB✓ fits
pplx-computer-qwen-3-8-27b-dflash2-20260824Open weightsPerplexity AIestimated: 18.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens18.8B4bit11.3 GB+114.7 GB✓ fits
pplx-qwen-3-8-27b-dflash2-20260819Open weightsPerplexity AIestimated: 18.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens18.8B4bit11.3 GB+114.7 GB✓ fits
nvidia/Qwen3.6-35B-A3B-NVFP4Open weightsNVIDIAestimated: 18.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens18.7B4bit11.2 GB+114.8 GB✓ fits
nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4Open weightsNVIDIAestimated: 18.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens18.3B4bit11.0 GB+115.0 GB✓ fits
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4Open weightsNVIDIAestimated: 17.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens17.8B4bit10.8 GB+115.3 GB✓ fits
Kimi-VL-A3B-Thinking-2506Open weightsMoonshot AIestimated: 16.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens16.4B4bit9.9 GB+116.1 GB✓ fits
Kimi-VL-A3B-ThinkingOpen weightsMoonshot AIestimated: 16.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens16.4B4bit9.9 GB+116.1 GB✓ fits
Kimi-VL-A3B-InstructOpen weightsMoonshot AIestimated: 16.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens16.4B4bit9.9 GB+116.1 GB✓ fits
Moonlight-16B-A3BOpen weightsMoonshot AIestimated: 16.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens16B4bit9.7 GB+116.3 GB✓ fits
Moonlight-16B-A3B-InstructOpen weightsMoonshot AIestimated: 16.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens16B4bit9.7 GB+116.3 GB✓ fits
DeepSeek-Coder-V2-Lite-InstructOpen weightsDeepSeekestimated: 15.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens15.7B4bit9.5 GB+116.5 GB✓ fits
DeepSeek-V2-LiteOpen weightsDeepSeekestimated: 15.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens15.7B4bit9.5 GB+116.5 GB✓ fits
bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUFOpen weightsbartowskiestimated: 15.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens15.7B4bit9.5 GB+116.5 GB✓ fits
amd/Qwen3.8-27B-Quark-AWQ-MXFP4Open weightsAMDestimated: 15.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens15.6B4bit9.5 GB+116.5 GB✓ fits
BAGEL-7B-MoTOpen weightsByteDanceestimated: 14.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens14.7B4bit8.9 GB+117.0 GB✓ fits
Phi 4Open weightsMicrosoftestimated: 14.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens14.7B4bit8.9 GB+117.1 GB✓ fits
nvidia/Gemma-4-26B-A4B-NVFP4Open weightsNVIDIAestimated: 14.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens14.4B4bit8.8 GB+117.2 GB✓ fits
Ministral 3 14B 2512Open weightsMistral AIestimated: 13.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens13.9B4bit8.5 GB+117.5 GB✓ fits
Ministral-3-14B-Reasoning-2512Open weightsMistral AIestimated: 13.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens13.9B4bit8.5 GB+117.5 GB✓ fits
FireLLaVA-13bRestrictedFireworks AIestimated: 13.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens13.3B4bit8.2 GB+117.8 GB✓ fits
gemma-4-12B-it-qat-w4a16-ctOpen weightsGoogleestimated: 13.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens13.3B4bit8.2 GB+117.8 GB✓ fits
Llama-2-13b-chat-hfRestrictedMeta AIestimated: 13.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens13B4bit8.0 GB+118.0 GB✓ fits
aya-101Open weightsCohereestimated: 12.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens12.9B4bit7.9 GB+118.1 GB✓ fits
Mistral-Nemo-Base-2407Open weightsMistral AIestimated: 12.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens12.3B4bit7.5 GB+118.5 GB✓ fits
Mistral NemoOpen weightsMistral AIestimated: 12.2B params × 0.5 B × 1.15 + KV cache for 8192 tokens12.3B4bit7.5 GB+118.5 GB✓ fits
bartowski/kai-os_Grug-12B-GGUFOpen weightsbartowskiestimated: 12.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens12B4bit7.4 GB+118.6 GB✓ fits
gemma-4-12B-itOpen weightsGoogleestimated: 12.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens12B4bit7.4 GB+118.6 GB✓ fits
unsloth/gemma-4-12B-it-qat-GGUFOpen weightsUnslothestimated: 11.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens11.9B4bit7.3 GB+118.7 GB✓ fits
FLUX.1-Fill-devRestrictedBlack Forest Labsestimated: 11.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens11.9B4bit7.3 GB+118.7 GB✓ fits
FLUX.1-Kontext-devRestrictedBlack Forest Labsestimated: 11.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens11.9B4bit7.3 GB+118.7 GB✓ fits
FLUX.1-devRestrictedBlack Forest Labsestimated: 11.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens11.9B4bit7.3 GB+118.7 GB✓ fits
FLUX.1-Krea-devRestrictedBlack Forest Labsestimated: 11.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens11.9B4bit7.3 GB+118.7 GB✓ fits
FLUX.1-schnellRestrictedBlack Forest Labsestimated: 11.9B params × 0.5 B × 1.15 + KV cache for 8192 tokens11.9B4bit7.3 GB+118.7 GB✓ fits
Intel/Qwen3-Coder-Next-int4-AutoRoundOpen weightsIntelestimated: 11.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens11.8B4bit7.3 GB+118.7 GB✓ fits
KaLM-Embedding-Gemma3-12B-2511Open weightsTencentestimated: 11.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens11.8B4bit7.3 GB+118.7 GB✓ fits
Nous-Hermes-2-SOLAR-10.7BOpen weightsNous Researchestimated: 10.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens10.7B4bit6.7 GB+119.3 GB✓ fits
Llama-3.2-11B-Vision-InstructRestrictedMeta AIestimated: 10.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens10.7B4bit6.6 GB+119.4 GB✓ fits
GLM-4.1V-9B-ThinkingOpen weightsZ.ai (Zhipu AI)estimated: 10.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens10.3B4bit6.4 GB+119.6 GB✓ fits
GLM-4.6V-FlashOpen weightsZ.ai (Zhipu AI)estimated: 10.3B params × 0.5 B × 1.15 + KV cache for 8192 tokens10.3B4bit6.4 GB+119.6 GB✓ fits
Kimi-Audio-7B-InstructOpen weightsMoonshot AIestimated: 9.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens9.77B4bit6.1 GB+119.9 GB✓ fits
Kimi-Audio-7BOpen weightsMoonshot AIestimated: 9.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens9.77B4bit6.1 GB+119.9 GB✓ fits
Qwen3.5-9BOpen weightsQwenestimated: 9.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens9.65B4bit6.0 GB+120.0 GB✓ fits
glm-4-9b-chatOpen weightsZ.ai (Zhipu AI)estimated: 9.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens9.4B4bit5.9 GB+120.1 GB✓ fits
academic-ds-9BOpen weightsByteDanceestimated: 9.4B params × 0.5 B × 1.15 + KV cache for 8192 tokens9.37B4bit5.9 GB+120.1 GB✓ fits
FLUX.2-klein-9BRestrictedBlack Forest Labsestimated: 9.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens9.08B4bit5.7 GB+120.3 GB✓ fits
FLUX.2-klein-9b-kvRestrictedBlack Forest Labsestimated: 9.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens9.08B4bit5.7 GB+120.3 GB✓ fits
FLUX.2-klein-base-9BRestrictedBlack Forest Labsestimated: 9.1B params × 0.5 B × 1.15 + KV cache for 8192 tokens9.08B4bit5.7 GB+120.3 GB✓ fits
bartowski/Fara1.5-9B-GGUFOpen weightsbartowskiestimated: 9.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens9B4bit5.7 GB+120.3 GB✓ fits
black-forest-labs/FLUX.2-klein-9b-fp8RestrictedBlack Forest Labsestimated: 9.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens9B4bit5.7 GB+120.3 GB✓ fits
bartowski/tencent_UI-Mate-9B-GGUFOpen weightsbartowskiestimated: 9.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens9B4bit5.7 GB+120.3 GB✓ fits
black-forest-labs/FLUX.2-klein-9b-kv-fp8Open weightsBlack Forest Labsestimated: 9.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens9B4bit5.7 GB+120.3 GB✓ fits
black-forest-labs/FLUX.2-klein-base-9b-fp8RestrictedBlack Forest Labsestimated: 9.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens9B4bit5.7 GB+120.3 GB✓ fits
unsloth/Qwen3.5-9B-GGUFOpen weightsUnslothestimated: 9.0B params × 0.5 B × 1.15 + KV cache for 8192 tokens8.95B4bit5.7 GB+120.3 GB✓ fits
Qianfan-VL-8BOpen weightsBaiduestimated: 8.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens8.81B4bit5.6 GB+120.4 GB✓ fits
internlm3-8b-instructOpen weightsInternLM (Shanghai AI Laboratory)estimated: 8.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens8.8B4bit5.6 GB+120.4 GB✓ fits
granite-4.1-8bOpen weightsIBMestimated: 8.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens8.79B4bit5.6 GB+120.4 GB✓ fits
amd/Qwen3-VL-8B-Instruct-w8a8-llmcompressorOpen weightsAMDestimated: 8.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens8.77B4bit5.5 GB+120.5 GB✓ fits
Qwen3 VL 8B InstructOpen weightsQwenestimated: 8.8B params × 0.5 B × 1.15 + KV cache for 8192 tokens8.77B4bit5.5 GB+120.5 GB✓ fits
VibeVoice-ASROpen weightsMicrosoftestimated: 8.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens8.67B4bit5.5 GB+120.5 GB✓ fits
Molmo2-8BOpen weightsAllen Institute for AIestimated: 8.7B params × 0.5 B × 1.15 + KV cache for 8192 tokens8.66B4bit5.5 GB+120.5 GB✓ fits
aya-vision-8bRestrictedCohereestimated: 8.6B params × 0.5 B × 1.15 + KV cache for 8192 tokens8.63B4bit5.5 GB+120.5 GB✓ fits

Headroom = device memory − 2 GB reserve − estimated need. Sorted by the API (fitting models first). Compare the models you shortlist with the + buttons. Devices with these memory sizes: hardware listing.