Skip to content
AI Atlas
Local AI Explorer

Run locally

Describe the machine — memory, number of devices, platform, quantization, context, batch — and the atlas lists the downloadable models estimated to fit, with the breakdown behind each number and the GGUF / MLX artifacts recorded for them. Estimates, never measurements.

EstimatedAll fit figures are estimates, not measurements.454 of 529 evaluated models fit

Every figure is an ESTIMATE: weights = params × bytes/param (× 1.15 overhead) unless an artifact's observed file size is available; KV cache uses architecture metadata when known, else 0.5 GB per 8K tokens × batch.

Assumptions (7)
  • Estimated, not measured: weights = parameters × bytes/param × 1.15 runtime overhead (or the observed artifact file size when one is recorded).
  • bytes/param: 4bit = 0.5, 8bit = 1.0, fp16 = 2.0 (uniform quantization, no per-layer exceptions).
  • KV cache: 2 × layers × kv_heads × head_dim × 2 bytes × context × batch when the architecture is known; otherwise 0.5 GB per 8 192 tokens (× batch), independent of architecture (GQA/MLA models need less).
  • A model 'fits' when the estimate is at most the device memory minus 2 GB reserved for the OS and framework.
  • Mixture-of-experts models are estimated on total parameters (all experts must be resident); active parameters are ignored.
  • Device memory uses the largest configuration when several are listed (e.g. Apple silicon tiers).
  • Multi-GPU: device memories are summed; interconnect bandwidth, tensor-parallel replication and pipeline bubbles are not modelled.

64 GB total1 × 64 GB4-bit8K contextbatch 1Show all evaluated models

ModelParamsEst. memoryHeadroomFitsBreakdownArtifacts
Llama-3.2-90B-Vision-Instructrestricted-weightsMeta AILlama-3.2-Community88.6B51.4 GB estimated+10.6 GB✓ fits
  • weights 44.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 6.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
HunyuanImage-3.0-InstructOpen weightsTencentOther83B48.2 GB estimated+13.8 GB✓ fits
  • weights 41.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 6.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
cerebras/GLM-4.5-Air-REAP-82B-A12BOpen weightsCerebras SystemsMIT81.9B47.6 GB estimated+14.4 GB✓ fits
  • weights 41.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 6.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Hunyuan A13B InstructOpen weightsTencentOther80.4B46.7 GB estimated+15.3 GB✓ fits
  • weights 40.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 6.0 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
UI-TARS-72B-DPOOpen weightsByteDanceApache-2.073.4B42.7 GB estimated+19.3 GB✓ fits
  • weights 36.7 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 5.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Kimi-Dev-72BOpen weightsMoonshot AIMIT72.7B42.3 GB estimated+19.7 GB✓ fits
  • weights 36.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 5.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Qwen2.5restricted-weightsQwen TeamResearch-Only72.7B42.3 GB estimated+19.7 GB✓ fits
  • weights 36.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 5.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
QVQ-72B-PreviewOpen weightsQwen Team72B41.9 GB estimated+20.1 GB✓ fits
  • weights 36.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 5.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Hermes-4-70Brestricted-weightsNous ResearchLlama-3-Community70.5B41.1 GB estimated+20.9 GB✓ fits
  • weights 35.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 5.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Llama 3.1 70B Instructrestricted-weightsMeta AILlama-3.1-Community70.5B41.1 GB estimated+20.9 GB✓ fits
  • weights 35.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 5.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
NousResearch/Meta-Llama-3.1-70B-Instructrestricted-weightsNous ResearchLlama-3.1-Community70.5B41.1 GB estimated+20.9 GB✓ fits
  • weights 35.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 5.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Meta-Llama-3-70Brestricted-weightsMeta AILlama-3-Community70.5B41.1 GB estimated+20.9 GB✓ fits
  • weights 35.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 5.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
fireworks-ai/llama-3-firefunction-v2restricted-weightsFireworks AILlama-3-Community70.5B41.1 GB estimated+20.9 GB✓ fits
  • weights 35.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 5.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Llama 3.3 70B Instructrestricted-weightsMeta AILlama-3.3-Community70.5B41.1 GB estimated+20.9 GB✓ fits
  • weights 35.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 5.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
2 recorded
amd/Llama-3.3-70B-Instruct-FP8-KVfp872.7 GB observed84.1 GB−22.1 GB✗ too large
  • weights 72.7 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 10.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
amd/Llama-3.3-70B-Instruct-FP8-KVfp872.7 GB observed84.1 GB−22.1 GB✗ too large
  • weights 72.7 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 10.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
NousResearch/Meta-Llama-3-70B-InstructOpen weightsNous ResearchOther70.5B41.1 GB estimated+20.9 GB✓ fits
  • weights 35.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 5.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Jamba-v0.1Open weightsAI21 LabsApache-2.051.6B30.2 GB estimated+31.8 GB✓ fits
  • weights 25.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 3.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
AI21-Jamba-Mini-1.5Open weightsAI21 LabsOther51.6B30.2 GB estimated+31.8 GB✓ fits
  • weights 25.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 3.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
AI21-Jamba-Mini-1.7Open weightsAI21 LabsOther51.6B30.2 GB estimated+31.8 GB✓ fits
  • weights 25.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 3.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
1 recorded
ai21labs/AI21-Jamba-Mini-1.7-FP855.1 GB observed63.8 GB−1.8 GB✗ too large
  • weights 55.1 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 8.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
AI21-Jamba-Mini-1.6Open weightsAI21 LabsOther51.6B30.2 GB estimated+31.8 GB✓ fits
  • weights 25.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 3.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
AI21-Jamba2-MiniOpen weightsAI21 LabsApache-2.051.6B30.2 GB estimated+31.8 GB✓ fits
  • weights 25.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 3.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
1 recorded
ai21labs/AI21-Jamba2-Mini-FP855.1 GB observed63.8 GB−1.8 GB✗ too large
  • weights 55.1 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 8.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
Kimi-Linear-48B-A3B-InstructOpen weightsMoonshot AIMIT49.1B28.7 GB estimated+33.3 GB✓ fits
  • weights 24.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 3.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Kimi-Linear-48B-A3B-BaseOpen weightsMoonshot AIMIT49.1B28.7 GB estimated+33.3 GB✓ fits
  • weights 24.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 3.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
firefunction-v1Open weightsFireworks AIApache-2.046.7B27.4 GB estimated+34.6 GB✓ fits
  • weights 23.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 3.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
function-calling-v1Open weightsFireworks AI46.7B27.4 GB estimated+34.6 GB✓ fits
  • weights 23.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 3.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Nous-Hermes-2-Mixtral-8x7B-DPOOpen weightsNous ResearchApache-2.046.7B27.4 GB estimated+34.6 GB✓ fits
  • weights 23.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 3.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Mixtral-8x7B-Instruct-v0.1Open weightsMistral AIApache-2.046.7B27.4 GB estimated+34.6 GB✓ fits
  • weights 23.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 3.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
DFN2B-CLIP-ViT-L-14-39Brestricted-weightsAppleApple-AMLR39B22.9 GB estimated+39.1 GB✓ fits
  • weights 19.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Seed-OSS-36B-BaseOpen weightsByteDanceApache-2.036.1B21.3 GB estimated+40.7 GB✓ fits
  • weights 18.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
1 recorded
NousResearch/Hermes-4.3-36B-GGUFgguf estimated21.3 GB+40.7 GB✓ fits
  • weights 18.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
Seed-OSS-36B-InstructOpen weightsByteDanceApache-2.036.1B21.3 GB estimated+40.7 GB✓ fits
  • weights 18.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Qwen3.6 35B A3BOpen weightsQwenApache-2.036B21.2 GB estimated+40.8 GB✓ fits
  • weights 18.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
11 recorded
Intel/Qwen3.6-35B-A3B-int4-mixed-AutoRound21.5 GB observed25.2 GB+36.8 GB✓ fits
  • weights 21.5 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 3.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
nvidia/Qwen3.6-35B-A3B-NVFP423.4 GB observed27.5 GB+34.5 GB✓ fits
  • weights 23.4 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 3.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
unsloth/Qwen3.6-35B-A3B-NVFP426.5 GB observed31.0 GB+31.0 GB✓ fits
  • weights 26.5 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 4.0 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
Qwen/Qwen3.6-35B-A3B-FP8fp837.5 GB observed43.6 GB+18.4 GB✓ fits
  • weights 37.5 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 5.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
+4 more artifacts on the model page.
cerebras/Kimi-Linear-REAP-35B-A3B-InstructOpen weightsCerebras SystemsMIT35.1B20.7 GB estimated+41.3 GB✓ fits
  • weights 17.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
c4ai-command-r-v01restricted-weightsCohereCC-BY-NC-4.035B20.6 GB estimated+41.4 GB✓ fits
  • weights 17.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Nous-Hermes-2-Yi-34BOpen weightsNous ResearchApache-2.034.4B20.3 GB estimated+41.7 GB✓ fits
  • weights 17.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
GKA-primed-HQwen3-32B-ReasonerOpen weightsAmazon Web ServicesApache-2.034.1B20.1 GB estimated+41.9 GB✓ fits
  • weights 17.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
MiniMax-H3Open weightsMiniMaxOther33.1B19.5 GB estimated+42.5 GB✓ fits
  • weights 16.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
2 recorded
unsloth/MiniMax-H3-GGUFgguf estimated19.5 GB+42.5 GB✓ fits
  • weights 16.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
unsloth/MiniMax-H3-GGUFgguf estimated19.5 GB+42.5 GB✓ fits
  • weights 16.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
DeepSeek-R1-Distill-Qwen-32BOpen weightsDeepSeekMIT32.8B19.3 GB estimated+42.7 GB✓ fits
  • weights 16.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
SynLogic-32BOpen weightsMiniMaxMIT32.8B19.3 GB estimated+42.7 GB✓ fits
  • weights 16.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
SynLogic-Mix-3-32BOpen weightsMiniMaxMIT32.8B19.3 GB estimated+42.7 GB✓ fits
  • weights 16.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Qwen3 32BOpen weightsQwenApache-2.032.8B19.3 GB estimated+42.7 GB✓ fits
  • weights 16.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Qwen2.5-Coderrestricted-weightsQwen TeamResearch-Only32.5B19.2 GB estimated+42.8 GB✓ fits
  • weights 16.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
aya-expanse-32brestricted-weightsCohereCC-BY-NC-4.032.3B19.1 GB estimated+42.9 GB✓ fits
  • weights 16.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
c4ai-command-r-08-2024restricted-weightsCohereCC-BY-NC-4.032.3B19.1 GB estimated+42.9 GB✓ fits
  • weights 16.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
OLMo-2-0325-32B-InstructOpen weightsAllen Institute for AIApache-2.032.2B19.0 GB estimated+43.0 GB✓ fits
  • weights 16.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
FLUX.2-devOpen weightsBlack Forest LabsOther32.2B19.0 GB estimated+43.0 GB✓ fits
  • weights 16.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
1 recorded
black-forest-labs/FLUX.2-dev-NVFP4 estimated19.0 GB+43.0 GB✓ fits
  • weights 16.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
QwQ-32BOpen weightsAlibaba GroupApache-2.032B18.9 GB estimated+43.1 GB✓ fits
  • weights 16.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
QwQ-32B-PreviewOpen weightsAlibaba Group32B18.9 GB estimated+43.1 GB✓ fits
  • weights 16.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Qwen2.5-VL-32B-InstructOpen weightsQwen TeamApache-2.032B18.9 GB estimated+43.1 GB✓ fits
  • weights 16.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Nemotron 3 Nano 30B A3BOpen weightsNVIDIAOther31.6B18.7 GB estimated+43.3 GB✓ fits
  • weights 15.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Gemma 4 31BOpen weightsGoogleApache-2.031.3B18.5 GB estimated+43.5 GB✓ fits
  • weights 15.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
5 recorded
mlx-community/gemma-4-31b-it-4bitmlx18.4 GB observed21.7 GB+40.3 GB✓ fits
  • weights 18.4 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 2.8 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
mlx-community/gemma-4-31b-it-4bitmlx18.4 GB observed21.7 GB+40.3 GB✓ fits
  • weights 18.4 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 2.8 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
Intel/gemma-4-31B-it-int4-AutoRound19.2 GB observed22.6 GB+39.4 GB✓ fits
  • weights 19.2 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 2.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
nvidia/Gemma-4-31B-IT-NVFP432.6 GB observed38.0 GB+24.0 GB✓ fits
  • weights 32.6 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 4.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
+1 more artifacts on the model page.
GLM 4.7 FlashOpen weightsZ.ai (Zhipu AI)MIT31.2B18.4 GB estimated+43.5 GB✓ fits
  • weights 15.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
browsesafeOpen weightsPerplexity AIMIT30.5B18.1 GB estimated+43.9 GB✓ fits
  • weights 15.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
North Mini Code (free)Open weightsCohereApache-2.030.5B18.0 GB estimated+44.0 GB✓ fits
  • weights 15.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Hy-MT2-30B-A3BOpen weightsTencentApache-2.030.1B17.8 GB estimated+44.2 GB✓ fits
  • weights 15.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
ERNIE-4.5-VL-28B-A3B-ThinkingOpen weightsBaiduApache-2.029.7B17.6 GB estimated+44.5 GB✓ fits
  • weights 14.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
ERNIE-4.5-VL-28B-A3B-PTOpen weightsBaiduApache-2.029.4B17.4 GB estimated+44.6 GB✓ fits
  • weights 14.7 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
granite-4.1-30bOpen weightsIBMApache-2.028.9B17.1 GB estimated+44.9 GB✓ fits
  • weights 14.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Qwen3.8 27BOpen weightsQwenApache-2.027.8B16.5 GB estimated+45.5 GB✓ fits
  • weights 13.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
16 recorded
mlx-community/Qwen3.8-27B-4bitmlx16.1 GB observed19.0 GB+43.0 GB✓ fits
  • weights 16.1 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 2.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
mlx-community/Qwen3.8-27B-4bitmlx16.1 GB observed19.0 GB+43.0 GB✓ fits
  • weights 16.1 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 2.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
amd/Qwen3.8-27B-Quark-AWQ-INT4-W4A16awq19.5 GB observed22.9 GB+39.1 GB✓ fits
  • weights 19.5 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 2.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
amd/Qwen3.8-27B-Quark-AWQ-INT4-W4A16awq19.5 GB observed22.9 GB+39.1 GB✓ fits
  • weights 19.5 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 2.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
+4 more artifacts on the model page.
Qwen3.6 27BOpen weightsQwenApache-2.027.8B16.5 GB estimated+45.5 GB✓ fits
  • weights 13.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 2.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
8 recorded
Intel/Qwen3.6-27B-int4-AutoRound19.0 GB observed22.4 GB+39.6 GB✓ fits
  • weights 19.0 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 2.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
unsloth/Qwen3.6-27B-NVFP423.4 GB observed27.4 GB+34.6 GB✓ fits
  • weights 23.4 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 3.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
Qwen/Qwen3.6-27B-FP8fp830.9 GB observed36.0 GB+26.0 GB✓ fits
  • weights 30.9 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 4.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
Qwen/Qwen3.6-27B-FP8fp830.9 GB observed36.0 GB+26.0 GB✓ fits
  • weights 30.9 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 4.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
+4 more artifacts on the model page.
Gemma 4 26B A4BOpen weightsGoogleApache-2.025.8B15.3 GB estimated+46.7 GB✓ fits
  • weights 12.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4 recorded
nvidia/Gemma-4-26B-A4B-NVFP418.8 GB observed22.1 GB+39.9 GB✓ fits
  • weights 18.8 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 2.8 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
unsloth/gemma-4-26B-A4B-it-qat-GGUFgguf estimated15.3 GB+46.7 GB✓ fits
  • weights 12.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
unsloth/gemma-4-26B-A4B-it-GGUFgguf estimated15.3 GB+46.7 GB✓ fits
  • weights 12.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
unsloth/gemma-4-26B-A4B-it-GGUFgguf estimated15.3 GB+46.7 GB✓ fits
  • weights 12.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
cerebras/Qwen3-Coder-REAP-25B-A3BOpen weightsCerebras SystemsApache-2.024.9B14.8 GB estimated+47.2 GB✓ fits
  • weights 12.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Voxtral Small 24B 2507Open weightsMistral AIApache-2.024.3B14.4 GB estimated+47.5 GB✓ fits
  • weights 12.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.8 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Devstral-Small-2-24B-Instruct-2512Open weightsMistral AIApache-2.024B14.3 GB estimated+47.7 GB✓ fits
  • weights 12.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.8 GB
  • reserved 2 GB
  • context 8,192 × batch 1
2 recorded
mlx-community/Devstral-Small-2-24B-Instruct-2512-4bitmlx15.1 GB observed17.9 GB+44.1 GB✓ fits
  • weights 15.1 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 2.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
mlx-community/Devstral-Small-2-24B-Instruct-2512-4bitmlx15.1 GB observed17.9 GB+44.1 GB✓ fits
  • weights 15.1 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 2.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
Mistral Small 3.1 24BOpen weightsMistral AIApache-2.024B14.3 GB estimated+47.7 GB✓ fits
  • weights 12.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.8 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Mistral Small 3.2 24BOpen weightsMistral AIApache-2.024B14.3 GB estimated+47.7 GB✓ fits
  • weights 12.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.8 GB
  • reserved 2 GB
  • context 8,192 × batch 1
1 recorded
Intel/Mistral-Small-3.2-24B-Instruct-2506-int4-AutoRound15.1 GB observed17.9 GB+44.1 GB✓ fits
  • weights 15.1 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 2.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
cerebras/GLM-4.7-Flash-REAP-23B-A3BOpen weightsCerebras SystemsMIT23B13.7 GB estimated+48.3 GB✓ fits
  • weights 11.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
ERNIE-4.5-21B-A3B-PTOpen weightsBaiduApache-2.021.9B13.1 GB estimated+48.9 GB✓ fits
  • weights 11.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
ERNIE-4.5-21B-A3B-Base-PTOpen weightsBaiduApache-2.021.8B13.1 GB estimated+49.0 GB✓ fits
  • weights 10.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
ERNIE-4.5-21B-A3B-ThinkingOpen weightsBaiduApache-2.021.8B13.1 GB estimated+49.0 GB✓ fits
  • weights 10.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
gpt-oss-safeguard-20bOpen weightsOpenAIApache-2.021.5B12.9 GB estimated+49.1 GB✓ fits
  • weights 10.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
gpt-oss-20bOpen weightsOpenAIApache-2.020.9B12.5 GB estimated+49.5 GB✓ fits
  • weights 10.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
2 recorded
mlx-community/gpt-oss-20b-MXFP4-Q8mlx12.1 GB observed14.4 GB+47.6 GB✓ fits
  • weights 12.1 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 1.8 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
mlx-community/gpt-oss-20b-MXFP4-Q8mlx12.1 GB observed14.4 GB+47.6 GB✓ fits
  • weights 12.1 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 1.8 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
gpt-neox-20bOpen weightsEleutherAIApache-2.020.7B12.4 GB estimated+49.6 GB✓ fits
  • weights 10.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
GPT-NeoXT-Chat-Base-20BOpen weightsTogether AIApache-2.020B12.0 GB estimated+50.0 GB✓ fits
  • weights 10.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
internlm2-base-20bOpen weightsInternLM (Shanghai AI Laboratory)Other20B12.0 GB estimated+50.0 GB✓ fits
  • weights 10.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
internlm2-20bOpen weightsInternLM (Shanghai AI Laboratory)Other20B12.0 GB estimated+50.0 GB✓ fits
  • weights 10.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
internlm2-chat-20bOpen weightsInternLM (Shanghai AI Laboratory)Other19.9B11.9 GB estimated+50.1 GB✓ fits
  • weights 9.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
pplx-qwen-3-8-27b-dflash2-20260819Open weightsPerplexity AIOther18.8B11.3 GB estimated+50.7 GB✓ fits
  • weights 9.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
pplx-computer-qwen-3-8-27b-dflash2-20260824Open weightsPerplexity AIOther18.8B11.3 GB estimated+50.7 GB✓ fits
  • weights 9.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Kimi-VL-A3B-ThinkingOpen weightsMoonshot AIMIT16.4B9.9 GB estimated+52.1 GB✓ fits
  • weights 8.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Kimi-VL-A3B-InstructOpen weightsMoonshot AIMIT16.4B9.9 GB estimated+52.1 GB✓ fits
  • weights 8.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Kimi-VL-A3B-Thinking-2506Open weightsMoonshot AIMIT16.4B9.9 GB estimated+52.1 GB✓ fits
  • weights 8.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Moonlight-16B-A3B-InstructOpen weightsMoonshot AIMIT16B9.7 GB estimated+52.3 GB✓ fits
  • weights 8.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Moonlight-16B-A3BOpen weightsMoonshot AIMIT16B9.7 GB estimated+52.3 GB✓ fits
  • weights 8.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
DeepSeek-Coder-V2-Lite-InstructOpen weightsDeepSeekOther15.7B9.5 GB estimated+52.5 GB✓ fits
  • weights 7.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
2 recorded
bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUFgguf estimated9.5 GB+52.5 GB✓ fits
  • weights 7.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUFgguf estimated9.5 GB+52.5 GB✓ fits
  • weights 7.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
DeepSeek-V2-LiteOpen weightsDeepSeekOther15.7B9.5 GB estimated+52.5 GB✓ fits
  • weights 7.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
BAGEL-7B-MoTOpen weightsByteDanceApache-2.014.7B8.9 GB estimated+53.0 GB✓ fits
  • weights 7.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Phi 4Open weightsMicrosoftMIT14.7B8.9 GB estimated+53.1 GB✓ fits
  • weights 7.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Ministral 3 14B 2512Open weightsMistral AIApache-2.013.9B8.5 GB estimated+53.5 GB✓ fits
  • weights 7.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Ministral-3-14B-Reasoning-2512Open weightsMistral AIApache-2.013.9B8.5 GB estimated+53.5 GB✓ fits
  • weights 7.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
FireLLaVA-13brestricted-weightsFireworks AILlama-2-Community13.3B8.2 GB estimated+53.8 GB✓ fits
  • weights 6.7 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.0 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Llama-2-13b-chat-hfrestricted-weightsMeta AILlama-2-Community13B8.0 GB estimated+54.0 GB✓ fits
  • weights 6.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.0 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
aya-101Open weightsCohereApache-2.012.9B7.9 GB estimated+54.1 GB✓ fits
  • weights 6.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 1.0 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Mistral-Nemo-Base-2407Open weightsMistral AIApache-2.012.3B7.5 GB estimated+54.5 GB✓ fits
  • weights 6.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Mistral NemoOpen weightsMistral AIApache-2.012.3B7.5 GB estimated+54.5 GB✓ fits
  • weights 6.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
gemma-4-12B-itOpen weightsGoogleApache-2.012B7.4 GB estimated+54.6 GB✓ fits
  • weights 6.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
FLUX.1-Fill-devOpen weightsBlack Forest LabsOther11.9B7.3 GB estimated+54.7 GB✓ fits
  • weights 6.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
FLUX.1-Kontext-devOpen weightsBlack Forest LabsOther11.9B7.3 GB estimated+54.7 GB✓ fits
  • weights 6.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
FLUX.1-Krea-devOpen weightsBlack Forest LabsOther11.9B7.3 GB estimated+54.7 GB✓ fits
  • weights 6.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
FLUX.1-devOpen weightsBlack Forest LabsOther11.9B7.3 GB estimated+54.7 GB✓ fits
  • weights 6.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
FLUX.1-schnellOpen weightsBlack Forest LabsApache-2.011.9B7.3 GB estimated+54.7 GB✓ fits
  • weights 6.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
KaLM-Embedding-Gemma3-12B-2511Open weightsTencentOther11.8B7.3 GB estimated+54.7 GB✓ fits
  • weights 5.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.9 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Nous-Hermes-2-SOLAR-10.7BOpen weightsNous ResearchApache-2.010.7B6.7 GB estimated+55.3 GB✓ fits
  • weights 5.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.8 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Llama-3.2-11B-Vision-Instructrestricted-weightsMeta AILlama-3.2-Community10.7B6.6 GB estimated+55.4 GB✓ fits
  • weights 5.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.8 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
GLM-4.6V-FlashOpen weightsZ.ai (Zhipu AI)MIT10.3B6.4 GB estimated+55.6 GB✓ fits
  • weights 5.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.8 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
GLM-4.1V-9B-ThinkingOpen weightsZ.ai (Zhipu AI)MIT10.3B6.4 GB estimated+55.6 GB✓ fits
  • weights 5.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.8 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Kimi-Audio-7BOpen weightsMoonshot AIMIT9.77B6.1 GB estimated+55.9 GB✓ fits
  • weights 4.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Kimi-Audio-7B-InstructOpen weightsMoonshot AIMIT9.77B6.1 GB estimated+55.9 GB✓ fits
  • weights 4.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Qwen3.5-9BOpen weightsQwenApache-2.09.65B6.0 GB estimated+56.0 GB✓ fits
  • weights 4.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
3 recorded
Intel/Qwen3.5-9B-int4-AutoRound9.0 GB observed10.8 GB+51.2 GB✓ fits
  • weights 9.0 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 1.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
unsloth/Qwen3.5-9B-GGUFgguf estimated6.0 GB+56.0 GB✓ fits
  • weights 4.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
unsloth/Qwen3.5-9B-GGUFgguf estimated6.0 GB+56.0 GB✓ fits
  • weights 4.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
glm-4-9b-chatOpen weightsZ.ai (Zhipu AI)Other9.4B5.9 GB estimated+56.1 GB✓ fits
  • weights 4.7 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
academic-ds-9BOpen weightsByteDanceApache-2.09.37B5.9 GB estimated+56.1 GB✓ fits
  • weights 4.7 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
FLUX.2-klein-9BOpen weightsBlack Forest LabsOther9.08B5.7 GB estimated+56.3 GB✓ fits
  • weights 4.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
FLUX.2-klein-base-9BOpen weightsBlack Forest LabsOther9.08B5.7 GB estimated+56.3 GB✓ fits
  • weights 4.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
FLUX.2-klein-9b-kvOpen weightsBlack Forest LabsOther9.08B5.7 GB estimated+56.3 GB✓ fits
  • weights 4.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
1 recorded
black-forest-labs/FLUX.2-klein-9b-kv-fp8 estimated5.7 GB+56.3 GB✓ fits
  • weights 4.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
Qianfan-VL-8BOpen weightsBaiduOther8.81B5.6 GB estimated+56.4 GB✓ fits
  • weights 4.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
internlm3-8b-instructOpen weightsInternLM (Shanghai AI Laboratory)Apache-2.08.8B5.6 GB estimated+56.4 GB✓ fits
  • weights 4.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
2 recorded
internlm/internlm3-8b-instruct-ggufgguf estimated5.6 GB+56.4 GB✓ fits
  • weights 4.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
internlm/internlm3-8b-instruct-ggufgguf estimated5.6 GB+56.4 GB✓ fits
  • weights 4.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
granite-4.1-8bOpen weightsIBMApache-2.08.79B5.6 GB estimated+56.4 GB✓ fits
  • weights 4.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Qwen3 VL 8B InstructOpen weightsQwenApache-2.08.77B5.5 GB estimated+56.5 GB✓ fits
  • weights 4.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
1 recorded
amd/Qwen3-VL-8B-Instruct-w8a8-llmcompressor10.6 GB observed12.7 GB+49.3 GB✓ fits
  • weights 10.6 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 1.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
VibeVoice-ASROpen weightsMicrosoftMIT8.67B5.5 GB estimated+56.5 GB✓ fits
  • weights 4.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Molmo2-8BOpen weightsAllen Institute for AIApache-2.08.66B5.5 GB estimated+56.5 GB✓ fits
  • weights 4.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
aya-vision-8brestricted-weightsCohereCC-BY-NC-4.08.63B5.5 GB estimated+56.5 GB✓ fits
  • weights 4.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Intern-S1-miniOpen weightsInternLM (Shanghai AI Laboratory)Apache-2.08.54B5.4 GB estimated+56.6 GB✓ fits
  • weights 4.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
GKA-primed-HQwen3-8B-ReasonerOpen weightsAmazon Web ServicesApache-2.08.5B5.4 GB estimated+56.6 GB✓ fits
  • weights 4.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
GDN-primed-HQwen3-8B-InstructOpen weightsAmazon Web ServicesApache-2.08.5B5.4 GB estimated+56.6 GB✓ fits
  • weights 4.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
LFM2.5-8B-A1BOpen weightsLiquid AIOther8.47B5.4 GB estimated+56.6 GB✓ fits
  • weights 4.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
2 recorded
LiquidAI/LFM2.5-8B-A1B-GGUFgguf estimated5.4 GB+56.6 GB✓ fits
  • weights 4.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
LiquidAI/LFM2.5-8B-A1B-GGUFgguf estimated5.4 GB+56.6 GB✓ fits
  • weights 4.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
olmOCR-2-7B-1025Open weightsAllen Institute for AIApache-2.08.29B5.3 GB estimated+56.7 GB✓ fits
  • weights 4.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
1 recorded
allenai/olmOCR-2-7B-1025-FP810.1 GB observed12.1 GB+49.9 GB✓ fits
  • weights 10.1 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 1.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
Qwen2.5-VL-7B-InstructOpen weightsQwenApache-2.08.29B5.3 GB estimated+56.7 GB✓ fits
  • weights 4.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
UI-TARS 7BOpen weightsByteDanceApache-2.08.29B5.3 GB estimated+56.7 GB✓ fits
  • weights 4.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
UI-TARS-7B-DPOOpen weightsByteDanceApache-2.08.29B5.3 GB estimated+56.7 GB✓ fits
  • weights 4.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
UI-TARS-7B-SFTOpen weightsByteDanceApache-2.08.29B5.3 GB estimated+56.7 GB✓ fits
  • weights 4.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Seed-Coder-8B-InstructOpen weightsByteDanceMIT8.25B5.3 GB estimated+56.8 GB✓ fits
  • weights 4.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Seed-Coder-8B-BaseOpen weightsByteDanceMIT8.25B5.3 GB estimated+56.8 GB✓ fits
  • weights 4.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Seed-Coder-8B-ReasoningOpen weightsByteDanceMIT8.25B5.3 GB estimated+56.8 GB✓ fits
  • weights 4.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Stable-DiffCoder-8B-BaseOpen weightsByteDanceMIT8.25B5.3 GB estimated+56.8 GB✓ fits
  • weights 4.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Stable-DiffCoder-8B-InstructOpen weightsByteDanceMIT8.25B5.3 GB estimated+56.8 GB✓ fits
  • weights 4.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
DeepSeek-R1-0528-Qwen3-8BOpen weightsDeepSeekMIT8.19B5.2 GB estimated+56.8 GB✓ fits
  • weights 4.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Qwen3 8BOpen weightsQwenApache-2.08.19B5.2 GB estimated+56.8 GB✓ fits
  • weights 4.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
2 recorded
mlx-community/Qwen3-8B-4bitmlx4.6 GB observed5.8 GB+56.2 GB✓ fits
  • weights 4.6 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
mlx-community/Qwen3-8B-4bitmlx4.6 GB observed5.8 GB+56.2 GB✓ fits
  • weights 4.6 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
granite-guardian-3.3-8bOpen weightsIBMApache-2.08.17B5.2 GB estimated+56.8 GB✓ fits
  • weights 4.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
granite-3.0-8b-instructOpen weightsIBMApache-2.08.17B5.2 GB estimated+56.8 GB✓ fits
  • weights 4.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
stable-diffusion-3.5-largeOpen weightsStability AIOther8.15B5.2 GB estimated+56.8 GB✓ fits
  • weights 4.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
ERNIE-Image-TurboOpen weightsBaiduApache-2.08.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
ERNIE-ImageOpen weightsBaiduApache-2.08.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Meta-Llama-3-8B-Instructrestricted-weightsMeta AILlama-3-Community8.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Llama 3.1 8Brestricted-weightsMeta AILlama-3.1-Community8.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
1 recorded
NousResearch/Hermes-3-Llama-3.1-8B-GGUFgguf estimated5.1 GB+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
NousResearch/Meta-Llama-3.1-8Brestricted-weightsNous ResearchLlama-3.1-Community8.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
NousResearch/Meta-Llama-3.1-8B-Instructrestricted-weightsNous ResearchLlama-3.1-Community8.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Hermes-2-Theta-Llama-3-8BOpen weightsNous ResearchApache-2.08.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Llama 3.1 8B Instructrestricted-weightsMeta AILlama-3.1-Community8.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
2 recorded
mlx-community/Llama-3.1-8B-Instruct-4bitmlx4.5 GB observed5.7 GB+56.3 GB✓ fits
  • weights 4.5 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
bartowski/Meta-Llama-3.1-8B-Instruct-GGUFgguf estimated5.1 GB+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
NousResearch/Meta-Llama-3-8BOpen weightsNous ResearchOther8.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Meta-Llama-3-8Brestricted-weightsMeta AILlama-3-Community8.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
NousResearch/Meta-Llama-3-8B-InstructOpen weightsNous ResearchOther8.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Hermes-3-Llama-3.1-8Brestricted-weightsNous ResearchLlama-3-Community8.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
1 recorded
NousResearch/Hermes-3-Llama-3.1-8B-GGUFgguf estimated5.1 GB+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
Salesforce/Llama-xLAM-2-8b-fc-rrestricted-weightsSalesforceCC-BY-NC-4.08.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Hy-MT2-7BOpen weightsTencentApache-2.08.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
2 recorded
tencent/Hy-MT2-7B-GGUFgguf estimated5.1 GB+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
tencent/Hy-MT2-7B-GGUFgguf estimated5.1 GB+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
c4ai-command-r7b-12-2024restricted-weightsCohereCC-BY-NC-4.08.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
aya-expanse-8brestricted-weightsCohereCC-BY-NC-4.08.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
aya-23-8Brestricted-weightsCohereCC-BY-NC-4.08.03B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Ministral-8B-Instruct-2410Open weightsMistral AIOther8.02B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
gemma-4-E4B-itOpen weightsGoogleApache-2.08B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
ERNIE-Image-AesOpen weightsBaiduApache-2.07.94B5.1 GB estimated+56.9 GB✓ fits
  • weights 4.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
internlm2_5-7b-chatOpen weightsInternLM (Shanghai AI Laboratory)Other7.74B5.0 GB estimated+57.0 GB✓ fits
  • weights 3.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
internlm2-chat-7bOpen weightsInternLM (Shanghai AI Laboratory)Other7.74B5.0 GB estimated+57.0 GB✓ fits
  • weights 3.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
StripedHyena-Hessian-7BOpen weightsTogether AIApache-2.07.65B4.9 GB estimated+57.1 GB✓ fits
  • weights 3.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
StripedHyena-Nous-7BOpen weightsTogether AIApache-2.07.65B4.9 GB estimated+57.1 GB✓ fits
  • weights 3.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
SynLogic-7BOpen weightsMiniMaxMIT7.62B4.9 GB estimated+57.1 GB✓ fits
  • weights 3.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Qwen2.5 7B InstructOpen weightsQwenApache-2.07.62B4.9 GB estimated+57.1 GB✓ fits
  • weights 3.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
2 recorded
bartowski/Qwen2.5-7B-Instruct-GGUFgguf estimated4.9 GB+57.1 GB✓ fits
  • weights 3.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
bartowski/Qwen2.5-7B-Instruct-GGUFgguf estimated4.9 GB+57.1 GB✓ fits
  • weights 3.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
BFS-Prover-V2-7BOpen weightsByteDanceApache-2.07.62B4.9 GB estimated+57.1 GB✓ fits
  • weights 3.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
DeepSeek-R1-Distill-Qwen-7BOpen weightsDeepSeekMIT7.62B4.9 GB estimated+57.1 GB✓ fits
  • weights 3.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Seed-X-PPO-7BOpen weightsByteDanceOther7.51B4.8 GB estimated+57.2 GB✓ fits
  • weights 3.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Hunyuan-7B-InstructOpen weightsTencent7.5B4.8 GB estimated+57.2 GB✓ fits
  • weights 3.8 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
OLMo-2-1124-7B-InstructOpen weightsAllen Institute for AIApache-2.07.3B4.7 GB estimated+57.3 GB✓ fits
  • weights 3.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
OLMo-2-1124-7BOpen weightsAllen Institute for AIApache-2.07.3B4.7 GB estimated+57.3 GB✓ fits
  • weights 3.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Olmo-3-7B-InstructOpen weightsAllen Institute for AIApache-2.07.3B4.7 GB estimated+57.3 GB✓ fits
  • weights 3.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Olmo-3-7B-ThinkOpen weightsAllen Institute for AIApache-2.07.3B4.7 GB estimated+57.3 GB✓ fits
  • weights 3.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Olmo-3-1025-7BOpen weightsAllen Institute for AIApache-2.07.3B4.7 GB estimated+57.3 GB✓ fits
  • weights 3.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.6 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
wildguardOpen weightsAllen Institute for AIApache-2.07.25B4.7 GB estimated+57.3 GB✓ fits
  • weights 3.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Mistral-7B-v0.3Open weightsMistral AIApache-2.07.25B4.7 GB estimated+57.3 GB✓ fits
  • weights 3.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Mistral-7B-Instruct-v0.3Open weightsMistral AIApache-2.07.25B4.7 GB estimated+57.3 GB✓ fits
  • weights 3.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Mistral-7B-Instruct-v0.1Open weightsMistral AIApache-2.07.24B4.7 GB estimated+57.3 GB✓ fits
  • weights 3.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Mistral-7B-v0.1Open weightsMistral AIApache-2.07.24B4.7 GB estimated+57.3 GB✓ fits
  • weights 3.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
1 recorded
NousResearch/Hermes-2-Pro-Mistral-7B-GGUFgguf estimated4.7 GB+57.3 GB✓ fits
  • weights 3.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
Mistral-7B-Instruct-v0.2Open weightsMistral AIApache-2.07.24B4.7 GB estimated+57.3 GB✓ fits
  • weights 3.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
SFR-Embedding-2_Rrestricted-weightsSalesforceCC-BY-NC-4.07.11B4.6 GB estimated+57.4 GB✓ fits
  • weights 3.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
SFR-Embedding-Mistralrestricted-weightsSalesforceCC-BY-NC-4.07.11B4.6 GB estimated+57.4 GB✓ fits
  • weights 3.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
SeedVR-7BOpen weightsByteDanceApache-2.07B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Qwen2-AudioOpen weightsQwen Team7B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
internlm-xcomposer2-7bOpen weightsInternLM (Shanghai AI Laboratory)Other7B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
fireworks-ai/mistral-7b-eagle-head-experimentalOpen weightsFireworks AIApache-2.07B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
internlm2-base-7bOpen weightsInternLM (Shanghai AI Laboratory)Other7B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
SeedVR2-7BOpen weightsByteDanceApache-2.07B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Qwen2.5-OmniOpen weightsQwen Team7B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
togethercomputer/Llama-2-7B-32K-Instructrestricted-weightsTogether AILlama-2-Community7B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
internlm-xcomposer-7bOpen weightsInternLM (Shanghai AI Laboratory)Apache-2.07B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
togethercomputer/LLaMA-2-7B-32Krestricted-weightsTogether AILlama-2-Community7B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
internlm2-7bOpen weightsInternLM (Shanghai AI Laboratory)Other7B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
internlm-chat-7bOpen weightsInternLM (Shanghai AI Laboratory)7B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Pythia-Chat-Base-7BOpen weightsTogether AIApache-2.07B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
RedPajama-INCITE-7B-BaseOpen weightsTogether AIApache-2.07B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
RedPajama-INCITE-7B-ChatOpen weightsTogether AIApache-2.07B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
RedPajama-INCITE-7B-InstructOpen weightsTogether AIApache-2.07B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
pythia-6.9bOpen weightsEleutherAIApache-2.06.99B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
OLMoE-1B-7B-0924Open weightsAllen Institute for AIApache-2.06.92B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
OLMoE-1B-7B-0125-InstructOpen weightsAllen Institute for AIApache-2.06.92B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
deepseek-coder-7b-instruct-v1.5Open weightsDeepSeekOther6.91B4.5 GB estimated+57.5 GB✓ fits
  • weights 3.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
deepseek-coder-6.7b-instructOpen weightsDeepSeekOther6.74B4.4 GB estimated+57.6 GB✓ fits
  • weights 3.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
NousResearch/Llama-2-7b-hfOpen weightsNous Research6.74B4.4 GB estimated+57.6 GB✓ fits
  • weights 3.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
NousResearch/Llama-2-7b-chat-hfOpen weightsNous Research6.74B4.4 GB estimated+57.6 GB✓ fits
  • weights 3.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Llama-2-7b-hfrestricted-weightsMeta AILlama-2-Community6.74B4.4 GB estimated+57.6 GB✓ fits
  • weights 3.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Llama-2-7b-chat-hfrestricted-weightsMeta AILlama-2-Community6.74B4.4 GB estimated+57.6 GB✓ fits
  • weights 3.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Nous-Hermes-llama-2-7bOpen weightsNous ResearchMIT6.74B4.4 GB estimated+57.6 GB✓ fits
  • weights 3.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
evo-1-8k-baseOpen weightsTogether AIApache-2.06.45B4.2 GB estimated+57.8 GB✓ fits
  • weights 3.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
evo-1-131k-baseOpen weightsTogether AIApache-2.06.45B4.2 GB estimated+57.8 GB✓ fits
  • weights 3.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
chatglm3-6bOpen weightsZ.ai (Zhipu AI)6.24B4.1 GB estimated+57.9 GB✓ fits
  • weights 3.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
GPT-JT-6B-v1Open weightsTogether AIApache-2.06B4.0 GB estimated+58.0 GB✓ fits
  • weights 3.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
gpt-j-6bOpen weightsEleutherAIApache-2.06B4.0 GB estimated+58.0 GB✓ fits
  • weights 3.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
chatglm2-6bOpen weightsZ.ai (Zhipu AI)6B4.0 GB estimated+58.0 GB✓ fits
  • weights 3.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
gemma-4-E2B-itOpen weightsGoogleApache-2.05.12B3.4 GB estimated+58.6 GB✓ fits
  • weights 2.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Molmo2-4BOpen weightsAllen Institute for AIApache-2.04.85B3.3 GB estimated+58.7 GB✓ fits
  • weights 2.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Qianfan-OCROpen weightsBaiduApache-2.04.74B3.2 GB estimated+58.8 GB✓ fits
  • weights 2.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.4 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Voxtral-Mini-3B-2507Open weightsMistral AIApache-2.04.68B3.2 GB estimated+58.8 GB✓ fits
  • weights 2.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Qwen3.5-4BOpen weightsQwenApache-2.04.66B3.2 GB estimated+58.8 GB✓ fits
  • weights 2.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
5 recorded
Intel/Qwen3.5-4B-int4-AutoRound4.6 GB observed5.8 GB+56.2 GB✓ fits
  • weights 4.6 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 0.7 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
unsloth/Qwen3.5-4B-GGUFgguf estimated3.2 GB+58.8 GB✓ fits
  • weights 2.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
bartowski/Qwen_Qwen3.5-4B-GGUFgguf estimated3.2 GB+58.8 GB✓ fits
  • weights 2.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
bartowski/Qwen_Qwen3.5-4B-GGUFgguf estimated3.2 GB+58.8 GB✓ fits
  • weights 2.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
+1 more artifacts on the model page.
Qwen3-VL-4B-InstructOpen weightsQwenApache-2.04.44B3.0 GB estimated+59.0 GB✓ fits
  • weights 2.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Voxtral-Mini-4B-Realtime-2602Open weightsMistral AIApache-2.04.43B3.0 GB estimated+59.0 GB✓ fits
  • weights 2.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Gemma 3 4Brestricted-weightsGoogleGemma-Terms4.3B3.0 GB estimated+59.0 GB✓ fits
  • weights 2.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
1 recorded
mlx-community/gemma-3-4b-it-qat-4bitmlx3.0 GB observed4.0 GB+58.0 GB✓ fits
  • weights 3.0 GB observed
  • KV cache 0.50 GB (heuristic)
  • overhead 0.5 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
Ministral-3-3B-Reasoning-2512Open weightsMistral AIApache-2.04.25B3.0 GB estimated+59.0 GB✓ fits
  • weights 2.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Phi-3.5-vision-instructOpen weightsMicrosoftMIT4.15B2.9 GB estimated+59.1 GB✓ fits
  • weights 2.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
instructblip-flan-t5-xlOpen weightsSalesforceMIT4.02B2.8 GB estimated+59.2 GB✓ fits
  • weights 2.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
pplx-embed-context-v1-4bOpen weightsPerplexity AIMIT4.02B2.8 GB estimated+59.2 GB✓ fits
  • weights 2.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
pplx-embed-v1-4bOpen weightsPerplexity AIMIT4.02B2.8 GB estimated+59.2 GB✓ fits
  • weights 2.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Qwen3-4BOpen weightsQwenApache-2.04.02B2.8 GB estimated+59.2 GB✓ fits
  • weights 2.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
TRELLIS.2-4BOpen weightsMicrosoftMIT4B2.8 GB estimated+59.2 GB✓ fits
  • weights 2.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
granite-vision-4.1-4bOpen weightsIBMApache-2.04B2.8 GB estimated+59.2 GB✓ fits
  • weights 2.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
blip2-flan-t5-xlOpen weightsSalesforceMIT3.94B2.8 GB estimated+59.2 GB✓ fits
  • weights 2.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
FLUX.2-klein-4BOpen weightsBlack Forest LabsApache-2.03.88B2.7 GB estimated+59.3 GB✓ fits
  • weights 1.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
FLUX.2-klein-base-4BOpen weightsBlack Forest LabsApache-2.03.88B2.7 GB estimated+59.3 GB✓ fits
  • weights 1.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Cosmos3-EdgeOpen weightsNVIDIAOther3.86B2.7 GB estimated+59.3 GB✓ fits
  • weights 1.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Ministral 3 3B 2512Open weightsMistral AIApache-2.03.85B2.7 GB estimated+59.3 GB✓ fits
  • weights 1.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
3b-zh-pretrain-research_releaserestricted-weightsCanopy LabsLlama-3.2-Community3.78B2.7 GB estimated+59.3 GB✓ fits
  • weights 1.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
orpheus-3b-0.1-ftOpen weightsCanopy LabsApache-2.03.78B2.7 GB estimated+59.3 GB✓ fits
  • weights 1.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
3b-ko-pretrain-research_releaseOpen weightsCanopy LabsApache-2.03.78B2.7 GB estimated+59.3 GB✓ fits
  • weights 1.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
3b-fr-pretrain-research_releaserestricted-weightsCanopy LabsLlama-3.2-Community3.78B2.7 GB estimated+59.3 GB✓ fits
  • weights 1.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
3b-hi-pretrain-research_releaserestricted-weightsCanopy LabsLlama-3.2-Community3.78B2.7 GB estimated+59.3 GB✓ fits
  • weights 1.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
3b-de-pretrain-research_releaserestricted-weightsCanopy LabsLlama-3.2-Community3.78B2.7 GB estimated+59.3 GB✓ fits
  • weights 1.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
3b-es_it-pretrain-research_releaserestricted-weightsCanopy LabsLlama-3.2-Community3.78B2.7 GB estimated+59.3 GB✓ fits
  • weights 1.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
orpheus-3b-0.1-pretrainedOpen weightsCanopy LabsApache-2.03.78B2.7 GB estimated+59.3 GB✓ fits
  • weights 1.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
blip2-opt-2.7bOpen weightsSalesforceMIT3.74B2.6 GB estimated+59.4 GB✓ fits
  • weights 1.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Qianfan-VL-3BOpen weightsBaiduOther3.71B2.6 GB estimated+59.4 GB✓ fits
  • weights 1.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
M1-3BOpen weightsTogether AIMIT3.45B2.5 GB estimated+59.5 GB✓ fits
  • weights 1.7 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
granite-4.1-3bOpen weightsIBMApache-2.03.4B2.5 GB estimated+59.5 GB✓ fits
  • weights 1.7 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
DeepSeek-OCR-2Open weightsDeepSeekApache-2.03.39B2.4 GB estimated+59.6 GB✓ fits
  • weights 1.7 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
tiny-aya-globalrestricted-weightsCohereCC-BY-NC-4.03.35B2.4 GB estimated+59.6 GB✓ fits
  • weights 1.7 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
tiny-aya-baserestricted-weightsCohereCC-BY-NC-4.03.35B2.4 GB estimated+59.6 GB✓ fits
  • weights 1.7 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
DeepSeek-OCROpen weightsDeepSeekMIT3.34B2.4 GB estimated+59.6 GB✓ fits
  • weights 1.7 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Unlimited-OCROpen weightsBaiduMIT3.34B2.4 GB estimated+59.6 GB✓ fits
  • weights 1.7 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
3b-fr-ft-research_releaseOpen weightsCanopy LabsApache-2.03.3B2.4 GB estimated+59.6 GB✓ fits
  • weights 1.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
3b-hi-ft-research_releaseOpen weightsCanopy LabsApache-2.03.3B2.4 GB estimated+59.6 GB✓ fits
  • weights 1.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
3b-ko-ft-research_releaseOpen weightsCanopy LabsApache-2.03.3B2.4 GB estimated+59.6 GB✓ fits
  • weights 1.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
3b-es_it-ft-research_releaseOpen weightsCanopy LabsApache-2.03.3B2.4 GB estimated+59.6 GB✓ fits
  • weights 1.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
3b-de-ft-research_releaseOpen weightsCanopy LabsApache-2.03.3B2.4 GB estimated+59.6 GB✓ fits
  • weights 1.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
3b-zh-ft-research_releaseOpen weightsCanopy LabsApache-2.03.3B2.4 GB estimated+59.6 GB✓ fits
  • weights 1.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.3 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Llama 3.2 3B Instructrestricted-weightsMeta AILlama-3.2-Community3.21B2.4 GB estimated+59.6 GB✓ fits
  • weights 1.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Llama-3.2-3Brestricted-weightsMeta AILlama-3.2-Community3.21B2.4 GB estimated+59.6 GB✓ fits
  • weights 1.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
AI21-Jamba-Reasoning-3BOpen weightsAI21 LabsApache-2.03.2B2.3 GB estimated+59.7 GB✓ fits
  • weights 1.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
1 recorded
ai21labs/AI21-Jamba-Reasoning-3B-GGUFgguf estimated2.3 GB+59.7 GB✓ fits
  • weights 1.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
LFM2.5-VL-3BOpen weightsLiquid AIOther3.12B2.3 GB estimated+59.7 GB✓ fits
  • weights 1.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
2 recorded
LiquidAI/LFM2.5-VL-3B-GGUFgguf estimated2.3 GB+59.7 GB✓ fits
  • weights 1.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
LiquidAI/LFM2.5-VL-3B-GGUFgguf estimated2.3 GB+59.7 GB✓ fits
  • weights 1.6 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
Qwen2.5-3B-InstructOpen weightsQwenOther3.09B2.3 GB estimated+59.7 GB✓ fits
  • weights 1.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
SmolLM3-3BOpen weightsHugging FaceApache-2.03.08B2.3 GB estimated+59.7 GB✓ fits
  • weights 1.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
SmolLM3-3B-BaseOpen weightsHugging FaceApache-2.03.08B2.3 GB estimated+59.7 GB✓ fits
  • weights 1.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
AI21-Jamba2-3BOpen weightsAI21 LabsApache-2.03.03B2.2 GB estimated+59.8 GB✓ fits
  • weights 1.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
RedPajama-INCITE-Chat-3B-v1Open weightsTogether AIApache-2.03B2.2 GB estimated+59.8 GB✓ fits
  • weights 1.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
SeedVR2-3BOpen weightsByteDanceApache-2.03B2.2 GB estimated+59.8 GB✓ fits
  • weights 1.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
SeedVR-3BOpen weightsByteDanceApache-2.03B2.2 GB estimated+59.8 GB✓ fits
  • weights 1.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
RedPajama-INCITE-Instruct-3B-v1Open weightsTogether AIApache-2.03B2.2 GB estimated+59.8 GB✓ fits
  • weights 1.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
RedPajama-INCITE-Base-3B-v1Open weightsTogether AIApache-2.03B2.2 GB estimated+59.8 GB✓ fits
  • weights 1.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
granite-vision-3.3-2bOpen weightsIBMApache-2.02.98B2.2 GB estimated+59.8 GB✓ fits
  • weights 1.5 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
pythia-2.8bOpen weightsEleutherAIApache-2.02.91B2.2 GB estimated+59.8 GB✓ fits
  • weights 1.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
stablelm-3b-4e1tOpen weightsStability AICC-BY-SA-4.02.8B2.1 GB estimated+59.9 GB✓ fits
  • weights 1.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
phi-2Open weightsMicrosoftMIT2.78B2.1 GB estimated+59.9 GB✓ fits
  • weights 1.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
WeMM-Embedding-2BOpen weightsTencentOther2.72B2.1 GB estimated+59.9 GB✓ fits
  • weights 1.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
LFM2.5-2.6B (free)Open weightsLiquid AIOther2.7B2.0 GB estimated+60.0 GB✓ fits
  • weights 1.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
1 recorded
LiquidAI/LFM2.5-2.6B-GGUFgguf estimated2.0 GB+60.0 GB✓ fits
  • weights 1.4 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
stable-diffusion-xl-base-1.0restricted-weightsStability AIOpenRAIL++-M2.57B2.0 GB estimated+60.0 GB✓ fits
  • weights 1.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
sdxl-turboOpen weightsStability AIOther2.57B2.0 GB estimated+60.0 GB✓ fits
  • weights 1.3 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
North-Micro-Vision-InstructOpen weightsCohereApache-2.02.48B1.9 GB estimated+60.1 GB✓ fits
  • weights 1.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
stable-diffusion-3.5-mediumOpen weightsStability AIOther2.47B1.9 GB estimated+60.1 GB✓ fits
  • weights 1.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
UI-TARS-2B-SFTOpen weightsByteDanceApache-2.02.44B1.9 GB estimated+60.1 GB✓ fits
  • weights 1.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Cosmos-Reason2-2BOpen weightsNVIDIAOther2.44B1.9 GB estimated+60.1 GB✓ fits
  • weights 1.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
MiniMax-Music3Open weightsMiniMax2.43B1.9 GB estimated+60.1 GB✓ fits
  • weights 1.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
granite-speech-4.1-2bOpen weightsIBMApache-2.02.31B1.8 GB estimated+60.2 GB✓ fits
  • weights 1.2 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
stable-audio-3-mediumOpen weightsStability AIOther2.31B1.8 GB estimated+60.2 GB✓ fits
  • weights 1.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
stable-diffusion-xl-refiner-1.0restricted-weightsStability AIOpenRAIL++-M2.26B1.8 GB estimated+60.2 GB✓ fits
  • weights 1.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
GLM-ASR-Nano-2512Open weightsZ.ai (Zhipu AI)MIT2.26B1.8 GB estimated+60.2 GB✓ fits
  • weights 1.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
granite-speech-4.1-2b-narOpen weightsIBMApache-2.02.25B1.8 GB estimated+60.2 GB✓ fits
  • weights 1.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
SmolVLM2-2.2B-InstructOpen weightsHugging FaceApache-2.02.25B1.8 GB estimated+60.2 GB✓ fits
  • weights 1.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
SmolVLM-InstructOpen weightsHugging FaceApache-2.02.25B1.8 GB estimated+60.2 GB✓ fits
  • weights 1.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
granite-speech-4.1-2b-plusOpen weightsIBMApache-2.02.11B1.7 GB estimated+60.3 GB✓ fits
  • weights 1.1 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
stable-diffusion-3-medium-diffusersOpen weightsStability AIOther2.08B1.7 GB estimated+60.3 GB✓ fits
  • weights 1.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.2 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
cohere-transcribe-03-2026Open sourceCohereApache-2.02.07B1.7 GB estimated+60.3 GB✓ fits
  • weights 1.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
cohere-transcribe-arabic-07-2026Open weightsCohereApache-2.02.07B1.7 GB estimated+60.3 GB✓ fits
  • weights 1.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
Hy-MT2-1.8BOpen weightsTencentApache-2.02.04B1.7 GB estimated+60.3 GB✓ fits
  • weights 1.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
2 recorded
tencent/Hy-MT2-1.8B-GGUFgguf estimated1.7 GB+60.3 GB✓ fits
  • weights 1.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
tencent/Hy-MT2-1.8B-GGUFgguf estimated1.7 GB+60.3 GB✓ fits
  • weights 1.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
4bit
Youtu-LLM-2BOpen weightsTencentOther1.96B1.6 GB estimated+60.4 GB✓ fits
  • weights 1.0 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
FastVLM-1.5Brestricted-weightsAppleApple-AMLR1.91B1.6 GB estimated+60.4 GB✓ fits
  • weights 0.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
internlm2_5-1_8b-chatOpen weightsInternLM (Shanghai AI Laboratory)Other1.89B1.6 GB estimated+60.4 GB✓ fits
  • weights 0.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
siglip2-giant-opt-patch16-384Open weightsGoogleApache-2.01.87B1.6 GB estimated+60.4 GB✓ fits
  • weights 0.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded
amazon/GPT-OSS-20B-P-EAGLEOpen weightsAmazon Web ServicesApache-2.01.8B1.5 GB estimated+60.5 GB✓ fits
  • weights 0.9 GB estimated
  • KV cache 0.50 GB (heuristic)
  • overhead 0.1 GB
  • reserved 2 GB
  • context 8,192 × batch 1
none recorded

Headroom = total device memory − reserve − estimate. 34 of 300 models have quantized artifacts recorded; artifact rows use the observed file size when a source publishes it, otherwise the estimate. Sorted by the API (fitting models first). Compare shortlisted models with the + buttons.

MethodEvery figure is an ESTIMATE: weights = params × bytes/param (× 1.15 overhead) unless an artifact's observed file size is available; KV cache uses architecture metadata when known, else 0.5 GB per 8K tokens × batch. /methodology