Self-host

RAM and whether you have a GPU are enough. Everything else is optional. The ranking is arithmetic on sourced parameter counts — NVIDIA does not pick the winner.

25 catalogue rows marked open-weight. Nothing else is a candidate.
25 of those have a sourced parameter count. The rest are listed as unknown, not guessed.
1 GiB runtime allowance, named, not measured on your machine.
GPU
OS (optional — software only)
Context to budget
KV cache grows with this. We do not assume the model card maximum.

7 of 25 profiled open-weight models fit 24 GiB (vram) at 8192 context. Best rated that fits: Gemma 4 31B (Q5_K_M).

Weights 21.106 GiB KV not estimated Runtime 1 GiB Headroom 1.9 GiB / 24 GiB

VRAM as entered. Driver reservation is inside the 1 GiB overhead, not extra.

48 GiB raises the quant of Gemma 4 31B to Q8_0. Same model, not a higher intelligence score.

Among models that fit

Highest LMArena Elo among models that fit: Gemma 4 31B 1443.4. Elo is a human-preference scale, not the intelligence index.

The arena boards are separate scales and are never blended. A missing cell is unmeasured, not zero. A model’s highest axis is relative to its own scores, not a claim it leads the catalogue.

Rated models that fit

ModelHuman preferenceQuant that fitsEstimated GiBHeadroom
Gemma 4 31B
visionvideoreasoningtools
1443.4Q5_K_M22.1061.894
Qwen3.8 27B
visionvideoreasoningtools
1440.7Q5_K_M20.0633.938
Phi 4 1216.7Q8_017.4386.563
  • KV heads and head dim are not in the public size table we used. Cache is not estimated.
  • KV cache not estimated — layer/head geometry is not in the sourced profile. Longer context may not fit.

Fits, unrated

What an upgrade actually buys

The fit bottleneck is GPU memory. More system RAM will not load extra weights onto this card.

Estimated GiBUnlocksExampleList price
48 GiBGemma 4 31B — Q8_0same model, higher quantrtx a6000, rtx 6000 ada, rtx pro 6000 blackwellNo consumer list in this file

The next higher-rated catalogue model is Kimi K3, and even Q4_K_M wants 1691.5 GiB. That is not a desktop card.

NVIDIA 'Starting at' list prices as published 2026-08-25, not a store quote and not a used-market index. Street prices in August 2026 are often much higher (Tom's Hardware Newegg median RTX 5090 $4,699 vs $1,999 list). Last-gen 24 GB cards (4090, 3090) have no current NVIDIA starting-at in this file. Datacenter cards have no consumer list here. RAM DIMMs are not priced — they move weekly.

Upgrades use the same fit formula as the ranking. List prices are NVIDIA starting-at, dated 2026-08-25. Street prices are often higher. Used 24 GB cards and datacenter GPUs have no list here. RAM DIMMs are not priced.

Does not fit, or not profiled

  • cohere/command-a-plus — even Q4_K_M wants 132.618 GiB; pool is 24 GiB
  • deepseek/deepseek-v4-flash-vision-exp — even Q4_K_M wants 184.902 GiB; pool is 24 GiB
  • z-ai/glm-5.3 — even Q4_K_M wants 455.624 GiB; pool is 24 GiB
  • qwen/qwen3.8-2.4t-a95b — even Q4_K_M wants 1450.719 GiB; pool is 24 GiB
  • moonshotai/kimi-k3 — even Q4_K_M wants 1691.5 GiB; pool is 24 GiB
  • z-ai/glm-5.2 — even Q4_K_M wants 455.624 GiB; pool is 24 GiB
  • moonshotai/kimi-k2.7-code — even Q4_K_M wants 604.75 GiB; pool is 24 GiB
  • nvidia/nemotron-3-ultra-550b-a55b — even Q4_K_M wants 333.063 GiB; pool is 24 GiB
  • minimax/minimax-m3 — even Q4_K_M wants 259.405 GiB; pool is 24 GiB
  • deepseek/deepseek-v4-pro — even Q4_K_M wants 967 GiB; pool is 24 GiB
  • deepseek/deepseek-v4-flash — even Q4_K_M wants 172.465 GiB; pool is 24 GiB
  • mistralai/mistral-small-2603 — even Q4_K_M wants 73.2 GiB; pool is 24 GiB
  • mistralai/mistral-large-2512 — even Q4_K_M wants 408.531 GiB; pool is 24 GiB
  • deepseek/deepseek-v3.2 — even Q4_K_M wants 406.116 GiB; pool is 24 GiB
  • openai/gpt-oss-120b — even Q4_K_M wants 72.201 GiB; pool is 24 GiB
  • meta-llama/llama-4-maverick — even Q4_K_M wants 242.5 GiB; pool is 24 GiB
  • meta-llama/llama-4-scout — even Q4_K_M wants 66.809 GiB; pool is 24 GiB
  • cohere/command-a — even Q4_K_M wants 68.016 GiB; pool is 24 GiB

Extra details (optional)

Does not change the ranking. Optional AI analysis of your notes against the numbers above, using NVIDIA NIM’s strongest models.

How the number is made

Weights: a measured GGUF size when we have one, otherwise total parameters × bits per weight ÷ 8. Q4_K_M is treated as 4.83 bits/param. MoE memory uses total parameters, not active parameters.

KV cache, when layer and head geometry is sourced: 2 × kvLayers × kvHeads × headDim × context × 2 bytes (FP16, batch 1). Hybrid models use attention-layer count, not every layer. If geometry is missing, cache is omitted and that is stated.

1 GiB is added for CUDA/runtime. Apple unified memory uses RAM as the pool. CPU-only is a memory fit, not a speed claim. DDR generation is ignored for fit — it changes bandwidth, not whether the weights sit in memory.

Only catalogue rows with openWeights=true are candidates. 25 such rows exist; 25 have sourced sizes. Unrated models that fit are listed separately. A higher score does not mean a drop-in replacement.

RAM type (DDR4/DDR5/LPDDR) is not a field. It does not change whether a model fits.

Evidence & Ask