Self-host

8 GiB GPU

8 GiB is the vendor-published framebuffer on RTX 5060 Ti 8, RTX 5060, RTX 5050, RTX 4060 Ti 8, RTX 4060, RTX 3070 Ti, RTX 3070, RTX 3060 Ti, RTX 3060 8. Same pool on every card in that list. Ranking is arithmetic on sourced parameter counts.

8 GiB framebuffer, taken from the sourced cards listed below. Not a free-after-driver figure.

2 of 25 profiled open-weight models fit 8 GiB (vram) at Q4_K_M, 8192 context. Best rated that fits: Ministral 3 8B 2512.

Weights 5.313 GiB KV cache 1.063 GiB Runtime 1 GiB Headroom 0.6 GiB / 8 GiB

VRAM as entered. Driver reservation is inside the 1 GiB overhead, not extra.

Fits, unrated

Does not fit, or not profiled

  • cohere/command-a-plus — even Q4_K_M wants 132.618 GiB; pool is 8 GiB
  • deepseek/deepseek-v4-flash-vision-exp — even Q4_K_M wants 184.902 GiB; pool is 8 GiB
  • z-ai/glm-5.3 — even Q4_K_M wants 455.624 GiB; pool is 8 GiB
  • qwen/qwen3.8-27b — even Q4_K_M wants 18 GiB; pool is 8 GiB
  • qwen/qwen3.8-2.4t-a95b — even Q4_K_M wants 1450.719 GiB; pool is 8 GiB
  • meta/muse-glimmer-30b — even Q4_K_M wants 19.277 GiB; pool is 8 GiB
  • moonshotai/kimi-k3 — even Q4_K_M wants 1691.5 GiB; pool is 8 GiB
  • z-ai/glm-5.2 — even Q4_K_M wants 455.624 GiB; pool is 8 GiB
  • moonshotai/kimi-k2.7-code — even Q4_K_M wants 604.75 GiB; pool is 8 GiB
  • nvidia/nemotron-3-ultra-550b-a55b — even Q4_K_M wants 333.063 GiB; pool is 8 GiB
  • minimax/minimax-m3 — even Q4_K_M wants 259.405 GiB; pool is 8 GiB
  • deepseek/deepseek-v4-pro — even Q4_K_M wants 967 GiB; pool is 8 GiB
  • deepseek/deepseek-v4-flash — even Q4_K_M wants 172.465 GiB; pool is 8 GiB
  • google/gemma-4-31b-it — even Q4_K_M wants 20.6 GiB; pool is 8 GiB
  • mistralai/mistral-small-2603 — even Q4_K_M wants 73.2 GiB; pool is 8 GiB
  • mistralai/ministral-14b-2512 — even Q4_K_M wants 10.642 GiB; pool is 8 GiB
  • mistralai/mistral-large-2512 — even Q4_K_M wants 408.531 GiB; pool is 8 GiB
  • deepseek/deepseek-v3.2 — even Q4_K_M wants 406.116 GiB; pool is 8 GiB
  • openai/gpt-oss-120b — even Q4_K_M wants 72.201 GiB; pool is 8 GiB
  • meta-llama/llama-4-maverick — even Q4_K_M wants 242.5 GiB; pool is 8 GiB
  • meta-llama/llama-4-scout — even Q4_K_M wants 66.809 GiB; pool is 8 GiB
  • cohere/command-a — even Q4_K_M wants 68.016 GiB; pool is 8 GiB
  • microsoft/phi-4 — even Q4_K_M wants 11.015 GiB; pool is 8 GiB

GPU

How the number is made

Weights: a measured GGUF size when we have one, otherwise total parameters × bits per weight ÷ 8. Q4_K_M is treated as 4.83 bits/param. MoE memory uses total parameters, not active parameters.

KV cache, when layer and head geometry is sourced: 2 × kvLayers × kvHeads × headDim × context × 2 bytes (FP16, batch 1). Hybrid models use attention-layer count, not every layer. If geometry is missing, cache is omitted and that is stated.

1 GiB is added for CUDA/runtime. Apple unified memory uses RAM as the pool. CPU-only is a memory fit, not a speed claim. DDR generation is ignored for fit — it changes bandwidth, not whether the weights sit in memory.

RAM type (DDR4/DDR5/LPDDR) is not a field. It does not change whether a model fits.

Self-host

Questions this page answers

What LLM can I run on a 8GB card?

2 of 25 profiled open-weight models fit 8 GiB (vram) at Q4_K_M, 8192 context. Best rated that fits: Ministral 3 8B 2512.

Which sourced cards publish 8 GiB?

RTX 5060 Ti 8 (8 GiB), RTX 5060 (8 GiB), RTX 5050 (8 GiB), RTX 4060 Ti 8 (8 GiB), RTX 4060 (8 GiB), RTX 3070 Ti (8 GiB), RTX 3070 (8 GiB), RTX 3060 Ti (8 GiB), RTX 3060 8 (8 GiB).

Does NVIDIA pick the winner?

No. Ranking is arithmetic on sourced parameter counts. Quality leads. Unrated models that fit are listed separately, never scored zero.

Evidence & Ask