---
title: "96 GiB GPU: open-weight models that fit · Undominated.ai"
canonical: https://undominated.ai/self-host/96gb/
description: "96 GiB is the vendor-published framebuffer on RTX PRO 6000 Blackwell. Catalogue open-weight models that fit that pool, ranked by measured quality. Unrated is not zero. NVIDIA does not pick the winner."
---

# 96 GiB GPU: open-weight models that fit · Undominated.ai

> 96 GiB is the vendor-published framebuffer on RTX PRO 6000 Blackwell. Catalogue open-weight models that fit that pool, ranked by measured quality. Unrated is not zero. NVIDIA does not pick the winner.

[Self-host](/self-host/)

# 96 GiB GPU

96 GiB is the vendor-published framebuffer on RTX PRO 6000 Blackwell. Same pool on every card in that list. Ranking is arithmetic on sourced parameter counts.

96 GiB framebuffer, taken from the sourced cards listed below. Not a free-after-driver figure.

11 of 25 profiled open-weight models fit 96 GiB (vram) at F16, 8192 context. Best rated that fits: Gemma 4 31B.

 ** ** **
 Weights 61.41 GiB KV not estimated Runtime 1 GiB Headroom 33.6 GiB / 96 GiB

VRAM as entered. Driver reservation is inside the 1 GiB overhead, not extra.

## Among models that fit

Highest LMArena Elo among models that fit: Gemma 4 31B 1443.4. Elo is a human-preference scale, not the intelligence index.

The arena boards are separate scales and are never blended. A missing cell is unmeasured, not zero. A model’s highest axis is relative to its own scores, not a claim it leads the catalogue.

## Rated models that fit

| Model | Human preference | Quant that fits | Estimated GiB | Headroom |
| --- | --- | --- | --- | --- |
| [Gemma 4 31B](/models/google__gemma-4-31b-it/) vision video reasoning tools Fit assumptions KV heads and head dim are not in the public size table we used. Cache is not estimated. KV cache not estimated — layer/head geometry is not in the sourced profile. Longer context may not fit. | 1443.4 | F16 | 62.41 | 33.59 |
| [Qwen3.8 27B](/models/qwen__qwen3.8-27b/) vision video reasoning tools Fit assumptions Hybrid: only 16 of 64 layers are Gated Attention. KV uses kvLayers=16, not nLayers=64. Measured Q4 is Unsloth UD-Q4_K_M, not a vanilla Q4_K_M. | 1440.7 | F16 | 55.5 | 40.5 |
| [gpt-oss-120b](/models/openai__gpt-oss-120b/) reasoning tools Fit assumptions MoE: memory follows 117B total, not 5.1B active. Vendor 80GB claim is for MXFP4, not Q4_K_M. KV from the published config: 36 layers, 8 KV heads, head dim 64. MoE: memory follows 117B total parameters, not 5.1B active. | 1365.4 | Q5_K_M | 82 | 14 |
| [Phi 4](/models/microsoft__phi-4/) Fit assumptions Hub safetensors listing says 15B params; the card states 14B. We use the card. KV from config.json: 40 layers, 10 KV heads, hidden 5120 → head dim 128. | 1216.7 | F16 | 30.563 | 65.438 |

## Fits, unrated

 - [Mistral Small 4](/models/mistralai__mistral-small-2603/) — Q5_K_M, 82.813 GiB unrated
- [Command A](/models/cohere__command-a/) — Q5_K_M, 77.313 GiB unrated
- [Llama 4 Scout](/models/meta-llama__llama-4-scout/) — Q5_K_M, 75.938 GiB unrated
- [Muse Glimmer 30B](/models/meta__muse-glimmer-30b/) — F16, 60.606 GiB unrated
- [Ministral 3 14B 2512](/models/mistralai__ministral-14b-2512/) — F16, 30.05 GiB unrated
- [Ministral 3 8B 2512](/models/mistralai__ministral-8b-2512/) — F16, 19.663 GiB unrated
- [Ministral 3 3B 2512](/models/mistralai__ministral-3b-2512/) — F16, 9.413 GiB unrated

## Does not fit, or not profiled

 - cohere/command-a-plus — even Q4_K_M wants 132.618 GiB; pool is 96 GiB
- deepseek/deepseek-v4-flash-vision-exp — even Q4_K_M wants 184.902 GiB; pool is 96 GiB
- z-ai/glm-5.3 — even Q4_K_M wants 455.624 GiB; pool is 96 GiB
- qwen/qwen3.8-2.4t-a95b — even Q4_K_M wants 1450.719 GiB; pool is 96 GiB
- moonshotai/kimi-k3 — even Q4_K_M wants 1691.5 GiB; pool is 96 GiB
- z-ai/glm-5.2 — even Q4_K_M wants 455.624 GiB; pool is 96 GiB
- moonshotai/kimi-k2.7-code — even Q4_K_M wants 604.75 GiB; pool is 96 GiB
- nvidia/nemotron-3-ultra-550b-a55b — even Q4_K_M wants 333.063 GiB; pool is 96 GiB
- minimax/minimax-m3 — even Q4_K_M wants 259.405 GiB; pool is 96 GiB
- deepseek/deepseek-v4-pro — even Q4_K_M wants 967 GiB; pool is 96 GiB
- deepseek/deepseek-v4-flash — even Q4_K_M wants 172.465 GiB; pool is 96 GiB
- mistralai/mistral-large-2512 — even Q4_K_M wants 408.531 GiB; pool is 96 GiB
- deepseek/deepseek-v3.2 — even Q4_K_M wants 406.116 GiB; pool is 96 GiB
- meta-llama/llama-4-maverick — even Q4_K_M wants 242.5 GiB; pool is 96 GiB

## GPU

 [RTX PRO 6000 Blackwell](/self-host/rtx-pro-6000-blackwell/)

## How the number is made

Weights: a measured GGUF size when we have one, otherwise total parameters × bits per weight ÷ 8. Q4_K_M is treated as 4.83 bits/param. MoE memory uses total parameters, not active parameters.

KV cache, when layer and head geometry is sourced: 2 × kvLayers × kvHeads × headDim × context × 2 bytes (FP16, batch 1). Hybrid models use attention-layer count, not every layer. If geometry is missing, cache is omitted and that is stated.

1 GiB is added for CUDA/runtime. Apple unified memory uses RAM as the pool. CPU-only is a memory fit, not a speed claim. DDR generation is ignored for fit — it changes bandwidth, not whether the weights sit in memory.

RAM type (DDR4/DDR5/LPDDR) is not a field. It does not change whether a model fits.

[Self-host](/self-host/)

## Questions this page answers

 What LLM can I run on a 96GB card?

11 of 25 profiled open-weight models fit 96 GiB (vram) at F16, 8192 context. Best rated that fits: Gemma 4 31B.

 Which sourced cards publish 96 GiB?

RTX PRO 6000 Blackwell (96 GiB).

 Does NVIDIA pick the winner?

No. Ranking is arithmetic on sourced parameter counts. Quality leads. Unrated models that fit are listed separately, never scored zero.

## Continue your investigation

 - [Inspect open-weight models](/open-weights/)
- [Compare model variants](/families/)
- [Review requirements](/compare/)
- [Find serving tools](/tools/)
