Blog ·

RTX PRO 6000 vs H100 for LLM inference

By the NEOXIS team · 4 min read

Both cards serve large language models in production, but they are built very differently. The RTX PRO 6000 Blackwell has 96 GB of GDDR7 memory and runs FP4 natively. The H100 has 80 GB of much faster HBM3 and NVLink for multi-GPU work. Which one is better for you depends mostly on how big your model is and how many GPUs it has to be split across.

The short answer

  • Models that fit on one GPU, or on two to four: the RTX PRO 6000 usually gives more throughput per dollar.
  • One model split across eight GPUs, and training: the H100 and H200 pull ahead, thanks to NVLink.

The specifications that matter for inference

RTX PRO 6000 Server Edition H100 SXM H100 NVL
GPU memory 96 GB GDDR7 80 GB HBM3 94 GB HBM3
Memory bandwidth 1,597 GB/s 3.35 TB/s 3.9 TB/s
FP4 support Yes No No
GPU-to-GPU link PCIe 5.0 x16 NVLink, 900 GB/s NVLink bridge, 600 GB/s
Maximum power up to 600 W up to 700 W 350 to 400 W
MIG instances up to 4 up to 7 up to 7
Architecture Blackwell (2025) Hopper (2022) Hopper

Memory capacity decides what fits

For inference, the first question is whether the model and its KV cache fit on the GPUs. A 70B model in FP8 needs about 70 GB for its weights alone. On a 96 GB RTX PRO 6000 that leaves room for the KV cache on a single card. On an 80 GB H100 SXM it leaves about 10 GB, so many teams spread the same model over two cards.

FP4 halves the weights again, to about 35 GB for a 70B model, and only Blackwell runs it natively. On the H100 the smallest format with hardware support is FP8.

Memory bandwidth decides the speed of one conversation

When a model writes an answer for one user, the GPU reads the weights again for every new token, so memory bandwidth sets the pace. The H100 SXM has about twice the bandwidth of the RTX PRO 6000 Server Edition, so a single conversation runs faster on it.

With many requests in parallel, the GPU reuses each read across the whole batch. Compute and memory capacity then matter more, and the gap narrows. FP4 weights also halve the data the GPU has to read per token.

NVLink decides how far you can scale one model

The RTX PRO 6000 talks to other GPUs over PCIe 5.0 x16, at 128 GB/s in both directions combined. The H100 SXM uses NVLink at 900 GB/s. In November 2025 CloudRift benchmarked both with vLLM, and the results show what this means in practice:

Setup (CloudRift, vLLM) RTX PRO 6000 H100 SXM H200
1 GPU, GLM-4.5-Air, 4-bit 3,140 tokens/s 2,987 tokens/s not tested
4 GPUs, tensor parallel, Qwen3-Coder-480B, 4-bit about 1,600 tokens/s about 2,100 tokens/s about 3,400 tokens/s
8 GPUs, tensor parallel, GLM-4.6, FP8 about 900 tokens/s about 2,700 tokens/s about 3,600 tokens/s

On a single GPU the RTX PRO 6000 was slightly faster than the H100. Splitting one large model across four GPUs left it about 25% behind. Across eight GPUs, the H100 was three times faster.

In the same test, eight independent copies of the small model on eight RTX PRO 6000 cards reached 15,713 tokens per second. For serving many users, running copies side by side scales much better than splitting one model across all the GPUs.

What it costs to rent

Rental prices vary a lot from provider to provider. In September 2026, the price tracker getdeploying listed the RTX PRO 6000 from $0.65 per GPU-hour, with a median of $2.18 across 53 providers. The H100 started at $1.30, with a median of $3.39 across 57 providers. Per hour of GPU time, the RTX PRO 6000 is roughly 35 to 50% cheaper. Our own server with eight RTX PRO 6000 cards rents from $1.70 per GPU-hour, as a whole machine on a monthly contract with power, cooling, internet and rack space included.

For steady workloads, renting a whole server for a month at a fixed price removes the hourly price risk altogether.

Which one should you choose?

Choose the RTX PRO 6000 if:

  • your models fit on one, two or four GPUs, that is up to about 380 GB of weights and cache,
  • you serve many users at once and care about the cost per token,
  • you want to run FP4-quantized models.

Choose the H100 or H200 if:

  • one model has to span all eight GPUs,
  • the speed of a single answer matters more than the cost,
  • you train or fine-tune large models across several GPUs.

How we run it

Our server has eight RTX PRO 6000 Server Edition cards. For inference we recommend running each model on one, two or four GPUs and adding copies, instead of splitting a single model across all eight. See the server specifications or read how teams use it for private LLM hosting.

Sources