LLM Enhancements

H100 vs H200: Memory, Bandwidth and What It Means for LLM Inference
Short answer: The H100 and H200 share the same NVIDIA Hopper architecture and the same Tensor Core compute. The difference is memory: the H200 has 141 GB of HBM3e at 4.8 TB/s, versus 80 GB of HBM3 at 3.35 TB/s on the H100 SXM. For LLM inference, that usually means larger batches, longer contexts and fewer GPUs per model.
If you are choosing between the two for inference, or deciding whether an existing H100 fleet is still the right foundation, the spec sheet tells a simpler story than most comparisons suggest. This guide puts the official NVIDIA numbers side by side, explains why memory matters more than FLOPS for serving LLMs, and gives a practical way to decide.
H100 vs H200 specs at a glance
All figures below come from NVIDIA's product pages (H100 and H200, both accessed Oct 9, 2026). NVIDIA marks H200 figures as preliminary, and Tensor Core figures are quoted with sparsity.
| Spec | H100 SXM | H200 SXM | H100 NVL | H200 NVL |
|---|---|---|---|---|
| GPU memory | 80 GB | 141 GB | 94 GB | 141 GB |
| Memory type | HBM3 | HBM3e | Not listed | HBM3e |
| Memory bandwidth | 3.35 TB/s | 4.8 TB/s | 3.9 TB/s | 4.8 TB/s |
| FP8 Tensor Core | 3,958 TFLOPS | 3,958 TFLOPS | 3,341 TFLOPS | 3,341 TFLOPS |
| BF16 Tensor Core | 1,979 TFLOPS | 1,979 TFLOPS | 1,671 TFLOPS | 1,671 TFLOPS |
| Max TDP | Up to 700 W | Up to 700 W | 350 to 400 W | Up to 600 W |
| NVLink | 900 GB/s | 900 GB/s | 600 GB/s | 900 GB/s per GPU (bridge) |
| MIG | Up to 7 @ 10 GB | Up to 7 @ 18 GB | Up to 7 @ 12 GB | Up to 7 @ 16.5 GB |
| Form factor | SXM | SXM | PCIe dual-slot | PCIe dual-slot |
The H100 SXM's HBM3 memory type is confirmed in NVIDIA's architecture deep dive, which describes the H100 SXM5 as the first GPU with HBM3, configured with 80 GB across five HBM3 stacks (NVIDIA Technical Blog, "NVIDIA Hopper Architecture In-Depth", Mar 22, 2022).
What is the same
Both GPUs are Hopper parts. In the SXM form factor, NVIDIA lists identical FP64, FP32, TF32, BF16, FP16, FP8 and INT8 Tensor Core figures, the same 900 GB/s NVLink, and the same configurable power ceiling of up to 700 W. NVIDIA also states that HGX H200 boards in four-way and eight-way configurations are compatible with both the hardware and software of HGX H100 systems (NVIDIA Newsroom, Nov 13, 2023). In other words, the H200 is not a new architecture. Software, kernels and serving stacks that run on H100 carry over.
What is different: memory capacity and bandwidth
NVIDIA describes the H200 as the first GPU with HBM3e, offering 141 GB at 4.8 TB/s, "nearly double the capacity of the NVIDIA H100 GPU with 1.4X more memory bandwidth" (NVIDIA H200 product page, accessed Oct 9, 2026). Against the H100 SXM, that works out to about 76% more memory (141 GB vs 80 GB) and about 43% more bandwidth (4.8 TB/s vs 3.35 TB/s).
At node level, an eight-way HGX H200 provides 1.1 TB of aggregate high-bandwidth memory and over 32 petaflops of FP8 compute (NVIDIA Newsroom, Nov 13, 2023). Eight H100 SXM GPUs at 80 GB each total 640 GB.
Why memory matters more than FLOPS for inference
LLM inference has two phases. Prefill processes the prompt in parallel and keeps the GPU busy. Decode generates one token at a time, and NVIDIA describes it as memory-bound: the speed of moving weights, keys and values from memory dominates latency (NVIDIA Technical Blog, Nov 17, 2023). If you are new to these concepts, our guide to what LLM inference is walks through them.
That has two consequences for H100 vs H200.
Bandwidth sets the decode ceiling
Each decode step reads the weights. With 43% more bandwidth, the H200 can, in principle, move the same weights faster, which lowers inter-token latency when decode is the bottleneck.
Capacity sets batch size and context length
Memory holds the weights plus the KV cache, which grows linearly with batch size and sequence length (NVIDIA, Nov 17, 2023). Using NVIDIA's rule of thumb of about two bytes per parameter in 16-bit precision, a 70B-parameter model needs roughly 140 GB for weights alone. That exceeds one H100's 80 GB and nearly fills one H200's 141 GB. Quantized to 8-bit, the same weights take about 70 GB: on an H100 SXM that leaves around 10 GB for KV cache and activations, while on an H200 it leaves around 70 GB. More free memory means more concurrent sequences per GPU, and more sequences per pass over the weights means more throughput. These are approximations that ignore framework overhead.
This is also why NVIDIA's own H200 comparison should be read carefully. Its Llama 2 70B inference result of 1.9X faster compares an H100 SXM at batch size 8 with an H200 SXM at batch size 32, and its GPT-3 175B result of 1.6X uses batch size 64 vs 128 on eight GPUs (NVIDIA H200 product page, accessed Oct 9, 2026). A large part of the gain comes from fitting bigger batches into memory. For the PCIe parts, NVIDIA cites up to 1.7x faster LLM inference for H200 NVL over H100 NVL (same source).
Power and efficiency
NVIDIA states that the H200 delivers its performance "all within the same power profile as the H100", and both SXM parts list a configurable maximum of up to 700 W (NVIDIA H200 product page; NVIDIA H100 product page, accessed Oct 9, 2026). If the H200 serves more tokens within the same power envelope, energy per token falls, but the size of that effect depends on your model and batch profile.
H100 or H200: how to choose for inference
The H100 is often sufficient when:
- Your models fit comfortably in 80 GB per GPU with room for KV cache, for example smaller models or quantized mid-size models.
- Prompts and outputs are short and concurrency is moderate.
- You already own H100 nodes and the bottleneck is software utilization rather than memory.
The H200 tends to pay off when:
- You serve 70B-class or larger models and want fewer GPUs per replica.
- You run long contexts, RAG with large retrieved passages, or agent loops that keep long histories.
- You need high concurrency, where KV-cache capacity limits batch size.
- You want to drop H200 boards into existing HGX H100-compatible infrastructure.
In practice, benchmark your own model, prompt lengths, output lengths and concurrency on both. Published multipliers are tied to specific configurations.
Where B200 fits
The next step up is Blackwell. NVIDIA lists DGX B200 with eight Blackwell GPUs, 1,440 GB of total GPU memory and 64 TB/s of HBM3e bandwidth (NVIDIA DGX B200 page, accessed Oct 9, 2026). NVIDIA's H100 page also cites SemiAnalysis InferenceX figures from April 2026 of about $0.02 per million tokens on B200 with TensorRT-LLM versus about $0.09 on H100 with vLLM for GPT-OSS-120B (NVIDIA H100 product page, accessed Oct 9, 2026). Note that the two figures use different serving software, a reminder that the software stack moves cost per token as much as the silicon does.
Before you buy: get more from the nodes you have
Whether you run H100 or H200, the gap between the spec-sheet ceiling and delivered throughput is set by software: batching, scheduling, memory management and decoding. EucaX Cortex installs on top of existing HGX H100, HGX H200 and B200 nodes. EucaX states that operators typically unlock 3x or more serving capacity, up to 6x on HGX-class systems, depending on model, workload and configuration, and measures the uplift against your own baseline. Before adding another hardware cycle, you can watch the technical demo or request a live test on your configuration.
FAQ
Is the H200 just an H100 with more memory?
Essentially, yes. Both use the Hopper architecture with identical Tensor Core compute in the SXM form factor. The H200 upgrades memory to 141 GB of HBM3e at 4.8 TB/s, compared with 80 GB of HBM3 at 3.35 TB/s on the H100 SXM.
How much faster is the H200 than the H100 for LLM inference?
NVIDIA reports 1.9X on Llama 2 70B and 1.6X on GPT-3 175B, both preliminary and both using larger batch sizes on the H200. Real gains depend on whether your workload is limited by memory capacity, bandwidth or software.
Can H200 GPUs go into existing H100 servers?
NVIDIA states that HGX H200 boards are compatible with the hardware and software of HGX H100 systems, and that partner server makers can update existing systems with an H200. Confirm specifics with your server vendor.
Do the H100 and H200 use the same power?
Both SXM versions list a configurable maximum TDP of up to 700 W, and NVIDIA says the H200 operates within the same power profile as the H100. The PCIe H200 NVL lists up to 600 W versus 350 to 400 W for the H100 NVL.
Sources
- [1]NVIDIA H100 GPU product page and specifications, accessed Oct 9, 2026
- [2]NVIDIA H200 GPU product page and specifications (preliminary), accessed Oct 9, 2026
- [3]NVIDIA Technical Blog, "NVIDIA Hopper Architecture In-Depth", published Mar 22, 2022
- [4]NVIDIA Newsroom, "NVIDIA Supercharges Hopper, the World's Leading AI Computing Platform", Nov 13, 2023
- [5]NVIDIA Technical Blog, "Mastering LLM Techniques: Inference Optimization", published Nov 17, 2023
- [6]NVIDIA DGX B200 product page and specifications, accessed Oct 9, 2026
See what this means for your infrastructure.
Book a discovery call and get a benchmark projection for your hardware and model.
SCHEDULE A CALL