GPU MODELS / ALTERNATIVES TO H800

Alternatives to the H800

The H800 was the first export-compliant Hopper for the China market: near-H100 tensor throughput but NVLink cut from 900 to 400 GB/s and FP64 essentially removed. It trains well inside a server but scales out worse than an H100 because the intra-node fabric is halved. Where the panel lists it, it sits between the H100 and the later, more restricted H20.

People usually compare the H800 against H100 SXM, H20. Prices below are live on-demand published rates from the panel.

LOADING LIVE DATA...

AlternativeMemoryBandwidthCheapest live rateHead to head
H100 SXM80 GB HBM33.35 TB/s-H800 vs H100 SXM
H2096 GB HBM34.0 TB/s-H800 vs H20

WHEN EACH ALTERNATIVE WINS

H100 SXM
The H800 keeps about 76% of the H100's FP16 tensor throughput (1,513 vs 1,979 TFLOPS) but halves NVLink to 400 GB/s and drops FP64. Single-node training is close; scale-out training and HPC are not.
H20
The H20 is the later, tighter compliance cut: 96 GB of memory but only 148 TFLOPS FP16. The H800 has roughly 10x the compute; the H20 has the newer memory system. Training says H800, memory-per-dollar inference says H20.

H800 ALTERNATIVES - FAQ

What is the best alternative to the H800?
It depends on the workload. H100 SXM: The H100 SXM is the chip the current generation of frontier models was trained on, and the benchmark every other datacenter GPU is priced against. H20: The H20 is the China-market Hopper: full 96 GB of HBM3 and 4.0 TB/s of memory bandwidth, but tensor compute capped at 148 TFLOPS of FP16 - a fraction of the H100's. Live rates for each are in the table above.
Why look for H800 alternatives?
FP64 scientific computing - the double-precision units are fused off to roughly 1 TFLOPS.
How much does the H800 cost compared to its alternatives?
Each alternative's current cheapest published rate is listed in the table above; figures recompute hourly from provider-published prices.

H800 pricing · all comparisons