GPU MODELS / H20

H20 cloud GPU pricing

3 listings across 3 providers. Cheapest published rate: $1.010/GPU/hr at PPIO. Sorted by per-GPU hourly price. Every row links to the provider's own published price.

WHAT IS THE H20

The H20 is the China-market Hopper: full 96 GB of HBM3 and 4.0 TB/s of memory bandwidth, but tensor compute capped at 148 TFLOPS of FP16 - a fraction of the H100's. That inversion flips the usual trade: per dollar of compute it is weak, but per gigabyte of memory it is one of the cheapest ways to hold a large model. It exists because export rules limit compute, not memory, and the market prices it accordingly.

MAKER
NVIDIA
ARCHITECTURE
Hopper
LAUNCHED
2024
MEMORY
96 GB HBM3
MEMORY BANDWIDTH
4.0 TB/s
TENSOR COMPUTE
148 TFLOPS FP16/BF16 tensor (296 TFLOPS FP8, 44 TFLOPS FP32)
POWER DRAW
400 W
INTERCONNECT
SXM5 (HGX H20, NVLink)

WHAT IT IS GOOD AT

Large-model inference
The design target in practice. 96 GB and 4.0 TB/s feed decode well, and inference is far less compute-bound than training. Panels in China price it as the default large-model serving chip.
Memory-bound HPC and analytics
Workloads limited by bandwidth rather than FLOPS get most of an H100's memory system at a lower price.
LLM training
Weak. 148 TFLOPS of FP16 is under 8% of an H100 SXM's tensor throughput; training at scale on H20 is a memory purchase, not a compute one.

WHO SHOULD NOT RENT THE H20

Compute-bound training: the FP16 cap is the entire point of the SKU. Buy H100-class silicon where rules allow.
Deployments needing NVLink-scale multi-GPU training - the HGX H20 interconnect exists, but the compute per GPU makes large training clusters uneconomic.
ProviderConfigVRAMLocation$/hr$/GPU/hrSource
PPIOPPIOH20 × 196 GBChina$1.010$1.010receipt
CloudGPUCloudGPUH20 × 196 GBNot published$1.290$1.290receipt
AutoDLAutoDLH20 × 196 GBChina$1.560$1.560receipt

THE H20 MARKET, BY THE NUMBERS

PROVIDERS
3
CONFIGURATIONS
3
MEDIAN $/GPU/HR
$1.29
SPREAD
1.5x

3 providers publish 3 H20 configurations on the panel today. Published per-GPU hourly rates run from $1.010 (PPIO) to $1.56 (AutoDL), a 2x spread between the cheapest and the most expensive published rate for the same chip. The median listing sits at $1.29/GPU/hr. 1 rows were reported in stock by the provider at last observation; the rest are published catalog rates. All rates are on-demand - committed-use and reserved pricing is excluded.

H20 PRICE TREND, DAILY PANEL SNAPSHOTS

Across every H20 row on the panel, the daily floor moved from $1.290 to $1.010/GPU/hr over 10 days of tracking (-21.7%). Snapshots run daily; the full history is public in the repo.

H20 PROVIDER NOTES
PPIO · 1 config · 1 region
Sets the panel floor for this model.
from $1.010/hr
CloudGPU · 1 config · 1 in stock
1 of 1 configs reported in stock at last check.
from $1.290/hr
AutoDL · 1 config · 1 region
Published catalog rates.
from $1.560/hr
BUYING H20: WHAT THE PANEL SAYS

At the current floor, one H20 running around the clock costs about $738 a month (730 hours of on-demand arithmetic). Compare the alternatives below before committing - the same budget often buys more than one chip class.

H20 VS THE ALTERNATIVES

H20 vs H100 SXM
Same Hopper die family and similar memory system (96 GB vs 80 GB, 4.0 vs 3.35 TB/s), but the H20's FP16 tensor throughput is about 7% of the H100's. For inference holding a big model, compare the panel prices; for training, there is no comparison.
H20 vs H800
The H800 is the earlier export-compliant Hopper: far more compute (about 10x the H20's FP16) but 80 GB and lower interconnect. For training under the rules, the H800 wins; for pure inference-per-dollar, the H20's newer memory system competes.

Compare live rates: H20 vs H100 SXM · H20 vs H800 · H20 vs L20 · H20 vs Ascend 910B

H20 PRICING FAQ

What is the cheapest H20 cloud GPU right now?
The cheapest published H20 rate on the panel is $1.010 per GPU-hour at PPIO (H20 x 1, China). Every price links to the provider's own published rate.
How much does a H20 cost per hour?
Across 3 providers tracked today, published per-GPU hourly rates for H20 run from $1.01 to $1.56, with a median of $1.29. Committed-use and reserved rates are excluded; everything shown is on-demand.
How many providers rent H20 GPUs?
3 providers publish 3 distinct H20 configurations on Compute Cafe. 1 of those rows were reported in stock by the provider's live endpoint at last observation.
Are H20 prices going up or down?
Over the 10 days Compute Cafe has tracked the panel, the H20 floor is falling: $1.290 to $1.010/GPU/hr (-21.7%). Daily snapshots are public in the repo, so the trend is auditable.
What does a H20 cost per month?
At the current floor of $1.010/GPU/hr, a H20 running 24/7 costs about $738 per month (730 hours of on-demand arithmetic on the cheapest published rate). Committed-use contracts can price lower; the panel tracks on-demand rates only.
Is the H20 good for LLM inference?
Yes - it is effectively an inference SKU. The 96 GB HBM3 and 4.0 TB/s bandwidth are close to H100-class memory, which is what decode speed depends on. The FP16 cap (148 TFLOPS) hurts prefill throughput on very large batches, but for most serving shapes it is the China market's default large-model chip.
Why is the H20 so much weaker on paper than the H100?
Export compliance. The rules cap compute density, not memory, so NVIDIA shipped a Hopper with the tensor cores cut down and the memory system intact. Compare it on memory per dollar, not FLOPS per dollar.

See also alternatives to the H20 · cheapest GPU-hour per model · inference chips · training chips · by location