GPU MODELS / A40

A40 cloud GPU pricing

19 listings across 14 providers. Cheapest published rate: $0.090/GPU/hr at SpotGPUs. Sorted by per-GPU hourly price. Every row links to the provider's own published price.

WHAT IS THE A40

The A40 is Ampere's big-memory visualization card: 48 GB of GDDR6, RT cores, and vGPU support. Predating the L40/L40S, it still serves and renders competently, and often rents at a discount that makes it interesting for 13B-34B inference.

MAKER
NVIDIA
ARCHITECTURE
Ampere
LAUNCHED
2020
MEMORY
48 GB GDDR6
MEMORY BANDWIDTH
696 GB/s
TENSOR COMPUTE
150 TFLOPS FP16 tensor dense (300 sparse)
POWER DRAW
300 W
INTERCONNECT
PCIe 4.0 x16 (NVLink bridge on pairs)

WHAT IT IS GOOD AT

Mid-size inference
48 GB serves 34B quantized.
Rendering and VDI
Its design role.

WHO SHOULD NOT RENT THE A40

AI serving where the L40S rents close on price - 2.4x tensor throughput gap.
Memory-bandwidth-bound LLM decoding - GDDR6 caps it.
ProviderConfigVRAMLocation$/hr$/GPU/hrSource
SpotGPUsSpotGPUsA40 × 148 GBNot published$0.090$0.090receipt
AI GalaxyAI GalaxyA40 × 148 GBChina$0.327$0.327receipt
RunPodRunPodA40 × 148 GBGlobal$0.350$0.350receipt
MatPoolMatPoolA40 × 147.5 GBChina$0.423$0.423receipt
AutoDLAutoDLA40 × 148 GBChina$0.467$0.467receipt
gpu.aigpu.aiA40 × 148 GBNot published$0.490$0.490receipt
GPUBrazilGPUBrazilA40 × 145 GBNorth America$0.537$0.537receipt
Hostnot GPUHostnot GPUA40 × 148 GBCanada Central$0.540$0.540receipt
Hostnot GPUHostnot GPUA40 × 148 GBEU West$0.540$0.540receipt
GPUBrazilGPUBrazilA40 × 145 GBAsia-Pacific$0.649$0.649receipt
Denvr DataworksDenvr DataworksA40 × 448 GBCanada$2.600$0.650receipt
E2E NetworksE2E NetworksA40 × 148 GBIndia$1.001$1.001receipt
InHostedInHostedA40 × 148 GBIndia$1.001$1.001receipt
ExoscaleExoscaleA40 × 148 GBEU$1.045$1.045receipt
ExoscaleExoscaleA40 × 448 GBEU$4.181$1.045receipt
ExoscaleExoscaleA40 × 848 GBEU$8.363$1.045receipt
ExoscaleExoscaleA40 × 248 GBEU$2.091$1.045receipt
ACCACCA40 × 148 GBSingapore$1.220$1.220receipt
VultrVultrA40 × 148 GBNot published$1.712$1.712receipt

THE A40 MARKET, BY THE NUMBERS

PROVIDERS
14
CONFIGURATIONS
19
MEDIAN $/GPU/HR
$0.65
SPREAD
19.0x

14 providers publish 19 A40 configurations on the panel today. Published per-GPU hourly rates run from $0.090 (SpotGPUs) to $1.71 (Vultr), a 19x spread between the cheapest and the most expensive published rate for the same chip. The median listing sits at $0.65/GPU/hr. 16 rows were reported in stock by the provider at last observation; the rest are published catalog rates. All rates are on-demand - committed-use and reserved pricing is excluded.

CHEAPEST A40 BY GEOGRAPHY
chinaAI Galaxy$0.327/hrunited statesGPUBrazil$0.537/hrcanadaHostnot GPU$0.540/hrindiaE2E Networks$1.001/hrsingaporeACC$1.220/hr
A40 PRICE TREND, DAILY PANEL SNAPSHOTS

Across every A40 row on the panel, the daily floor moved from $0.350 to $0.090/GPU/hr over 11 days of tracking (-74.3%). The median continuously-tracked listing held flat at $0.65. Snapshots run daily; the full history is public in the repo.

A40 PROVIDER NOTES
SpotGPUs · 1 config · 1 in stock
Sets the panel floor for this model.
from $0.090/hr
AI Galaxy · 1 config · 1 region
Published catalog rates.
from $0.327/hr
RunPod · 1 config · 1 region · 1 in stock
1 of 1 configs reported in stock at last check.
from $0.350/hr
MatPool · 1 config · 1 region · 1 in stock
1 of 1 configs reported in stock at last check.
from $0.423/hr
AutoDL · 1 config · 1 region
Published catalog rates.
from $0.467/hr
gpu.ai · 1 config · 1 in stock
1 of 1 configs reported in stock at last check.
from $0.490/hr
GPUBrazil · 2 configs · 2 regions · 2 in stock
2 of 2 configs reported in stock at last check.
from $0.537/hr
Hostnot GPU · 2 configs · 2 regions · 2 in stock
2 of 2 configs reported in stock at last check.
from $0.540/hr
BUYING A40: WHAT THE PANEL SAYS

At the current floor, one A40 running around the clock costs about $66 a month (730 hours of on-demand arithmetic). Node shape matters: the single-GPU floor is $0.090/GPU/hr against $1.045/GPU/hr on an 8-GPU node (+1061.5%). Per-GPU floors by config size: 1x $0.09 · 2x $1.05 · 4x $0.65 · 8x $1.05. Compare the alternatives below before committing - the same budget often buys more than one chip class.

A40 VS THE ALTERNATIVES

A40 vs L40S
The L40S has 2.4x the tensor compute and FP8. The A40 only wins when the panel prices it well below - check the spread.

Compare live rates: A40 vs L40S

A40 PRICING FAQ

What is the cheapest A40 cloud GPU right now?
The cheapest published A40 rate on the panel is $0.090 per GPU-hour at SpotGPUs (A40 x 1, Not published). Every price links to the provider's own published rate.
How much does a A40 cost per hour?
Across 14 providers tracked today, published per-GPU hourly rates for A40 run from $0.09 to $1.71, with a median of $0.65. Committed-use and reserved rates are excluded; everything shown is on-demand.
How many providers rent A40 GPUs?
14 providers publish 19 distinct A40 configurations on Compute Cafe. 16 of those rows were reported in stock by the provider's live endpoint at last observation.
Where is A40 cheapest?
By published per-GPU hourly rate, the cheapest geography for A40 today is china at $0.33/hr (AI Galaxy). The spread across 5 tracked geographies runs to $1.22/hr in singapore.
Are A40 prices going up or down?
Over the 11 days Compute Cafe has tracked the panel, the A40 floor is falling: $0.350 to $0.090/GPU/hr (-74.3%). The median continuously-tracked listing moved from $0.65 to $0.65 (+0.0%). Daily snapshots are public in the repo, so the trend is auditable.
What does a A40 cost per month?
At the current floor of $0.090/GPU/hr, a A40 running 24/7 costs about $66 per month (730 hours of on-demand arithmetic on the cheapest published rate). Committed-use contracts can price lower; the panel tracks on-demand rates only.
Is a full 8-GPU A40 node cheaper than a single card?
On today's panel: the single-GPU floor is $0.090/GPU/hr and the 8-GPU node floor is $1.045/GPU/hr - more expensive per GPU on a full node (+1061.5%). Check the table for the exact configs behind each floor.
Is the A40 good for AI?
Adequate, not fast: it lacks FP8 and Ada's tensor throughput. At the right hourly rate it serves 13B-34B models economically.

See also alternatives to the A40 · cheapest GPU-hour per model · inference chips · training chips · by location