USE CASES / BATCH-INFERENCE
GPU pricing for batch inference
Batch inference (offline scoring, embeddings, evaluations) has no latency budget, so the cheapest chip that fits the model wins. Small cards with good bandwidth dominate; memory bandwidth per dollar is the number that matters, and spot-priced legacy parts are often the best deal on the panel.
LOADING LIVE DATA...
Serving instead? Inference-chip pricing.