USE CASES / FINE-TUNING

GPU pricing for fine-tuning

Fine-tuning sits between serving and pre-training: you need enough VRAM to hold the model plus optimizer state, but not a training cluster. QLoRA and LoRA stretch smaller cards a long way - a 24 GB consumer card handles 7B-13B adapters, 48 GB workstation parts reach 34B, and 80 GB HBM parts take 70B-class models. These are the chips on the panel that fit that profile.

LOADING LIVE DATA...

Serving instead? Inference-chip pricing.