The rounding tax: a $2.00 H100 that costs more than a $3.24 one

Every GPU price on the panel is a rate, and every rate hides a second number: the smallest unit of time the provider will bill you for. Rent a machine for a 20-minute experiment and the meter does not care that you used 20 minutes - it cares about the quantum. Billed by the second, you pay for 20 minutes. Billed by the hour, you pay for 60. The hourly rate on the pricing page is the same either way, which is exactly why the quantum never makes the headline.
Tonight the panel holds verified billing terms at both ends of the market. SpotGPUs publishes an on-demand H100 SXM at $0.89 per GPU-hour, billed per minute, never reclaimed. gpu.ai lists the same chip at $3.24 on its community tier, billed per second. VoltageGPU's H100 is $6.95 with per-second billing, AWS's p5.4xlarge is $6.88 with partial hours billed per second, and Google's A3 block works out to $11.06 per GPU-hour with a one-minute minimum followed by per-second increments - the last two verified against the clouds' own pricing pages. Most of the panel's 128 providers publish no quantum at all, and the safe assumption for silence is the hour.
The tax, in dollars
Run the arithmetic on two ordinary workloads. A 20-minute experiment - a fine-tune smoke test, a batch of evals, a render check - costs $0.30 at the per-minute floor, $1.08 on gpu.ai's per-second machine, and $2.29 at AWS. Now picture an hourly-billed provider charging $2.00, a price that looks like the second-cheapest machine on this list. The 20-minute job there costs the full $2.00 - nearly seven times the floor, and more than the per-second machine at triple its sticker. The interactive day is worse: eight sessions of 12 minutes each is 96 minutes of actual compute. Per-second and per-minute providers bill 96 minutes. The hourly provider bills eight hours, $16.00, more than AWS's $11.01 for the same work on a machine with triple the sticker price.
| Provider | Rate /GPU-hr | Billed by | 20-min job | 8 x 12-min day |
|---|---|---|---|---|
| SpotGPUs | $0.89 | Minute | $0.30 | $1.42 |
| gpu.ai (community) | $3.24 | Second | $1.08 | $5.18 |
| Machine0 | $4.85 | Minute | $1.62 | $7.76 |
| AWS p5.4xlarge | $6.88 | Second | $2.29 | $11.01 |
| VoltageGPU | $6.95 | Second | $2.32 | $11.12 |
| Hypothetical hourly provider | $2.00 | Hour | $2.00 | $16.00 |
The last row is the whole argument. A $2.00 hourly-billed H100 is not cheaper than a $3.24 per-second one. For anything under about 40 minutes of actual use per started hour, it is more expensive - and the gap grows with every start-stop cycle, because the tax is collected per session, not per day. The sticker comparison and the real comparison part ways precisely where exploratory work lives.
When the tax doesn't matter
Long runs amortize rounding into noise. A three-day training job that ends 4 minutes into its final hour overpays by 56 minutes on an hourly meter - 1.3% of the bill, worth knowing but not worth switching providers over. This is why the quantum is invisible in most comparisons: the workloads that dominate GPU spending (multi-day training, always-on inference) barely feel it. The workloads that feel it are the ones that dominate GPU usage - development, debugging, evaluation, experiments - the hundreds of short sessions that never make it into a capacity plan. If your work is steady-state, shop on rate. If your work is a loop, shop on the quantum first and the rate second.
The buyer's checklist
- Estimate your job shape before comparing rates: many short starts make the quantum the price; one long run makes it a rounding error.
- Compute effective cost, not sticker cost: rate times ceil(runtime divided by quantum). A 20-minute job at an hourly meter is a 3x multiplier whatever the rate says.
- Read minimums as part of the quantum: Google's per-second billing starts after a one-minute minimum and AWS bills partial hours per second - both harmless except for sub-minute jobs.
- Treat silence as hourly. Most of the panel publishes no billing quantum; assume the most expensive interpretation until the provider says otherwise.
- Do not confuse granularity with availability: a per-second machine that is reclaimed mid-job costs more than a per-hour one that stays up. Quantum and reclaim policy are separate checks.
The GPU market trains buyers to read one number per row. The row has always had two: what the hour costs, and how little of the hour you are allowed to buy. The second number decides which of the two you actually pay.