MODULE 10  ยท  GPU & Hardware

Cloud GPU Pricing: Where to Train and Inference Without Going Broke

You need GPU compute. Cloud providers all want your money. Here's the real pricing breakdown so you don't go broke.

๐Ÿ“… Mar 2026
โฑ 10 min read
๐ŸŽฏ Episode 68 of 98
Cloud GPUPricingAWSCost
In this episode

You need to train a 7B parameter model. Or run inference at scale. Or fine-tune on your own data.

Your laptop isn't cutting it. You need GPUs. Real GPUs. The kind that cost more than a car.

But where do you get them? AWS? GCP? Some random website with "vast" in the name? The prices vary by 10x. The reliability varies even more. And if you pick wrong, you'll spend a month debugging why your training run keeps failing instead of actually training.

This episode is your field guide to cloud GPU providers. What they cost. What they're good for. And where the hidden traps are.

The Big Three: AWS, GCP, Azure

AWS (EC2 + SageMaker)

The safe choice. The expensive choice. The "nobody ever got fired for choosing AWS" choice.

Instance Type GPU VRAM On-Demand Spot

g4dn.xlargeT416 GB$0.526/hr~$0.16/hr
g5.xlargeA10G24 GB$1.006/hr~$0.30/hr
p3.2xlargeV10016 GB$3.06/hr~$0.92/hr
p4d.24xlargeA100 ร— 8320 GB$32.77/hr~$9.83/hr
p5.48xlargeH100 ร— 8640 GB$98.32/hr~$29.50/hr

AWS Pros:

Reliability is top-tier

SageMaker makes training jobs easy

Spot instances are actually available (unlike GCP)

Integration with the rest of AWS ecosystem

AWS Cons:

Most expensive option

GPU availability can be spotty for A100/H100

Complex pricing (egress, storage, IP addresses all add up)

Google Cloud Platform (GCP)

The ML-native choice. TensorFlow was built here. TPUs live here. But their GPU availability is... inconsistent.

Instance Type GPU VRAM On-Demand Spot

n1-standard-4 + T4T416 GB$0.35/hr~$0.10/hr
g2-standard-4L424 GB$0.70/hr~$0.21/hr
a2-highgpu-1gA10040 GB$3.67/hr~$1.10/hr
a2-ultragpu-1gA100 80GB80 GB$5.02/hr~$1.51/hr
a3-highgpu-8gH100 ร— 8640 GB~$40/hrRarely available

GCP Pros:

Best TPU support if you're in that ecosystem

Good integration with Vertex AI

Preemptible pricing can be very cheap

GCP Cons:

GPU availability is the worst of the big three

Spot/preemptible instances get reclaimed constantly

Support can be... Google-y (documentation > humans)

Azure

The enterprise choice. The "we already pay for Microsoft" choice. Actually quite good for AI workloads.

Instance Type GPU VRAM On-Demand Spot

NC4as T4 v3 T4 16 GB $0.526/hr ~$0.08/hr

NC24ads A100 v4 A100 80 GB $3.60/hr ~$1.08/hr

ND96asr A100 v4 A100 ร— 8 640 GB $36.29/hr ~$10.89/hr

ND96isr H100 v5 H100 ร— 8 640 GB ~$90/hr ~$27/hr

Azure Pros:

OpenAI partnership means easy GPT-4 access

Good availability for A100s

Spot pricing is aggressive

Best for Windows/enterprise workflows

Azure Cons:

Portal UX is... not great

Documentation can be fragmented

Less community content than AWS

โšก

The Challengers: Lambda Labs, RunPod, Vast.ai

These providers rent GPUs cheaper than the big three. Much cheaper. But with tradeoffs.

Lambda Labs

The "researcher-friendly" alternative. Simple pricing, good support, focused on ML workloads.

GPU VRAM Price/hr Notes

RTX A6000 48 GB $0.80 Great for inference

A100 40GB40 GB$1.10Training workhorse
A100 80GB80 GB$1.60Large model training
H100 80GB80 GB$2.40Latest and greatest
H100 ร— 8640 GB$19.20Multi-node training

Lambda Pros:

2-3x cheaper than AWS for same GPUs

Persistent storage (your data stays between sessions)

Jupyter notebooks built-in

Good customer support (actual humans)

Lambda Cons:

Smaller ecosystem than AWS/GCP

Occasional availability issues

Fewer regions (mainly US)

RunPod

The "serverless GPU" platform. Pay by the second. Good for bursty workloads.

GPU VRAM Community Cloud Secure Cloud

RTX 4090 24 GB $0.44/hr $0.69/hr

RTX A6000 48 GB $0.80/hr $1.19/hr

A100 40GB40 GB$1.19/hr$1.69/hr
A100 80GB80 GB$1.79/hr$2.49/hr
H100 80GB80 GB$2.69/hr$3.59/hr

RunPod Pros:

Cheapest option for many GPUs

Serverless endpoints (scale to zero)

Community cloud is very cheap (but less reliable)

Good for inference endpoints

RunPod Cons:

Community cloud can be unreliable (random failures)

Network storage costs add up

Support is community-based for lower tiers

Vast.ai

The "eBay of GPUs." Individuals rent out their personal GPUs. Cheapest option, highest variance.

GPU VRAM Price Range

RTX 3090 24 GB $0.20 - $0.50/hr

RTX 4090 24 GB $0.40 - $0.80/hr

RTX A6000 48 GB $0.60 - $1.20/hr

A100 40GB40 GB$0.80 - $1.50/hr
A100 80GB80 GB$1.20 - $2.50/hr

Vast.ai Pros:

Cheapest GPUs anywhere

Massive selection

Can find deals 5x cheaper than AWS

Vast.ai Cons:

Reliability is a coin flip

No support (you're renting from random people)

Data security concerns (who owns that machine?)

Machines can disappear mid-training

Setup is manual (SSH, configure yourself)

When to use Vast.ai:

Experiments and prototyping

When cost matters more than reliability

When you can checkpoint frequently

Never for production or sensitive data

๐Ÿ”ง

Spot vs On-Demand: The 70% Discount

Spot instances (AWS) / Preemptible (GCP) / Spot (Azure) are excess capacity sold at discount. The catch: they can be reclaimed with 2 minutes notice.

Provider Discount Interruption Rate Best For

AWS Spot Up to 90% 5-10% Training with checkpointing

GCP Preemptible Up to 80% 10-20% Fault-tolerant batch jobs

Azure SpotUp to 90%5-10%Similar to AWS
Lambda SpotN/AN/ALambda doesn't do spot

Spot Strategy for Training

python

Pseudo-code for spot-aware training

import signal

import sys

snippet
code
checkpoint_path = "/workspace/checkpoint.pt"
def save_checkpoint():
torch.save({
        'epoch': epoch,
        'model_state_dict': model.state_dict(),
        'optimizer_state_dict': optimizer.state_dict(),
    }, checkpoint_path)
    print(f"Checkpoint saved to {checkpoint_path}")

def handle_sigterm(signum, frame):

example
code
"""AWS/GCP send SIGTERM before reclaiming spot instance"""
    print("Spot interruption detected! Saving checkpoint...")
    save_checkpoint()
    sys.exit(0)

signal.signal(signal.SIGTERM, handle_sigterm)

Also checkpoint regularly

for epoch in range(num_epochs):

example
code
train_epoch()
    if epoch % 5 == 0:
        save_checkpoint()

Rule of thumb: If your training job can checkpoint and resume, use spot. If it can't, pay for on-demand.

๐Ÿ“Š

Cost Comparison: Training a 7B Model

Let's say you're fine-tuning Llama 3.1 8B on a custom dataset. You need an A100 80GB.

ProviderPrice/hr1 Week Cost1 Month Cost
AWS On-Demand$5.02$843$3,614
AWS Spot$1.51$254$1,087
GCP On-Demand$5.02$843$3,614
GCP Spot$1.51$254$1,087
Lambda$1.60$269$1,152
RunPod Secure$2.49$418$1,792
RunPod Community$1.79$301$1,289
Vast.ai (avg)$1.50$252$1,080

The lesson: Spot instances and alternative providers can cut your training costs by 70%. For a month-long training run, that's the difference between $3,600 and $1,000.

Hidden Costs

The hourly GPU price is just the start. Watch out for:

Cost AWS GCP Lambda RunPod

Egress (data out)$0.09/GB$0.12/GB$0.01/GBVaries
Storage$0.10/GB/mo$0.04/GB/mo$0.20/GB/mo$0.10/GB/mo
IP address$0.005/hr$0.004/hrFreeFree
Load balancer$0.022/hr$0.025/hrN/AN/A

Example: Training a model that generates 100GB of checkpoints weekly:

AWS storage: $10/month

Downloading those checkpoints: $9 ร— 4 = $36/month

Suddenly your "cheap" spot instance has $46 in extra costs

๐Ÿ’ก

Choosing a Provider: Decision Matrix

What's your priority?

โ”‚

โ”œโ”€โ”€ Reliability is critical (production)

โ”‚ โ”œโ”€โ”€ Need managed ML platform โ†’ AWS SageMaker or GCP Vertex

โ”‚ โ””โ”€โ”€ Just need reliable GPUs โ†’ AWS On-Demand or Azure

โ”‚

โ”œโ”€โ”€ Cost is critical (research, experiments)

โ”‚ โ”œโ”€โ”€ Can handle interruptions โ†’ AWS/GCP Spot

โ”‚ โ”œโ”€โ”€ Need persistence โ†’ Lambda Labs

โ”‚ โ”œโ”€โ”€ Need serverless scaling โ†’ RunPod

โ”‚ โ””โ”€โ”€ Maximum cheap, can tolerate risk โ†’ Vast.ai

โ”‚

โ”œโ”€โ”€ Training large models (70B+)

โ”‚ โ””โ”€โ”€ Need multi-node โ†’ AWS p4d/p5 or Lambda 8xH100

โ”‚

โ””โ”€โ”€ Inference at scale

example
code
โ”œโ”€โ”€ Steady load โ†’ Lambda or RunPod dedicated
    โ””โ”€โ”€ Bursty load โ†’ RunPod serverless or AWS SageMaker endpoints

๐Ÿ”ฌ

Practical Takeaways

Spot instances save 70% โ€” if your workload can handle interruptions, always use spot

Lambda Labs is the sweet spot โ€” 2-3x cheaper than AWS with better reliability than Vast

Vast.ai for experiments only โ€” cheapest prices, but don't run production there

Watch hidden costs โ€” egress and storage can add 20-30% to your bill

Multi-node training = AWS or Lambda โ€” few providers offer reliable 8xGPU instances

Start with on-demand, move to spot โ€” prove your code works, then optimize for cost

๐Ÿ›ก๏ธ

What's Next?

Episode 69: TPUs vs GPUs โ€” Google bet on TPUs. NVIDIA owns GPUs. Which is better for your workload? We'll dive into systolic arrays, XLA compilation, JAX vs PyTorch, and when Google's custom chips actually win.

โ† Previous

Ep 67: Apple Silicon for AI

Next โ†’

Ep 69: TPUs vs GPUs

Next: Episode 69 โ€” TPUs vs GPUs

NVIDIA GPUs power virtually every AI model you've heard of. So why did Google build its own chip? Here's the TPU story.

This is part of a 98-episode series covering AI engineering from tokens to production deployment.

โ† Previous Ep 67: Apple Silicon for AI: Why Macs Are Legit for Inference