You need to train a 7B parameter model. Or run inference at scale. Or fine-tune on your own data.
Your laptop isn't cutting it. You need GPUs. Real GPUs. The kind that cost more than a car.
But where do you get them? AWS? GCP? Some random website with "vast" in the name? The prices vary by 10x. The reliability varies even more. And if you pick wrong, you'll spend a month debugging why your training run keeps failing instead of actually training.
This episode is your field guide to cloud GPU providers. What they cost. What they're good for. And where the hidden traps are.
The Big Three: AWS, GCP, Azure
AWS (EC2 + SageMaker)
The safe choice. The expensive choice. The "nobody ever got fired for choosing AWS" choice.
Instance Type GPU VRAM On-Demand Spot
| g4dn.xlarge | T4 | 16 GB | $0.526/hr | ~$0.16/hr |
|---|---|---|---|---|
| g5.xlarge | A10G | 24 GB | $1.006/hr | ~$0.30/hr |
| p3.2xlarge | V100 | 16 GB | $3.06/hr | ~$0.92/hr |
| p4d.24xlarge | A100 ร 8 | 320 GB | $32.77/hr | ~$9.83/hr |
| p5.48xlarge | H100 ร 8 | 640 GB | $98.32/hr | ~$29.50/hr |
AWS Pros:
Reliability is top-tier
SageMaker makes training jobs easy
Spot instances are actually available (unlike GCP)
Integration with the rest of AWS ecosystem
AWS Cons:
Most expensive option
GPU availability can be spotty for A100/H100
Complex pricing (egress, storage, IP addresses all add up)
Google Cloud Platform (GCP)
The ML-native choice. TensorFlow was built here. TPUs live here. But their GPU availability is... inconsistent.
Instance Type GPU VRAM On-Demand Spot
| n1-standard-4 + T4 | T4 | 16 GB | $0.35/hr | ~$0.10/hr |
|---|---|---|---|---|
| g2-standard-4 | L4 | 24 GB | $0.70/hr | ~$0.21/hr |
| a2-highgpu-1g | A100 | 40 GB | $3.67/hr | ~$1.10/hr |
| a2-ultragpu-1g | A100 80GB | 80 GB | $5.02/hr | ~$1.51/hr |
| a3-highgpu-8g | H100 ร 8 | 640 GB | ~$40/hr | Rarely available |
GCP Pros:
Best TPU support if you're in that ecosystem
Good integration with Vertex AI
Preemptible pricing can be very cheap
GCP Cons:
GPU availability is the worst of the big three
Spot/preemptible instances get reclaimed constantly
Support can be... Google-y (documentation > humans)
Azure
The enterprise choice. The "we already pay for Microsoft" choice. Actually quite good for AI workloads.
Instance Type GPU VRAM On-Demand Spot
NC4as T4 v3 T4 16 GB $0.526/hr ~$0.08/hr
NC24ads A100 v4 A100 80 GB $3.60/hr ~$1.08/hr
ND96asr A100 v4 A100 ร 8 640 GB $36.29/hr ~$10.89/hr
ND96isr H100 v5 H100 ร 8 640 GB ~$90/hr ~$27/hr
Azure Pros:
OpenAI partnership means easy GPT-4 access
Good availability for A100s
Spot pricing is aggressive
Best for Windows/enterprise workflows
Azure Cons:
Portal UX is... not great
Documentation can be fragmented
Less community content than AWS
The Challengers: Lambda Labs, RunPod, Vast.ai
These providers rent GPUs cheaper than the big three. Much cheaper. But with tradeoffs.
Lambda Labs
The "researcher-friendly" alternative. Simple pricing, good support, focused on ML workloads.
GPU VRAM Price/hr Notes
RTX A6000 48 GB $0.80 Great for inference
| A100 40GB | 40 GB | $1.10 | Training workhorse |
|---|---|---|---|
| A100 80GB | 80 GB | $1.60 | Large model training |
| H100 80GB | 80 GB | $2.40 | Latest and greatest |
| H100 ร 8 | 640 GB | $19.20 | Multi-node training |
Lambda Pros:
2-3x cheaper than AWS for same GPUs
Persistent storage (your data stays between sessions)
Jupyter notebooks built-in
Good customer support (actual humans)
Lambda Cons:
Smaller ecosystem than AWS/GCP
Occasional availability issues
Fewer regions (mainly US)
RunPod
The "serverless GPU" platform. Pay by the second. Good for bursty workloads.
GPU VRAM Community Cloud Secure Cloud
RTX 4090 24 GB $0.44/hr $0.69/hr
RTX A6000 48 GB $0.80/hr $1.19/hr
| A100 40GB | 40 GB | $1.19/hr | $1.69/hr |
|---|---|---|---|
| A100 80GB | 80 GB | $1.79/hr | $2.49/hr |
| H100 80GB | 80 GB | $2.69/hr | $3.59/hr |
RunPod Pros:
Cheapest option for many GPUs
Serverless endpoints (scale to zero)
Community cloud is very cheap (but less reliable)
Good for inference endpoints
RunPod Cons:
Community cloud can be unreliable (random failures)
Network storage costs add up
Support is community-based for lower tiers
Vast.ai
The "eBay of GPUs." Individuals rent out their personal GPUs. Cheapest option, highest variance.
GPU VRAM Price Range
RTX 3090 24 GB $0.20 - $0.50/hr
RTX 4090 24 GB $0.40 - $0.80/hr
RTX A6000 48 GB $0.60 - $1.20/hr
| A100 40GB | 40 GB | $0.80 - $1.50/hr |
|---|---|---|
| A100 80GB | 80 GB | $1.20 - $2.50/hr |
Vast.ai Pros:
Cheapest GPUs anywhere
Massive selection
Can find deals 5x cheaper than AWS
Vast.ai Cons:
Reliability is a coin flip
No support (you're renting from random people)
Data security concerns (who owns that machine?)
Machines can disappear mid-training
Setup is manual (SSH, configure yourself)
When to use Vast.ai:
Experiments and prototyping
When cost matters more than reliability
When you can checkpoint frequently
Never for production or sensitive data
๐ง
Spot vs On-Demand: The 70% Discount
Spot instances (AWS) / Preemptible (GCP) / Spot (Azure) are excess capacity sold at discount. The catch: they can be reclaimed with 2 minutes notice.
Provider Discount Interruption Rate Best For
AWS Spot Up to 90% 5-10% Training with checkpointing
GCP Preemptible Up to 80% 10-20% Fault-tolerant batch jobs
| Azure Spot | Up to 90% | 5-10% | Similar to AWS |
|---|---|---|---|
| Lambda Spot | N/A | N/A | Lambda doesn't do spot |
Spot Strategy for Training
python
Pseudo-code for spot-aware training
import signal
import sys
checkpoint_path = "/workspace/checkpoint.pt"
def save_checkpoint():
torch.save({
'epoch': epoch,
'model_state_dict': model.state_dict(),
'optimizer_state_dict': optimizer.state_dict(),
}, checkpoint_path)
print(f"Checkpoint saved to {checkpoint_path}")def handle_sigterm(signum, frame):
"""AWS/GCP send SIGTERM before reclaiming spot instance"""
print("Spot interruption detected! Saving checkpoint...")
save_checkpoint()
sys.exit(0)signal.signal(signal.SIGTERM, handle_sigterm)
Also checkpoint regularly
for epoch in range(num_epochs):
train_epoch()
if epoch % 5 == 0:
save_checkpoint()Rule of thumb: If your training job can checkpoint and resume, use spot. If it can't, pay for on-demand.
๐
Cost Comparison: Training a 7B Model
Let's say you're fine-tuning Llama 3.1 8B on a custom dataset. You need an A100 80GB.
| Provider | Price/hr | 1 Week Cost | 1 Month Cost |
|---|---|---|---|
| AWS On-Demand | $5.02 | $843 | $3,614 |
| AWS Spot | $1.51 | $254 | $1,087 |
| GCP On-Demand | $5.02 | $843 | $3,614 |
| GCP Spot | $1.51 | $254 | $1,087 |
| Lambda | $1.60 | $269 | $1,152 |
| RunPod Secure | $2.49 | $418 | $1,792 |
| RunPod Community | $1.79 | $301 | $1,289 |
| Vast.ai (avg) | $1.50 | $252 | $1,080 |
The lesson: Spot instances and alternative providers can cut your training costs by 70%. For a month-long training run, that's the difference between $3,600 and $1,000.
Hidden Costs
The hourly GPU price is just the start. Watch out for:
Cost AWS GCP Lambda RunPod
| Egress (data out) | $0.09/GB | $0.12/GB | $0.01/GB | Varies |
|---|---|---|---|---|
| Storage | $0.10/GB/mo | $0.04/GB/mo | $0.20/GB/mo | $0.10/GB/mo |
| IP address | $0.005/hr | $0.004/hr | Free | Free |
| Load balancer | $0.022/hr | $0.025/hr | N/A | N/A |
Example: Training a model that generates 100GB of checkpoints weekly:
AWS storage: $10/month
Downloading those checkpoints: $9 ร 4 = $36/month
Suddenly your "cheap" spot instance has $46 in extra costs
Choosing a Provider: Decision Matrix
What's your priority?
โ
โโโ Reliability is critical (production)
โ โโโ Need managed ML platform โ AWS SageMaker or GCP Vertex
โ โโโ Just need reliable GPUs โ AWS On-Demand or Azure
โ
โโโ Cost is critical (research, experiments)
โ โโโ Can handle interruptions โ AWS/GCP Spot
โ โโโ Need persistence โ Lambda Labs
โ โโโ Need serverless scaling โ RunPod
โ โโโ Maximum cheap, can tolerate risk โ Vast.ai
โ
โโโ Training large models (70B+)
โ โโโ Need multi-node โ AWS p4d/p5 or Lambda 8xH100
โ
โโโ Inference at scale
โโโ Steady load โ Lambda or RunPod dedicated
โโโ Bursty load โ RunPod serverless or AWS SageMaker endpoints๐ฌ
Practical Takeaways
Spot instances save 70% โ if your workload can handle interruptions, always use spot
Lambda Labs is the sweet spot โ 2-3x cheaper than AWS with better reliability than Vast
Vast.ai for experiments only โ cheapest prices, but don't run production there
Watch hidden costs โ egress and storage can add 20-30% to your bill
Multi-node training = AWS or Lambda โ few providers offer reliable 8xGPU instances
Start with on-demand, move to spot โ prove your code works, then optimize for cost
๐ก๏ธ
What's Next?
Episode 69: TPUs vs GPUs โ Google bet on TPUs. NVIDIA owns GPUs. Which is better for your workload? We'll dive into systolic arrays, XLA compilation, JAX vs PyTorch, and when Google's custom chips actually win.
โ Previous
Ep 67: Apple Silicon for AI
Next โ
Ep 69: TPUs vs GPUs
Next: Episode 69 โ TPUs vs GPUs
NVIDIA GPUs power virtually every AI model you've heard of. So why did Google build its own chip? Here's the TPU story.
This is part of a 98-episode series covering AI engineering from tokens to production deployment.