How Much Does It Cost to Rent a GPU in the Cloud? [Live Pricing Guide]

The answer: it depends.
Cloud GPU rental costs depend on the GPU model you're looking for, supply and demand across the marketplace, whether you're renting an on-demand or interruptible instance, and many other factors. Prices can vary from one day to the next.
To give you an idea of current marketplace rates on Vast.ai, below is a live pricing table that updates automatically with the rental costs of popular GPUs in real time.
In this guide, you'll also find a quick overview of the most in-demand GPUs on Vast.ai and tips for choosing the right hardware for your workload and budget.
Live Cloud GPU Pricing on Vast.ai
Updated automatically with current marketplace data, this card shows pricing, inventory, and availability across the GPUs featured or referenced in this guide.
- "From" hourly price is the current 25th-percentile marketplace rate. It provides a lower-market reference without relying on one unusually cheap listing.
- "Median" hourly price represents the typical hourly rental rate across current offers.
- "On platform" is the current number of GPUs represented in the marketplace snapshot.
The card is intentionally limited to GPUs covered in this guide. Inventory order is a live platform snapshot, not an annual rental-popularity ranking.
Live Vast.ai and Published Provider Price Comparison
For directional context, the Vast.ai row below updates automatically. Competitor rows use date-stamped rates from each provider's published pricing page.
| Provider | H100 SXM | H200 | B200 | B300 | Published basis |
|---|---|---|---|---|---|
| Vast.ai | Loading… | Loading… | Loading… | Loading… | Marketplace P25 “from” price |
| RunPod | $2.99/GPU/hr | $4.39/GPU/hr | $5.89/GPU/hr | $7.39/GPU/hr | Pod rate per GPU |
| Lambda | $3.99/GPU/hr | Not listed | $6.69/GPU/hr | Not listed | 8-GPU instance, price per GPU |
| CoreWeave | $6.16/GPU/hr | $6.31/GPU/hr | $8.60/GPU/hr | Contact sales | 8-GPU on-demand instance divided by 8 |
Vast.ai values refresh every five minutes. Competitor values are published references checked July 20, 2026; configurations and billing terms differ.
The Most Popular Cloud GPUs Right Now
So, which GPUs are people actually renting? Here's a look at ten of the most in-demand GPU options on Vast.ai today.
1.RTX 5090Live price
- 32GB GDDR7 memory
- 1.79 TB/s bandwidth
The NVIDIA GeForce RTX 5090 is one of the most sought-after consumer GPUs and boasts the largest inventory of all machines on Vast.ai. It's well suited for fine-tuning smaller models, inference, and demanding creative workloads.
2.RTX 4090Live price
- 24GB GDDR6X memory
- 1.01 TB/s bandwidth
Although now a legacy card, the RTX 4090 remains a popular option as a dependable workhorse for those who don't need cutting-edge specs. It continues to offer excellent performance for tasks like LLM inference and image generation.
3.H200Live price
- 141GB HBM3e memory
- 4.8 TB/s bandwidth
One of the highest-demand cards on Vast.ai, the H200 is a leading enterprise AI accelerator for serious model training and high-throughput inference.
4.H100 SXMLive price
- 80GB HBM3 memory
- 3.35 TB/s bandwidth
The enterprise-grade H100 SXM is well suited for workloads involving massive datasets, large-scale simulations, and advanced AI training for foundation models.
5.RTX PRO 6000 SandWSLive price
- 96GB GDDR7 memory
- 1.79 TB/s bandwidth
Both the RTX PRO 6000 Server and Workstation Editions are designed for high-performance professional applications in areas like engineering, simulation, and media production. These machines balance AI acceleration with visualization and rendering capabilities.
6.A100 SXM4Live price
- 80GB HBM2e memory
- 2.04 TB/s bandwidth
Built on the Ampere architecture, the A100 SXM4 has somewhat lower rental demand but continues to power production AI environments thanks to its mature software ecosystem and broad availability.
7.B200Live price
- 192GB HBM3e memory
- 8.0 TB/s bandwidth
Running at high utilization on Vast.ai, the B200 is an incredibly in-demand GPU. It serves as a unified AI platform for develop-to-deploy pipelines for organizations of any size.
8.L40SLive price
- 48GB GDDR6 memory
- 864 GB/s bandwidth
The L40S is best suited for inference and rendering-focused workloads rather than large-scale model training. With plenty of memory capacity for many production AI and graphics applications, it's a popular choice for those looking to balance performance with cost.
9.RTX 3090Live price
- 24GB GDDR6X memory
- 936 GB/s bandwidth
The RTX 3090 is a budget-friendly option for AI experimentation as well as gaming and professional tasks, delivering solid performance despite its age.
10.Tesla V100Live price
- 32GB HBM2 memory
- 900 GB/s bandwidth
Another older card, the Tesla V100 continues to offer a good fit for legacy AI workloads and research environments when accessing the latest hardware isn't a priority.
How to Choose the Right GPU for Your Budget
To get the most bang for your buck, first identify your workload. The amount of VRAM you need may eliminate some GPUs from consideration entirely. From there, look for the hardware that delivers the performance you need without paying for unnecessary processing power.
Below are some recommended GPUs based on different types of workloads:
| Workload | Recommended GPUs |
|---|---|
| LLM inference | L40S, H100, H200 |
| Fine-tuning smaller models | RTX 4090, RTX 5090, A100 |
| Training large language models | H100, H200, B200, B300 |
| 3D rendering | RTX PRO 6000, L40S |
| Scientific computing | A100, H100 |
Tip: Consider how quickly each GPU can complete your workload. A more powerful GPU may have a higher hourly rate, but if it finishes the job much sooner, your total compute cost may be lower.
If you're prioritizing value over maximum performance, compare the live from price, median, and availability in the pricing table at the top of this page.
Start Building with the Right GPU on Vast.ai
Once you've found the hardware that fits your workload and budget, the next step is spinning up an instance. Vast.ai offers flexible ways to deploy.
Need a fully managed inference endpoint? Vast.ai Serverless automatically provisions GPUs and scales capacity up or down using predictive optimization, with no tiers or pricing surcharges.
Want to get up and running quickly? The Vast.ai Model Library provides ready-to-run templates for popular open-weight models and frameworks, so you can get started in minutes instead of configuring everything from scratch.
Browse live GPU prices and launch your instance on Vast.ai today.


