August 2026 Product Update

The past couple of months have been busy at Vast.ai. From fresh GPU supply to an expanded Model Library, here are the latest updates.
Surging B200 Capacity in August and September
Hundreds of new NVIDIA B200 GPUs are already live on Vast.ai, with more capacity arriving through September. Whether you're training models or serving production inference, you can rent by the hour or reserve long-term without waiting months for hardware to become available.
Search "B200" in the Vast.ai console to grab capacity as it comes online, and check live pricing here anytime.
New Templates: Kimi K3, Gemma 4, and More
Kimi K3 templates are now available in the Vast.ai Model Library.
As Moonshot AI's most capable model to date, Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model featuring native multimodal capabilities and a massive one-million-token context window. It's the world's first open-weight 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and complex reasoning tasks.
Choose between the full-precision vLLM flagship on 8xB300, or single-node Unsloth 1-bit and 2-bit quantized variants that run on much smaller GPU configurations.
Alongside Kimi K3, we've added a wide range of new templates to the library, including Kimi K2.7 Code; Gemma 4 12B, E2B, and E4B; DeepSeek V4 Flash; Hunyuan HY3; and Krea 2 Turbo.
Platform Updates: Notifications, Webhooks, and Other Improvements
We've expanded our notification system with per-event controls and webhook delivery. Renters can now receive alerts for instance lifecycle events, interruptible instance outbids, disk warnings, and more. Notifications can be delivered by email or sent directly to an HTTPS endpoint you run, making it easy to trigger your own scripts.
You can now also connect your Hugging Face Storage Buckets directly to Vast.ai. Every GPU instance you rent will be able to pull datasets and model checkpoints straight from your bucket - and push results back when jobs finish - without manual transfers or re-uploads. This means you can keep your data on Hugging Face while your compute runs on Vast GPUs.
In addition, we've added support for ARM-based NVIDIA GPUs (including the GB10 found in DGX Spark systems), improved the Billing page with clearer autobilling information, updated serverless endpoint analytics with improved granularity and additional metrics, fixed earnings display issues, and rolled out UI improvements across the Search page and Template Editor.
Our Commitment
As always, our mission remains the same: making high-performance AI infrastructure accessible to everyone, everywhere. By expanding GPU availability and shipping new platform features and developer tools, we're continuing to help our users build, train, fine-tune, and deploy AI workloads faster and more affordably, at any scale.
Need help getting started? Contact us anytime at support@vast.ai or join our Discord server to connect with the community and discover our latest updates as they happen.
Change Log
New Features
- Expanded notification system for renters, with per-event controls and webhook delivery (enable Notifications and Webhooks on your Settings page; read the full list here):
- New notification events: instance lifecycle (e.g., stopped, offline, back online after scheduling), outbid alerts, disk warnings, and more. Enable them on the Settings page under Notifications & Webhooks.
- Notification Settings: a new section in the console under Settings lets you turn each event on or off individually.
- Webhooks: send any notification event to an HTTPS endpoint you run, and use it to trigger your own scripts: post instance alerts into Slack, pause automation on low balance, or run cleanup when an instance goes offline. Deliveries are signed and retried.
- API, CLI, and SDK support: notification preferences, the inbox, and webhook management are all available programmatically.
- Hugging Face Storage Buckets can be added as a Cloud Connection on your Settings page, available now for all users. Read more about HF Storage Buckets and Vast Cloud Storage docs here.
- GB10s (DGX Sparks) are now listed for rent on the platform (and Vast.ai now supports ARM-based NVIDIA GPUs).
- Improved Billing page header messaging around autobilling and made saved cards more visible.
- Updated serverless endpoint analytics with improved granularity and additional metrics. Opt in on Settings → Early Access.
- Updated Search Page UI: improved filter UX, page scrolling, and other general improvements
- Updated Template Editor UI: improved flow for private registry images, and other general improvements.
- Expanded B200 supply.
Issues Resolved
- Clear prior template extra_filters on template switch
- Fixed earnings display issue
- Notification email delivery timing
- Earnings PDF export now displays invoice number
- Various page crashing issues
New Templates
LLMs & Multimodal:
- Kimi K3 (including 1-bit and 2-bit quantized variants)
- Kimi K2.7 Code
- DeepSeek V4 Flash
- Hunyuan HY3
- Gemma 4 12B, Gemma 4 E2B, Gemma 4 E4B
- Granite 4.0 H Small
- Qwen3-VL 30B A3B Instruct
- Seed-OSS 36B Instruct
Image Generation:
New Guides
- Webhooks Example: Send notifications to Slack via webhooks
- Webhooks Example: Send notifications to Google Chat via webhooks
Top Blog Posts
- Build an AI Agent on Vast.ai in <100 Lines of Python
- Everything You Need to Know About the NVIDIA Blackwell Ultra B300
- NVIDIA B300 vs H200: The Blackwell Ultra Upgrade
- NVIDIA Rubin Platform: Everything We Know So Far
- How Much Does It Cost to Rent a GPU in the Cloud? Live Pricing Guide
- The Future of AI Inference in 2026
- Why AI Agents Cost More Than Chatbots
- Matryoshka Vector Embeddings
- Why Vast.ai Is Ideal for Mission-Critical Production Systems


