For auto-scaling inference endpoints to zero, platform choice depends on your infrastructure needs. RunPod and Modal provide excellent performance for specialized GPU workloads with fast cold starts and per-second billing. For development teams already embedded in AWS, Amazon SageMaker Serverless Inference offers native integration. Other specialized platforms like Baseten and SiliconFlow are also recommended for their ease of use and cost-optimized auto-scaling across varied AI workloads.
1Amazon SageMakerBest for enterprises using AWS. It provides a managed way to scale to zero while keeping your inference workflows within the AWS ecosystem. Ideal if you already rely on AWS for other infrastructure needs.69%
2RunPodBest for performance-focused GPU inference. It offers rapid scaling from zero and per-second billing, which helps minimize costs during idle periods. Preferred by teams needing high-speed hardware with flexible deployment.61%
If your main requirement is GPU inference endpoints that automatically scale with traffic, including down to zero when idle, I’d shortlist Modal, Replicate, and Google Cloud Run. AWS SageMaker is also viable, particularly if you’re already deep in AWS.
My pick: Modal
Modal Inference is probably the best fit if you're building a custom inference service rather than simply exposing an existing model.
For inference endpoints specifically, the choice usually comes down to cold-start tolerance, GPU support, and how much infrastructure you want to manage. I’ll compare the leading scale-to-zero serverless options on those dimensions.
Why Modal: its inference infrastructure explicitly supports elastic GPU autoscaling down to zero, and its autoscaler exposes controls such as minimum/maximum containers, target concurrency, scale-up windows, and scale-down windows. Modal also offers GPU snapshotting and model/weight caching aimed at reducing cold-start latency.
One important caveat: scale-to-zero means cold starts. Modal's serverless HTTP servers can return a 503 while the first container is starting after being at zero, so your client/load balancer needs appropriate retries.
When I'd choose the others
Replicate is attractive if you want the least infrastructure work. Deployments can automatically scale from zero to hundreds of instances based on traffic, with configurable minimum/maximum instances and hardware.
Cloud Run is particularly compelling if you already have a Dockerized inference server. GPU-backed Cloud Run services can scale to zero, and its autoscaler scales based on request concurrency/CPU rather than GPU utilization. That latter point means you'll need to tune concurrency carefully for GPU workloads.
SageMaker Serverless Inference is convenient if you're already on AWS. AWS manages the scaling and automatically brings serverless endpoints down to zero when idle. However, the serverless option has limits such as 4 MB payloads and 60-second processing time, which can make it unsuitable for larger/longer-running LLM inference. docs.aws.amazon.com AWS also supports scale-to-zero with asynchronous inference for workloads where requests can be queued.
Bottom line
For a new inference API, I'd generally evaluate them in this order:
Modal → Replicate → Cloud Run → SageMaker
The deciding variable is really cold-start tolerance. If your endpoint must respond in tens of milliseconds even after long idle periods, true scale-to-zero isn't compatible with that requirement; you'd keep at least one warm replica. If a few seconds of startup latency is acceptable for occasional traffic, scale-to-zero can substantially reduce idle GPU cost.
If you tell me what model you're serving (e.g. Llama, Whisper, Stable Diffusion), GPU size, expected requests/sec, and acceptable cold-start latency, I can do a concrete Modal vs. Replicate vs. Cloud Run cost/architecture comparison.
When choosing a serverless platform that scales to and from zero for AI inference, the decision usually comes down to whether you are running standard CPU-bound models or heavy GPU-accelerated deep learning models (like LLMs or stable diffusion).
The top-tier specialized platforms for serverless GPU/CPU inference with scale-to-zero capabilities balance cost-efficiency, cold-start times, and developer ergonomics:
Modal – Best for developer-first Python/GPU workflows. Modal is a general serverless compute platform that treats AI inference as a first-class citizen. It scales seamlessly to and from zero, bills precisely per second, boasts some of the lowest raw GPU rates on the market (e.g., A100/H100), and lets you define infrastructure purely via a Python SDK.
Baseten – Best for production-grade model serving. Baseten is purpose-built for high-performance ML inference, offering robust auto-scaling to zero, optimized model runtimes (like vLLM), fast cold starts, and explicit controls over concurrency and replica counts.
Replicate – Best for pre-packaged or fast-to-deploy models. If you want to run open-source models with zero infrastructure management or deploy custom weights into an established ecosystem with pay-per-second scaling to zero, Replicate abstracts away almost everything.
Cerebrium – Best for ultra-low latency cold starts. Cerebrium focuses heavily on optimizing serverless GPU infrastructure (using snapshot restores and optimized container caching) to minimize the dreaded cold-start lag when scaling up from zero.
Standard Hyperscalers (AWS Lambda / Google Cloud Functions) – Best if your inference is lightweight and strictly CPU-bound (e.g., scikit-learn, small NLP pipelines). For heavy LLMs or computer vision, they lack the native, cost-effective GPU scale-to-zero agility of dedicated AI platforms.
To help narrow down the ideal platform, tell me:
What framework or model are you deploying (e.g., Llama 3, a custom PyTorch model, or a small CPU model)?
What are your expectations for cold-start latency tolerance?
Are you looking for a pure API service or a custom container environment?
If you mean GPU inference endpoints that automatically scale from 0 → N → 0, my current shortlist is:
Platform
Scale to zero
Cold-start story
GPU flexibility
Best fit
Modal
Yes
Excellent, especially with snapshots
High
Python-first teams
RunPod Serverless
Yes
Very strong; FlashBoot
Very high
Custom containers + cost/control
Replicate Deployments
Yes
Good
High
Fastest path from model → API
Baseten
Yes
Good, with warm replicas available
High
Managed production model serving
Google Cloud Run GPU
Yes
~5s instance startup
Moderate
Teams already on GCP
My recommendation
For a new inference service, I'd start with Modal. It has a particularly clean model of "this Python function is an autoscaling service": you can configure minimum/maximum containers, GPUs, and concurrency, while scale-to-zero is natural. Current comparisons also highlight its per-second billing and memory-snapshot support for reducing cold starts.
Choose RunPod Serverless instead if GPU choice and infrastructure control are more important. It currently advertises 13 GPU tiers from 16–280 GB VRAM, per-second billing, configurable worker counts, and scaling from zero to hundreds of workers.
Choose Cloud Run GPU if you're already deeply invested in GCP and want a conventional container deployment. Google currently supports L4 and RTX PRO 6000 Blackwell GPUs, and GPU-enabled services automatically scale down to zero. docs.cloud.google.com The tradeoff is that Cloud Run's GPU autoscaling is primarily driven by request concurrency/CPU rather than GPU utilization, so you'll want to tune concurrency carefully for your model.
Replicate is attractive if you want the least infrastructure work. Its Deployments support custom models, configurable min/max instances, traffic-based autoscaling, and scale-to-zero.
The important caveat
"Scale to zero" doesn't necessarily mean zero latency. Your first request after an idle period has to provision a worker and potentially download/load model weights. Replicate explicitly notes that setup can take several seconds; RunPod and Modal have invested heavily in reducing this with worker/model caching and snapshots.
So I'd make the decision based on your traffic pattern:
Bursty traffic + custom Python model: Modal
Bursty traffic + maximum GPU/container control: RunPod Serverless
Want an API around an existing/custom model quickly: Replicate
Production serving with more managed infrastructure: Baseten
Already standardized on Google Cloud: Cloud Run GPU
Traffic is nearly continuous: consider a warm/dedicated GPU instead—the economics of scale-to-zero become less compelling as utilization rises.
If you tell me the model (e.g. vLLM/Llama, Whisper, Stable Diffusion), GPU size, expected requests/sec, and acceptable cold-start latency, I can narrow this down to a concrete architecture and estimate the monthly cost.
For auto-scaling inference endpoints to and from zero (especially if you are running custom machine learning models that require GPUs or specific container runtimes), the "best" platform depends on whether you are deploying raw custom containers or utilizing managed AI inference infrastructure.
The top serverless platforms for scaling to and from zero break down by use case:
Google Cloud Run: Best for standard / CPU-based inference containers. It allows you to package any custom model code into a Docker container, scales seamlessly from zero to thousands of instances based on incoming HTTP traffic/concurrency, and you pay only for the exact millisecond execution time.
Modal: Best for serverless GPU inference and heavy AI workloads. Modal allows you to run arbitrary Python code and heavy deep learning models on serverless GPUs (like A10Gs or H100s) that scale completely down to zero when idle, spinning back up in seconds when a request hits.
Cerebrium: Best specialized infrastructure for fine-tuned or custom ML models. It provides sub-second spin-up times for serverless GPU microservices, handling the complex orchestration of scaling weights and containers to/from zero with per-second billing.
Baseten or RunPod Serverless: Great alternatives for high-performance LLM/vLLM deployments. They let you deploy custom or open-source models with scale-to-zero configurations, minimizing idle costs while keeping cold-start latencies optimized via warm pool settings.
To help narrow down the ideal platform, tell me:
Are you running CPU-only or GPU-intensive models (like LLMs or Diffusion)?
What is your tolerance for cold start latency when scaling up from zero?
Are you tied to a specific cloud provider ecosystem (AWS, GCP, Azure)?
Choosing the "best" serverless platform with scale-to-zero for AI inference depends heavily on whether your models require GPUs , how much you care about cold-start latency , and your preferred developer workflow.
The top serverless platforms for auto-scaling to and from zero break down by use-case:
Modal (Modal Labs) – Best Developer Experience & Python Integration
Why it shines: Modal allows you to turn arbitrary Python functions and container environments into serverless web endpoints with a simple decorator. It scales from zero to hundreds of GPUs seamlessly and bills down to the millisecond of active compute time.
Trade-off: Proprietary ecosystem; tightly bound to their Python SDK workflow.
RunPod Serverless (RunPod) – Best for Speed and Cost-Efficiency
Why it shines: RunPod features FlashBoot , which dramatically minimizes cold-start times (sometimes down to ~500ms). It provides robust per-second billing and an extensive hardware catalog ranging from affordable consumer cards to heavy enterprise GPUs.
Trade-off: Lower-level configuration required compared to abstraction-heavy platforms like Modal.
Baseten (Baseten) – Best for Production-Grade Enterprise ML
Why it shines: Built around the open-source Truss packaging framework, Baseten is tailored for production model serving. It supports high-performance backends like vLLM and TensorRT-LLM out-of-the-box, alongside native scaling to zero, fine-grained batching controls, and robust observability dashboards.
Trade-off: Higher price point on premium hardware and slight migration friction when adopting Truss.
DigitalOcean / Koyeb (DigitalOcean / Koyeb) – Best for General-Purpose & Unified Stacks
Why it shines: If your inference endpoints sit alongside traditional microservices, databases, and standard CPU workloads, platforms like DigitalOcean and Koyeb offer clean, per-second scale-to-zero infrastructure without forcing you into a pure AI-only cloud.
To help narrow down the ideal fit, tell me:
Are you serving LLMs / vision models (needing GPUs) or lightweight CPU-based models?
What is your primary framework (vLLM, Hugging Face, custom PyTorch/ONNX)?
For GPU inference endpoints that need true scale-to-zero, I’d shortlist Modal, Runpod Serverless, and Baseten. Based on their current capabilities, Modal is the strongest default choice for a new inference service if you want a developer-friendly platform rather than building around a hyperscaler.
Platform
Scale to zero
GPU inference
Cold-start story
Best fit
Modal
✅
Excellent
Very strong
General-purpose production inference
Runpod Serverless
✅
Excellent
FlashBoot, sub-200ms claimed
Cost-sensitive GPU inference / containers
Baseten
✅
Excellent
Fast cold starts
Production LLM serving
SageMaker Serverless
✅
⚠️ More constrained
Cold starts expected
AWS-native CPU/serverless workloads
1. Modal — my default pick
Modal is specifically designed around the pattern you're describing: deploy an inference function, let the platform provision GPUs when requests arrive, and scale back to zero when idle. Its current inference offering advertises scaling to 1,000+ GPUs and back to zero, with support for custom models and inference engines.
The programming model is also unusually clean—you can treat the endpoint much more like application code than Kubernetes infrastructure.
I'd choose Modal if:
You have custom PyTorch/vLLM/TensorRT/etc. inference.
Traffic is bursty.
You care about minimizing idle GPU spend.
You want streaming/WebSockets as well as ordinary HTTP.
You don't want to operate Kubernetes.
2. Runpod Serverless — strong alternative
Runpod's current Serverless offering is explicitly built around containerized GPU inference and scale-to-zero. It advertises per-second billing, autoscaling from zero to hundreds of workers, and FlashBoot cold starts under 200 ms. It also offers a broad selection of GPUs.
This is particularly attractive if you already have a Dockerized inference server and want relatively direct control over the GPU environment.
I'd choose Runpod if:
You already have a containerized inference stack.
GPU price/availability is a major concern.
You want lots of GPU types.
You're comfortable with a somewhat more infrastructure-oriented model.
3. Baseten — worth considering for LLM serving
Baseten has explicit replica autoscaling, including scale-to-zero, configurable minimum/maximum replicas, concurrency targets, scaling windows, and scale-down delays.
I'd look particularly closely at Baseten if your workload is primarily LLM inference and you want a more specialized model-serving product rather than a general serverless compute platform.
What about AWS SageMaker?
SageMaker Serverless Inference genuinely does scale to zero automatically and is pay-per-use.
However, there is an important distinction: AWS's Serverless Inference has constraints such as 4 MB payloads and 60-second processing limits, and AWS positions it primarily for workloads that can tolerate cold starts.
SageMaker Asynchronous Inference can also scale to zero and supports much larger/longer-running jobs, but it introduces a queue-based asynchronous architecture.
So I wouldn't make SageMaker Serverless my first choice for a modern, low-latency GPU LLM endpoint.
If GPU cost is the dominant concern, benchmark the same container on Runpod Serverless.
If you're deploying large LLMs and want model-serving-specific tooling, benchmark Baseten as well.
The biggest thing I'd benchmark before committing isn't the nominal GPU hourly price—it is cold-start latency + model load time + scale-out behavior under your actual traffic pattern. A platform that is 20% cheaper per GPU-hour can easily become more expensive if it spends substantially longer loading your model or requires keeping warm replicas.
If you tell me which model (e.g. Llama/Qwen/Whisper/Stable Diffusion), GPU size, expected requests/sec, and whether requests are streaming, I can do a concrete Modal vs Runpod vs Baseten cost/latency comparison for your workload.
When it comes to automated scaling to and from zero for AI and ML inference endpoints , traditional FaaS platforms like AWS Lambda fall short because they lack native GPU support. Instead, you need a specialized Serverless GPU platform that handles container orchestration, fast model weight loading, and scale-to-zero infrastructure seamlessly.
The "best" platform depends on whether you prioritize developer experience, cost, or heavy production-grade operations. The top platforms on the market excel in different areas:
Modal – Best for Developer Experience & Speed
How it works: You write standard Python code, apply a simple decorator (@app.cls ), and deploy via CLI. No Kubernetes or complex Dockerfiles required.
Scale-to-zero & Cold Starts: Exceptional filesystem snapshotting makes its cold starts some of the fastest in the industry (2–5 seconds for small models, though larger 7B+ LLMs will naturally take 15–30 seconds to load weights).
Tradeoff: It is slightly more expensive per GPU-hour than raw infrastructure providers, as you are paying for the polished DX. Explore details on Modal.
Baseten – Best for Production-Grade Inference Ops
How it works: Built specifically for high-performance model serving with production monitoring, explicit autoscaling controls, and optimized Triton/vLLM backends.
Scale-to-zero & Cold Starts: Scales smoothly to zero when idle and warms up efficiently under incoming traffic spikes.
Tradeoff: Steeper operational setup than a pure code-first abstraction like Modal, but far more robust for complex enterprise pipelines. Check out Baseten.
RunPod (Serverless) – Best for Cost-Efficiency & Control
How it works: Offers serverless GPU endpoints utilizing a global network of distributed workers, allowing you to pay strictly per millisecond/second of actual inference execution.
Scale-to-zero & Cold Starts: Features fast initialization variants (like FlashBoot) and gives you broad infrastructure flexibility.
Tradeoff: Requires a bit more manual configuration of your worker environment and templates compared to code-first platforms. Learn more via RunPod.
Google Cloud Run (with GPUs) – Best if you are locked into GCP
How it works: Cloud Run supports scale-to-zero container execution with L4 GPU availability.
Scale-to-zero & Cold Starts: Native integration with Google Cloud's ecosystem, making it easy to tie into existing IAM, VPCs, and logging.
Tradeoff: Less specialized for heavy LLM fine-tuning or exotic hardware adjustments compared to dedicated AI clouds like Modal or Baseten. Look into Google Cloud Run.
To help narrow down the ideal platform for your specific stack, let me know:
What type and size of model are you deploying (e.g., a small embedding model, a 7B/70B LLM, or a stable diffusion pipeline)?
What is your target latency constraint for cold starts from zero?
Modal, Baseten, and Google Cloud Run (with GPUs) are the leading serverless platforms that offer seamless auto-scaling to and from zero for machine learning inference endpoints.
The choice of the "best" platform depends on whether you are deploying custom deep-learning models (requiring specialized GPU hardware) or lightweight CPU-bound models.
Top Serverless Platforms for Inference Auto-Scaling to Zero
Platform
Primary Strength
GPU Support
Scale-to-Zero Behavior
Modal
Developer experience & fast cold starts for Python/AI workloads
Yes (NVIDIA T4, A10G, A100, H100)
Scales to zero instantly; extremely rapid container boot times.
Baseten
Production ML infrastructure with high-throughput vLLM support
Yes (Dedicated and serverless GPU options)
Scales to zero efficiently with optimized model-loading strategies.
Standard container-based serverless with general multi-language support
Yes (NVIDIA L4 GPUs GA)
Scales to zero natively for both CPU and GPU container workloads.
Amazon SageMaker Serverless Inference
Native AWS ecosystem integration
No (CPU only, up to 6 GB memory)
Scales down to 0 when idle; billed by the millisecond of execution.
Key Considerations for Scale-to-Zero Inference
Cold Start Latency: Scaling to zero means the next incoming request triggers a cold start, which can take anywhere from 2 to 30+ seconds depending on model weight size and container image pulling. Platforms like Modal optimize this using network-attached filesystems and memory snapshotting.
Hardware Requirements: Standard function-as-a-service platforms like AWS Lambda do not support GPUs . If your inference payload requires heavy tensor math (e.g., LLMs or stable diffusion), you must use GPU-enabled serverless infrastructure like Modal, Google Cloud Run, or Baseten.
Concurrency and Queuing: Ensure the platform supports request queuing during scale-from-zero events so that sudden traffic bursts do not drop initial client connections.
To help narrow down the ideal option, tell me:
What kind of model are you deploying (e.g., small NLP classifier, large LLM like Llama 3, or computer vision)?
What is your preferred cloud provider ecosystem (AWS, GCP, or cloud-agnostic)?
For GPU inference endpoints that need to scale from zero → many replicas → zero, I’d shortlist Modal, RunPod Serverless, and Baseten. They all support the core pattern, but optimize for different things.
Platform
Scale to zero
Cold-start story
Best fit
Modal
✅
Very fast; optimized for inference
Best overall developer experience
RunPod Serverless
✅
FlashBoot claims sub-200ms for supported setups
Broad GPU choice + cost-sensitive workloads
Baseten
✅
Optimized model cold starts
Production ML serving / observability
AWS SageMaker
✅
Generally slower provisioning from zero
AWS-native workloads, especially async
My default choice: Modal
Modal is particularly attractive if you're building the inference service yourself in Python. It supports elastic scaling to 1,000+ GPUs and back to zero, and has infrastructure specifically optimized for real-time inference, streaming, WebSockets, and dynamically batched inference.
The programming model is also unusually clean: you deploy Python code rather than having to build a whole Kubernetes-style serving stack.
When I'd choose RunPod instead
RunPod Serverless is compelling if GPU selection and price flexibility matter more than having the nicest developer abstraction. It supports arbitrary containers, scales workers from zero based on request demand, and bills by the second. Its current Serverless offering advertises GPUs ranging from 16 GB to 280 GB VRAM.
That's especially useful if you're running things like vLLM, custom CUDA images, image generation, or unusual model-serving stacks.
When I'd choose Baseten
Baseten is worth considering when inference is becoming a serious production service and you care about model-specific autoscaling controls. It lets you configure minimum/maximum replicas, concurrency targets, autoscaling windows, and scale-to-zero behavior.
One important AWS caveat
If you're thinking "we're already on AWS, so let's use SageMaker," check the workload carefully. SageMaker's Serverless Inference scales to zero, but the serverless offering is CPU-oriented; SageMaker's asynchronous endpoints can scale GPU/other instances to zero, but waking from zero can take minutes and requests may queue during provisioning.
So, if I were starting a new GPU inference API today:
Modal → my first choice for most teams.
RunPod Serverless → if GPU cost/selection and arbitrary containers are paramount.
Baseten → if you're building a more sophisticated production model-serving platform.
SageMaker → mainly when deep AWS integration outweighs the scale-from-zero latency tradeoff.
If you tell me model size (e.g. 7B/70B), GPU requirement, expected requests/sec, and whether latency or cost matters more, I can narrow this down and give you an architecture + rough monthly cost comparison.