GLM 5.3 now online!Try it →
TokenGO

─── BLOG

Cold starts in elastic auto-scaling, and a solution

Aug 8, 2026 · TokenGO

Why "millisecond" wake times are not inference latency, and how suspension keeps weights in VRAM without starting from zero every time.

Teams building serverless GPUs will eventually face a common problem. Functions scale down to zero when idle, and machines spin up only when a request arrives. The waiting period from zero to being ready to work is the cold start.

Latency has two layers, often conflated

The latency experienced by users consists of two layers. One layer is the time it takes for the platform to spin up the container and load model weights into VRAM; this is the part the platform can optimize. The other layer is a hard physical constraint: the larger the model, the more data needs to be moved into VRAM. No matter how much you optimize, this cannot be compressed to zero. For a 70-billion parameter model, just moving the weights into VRAM takes tangible time.

Two layers of latency: platform spin-up versus the physical cost of loading weights

The "millisecond-level" latency touted in many marketing materials refers to waking up an empty container, not the model being truly ready. When customers mistake idle container latency for inference latency, things break in production. Actual tests reveal latency in the double-digit seconds, and complaints naturally follow.

Don't start from zero every time: suspend

Pushing down cold start times relies on not starting from scratch every time. One approach is to suspend the spun-up container instead of destroying it. The model weights remain in VRAM, merely yielding the compute resources temporarily. When the next request arrives, the container is awakened. The weights are still there, skipping the massive overhead of reloading.

autoscale:
  min_replicas: 0
  suspend_threshold_sec: 300   # Suspend after 5 minutes of no requests
  resume_target_p99_ms: 200    # Target latency for first token upon resume
  suspend_cost_model: shared   # VRAM cost during suspension is centrally scheduled/billed by the platform

This approach has a prerequisite: suspended containers still occupy VRAM, and this resource cannot be considered purely wasted. We've seen two ways to handle this: either whoever suspends the container bears the cost (keeping the billing straightforward), or the suspended state is treated as a tradable resource centrally scheduled by the platform, so teams yielding idle VRAM pay less. The latter is more flexible, but the accounting must be crystal clear.

Huge differences in real-world application

For a speech-to-text API, scaling to zero normally costs nothing. When a sudden burst of requests hits, waking from a suspended state returns the first transcription in the 200-millisecond range, rather than making the user wait over a dozen seconds. We benchmarked a mid-sized speech model: waking from suspension versus a cold start from zero showed an order-of-magnitude difference in first-token latency.

Conversely, for tasks like image generation that inherently take several seconds to execute, the cold start gap isn't as critical, so keeping them permanently suspended might not be worth it.

The cost: VRAM cannot be completely freed

The speed gained from suspension comes at the cost of not fully releasing VRAM. On an 80GB card, if you suspend just one 24GB model, the remaining 56GB cannot be allocated to anyone else. The platform has to strike a balance: which models are worth keeping suspended, and which can tolerate a slower start from zero.

An 80GB card with 24GB held by a suspended model

For steady traffic that is latency-sensitive, permanent suspension is cost-effective; for sparse traffic that doesn't care about a few extra seconds, scaling down to zero saves more.

Give the choice to the user

Our approach is to let the customer draw that line themselves, rather than the platform imposing a one-size-fits-all policy. The platform minimizes the switching cost between the two states; customers choose based on their business needs and can change it at any time.

We also provide defaults: latency-sensitive small to medium models default to suspend, while large-scale models default to scale-to-zero. However, they can override this at any time. For the first week a new API goes live, it runs on scale-to-zero. We observe its actual cold start ratio before recommending whether to enable suspend, ensuring decisions are based on data, not guesswork.

Default policy: suspend for small and medium models, scale-to-zero for large models

A set of our benchmark comparisons

For the same mid-sized speech model, we tested the initial response latency across three states:

StateFirst-token latency
Cold start from zero1800 ms
Wake from suspend180 ms
Always-on hot state40 ms

First-token latency across cold start, suspend, and always-on

Waking from suspension is an order of magnitude faster than a cold start from zero. While it's a few times slower than an always-on hot state, the VRAM cost is much lower.

Therefore, our default is neither suspend-all nor scale-to-zero-all; it is categorized by model profile. Latency-sensitive small to medium models default to suspend, while large-scale models that inherently boot slowly default to scale-to-zero. As mentioned, for the first week a new API goes live, it runs on scale-to-zero, and we observe its actual cold start ratio before recommending whether to enable suspension.

Cold starts inherently are not the enemy; they are an inevitable cost of elastic auto-scaling. The real mastery lies in what comes after: how to accurately draw the line between always-on and on-demand, ensuring most requests feel no wait time while not paying for idle resources unnecessarily.

We don't measure our success by how high our suspend rate is, but by whether the user-side p99 first-token latency remains within an acceptable range while keeping the idle cost on the bill sufficiently low. It only counts when both of these are true simultaneously.

─── NEXT

Try TokenGo's serverless inference API now →