GLM 5.3 now online!Try it →
TokenGO

─── BLOG

Decoupling the inference optimization layer from elastic scheduling

Aug 13, 2026 · TokenGO

The inference optimization layer should be decoupled from elastic scheduling, allowing each to iterate independently.

The inference optimization layer should be decoupled from elastic scheduling, allowing each to iterate independently.

When building inference services, there is a choice we simply cannot avoid. Should optimizing inference performance be bundled tightly with underlying resource scheduling, or should they be kept separate? We chose to separate them, and we are increasingly convinced this was the right call.

What does inference optimization involve?

It involves model quantization, trading lower precision for speed, like compressing from 16-bit float down to 8-bit, which saves both VRAM and compute. It involves speculative decoding, guessing a few steps ahead and then verifying, which reduces serial waiting so the user sees the first token faster. It involves batching, grouping multiple requests to execute together, driving up throughput per unit of time. And then there are various low-level kernel tweaks.

When you do these things well, the same GPU can process significantly more requests per second, driving down the cost per token. We benchmarked a typical workload: after proper optimization, the cost per million tokens dropped by over 30%.

Typical workload: cost per million tokens dropped by over 30% after optimization

The cost of tight coupling

If optimizations are hardcoded into a specific scheduling framework, users are paralyzed: if they want to change the scaling strategy, they can't touch the optimization; if they want a new optimization, they can't touch the scheduler. They are locked in on both ends. For instance, if you fine-tune a brilliant quantization scheme, you might have to rewrite it entirely just to switch scheduling frameworks.

We prefer decoupling. The optimization layer focuses solely on running the model as fast and cheaply as possible. The scheduling layer focuses solely on when to spin up machines, how many to spin up, and when to scale down. The two layers communicate through clean APIs.

Optimization runs the model. Scheduling runs the machines. They meet at an API.

class InferenceOptimizer(Protocol):
    def optimize(self, model: Model) -> OptimizedModel: ...
 
registry.register("speculative-v3", SpeculativeDecoder())
 
scheduler.use(registry.get("speculative-v3"))

The benefit: independent iterations without interference

If the scheduling team wants to add routing logic for a new region, they don't have to touch inference. If the inference team wants to ship a new speculative decoding algorithm, they don't need to worry about how the underlying infrastructure scales up or down.

In the first half of the year, our optimization team iterated through four versions of decoding strategies. The scheduling side didn't touch a single line of code, yet customers seamlessly enjoyed the faster versions. If they were tightly coupled, every one of those four iterations would have required rewriting the scheduler alongside it.

Four decoding versions shipped while the scheduler stayed unchanged

Two layers enabling each other

There is a detail worth mentioning here: making single-turn inference faster at the optimization layer allows the scheduling layer to scale down much more aggressively. Because each request is processed faster, the persistent baseline resources required for the same amount of traffic shrink, reducing idle time.

The two layers mutually enable each other. When tuning parameters, we often find that bumping up optimization by one tier allows the scheduler to drop a persistent node tier. The freed-up VRAM can then be allocated to other smaller tasks, pushing overall utilization up even further.

Conclusion: build it as an independent, pluggable layer

Inference optimization is a core competency that platforms must continuously invest in, but it should not be locked inside a specific scheduler. By building it as an independent, pluggable layer, a platform can run fast while remaining highly flexible. Our current optimization layer is plug-in based; when a new algorithm arrives, it connects via API, and the scheduling side doesn't change a single word.

How to define the APIs between the two layers, and how the scheduler provides a fallback if the optimization layer errors out, these are real, practical problems to solve, but we have never wavered on the direction.

A look at before and after optimization

Take the same model with the same traffic: after optimization, the GPUs burned in a single day were cut by more than a third compared to before. The customer didn't increase their budget, but their throughput went up. The compute they saved went toward running more experiments, which actually accelerated their model iteration. This is far more tangible than just slashing unit prices, because it relies on our own technological edge, not vendor discounts.

Before a new optimization algorithm goes live, we have a strict validation gate: we run it on shadow traffic first, comparing its latency, accuracy, and unit cost against the original algorithm. It only gets switched to the main path if all three metrics pass. If the optimization layer throws an error, the scheduler doesn't blindly wait; it identifies the issue and switches to a fallback path. We have never wavered on this architecture. Because of this decoupling, both sides of the platform can advance independently, ensuring technical progress on one side is never dragged down by the other.

Shadow-traffic gate: latency, accuracy, and unit cost must all pass

Our optimization team and scheduling team run on their own independent roadmaps. The customer doesn't need to understand this architectural split behind the scenes; they only need to know that the platform is getting cheaper and more stable. Our job is to nail both layers without them interfering with one another.

We have an internal saying: optimization saves money, scheduling saves headaches. Doing them separately, but delivering them together, is exactly what a platform should provide. The customer might not grasp how these two layers are sliced, but when they feel their bills dropping and their service remaining rock-solid, that is more than enough.

─── NEXT

Try TokenGo's serverless inference API now →