Agent workloads are inherently bursty; provisioning resources based on averages is a guaranteed recipe for failure.
We have spoken with quite a few teams building agents, and the consensus is becoming increasingly clear. The workload shape of an executing agent is a completely different beast compared to a single model inference call.
Single-turn inference is flat; agents are bursty
A single inference call is straightforward: a request comes in, runs, and exits. The traffic can be drawn as a smooth line, and you can provision resources based on the average. Agents are not like that. They plan first, then loop, and might execute a bunch of parallel tool calls mid-way. One specific branch might suddenly require a massive chunk of compute, and after it is done, it goes completely quiet again. Its workload consists of bursts.
We once plotted the daily compute curve for a team building a coding assistant. Usage floated around 0 most of the time, then spiked 20x during intense periods, within a 10 minute window.
Provisioning by average pleases no one
You cannot provision compute for agents based on averages. If you do, you waste money idling machines during downtime, and when a burst hits, you don't have enough. The queue piles up, and the user gets stuck waiting. We've seen too many teams with heartbreakingly low GPU utilization on average, only to completely drop the ball during peak business hours. Worse yet, some teams try to prevent bursts by keeping machines constantly provisioned at peak capacity. The result is that those machines are sleeping 90% of the time, but the bills keep rolling in anyway.
Different patterns, different compute requirements
Decoupled planning and execution. The orchestration layer is lightweight and persistent, while the truly compute-heavy tool calls are spun up elastically.
Step-by-step scheduling. Long tasks are broken down into many small steps, and each step is scheduled independently. Whichever step is busy gets more resources.
Reflection loops. The model decides for itself whether to retry or take a different path. This is the most compute-intensive but also the most stable pattern, suitable for tasks with zero tolerance for error. We saw a team that cranked their reflection loop too high, retrying every step three times, which instantly doubled their compute usage. They later changed it to trigger based on confidence scores; costs dropped back down, but accuracy didn't suffer.
Three requirements for underlying infrastructure
Fast scale-up. When a burst hits, machines must be added in seconds. You can't have the spike pass before the scale-up even finishes.
Clean scale-down. As soon as a task stops, resources must be released immediately. Many billing traps are hidden in forgotten, un-closed resources.
Dependency awareness. The scheduler must understand the dependencies between tasks. It needs to know what can run in parallel and what must wait in a queue. Otherwise, elasticity turns into chaos, and adding more machines is pointless.
The worst-case scenario: absorbing bursts with steady-state GPUs
Using dedicated GPUs designed for steady-state workloads to handle agent bursts guarantees that the cards will either idle to death or bottleneck completely. We met a team building a customer service agent who rented a fixed number of cards at the beginning of the month based on historical averages. On the day of a major promotion, query volume surged 5x. The cards couldn't keep up, and queues grew so long that users simply abandoned them.
Our direction: workload-shape-aware scheduling
We built the scheduler itself to be aware of the workload shape. Instead of a fixed-size pool, the resources follow the tasks. Consume nothing when idle, be available instantly when needed, and disappear immediately when done. Implementing this relies on resource profiling for each step of the task. By knowing roughly how much VRAM a step will take and how long it will run, we can provision the GPU half a step ahead of time. For example, if the orchestration layer predicts the next step involves calling an image generation tool, it pre-warms a GPU with the right amount of VRAM in a nearby region ahead of time.
What our elastic scaling policy looks like
In terms of configuration, we set two things for every agent task: trigger conditions for scaling up/down, and resource profiles for each step.
agent_autoscale:
scale_up: p95_latency > 800ms # Scale up in seconds if latency exceeds threshold
scale_down: idle > 60s # Release after one minute of idle time
min_replicas: 0 # Scale to zero during downtime
per_step_profile: true # Predict VRAM and duration per stepWith a resource profile for every step, the orchestration layer can prep the GPU half a step early. When the call actually hits, the machine is already waiting for it. All the user feels is that the agent is consistently fast. They will never know there's a scheduling system behind the scenes, frantically adding and dropping machines in lockstep with their task. To paraphrase Jordan Schlansky, if we perform our duties to the optimal level, you won't notice us.
This kind of elasticity is exactly what agents need. The agent itself doesn't know how much compute it will need next so the platform must absorb that uncertainty for it. We calculated the numbers for an onboarded client: handling the exact same peak traffic before and after the revamp, their machine costs dropped by 40%. Meanwhile, the user-perceived wait times stablilized because they no longer had to fight others for a fixed quota of GPUs.