Blog · Inference and Infrastructure · October 2, 2026

The Energy Cost of an Answer Includes the Queue

The electricity used to answer a prompt is not fixed. The answer's length matters, but so do the requests beside it, the capacity waiting for traffic, and the deadline the service promises.

Waiting for company

A request sits in a queue while a server gathers other requests to process alongside it. In this hypothetical service, the delay is deliberate: sharing execution may reduce the electricity assigned to each answer. The user sees a slower response. The operator may see better efficiency. Neither observation tells us whether the arrangement is worthwhile.

Now suppose the same queue forms because the service is overwhelmed. Waiting could grow without producing a useful energy saving. A spinning indicator cannot distinguish these situations. Our argument is that energy accounting needs the service's operating conditions, including its queues and response promises, alongside the model's identity.

This is a practical question for anyone comparing an immediate chat response with a deferred document-processing job. They may use the same model, yet purchase different arrangements of capacity and time. Asking which consumes less electricity requires specifying what each arrangement must deliver.

Where the meter stops

The MLPerf power rules provide a concrete starting point: measure the defined system at its electrical supply, including the host, accelerators, memory and fans activated by the benchmark. The document, marked January 2022 and consulted for this essay, aligns energy accounting with the performance run. That differs from reading only an accelerator's telemetry.

The MLPerf Power methodology paper, revised in February 2025, also distinguishes machine efficiency from building efficiency. It warns that shared cooling and coarse facility meters complicate attribution. A server measurement does not automatically cover the facility overhead needed to operate that server.

For a comparison, write down where the meter stops before discussing how small the number is. Device energy can help diagnose a kernel improvement. Server energy can compare complete machines. An estimate including facility overhead answers a broader operating question. These quantities can all be useful if the boundary stays attached.

Then specify the interval. A short execution measurement and an allocation of the service's electricity over a working day have different numerators. Adding them without checking overlap can count the same electricity twice. Conversely, presenting the execution interval as the complete cost of keeping a service available leaves the quiet hours unexplained.

What batching buys

Delavande, Pierrard and Luccioni's January 2026 preprint tested inference on H100 GPUs. Its experiments show batching can spread execution overhead across requests, while padding shorter inputs to match longer ones can waste work. The apparent optimum also changes with the denominator: energy per input token and per output token describe different quantities. The study's serving experiments vary arrival patterns as well as software configuration. Its methods describe CPU measurements and RAM estimates, while its limitations describe a GPU-only focus; we therefore avoid its headline savings estimate.

Our interpretation is that a batch is an agreement about sharing equipment. It offers a way to divide some costs across useful work, but it also asks requests to coexist. A hypothetical overnight summarization service can collect a backlog before processing it. A conversational assistant must keep answering while new requests arrive. An efficient arrangement for the former may be unacceptable for the latter.

That does not make slower delivery intrinsically greener. Waiting helps only if the resulting execution or capacity arrangement saves enough electricity. A delay that changes nothing about the work simply postpones completion. Likewise, allowing a larger batch should be evaluated against the actual mixture of short and long requests, rather than assuming every additional request improves efficiency.

A more useful product choice would be an explicit deadline: immediate, or ready by an agreed time. That is our proposal for exposing flexibility. The operator would still need to demonstrate a saving; permission to defer work is an opportunity, not evidence that electricity was saved.

Which clock counts?

In Poddar and colleagues' NAACL 2025 experiments, response time and energy were strongly correlated within the tested setup, and output length was an important energy driver. Most experiments ran on an A6000 GPU. Their software accounting combined component measurements and estimates, with facility overhead omitted. These results support watching generation length and runtime under controlled conditions; they do not turn a remote service's visible waiting time into an electricity meter.

The MLPerf inference rules consulted October 2, 2026 distinguish an Offline scenario, with inputs available together, from Server testing, with arriving requests and latency constraints. For language models, the rules also distinguish time to the first token from time per output token. Throughput alone cannot describe how promptly a user receives an answer.

The practical consequence is to retain several clocks. Record arrival, the start of execution, the first visible token, and completion. A service might begin streaming quickly yet take too long to finish. Another might process many answers per second while an unlucky request waits behind other work. The average experience can hide that delay.

We would therefore compare energy at an agreed quality and delivery target. Otherwise, a system can appear efficient by accepting waiting times the buyer never intended to tolerate. The site’s guardrail latency essay examines a different component of response time; the measurement here concerns the service surrounding the model.

Electricity between answers

Consider another hypothetical comparison. Two dedicated services execute identical requests with identical active energy. One receives steady traffic; the other keeps equipment ready through long quiet periods. If both allocate their entire operating electricity to completed requests, the quieter service assigns more overhead to each answer.

This is an accounting consequence, not evidence that its model performs more computation. It also separates average allocation from the extra electricity caused by adding a request. Dividing the day's electricity by the day's requests answers the first question. Estimating the second requires a comparison with what the service would have consumed without that request, including whether capacity would change.

Neither measure should impersonate the other. An average can help explain the cost of maintaining availability. A marginal estimate can help assess an additional workload. For shared equipment, allocating all concurrent power separately to every request would overcount; an allocation rule must reconcile with the measured total.

Our recommendation is to show idle allocation explicitly. State whether it is included, over what period, and how shared capacity is divided. A busy benchmark can establish useful execution efficiency while leaving open what the service consumes overnight. A quiet period can reveal that operating cost without disproving the benchmark.

A useful comparison

For an actual deployment decision, we propose testing the same representative workload under both quiet and busy traffic, with matching output requirements. Record input and output lengths separately, the serving software, batching policy, measurement boundary, total electricity, and the distribution of completion times. Preserve failed and cancelled attempts in the energy total.

Then add a denominator that reflects the job: accepted results. A hypothetical summarization system that produces brief but incomplete summaries may look economical per request while forcing repeated attempts. Conversely, a longer first answer may finish the assignment. Define acceptance before comparing systems, and include the measured electricity of retries or automated checking within the stated workflow boundary.

This proposed measure complements energy per token. Tokens help describe machine activity; accepted results help assess the service purchased. Neither should erase absolute electricity consumption, because total demand can change even when efficiency improves.

A useful claim would identify a workload, a quality requirement, a delivery target and an energy boundary, then report the result under stated traffic conditions. That lets a buyer decide whether waiting is an acceptable exchange and lets an operator receive credit for demonstrated savings.

Sources

Production: commissioned by the site operator; researched and drafted by GPT-6 Astra; editorial and source review by the coordinating AI assistant.


Return to Blog