Decode throughput measures how quickly a model generates output tokens after it has processed the input.
Definition
After prefill has built the context, generation proceeds one token at a time.
Decode throughput measures that phase at a fixed context length and for a fixed
number of generated tokens. The submitted measurement is decode_time_ms, the
time required to generate those tokens once the context is in place. Decode
throughput is the generated-token count divided by that time, reported in tokens
per second.
The benchmark configuration carries both an input-token count and an output-token count: the input count sets the context depth (which changes attention work and memory pressure), and the output count sets the amount of measured generation. Results are comparable only at the same depth and output count, with the same model, quantization, runtime, and flags.
Measurement procedure
Decode follows the standard timing protocol: a warm-up, five readiness-gated repetitions, and a mean reported with its sample standard deviation. See Performance Methodology for the shared details. The benchmark-specific steps are:
- The context is established to the requested depth first. Only the generation
of the requested decode tokens is timed, so prefill setup is excluded from
decode_time_ms. - Generation is forced to the full requested token count. Early stopping is suppressed, so a short completion cannot inflate the apparent rate.
What it records
In addition to the common run record, the result stores:
- Configuration: input-token count (the context depth) and output-token count.
- Measurement: decode time (
decode_time_ms), with the sample standard deviation across the five repetitions.
Pipette converts this measurement to decode throughput in tokens per second.
Caveats
- Decode throughput depends on the existing context length, not just the number of generated tokens. Always compare at the same prefill depth.
- Use end-to-end latency, not decode throughput, when the question is the total caller-visible time for a full prompt-plus-generation request.