This page explains how to compare Pipette results and interpret the metrics and configurations shown on the dashboard.
Choose the metric for the question
| Question | Metric | Direction |
|---|---|---|
| How quickly is the prompt processed? | Prefill throughput | ↑ Higher is better |
| How quickly are output tokens generated? | Decode throughput | ↑ Higher is better |
| How long does the complete request take? | End-to-end latency | ↓ Lower is better |
| How much memory does inference require? | Peak memory | ↓ Lower is better |
| How well does the model perform a specific task? | IFBench, GPQA Diamond, or MATH-500 | ↑ Higher is better |
The best deployment is usually a trade-off across quality, speed, and memory on the target hardware. A model that leads one metric may be unsuitable once the other constraints are considered.
Match the full configuration
A performance value is comparable only when these four contexts match:
Runtime and model flags are part of those contexts. A different flash-attention setting, thread count, GPU-layer count, thinking mode, quantization, or token configuration represents a different configuration.
See Result Publication Methodology for the complete list of fields that must match and the rules for selecting results.
Understand input and output tokens
Pipette benchmarks describe workloads using two token counts:
- Input tokens are the prompt tokens processed during prefill.
- Output tokens are generated during decode.
During generation, the active context contains the prompt and the output generated so far. Longer contexts increase the attention workload and may require more KV-cache memory, depending on how the runtime allocates its context.
For a matching benchmark configuration, divide the token count by the reported throughput to estimate the isolated phase time. At 5,000 tokens per second, processing 1,024 input tokens takes about 0.2 seconds. At 80 tokens per second, generating 100 output tokens takes about 1.25 seconds.
These are isolated phase estimates, not a prediction of end-to-end latency. End-to-end measurements combine prompt processing and generation and may also include tokenization and local request overhead, depending on the client path.
Interpret standard deviation and peak memory
Timing results report a mean and sample standard deviation across five measured repetitions within one benchmark run. The standard deviation describes repeat-to-repeat variation on that device; it is not calculated across devices or submissions, even when they share a configuration. Treat a standard deviation greater than 5% of the mean as evidence that the run was noisy or unstable, rather than relying on the rounded mean alone.
Peak memory is the highest memory use recorded during inference, not the device's total memory capacity. Platforms measure that peak with different counters, so use it to compare models and quantizations within the same device and runtime configuration rather than as an identical cross-platform measure. See Peak Memory for the exact counter used by each platform.
Combine performance and accuracy
Accuracy and device performance answer different questions and are measured in separate runs. Accuracy runs use the same dataset version, sampling policy, and task-specific scorer for each evaluation. For some models, answers may be generated on a different runtime or device from the one shown for performance. Performance runs measure latency, throughput, and memory on the selected device.
For example, the dashboard can match the GPQA score for Model A at Q4 with performance results for Model A at Q4 on a Mac or phone. Selecting a different device changes the performance results, while the shared GPQA score provides a consistent quality reference. It does not mean that GPQA ran on the selected device.
See Evaluation Methodology for how accuracy is measured and Result Publication Methodology for how the dashboard combines the values.