This page explains how Pipette makes performance and quality results comparable on the dashboard. It covers configuration matching, normalization, deduplication, and how the two types of results are combined.
How results are identified
Different stages identify a result with different fields:
| Layer | Fields |
|---|---|
| Deployment configuration | Model descriptor, quantization, runtime descriptor and version, device, and effective runtime flags |
| Eligible performance submission | Deployment configuration plus benchmark ID, metric, token configuration, accepted benchmark flags, and any platform-specific client or operating-system requirements |
| Displayed performance value | Model, quantization, displayed device, runtime name and version, projected runtime flags, and normalized metric key |
| Raw submission | Complete submitted record, including the exact descriptors and flags, measured value, standard deviation where reported, client ID and version, result and submission IDs, filename, and submission time |
| Displayed quality score | Model, quantization, canonical evaluation version, and thinking mode |
The eligibility fields determine whether a raw performance row can appear in the public results. For an eligible row, the normalized metric key identifies the benchmark, metric, and token configuration. The descriptor format is documented in the version-specific storage specification.
Performance and quality publication
Performance and quality originate from different benchmark pipelines and follow separate eligibility rules. Performance results are eligible when their model, device, runtime, benchmark, and metric match an approved public configuration. Quality results are eligible when both their benchmark ID and source label are approved.
The dashboard backend selects eligible results, removes duplicates, and combines each approved quality score with compatible performance results for the same model and quantization.
Collection and evaluation scoring happen before this publication flow. See Evaluation Methodology for an overview and the version-specific management storage documentation for the exact processing sequence.
Performance selection
A performance row is eligible only when all required fields match the public configuration. These fields include the model, quantization, device, runtime descriptor, runtime flags, benchmark ID, metric, benchmark flags, and token configuration, along with any platform-specific requirements.
Rows outside the public configuration remain available in raw history without entering the public charts.
Most performance configurations do not require an exact client version. A result from a compatible client build can qualify when the model, runtime, flags, device, benchmark, and metric meet the remaining requirements. The iOS configuration also requires an internal-build version pattern.
Quality result selection
The database can contain multiple scored rows for the same evaluation from different clients, runtimes, or collection periods. The dashboard publishes the latest approved accuracy result for each model and quantization. Approval is evaluated separately for each evaluation version and thinking mode.
The dashboard combines the selected quality value with compatible performance rows for every device that runs the same quantization. Changing the performance device changes speed and memory, but it does not select a different quality row. This is a dashboard publication policy, not a claim that generation is independent of runtime or hardware.
The full-precision reference is the selected row for the same model at its native full-precision weights. It can be FP16 or BF16, depending on the model's published weights.