This page explains how Pipette defines the models, weights, runtimes, devices, and benchmark configurations included in the public dashboard. Coverage is not comprehensive across every available configuration. Result Publication Methodology explains how eligible raw rows become displayed values.
Selection boundaries
A public configuration must identify:
- A model and quantization included in the public chart scope
- A declared public weight source
- A runtime descriptor, version, and effective flags
- A supported device path
- A benchmark ID, metric, and exact token configuration
Changing any of those fields creates a different deployment or benchmark configuration. The Datasheet can contain additional submissions that are not displayed in the device charts.
Models and weights
The same selection criteria apply to Liquid AI models and models from other providers. Availability, public artifacts, runtime support, and useful device coverage determine which configurations the charts can include. Provider identity does not change those requirements.
When available, Pipette uses files published by the model author in the format
required by the runtime. Otherwise it uses a declared public community
conversion: unsloth GGUF builds when no official GGUF is available, and
mlx-community builds for MLX. Server runtimes load upstream model weights
directly.
The source priority is documented in the version-specific pipette-clients selection policy.
Runtimes
Public performance coverage defines llama.cpp configurations for macOS, Android, iOS, and Windows.
Current public Android performance measurements are collected by the pipette
CLI binary running on the phone. The Pipette plan runner dispatches that binary
through ADB, as shown in the
Android plan example.
These rows use the llamacpp_cli_stock_tools runtime descriptor with the
android-arm64-v8a flavor; they are not measurements submitted by the native
Android app.
Benchmark and device coverage
The standard Pipette client catalog includes prompt sizes of 100, 256, 512, 1024, 2048, 4096, and 8192 tokens. Public chart coverage is narrower and device-specific:
| Device path | Public prompt configurations |
|---|---|
| macOS | 256, 512, 1024, 2048, 4096, and 8192 tokens |
| Android and iOS | 256, 512, 1024, 2048, and 4096 tokens |
| Windows | 256, 512, 1024, 2048, 4096, and 8192 tokens |
Decode throughput uses 100 generated tokens. End-to-end latency uses 256 generated tokens. The public dashboard does not currently show the 100-token prompt configuration from the client catalog.
The client catalog is defined in the version-specific standard benchmark source.
Runtime flags
Runtime flags are part of the configuration identity. The public charts use one flash-attention policy per platform so each displayed result uses one setting:
| Platform | Current public setting |
|---|---|
| macOS and Windows | Flash attention on |
| Android CPU | Flash attention off |
| iOS | In-app runtime default |
Where the runtime exposes the switch, benchmark collection measures flash attention both on and off. The current macOS and Windows GPU paths require it to be on because that setting is faster for both prefill and decode.
The current Android path is CPU-only, with zero layers offloaded to a GPU. Flash attention is disabled because it is slower on this Android CPU path. Requiring it to be off keeps every public Android cell on one stable configuration.
Thread count, GPU-layer count, memory mapping, and context sizing are also fixed when the platform allows them to be configured. These values are part of the configuration identity and must match when comparing or reproducing results.