Pipette at a Glance
Understanding what on-device intelligence can deliver requires measuring models together with the systems that run them. Pipette evaluates complete deployment configurations: model, quantization, runtime, device, and workload. It brings speed, memory use, and accuracy into one view, using transparent and reproducible methods to support fair comparisons.
The current public results come from verified1 benchmark runs in Liquid AI's optimized lab environment. Performance runs use the documented power and cooling setup for each device path. Before every measured timing repetition, a platform-specific check waits for the device to be cool and idle. These controls reduce environmental variation within a comparison. Read more about the readiness checks, power, and cooling setup in the device-conditions methodology.
Results from these controlled runs appear in three dashboard views. Use the Leaderboard to compare configurations, Results to inspect their measurements, and Submissions to trace those measurements to submitted benchmark rows. To run or reproduce a benchmark, use the Replication Guide.
Compare Performance and Accuracy Together
Performance is measured on the device shown in a chart. Quality evaluations run separately on an H100 or another designated NVIDIA GPU system. The dashboard combines the results when they use the same model and quantization, so a quality score shown beside phone performance does not mean the evaluation ran on the phone. Result Publication Methodology explains the matching rules.
What to read next
Choose a guide based on what you need:
- Read and compare results: Compare Pipette Results explains which configurations can be compared and how to interpret charts.
- Quantization covers the weight formats shown in Pipette and their quality, speed, and memory tradeoffs.
- Review how measurements are produced: Performance Methodology explains the timing and memory protocols. Device Conditions covers readiness, power, and cooling. Evaluation Methodology explains datasets, generation, and scoring.
- Understand what appears on the dashboard: Coverage and Selection defines which models, quantizations, runtimes, devices, and benchmarks are included. Result Publication Methodology covers configuration matching, normalization, duplicate handling, and how performance and quality results are combined.
- Check a benchmark definition: End-to-End Latency, Prefill Throughput, Decode Throughput, and Peak Memory define the performance measurements. IFBench, GPQA Diamond, and MATH-500 document the evaluation datasets and scoring rules.
- Understand current constraints: Limitations and Future Directions explains how to interpret the current results and where coverage is still limited.
- Run Pipette: The Replication Guide points to available client releases and covers device-specific benchmarking, evaluation scoring, and management-service setup.
- Analyze the data offline: Data Exports explains how to request a snapshot of the public benchmark data.
Footnotes
-
At this time, the verified results shown come from Liquid AI devices. If you are interested in contributing benchmark submissions in the future, let us know through the Feedback form. ↩