Pipette at a Glance

Understanding what on-device intelligence can deliver requires measuring models together with the systems that run them. Pipette evaluates complete deployment configurations: model, quantization, runtime, device, and workload. It brings speed, memory use, and accuracy into one view, using transparent and reproducible methods to support fair comparisons.

The current public results come from verified1 benchmark runs in Liquid AI's optimized lab environment. Performance runs use the documented power and cooling setup for each device path. Before every measured timing repetition, a platform-specific check waits for the device to be cool and idle. These controls reduce environmental variation within a comparison. Read more about the readiness checks, power, and cooling setup in the device-conditions methodology.

Results from these controlled runs appear in three dashboard views. Use the Leaderboard to compare configurations, Results to inspect their measurements, and Submissions to trace those measurements to submitted benchmark rows. To run or reproduce a benchmark, use the Replication Guide.

Compare Performance and Accuracy Together

Performance is measured on the device shown in a chart. Quality evaluations run separately on an H100 or another designated NVIDIA GPU system. The dashboard combines the results when they use the same model and quantization, so a quality score shown beside phone performance does not mean the evaluation ran on the phone. Result Publication Methodology explains the matching rules.

Choose a guide based on what you need:

Footnotes

  1. At this time, the verified results shown come from Liquid AI devices. If you are interested in contributing benchmark submissions in the future, let us know through the Feedback form.