Replication Guide

The performance data shown in Pipette was generated with the published Pipette v1 client snapshot. The immutable source snapshots for this release are:

Use these repositories to run the benchmark suite or inspect how a result is produced.

Note: Compare results only after matching the recorded configuration and device conditions.

Matching the workload does not guarantee the same number. Hardware revision, operating system, temperature, power state, and background activity can affect the measurement. See Limitations and Future Directions for the current replication boundaries.

The sections below are independent workflows. Choose the one that matches what you want to run.

Run performance benchmarks

Start with the Pipette CLI guide, then choose the llama.cpp or MLX runtime path.

The model and runtime reference explains the values accepted by --model, --runtime, and --runtime-flags.

Run evaluation scoring

Use the pipette-scores quick start to install the scoring workspace. Its dataset guide documents the evaluation inputs, and the scoring guide documents each scorer.

Run the management service

You do not need a server for local benchmarks. To host the benchmark catalog, collect submissions, and process results, follow the pipette-mgmt quick start and operations guide.