Replication Guide
Pipette publishes versioned releases of the software used to collect, process, and score benchmark results:
Use these release pages to choose a published version for running Pipette.
Reproduce a benchmark
Match the model artifact and revision, quantization, runtime artifact, runtime flags, benchmark configuration, client behavior, device and operating system, and documented device conditions.
Run physical benchmarks under the documented lab conditions, including the same power, cooling, readiness, and background-activity controls. Even with a matched configuration and test environment, normal hardware and operating-system variation can produce small differences between runs. Compare results only when their recorded configurations and device conditions match.
Run performance benchmarks
Choose our latest benchmark client release. The links below show the reviewed reference workflow for this methodology.
Start with the Pipette CLI guide, then choose the llama.cpp or MLX runtime path.
The
model and runtime reference
explains the values accepted by --model, --runtime, and --runtime-flags.
Run evaluation scoring
Choose an evaluation dataset and scoring release. The reviewed reference workflow uses the pipette-scores quick start to install the scoring workspace. Its dataset guide documents the evaluation inputs, and the scoring guide documents each scorer.
Run the management service
You do not need a server for local benchmarks. To host the benchmark catalog, collect submissions, and process results, choose a management service release. The reviewed reference workflow is documented in the pipette-mgmt quick start and operations guide.