Peak Memory

Peak memory usage is the largest host or GPU memory reading captured for one benchmark cell. Each client path uses a specific counter and measurement window, documented in the linked client methodology.

Definition

This dashboard covers two independent memory measurements:

  • Host memory: peak host or process memory attributed to the workload.
  • GPU memory: peak GPU memory when the runtime and platform provide a GPU probe.

Current clients submit these measurements in bytes as max_ram_bytes and, when available, max_vram_bytes, as documented in the client peak-memory methodology. The management service maps those fields to max_host_bytes and max_gpu_bytes. It then stores each measurement in a separate database row as max_host_usage or max_gpu_usage, as documented in the management benchmark specification. Each database row contains one metric and one value, not both client fields. Neither measurement is calculated from or subtracted from the other.

Measurement procedure

A single controlled workload covers model load, prompt processing, and at least one generation step. The reported peak therefore includes allocations from all three phases. One observed peak is reported per run. There is no standard deviation because allocator caches and high-water counters retain state in the same process, so later attempts would depend on earlier ones.

Unlike the timing benchmarks, peak memory uses one measured workload and does not run the thermal or CPU readiness check. Clients differ in which counters they provide and how they behave under memory pressure; see the linked client methodology.

The reviewed methodology for measurement windows, token-control behavior, what each counter includes, and platform caveats is the pipette-clients peak-memory methodology.

The dashboard maps both database metrics to the same peak-memory result. When both qualifying metrics exist, max_gpu_usage takes priority; otherwise the dashboard uses max_host_usage.

What it records

In addition to the common run record, the result stores:

  • Configuration: input-token count (the sequence length the peak is measured at).
  • Client payload: max_ram_bytes and, where supported, max_vram_bytes.
  • Database metric: one row containing max_host_usage or max_gpu_usage, with the value stored in bytes.
  • Dashboard value: the selected database value converted from bytes to MiB. Peak memory is a single observed high-water mark, so it has no standard deviation.

Caveats

  • On unified-memory devices, the host measurement already includes GPU allocations. Do not add the host and GPU measurements.
  • On devices with a discrete GPU, host and GPU measurements represent separate pools. Check the host figure against system RAM and the GPU figure against VRAM independently, never as a sum.
  • An absent max_gpu_usage database row means GPU memory was unmeasured or unused, which is expected on CPU-only paths. It is not a zero-memory result.
  • On the Android CLI path, host memory includes sampled swap so pages moved into zram remain counted. See the Android peak-memory methodology and swap-exclusion policy.
  • Counters cover different windows by client path. Compare memory only across the same platform and runtime path.