Peak memory usage is the largest host or GPU memory reading captured for one benchmark cell. Each client path uses a specific counter and measurement window, documented in the linked client methodology.
Definition
This dashboard covers two independent memory measurements:
- Host memory: peak host or process memory attributed to the workload.
- GPU memory: peak GPU memory when the runtime and platform provide a GPU probe.
Current clients submit these measurements in bytes as max_ram_bytes and, when
available, max_vram_bytes, as documented in the
client peak-memory methodology.
The management service maps those fields to max_host_bytes and
max_gpu_bytes. It then stores each measurement in a separate database row as
max_host_usage or max_gpu_usage, as documented in the
management benchmark specification.
Each database row contains one metric and one value, not both client fields.
Neither measurement is calculated from or subtracted from the other.
Measurement procedure
A single controlled workload covers model load, prompt processing, and at least one generation step. The reported peak therefore includes allocations from all three phases. One observed peak is reported per run. There is no standard deviation because allocator caches and high-water counters retain state in the same process, so later attempts would depend on earlier ones.
Unlike the timing benchmarks, peak memory uses one measured workload and does not run the thermal or CPU readiness check. Clients differ in which counters they provide and how they behave under memory pressure; see the linked client methodology.
The reviewed methodology for measurement windows, token-control behavior, what each counter includes, and platform caveats is the pipette-clients peak-memory methodology.
The dashboard maps both database metrics to the same peak-memory result. When
both qualifying metrics exist, max_gpu_usage takes priority; otherwise the
dashboard uses max_host_usage.
What it records
In addition to the common run record, the result stores:
- Configuration: input-token count (the sequence length the peak is measured at).
- Client payload:
max_ram_bytesand, where supported,max_vram_bytes. - Database metric: one row containing
max_host_usageormax_gpu_usage, with the value stored in bytes. - Dashboard value: the selected database value converted from bytes to MiB. Peak memory is a single observed high-water mark, so it has no standard deviation.
Caveats
- On unified-memory devices, the host measurement already includes GPU allocations. Do not add the host and GPU measurements.
- On devices with a discrete GPU, host and GPU measurements represent separate pools. Check the host figure against system RAM and the GPU figure against VRAM independently, never as a sum.
- An absent
max_gpu_usagedatabase row means GPU memory was unmeasured or unused, which is expected on CPU-only paths. It is not a zero-memory result. - On the Android CLI path, host memory includes sampled swap so pages moved into zram remain counted. See the Android peak-memory methodology and swap-exclusion policy.
- Counters cover different windows by client path. Compare memory only across the same platform and runtime path.