Benchmark

Measure a model on your own Mac, the same way every time, so runs can be compared — across models, formats, settings and Macs.

Run one

Start the server, open Benchmark (in the menu or the Quail window), choose a model and click Run Benchmark. A small model takes about a minute. Whatever was loaded before is loaded again afterwards. From the terminal: quail bench <model>.

What it measures

The fixed quail-bench-1 suite: one warm-up, then three runs of each, reporting the median with the lowest and highest.

Prompt 512 / 4KHow fast it reads a prompt of 512 and 4,096 tokens (tokens a second).
GenerateHow fast it writes 256 tokens (tokens a second).
1st tokenTime to the first token of a reply to the 512-token prompt.
Next turnTime to the first token when a conversation continues, using what's already cached.
4 at onceTotal tokens a second with four requests at the same time.
LoadTime to load the model from disk.

Each result also records the Mac, the model file, the engine and its settings, and the conditions: thermal state, Low Power Mode, battery, and other models loaded.

Compare

Select two results to compare them: the older is the baseline, and each figure reads “1.23× faster” or “19% slower”. Right-click to set a different baseline. Copy as Markdown and Export JSON… share results. After a benchmark, the Models page shows the measured speed next to each model.

Fair numbers