Benchmark
Measure a model on your own Mac, the same way every time, so runs can be compared — across models, formats, settings and Macs.
Run one
Start the server, open Benchmark (in the menu or the Quail window), choose a model and click Run Benchmark. A small model takes about a minute. Whatever was loaded before is loaded again afterwards. From the terminal: quail bench <model>.
What it measures
The fixed quail-bench-1 suite: one warm-up, then three runs of each, reporting the median with the lowest and highest.
| Prompt 512 / 4K | How fast it reads a prompt of 512 and 4,096 tokens (tokens a second). |
|---|---|
| Generate | How fast it writes 256 tokens (tokens a second). |
| 1st token | Time to the first token of a reply to the 512-token prompt. |
| Next turn | Time to the first token when a conversation continues, using what's already cached. |
| 4 at once | Total tokens a second with four requests at the same time. |
| Load | Time to load the model from disk. |
Each result also records the Mac, the model file, the engine and its settings, and the conditions: thermal state, Low Power Mode, battery, and other models loaded.
Compare
Select two results to compare them: the older is the baseline, and each figure reads “1.23× faster” or “19% slower”. Right-click to set a different baseline. Copy as Markdown and Export JSON… share results. After a benchmark, the Models page shows the measured speed next to each model.
Fair numbers
- Plug a laptop in, and quit busy apps.
- Macs slow down as they heat up. Let the Mac cool between big runs, and alternate the order when comparing two things.
- Compare like with like: the same context size and KV cache, and nothing else loaded.