Models

Quail keeps every model in one folder and tells you, before you download, whether each one fits this Mac.

GGUF and MLX

Models come in two formats. GGUF files run on llama.cpp; MLX models run on Apple's MLX. Quail server, the default runtime, runs both. The llama.cpp runtime, kept as a fallback, runs GGUF only; choose the runtime on the Server page while the server is stopped.

On Apple silicon, MLX usually writes replies faster, and GGUF reads short prompts a little faster. Add Model suggests MLX when Quail server is the runtime.

Adding models

Models → Add Model… lists:

The catalog updates itself weekly. From the terminal, quail pull --list shows the catalog and quail pull <name>[:quant] downloads.

Models other apps downloaded — LM Studio, llama.cpp, the Hugging Face cache, oMLX — can be moved in rather than downloaded again: Quail offers them once, and Storage and downloads → ⋯ → Import Models from Other Apps… reopens the list.

Will it fit?

Every model gets a verdict for this Mac, worked out from its size and shape and the memory macOS lets the GPU use:

ComfortableRuns at the chosen context with room to spare.
TightFits at a smaller context, which Quail sets automatically. Other apps may be pushed out of memory.
Won't fitToo big for this Mac, even at a small context.

The Models page starts with a summary such as “This Mac: Apple M4 Pro · 64 GB · comfortable up to ~55B”. Speeds shown before a download are estimates from your chip's memory bandwidth; after a benchmark, the Models page shows the measured speed.

A model's settings

Click the sliders button on a model's row (or its “8K context” text) to open its settings.

Context size

How much text the model can work with at once — the conversation, files and instructions together. Automatic picks the largest of 32K, 16K and 8K that runs comfortably here, never more than the model was trained for. Coding agents need 32K or more. Each size is labelled with its fit. From the terminal: quail ctx <model> 32k.

KV cache

The memory that holds the context. 8-bit makes it about half the size and 4-bit about a quarter, so a longer context fits in the same memory, a little less accurately. On MLX models it also makes reading long prompts slower; on GGUF the cost is too small to measure.

Load when the server starts

The default model (the star on its row) loads as soon as the server starts. Others load on their first request. From the terminal: quail default <model>.

Context, KV cache and the default model apply the next time the server starts. So do models added or deleted while it runs: the Server page and the menu say when a restart is needed.

Choosing a model in a request

Clients pick a model by its id, which is its name on the Models page — send it as "model". Copy it from the model's settings. With Server → Models loaded at once above 1, several stay loaded; each keeps its weights in memory, and the least recently used is unloaded when there's no room.

Storage

Models live in ~/Library/Application Support/Quail/Models unless you move them. From Models → Storage and downloads: