Models
Quail keeps every model in one folder and tells you, before you download, whether each one fits this Mac.
GGUF and MLX
Models come in two formats. GGUF files run on llama.cpp; MLX models run on Apple's MLX. Quail server, the default runtime, runs both. The llama.cpp runtime, kept as a fallback, runs GGUF only; choose the runtime on the Server page while the server is stopped.
On Apple silicon, MLX usually writes replies faster, and GGUF reads short prompts a little faster. Add Model suggests MLX when Quail server is the runtime.
Adding models
Models → Add Model… lists:
- Recommended for this Mac: curated models that run comfortably here, best first.
- The curated catalog: Qwen, Gemma, gpt-oss, Llama, DeepSeek-R1 distills and others, marked with what each is good for (coding, agents and tools, reasoning, vision…).
- More MLX models, from Rapid-MLX's catalog.
- Any Hugging Face repo: paste its name.
The catalog updates itself weekly. From the terminal, quail pull --list shows the catalog and quail pull <name>[:quant] downloads.
Models other apps downloaded — LM Studio, llama.cpp, the Hugging Face cache, oMLX — can be moved in rather than downloaded again: Quail offers them once, and Storage and downloads → ⋯ → Import Models from Other Apps… reopens the list.
Will it fit?
Every model gets a verdict for this Mac, worked out from its size and shape and the memory macOS lets the GPU use:
| Comfortable | Runs at the chosen context with room to spare. |
|---|---|
| Tight | Fits at a smaller context, which Quail sets automatically. Other apps may be pushed out of memory. |
| Won't fit | Too big for this Mac, even at a small context. |
The Models page starts with a summary such as “This Mac: Apple M4 Pro · 64 GB · comfortable up to ~55B”. Speeds shown before a download are estimates from your chip's memory bandwidth; after a benchmark, the Models page shows the measured speed.
A model's settings
Click the sliders button on a model's row (or its “8K context” text) to open its settings.
Context size
How much text the model can work with at once — the conversation, files and instructions together. Automatic picks the largest of 32K, 16K and 8K that runs comfortably here, never more than the model was trained for. Coding agents need 32K or more. Each size is labelled with its fit. From the terminal: quail ctx <model> 32k.
KV cache
The memory that holds the context. 8-bit makes it about half the size and 4-bit about a quarter, so a longer context fits in the same memory, a little less accurately. On MLX models it also makes reading long prompts slower; on GGUF the cost is too small to measure.
Load when the server starts
The default model (the star on its row) loads as soon as the server starts. Others load on their first request. From the terminal: quail default <model>.
Choosing a model in a request
Clients pick a model by its id, which is its name on the Models page — send it as "model". Copy it from the model's settings. With Server → Models loaded at once above 1, several stay loaded; each keeps its weights in memory, and the least recently used is unloaded when there's no room.
Storage
Models live in ~/Library/Application Support/Quail/Models unless you move them. From Models → Storage and downloads:
- Show in Finder. Models you add or delete there show up in Quail by themselves.
- ⋯ → Move Folder… moves the whole store, to an external disk for example (with the server stopped).
- ⋯ → Clean Up… removes unfinished downloads and leftovers.
- Hugging Face token: only needed for gated models. It's kept in your Keychain.