A ternary model won’t load, or the quantization type is rejected
Three things cause this, depending on which file you have:
- Legacy
*-Q2_0.gguf (no g64 suffix) — pre-migration files stored under a ggml type id that now belongs to the official group-64 format. Current binaries refuse them with an error. Download the *-Q2_0_g64.gguf file instead.
*-PQ2_0.gguf — our group-128 packing, which needs the PrismML fork binaries (prism-b10658 or newer). Stock builds refuse it. Use the group-64 file on a stock build.
- Ternary Bonsai 2 (
PTQ1_0 / PQ2_0) — needs the fork for every backend, because the activation-side Hadamard transform isn’t upstream yet.
The plain ternary group-64 files run on mainline llama.cpp (CPU, Metal, Vulkan, CUDA), and ternary MLX weights run on stock MLX. See Formats & runtime support.
Never point stock llama.cpp at Ternary Bonsai 2. PQ2_0 and PTQ1_0 are refused outright, but a Q2_0 file loads without a warning and outputs gibberish.
Generation is much slower than the published numbers
The published speeds (131 tok/s on M4 Pro for 1-bit 8B, etc.) assume native low-bit kernels and full GPU offload. Check, in order:
- Is the runtime dequantizing? A runtime that recognizes the file but lacks the right kernels may fall back to higher precision: same output, none of the speed. If memory use during inference is far above the expected totals (e.g. multiple GB more than weights + KV cache), that’s the signature. Use the known-good binaries.
- Are layers on the GPU? With raw
llama-cli/llama-server, pass -ngl 99. The demo scripts set this automatically. The load log prints how many layers were offloaded.
- Is thinking eating the time? On the 27B, reasoning is usually the bulk of the wait. Pick a lower Reasoning effort in the chat UI, or cap it with
--reasoning-budget N.
- Right binary for your hardware? On x86 CPUs, the optimized build is significantly faster than the generic one; on macOS, make sure you’re using the Metal (Apple Silicon) build. See the binary variants.
- Long context? Decode slows as the KV cache grows; benchmark numbers are short-context. Compare against community benchmarks for your hardware.
Out of memory, or the machine freezes at startup
Total memory = weight file + KV cache, and the KV cache grows linearly with context (64 KiB/token on the 27B, 144 KiB/token on the 8B; see each model card).
Older revisions passed llama.cpp’s -c 0, which uses the model’s full training context (262K on the 27B) regardless of available memory. The scripts now pick a RAM-tiered context instead, and BONSAI_CTX=0 maps to that safe default rather than -c 0. If you still hit memory pressure:
- Pin a smaller context:
BONSAI_CTX=8192 ./scripts/start_llama_server.sh (or -c 8192 with raw llama-cli).
- Turn on the 4-bit KV cache:
BONSAI_KV4=1 cuts cache memory roughly 3.5x, or use llama.cpp’s own --cache-type-k q8_0 --cache-type-v q8_0 to roughly halve it.
- Free VRAM on tight cards with
BONSAI_MMPROJ_CPU=1, which keeps the vision projector in system RAM (~0.9 GiB back, slower image prompts only).
- Or step down: ternary → 1-bit, or one size smaller.
”GGUF model not found” from the scripts
The script’s BONSAI_MODEL/BONSAI_FAMILY don’t match what’s downloaded. Download the variant you asked for:
Also note BONSAI_MODEL=all / BONSAI_FAMILY=all are valid only for setup and download scripts, not for run/server scripts, which need exactly one model. If you installed with BONSAI_SKIP_GGUF=1 (MLX only), the llama.cpp scripts stop with an error pointing at both options: run run_mlx.sh / start_mlx_server.sh instead, or download the GGUF weights.
Port already in use
- The Bonsai server refuses to start if something is on 8080: stop it with
kill $(lsof -ti TCP:8080) (macOS/Linux), or start on another port by passing --port through the script.
- A standalone
open-webui serve also defaults to 8080. Use ./scripts/start_openwebui.sh (which picks 9090+) or pass --port. See Open WebUI.
Windows: “running scripts is disabled on this system”
PowerShell’s execution policy blocks setup.ps1. Allow it for the current session only:
On M5 (and A19) devices with certain macOS 26 point releases, ggml compiles its Metal library at runtime to enable the tensor API, and stricter MetalPerformancePrimitives headers break that compile:
This is an ecosystem-wide issue affecting every ggml-based project; M1–M4 are unaffected. Disable the tensor API — full Metal speed is kept, only the Neural Accelerator prefill boost is lost:
This is much faster than falling back to CPU (BONSAI_NGL=0). If out-of-memory errors persist on lower-memory machines, also pin a smaller context, e.g. -c 16384.
CUDA source build hangs or gets killed
Compiling CUDA kernels is memory-intensive — each parallel job can take several GB of VRAM and system RAM, so -j$(nproc) can exhaust it. build_cuda_linux.sh and build_cuda_windows.ps1 detect GPU VRAM and cap parallelism at -j 2 below 16 GB. Building by hand, reduce -j yourself or close other GPU-heavy applications.
MLX errors on 1-bit weights
1-bit support in stock MLX is pending (mlx#3161). Until it merges, install the fork (branch prism):
Ternary (2-bit) MLX weights, including Ternary Bonsai 2, work with stock mlx-lm.
The model emits <think> blocks or stalls before answering
Thinking-mode defaults differ by size. The 27B reasons by default; the 8B/4B/1.7B demo scripts disable it (--reasoning-budget 0 --reasoning-format none --chat-template-kwargs '{"enable_thinking": false}').
- To turn thinking off for 27B: start the server with
BONSAI_THINKING=0, or cap it with --reasoning-budget N.
- If you’re using the built-in chat UI, check the Reasoning effort picker (lightbulb icon) — it’s saved per browser and overrides the server default for every future chat, including new ones. If replies are slower than expected even after setting
BONSAI_THINKING=0, this is almost always why: set the picker itself to your desired effort, or clear site data for localhost:8080.
- If you launched
llama-server by hand on 8B/4B/1.7B without the disable flags above, add them, or strip reasoning client-side.
Still stuck?
Ask in the PrismML Discord or open an issue on the demo repo with your platform, binary variant, and the full load log.