Skip to main content

A ternary model won’t load, or the quantization type is rejected

Three things cause this, depending on which file you have:
  • Legacy *-Q2_0.gguf (no g64 suffix) — pre-migration files stored under a ggml type id that now belongs to the official group-64 format. Current binaries refuse them with an error. Download the *-Q2_0_g64.gguf file instead.
  • *-PQ2_0.gguf — our group-128 packing, which needs the PrismML fork binaries (prism-b10658 or newer). Stock builds refuse it. Use the group-64 file on a stock build.
  • Ternary Bonsai 2 (PTQ1_0 / PQ2_0) — needs the fork for every backend, because the activation-side Hadamard transform isn’t upstream yet.
The plain ternary group-64 files run on mainline llama.cpp (CPU, Metal, Vulkan, CUDA), and ternary MLX weights run on stock MLX. See Formats & runtime support.
Never point stock llama.cpp at Ternary Bonsai 2. PQ2_0 and PTQ1_0 are refused outright, but a Q2_0 file loads without a warning and outputs gibberish.

Generation is much slower than the published numbers

The published speeds (131 tok/s on M4 Pro for 1-bit 8B, etc.) assume native low-bit kernels and full GPU offload. Check, in order:
  1. Is the runtime dequantizing? A runtime that recognizes the file but lacks the right kernels may fall back to higher precision: same output, none of the speed. If memory use during inference is far above the expected totals (e.g. multiple GB more than weights + KV cache), that’s the signature. Use the known-good binaries.
  2. Are layers on the GPU? With raw llama-cli/llama-server, pass -ngl 99. The demo scripts set this automatically. The load log prints how many layers were offloaded.
  3. Is thinking eating the time? On the 27B, reasoning is usually the bulk of the wait. Pick a lower Reasoning effort in the chat UI, or cap it with --reasoning-budget N.
  4. Right binary for your hardware? On x86 CPUs, the optimized build is significantly faster than the generic one; on macOS, make sure you’re using the Metal (Apple Silicon) build. See the binary variants.
  5. Long context? Decode slows as the KV cache grows; benchmark numbers are short-context. Compare against community benchmarks for your hardware.

Out of memory, or the machine freezes at startup

Total memory = weight file + KV cache, and the KV cache grows linearly with context (64 KiB/token on the 27B, 144 KiB/token on the 8B; see each model card). Older revisions passed llama.cpp’s -c 0, which uses the model’s full training context (262K on the 27B) regardless of available memory. The scripts now pick a RAM-tiered context instead, and BONSAI_CTX=0 maps to that safe default rather than -c 0. If you still hit memory pressure:
  • Pin a smaller context: BONSAI_CTX=8192 ./scripts/start_llama_server.sh (or -c 8192 with raw llama-cli).
  • Turn on the 4-bit KV cache: BONSAI_KV4=1 cuts cache memory roughly 3.5x, or use llama.cpp’s own --cache-type-k q8_0 --cache-type-v q8_0 to roughly halve it.
  • Free VRAM on tight cards with BONSAI_MMPROJ_CPU=1, which keeps the vision projector in system RAM (~0.9 GiB back, slower image prompts only).
  • Or step down: ternary → 1-bit, or one size smaller.

”GGUF model not found” from the scripts

The script’s BONSAI_MODEL/BONSAI_FAMILY don’t match what’s downloaded. Download the variant you asked for:
Also note BONSAI_MODEL=all / BONSAI_FAMILY=all are valid only for setup and download scripts, not for run/server scripts, which need exactly one model. If you installed with BONSAI_SKIP_GGUF=1 (MLX only), the llama.cpp scripts stop with an error pointing at both options: run run_mlx.sh / start_mlx_server.sh instead, or download the GGUF weights.

Port already in use

  • The Bonsai server refuses to start if something is on 8080: stop it with kill $(lsof -ti TCP:8080) (macOS/Linux), or start on another port by passing --port through the script.
  • A standalone open-webui serve also defaults to 8080. Use ./scripts/start_openwebui.sh (which picks 9090+) or pass --port. See Open WebUI.

Windows: “running scripts is disabled on this system”

PowerShell’s execution policy blocks setup.ps1. Allow it for the current session only:

Apple M5: Metal compile errors, then out-of-memory

On M5 (and A19) devices with certain macOS 26 point releases, ggml compiles its Metal library at runtime to enable the tensor API, and stricter MetalPerformancePrimitives headers break that compile:
This is an ecosystem-wide issue affecting every ggml-based project; M1–M4 are unaffected. Disable the tensor API — full Metal speed is kept, only the Neural Accelerator prefill boost is lost:
This is much faster than falling back to CPU (BONSAI_NGL=0). If out-of-memory errors persist on lower-memory machines, also pin a smaller context, e.g. -c 16384.

CUDA source build hangs or gets killed

Compiling CUDA kernels is memory-intensive — each parallel job can take several GB of VRAM and system RAM, so -j$(nproc) can exhaust it. build_cuda_linux.sh and build_cuda_windows.ps1 detect GPU VRAM and cap parallelism at -j 2 below 16 GB. Building by hand, reduce -j yourself or close other GPU-heavy applications.

MLX errors on 1-bit weights

1-bit support in stock MLX is pending (mlx#3161). Until it merges, install the fork (branch prism):
Ternary (2-bit) MLX weights, including Ternary Bonsai 2, work with stock mlx-lm.

The model emits <think> blocks or stalls before answering

Thinking-mode defaults differ by size. The 27B reasons by default; the 8B/4B/1.7B demo scripts disable it (--reasoning-budget 0 --reasoning-format none --chat-template-kwargs '{"enable_thinking": false}').
  • To turn thinking off for 27B: start the server with BONSAI_THINKING=0, or cap it with --reasoning-budget N.
  • If you’re using the built-in chat UI, check the Reasoning effort picker (lightbulb icon) — it’s saved per browser and overrides the server default for every future chat, including new ones. If replies are slower than expected even after setting BONSAI_THINKING=0, this is almost always why: set the picker itself to your desired effort, or clear site data for localhost:8080.
  • If you launched llama-server by hand on 8B/4B/1.7B without the disable flags above, add them, or strip reasoning client-side.

Still stuck?

Ask in the PrismML Discord or open an issue on the demo repo with your platform, binary variant, and the full load log.