Skip to main content
llama.cpp is the reference runtime for Bonsai GGUF weights. It runs on CPU and every major GPU backend (CUDA, Metal, Vulkan, ROCm).
1-bit (Q1_0) and ternary (Q2_0) are merged into upstream llama.cpp, so recent upstream builds run them — for ternary, use the *-Q2_0_g64.gguf files. Ternary Bonsai 2 (PTQ1_0 / PQ2_0) still needs the PrismML fork (prism branch) or its pre-built binaries; on a stock build those files are refused, and a Q2_0 file loads silently and outputs gibberish. See Formats & runtime support.

Get a binary

The demo repo’s setup.sh downloads a compatible build automatically into bin/. To grab one yourself, use the PrismML llama.cpp release (prism-b10658 or newer), which covers all three bands on macOS (Apple Silicon and Intel), Linux (CPU / CUDA 12.4 / CUDA 12.8 / Vulkan / ROCm), Windows (CPU / CUDA / Vulkan / HIP), and iOS (XCFramework). The full variant list is in Formats & runtime support.

Run a chat

  • Always pass an explicit -c. Bare -c 0 uses the model’s full training context (262K on the 27B) regardless of available memory and can exhaust it; the demo scripts instead pick a RAM-tiered default. Memory-per-context numbers are on each model card.
  • The sampling flags are Bonsai’s recommended defaults.
  • -ngl 99 offloads all layers to the GPU (the wrapper script sets this automatically when a GPU is present).
The 27B needs two extra flags: --mmproj <path>.gguf for image input, and --jinja for native OpenAI-style tool_calls (the demo scripts already pass both).
The demo repo wraps all of this:

Start a server

This is the OpenAI-compatible endpoint described in Run the server, which also documents the recommended server flags.

Build from source

If the pre-built binaries don’t cover your platform or you want a specific backend:
The prism branch is required for Ternary Bonsai 2 and for the PQ2_0 ternary packing. For 1-bit or the group-64 ternary files, upstream llama.cpp works too. The demo repo also has ready-made build scripts: scripts/build_mac.sh, scripts/build_cpu_linux.sh, scripts/build_cuda_linux.sh, and scripts/build_cuda_windows.ps1.
Building with CUDA is memory-hungry — the build scripts cap parallelism at -j 2 when they detect less than 16 GB of GPU VRAM. Building by hand on a low-VRAM machine, reduce -j yourself.