Get a binary
The demo repo’ssetup.sh downloads a compatible build automatically into bin/. To grab one yourself, use the PrismML llama.cpp release (prism-b10658 or newer), which covers all three bands on macOS (Apple Silicon and Intel), Linux (CPU / CUDA 12.4 / CUDA 12.8 / Vulkan / ROCm), Windows (CPU / CUDA / Vulkan / HIP), and iOS (XCFramework). The full variant list is in Formats & runtime support.
Run a chat
- Always pass an explicit
-c. Bare-c 0uses the model’s full training context (262K on the 27B) regardless of available memory and can exhaust it; the demo scripts instead pick a RAM-tiered default. Memory-per-context numbers are on each model card. - The sampling flags are Bonsai’s recommended defaults.
-ngl 99offloads all layers to the GPU (the wrapper script sets this automatically when a GPU is present).
The 27B needs two extra flags:
--mmproj <path>.gguf for image input, and --jinja for native OpenAI-style tool_calls (the demo scripts already pass both).Start a server
Build from source
If the pre-built binaries don’t cover your platform or you want a specific backend:
The
prism branch is required for Ternary Bonsai 2 and for the PQ2_0 ternary packing. For 1-bit or the group-64 ternary files, upstream llama.cpp works too. The demo repo also has ready-made build scripts: scripts/build_mac.sh, scripts/build_cpu_linux.sh, scripts/build_cuda_linux.sh, and scripts/build_cuda_windows.ps1.
Building with CUDA is memory-hungry — the build scripts cap parallelism at
-j 2 when they detect less than 16 GB of GPU VRAM. Building by hand on a low-VRAM machine, reduce -j yourself.