GGUF
For llama.cpp and the ecosystem built on it. Cross-platform: CPU, CUDA, Metal, Vulkan, ROCm.
MLX
Apple’s array framework, tuned for Apple Silicon. Best raw performance on M-series Macs and the basis for the iPhone/iPad apps.
Quantization types
Neither Ternary Bonsai 2 packing is uniformly faster than the other; they trade footprint against unpack cost.
The three ternary GGUF files
Published Ternary-Bonsai files were deliberately not renamed when the format migrated, since too many things link to them. Three variants therefore sit in the current repos, and each needs the right binaries:
Future releases drop the transitional suffix.
setup.sh downloads PQ2_0 where the backend is optimized for it and the group-64 file otherwise.
Runtime support matrix
This is the part that matters most for a low-bit model: a runtime without native kernels either refuses the file or runs it dequantized, silently losing the speed and memory advantage.1-bit (Q1_0): merged upstream
Because
Q1_0 is upstream, any llama.cpp-based tool built from a recent enough version runs 1-bit Bonsai. When in doubt, use the PrismML llama.cpp binaries, which are known-good.
Ternary (Q2_0): merged upstream
Ternary support has landed in mainline llama.cpp, so the group-64 files run on a stock build with no fork needed.
Use the
*-Q2_0_g64.gguf file for a stock build. The PQ2_0 file is smaller and usually faster, but needs the fork binaries.
Ternary Bonsai 2 (PTQ1_0 / PQ2_0): PrismML fork for now
Ternary Bonsai 2 uses a rotated weight basis, which requires an activation-side Walsh–Hadamard transform at runtime. That transform is not upstream yet, so every band currently requires the demo’s binaries from the PrismML fork.
More will be added here as they go up.
The MLX package for Ternary Bonsai 2 uses a 2-bit, group-size-128 packing analogous to
PQ2_0 and runs on stock MLX.
Pre-built binaries
The PrismML llama.cpp release covers 1-bit, ternary, and Ternary Bonsai 2 on:
Use
prism-b10658 or newer: earlier builds predate the ternary group-64 migration and refuse the current files. setup.sh in the demo repo picks the right one automatically. Build-from-source instructions are on the llama.cpp page.