prism-ml org on Hugging Face. Names follow one convention:
{SIZE} is 27B, 8B, 4B, or 1.7B for the two earlier families. Ternary Bonsai 2 ships at 27B only. Not sure which variant you want? See the lineup for size trade-offs, and Formats & runtime support for GGUF vs MLX.
Ternary Bonsai 2 requires the demo’s llama.cpp binaries from the PrismML fork — stock llama.cpp cannot run these files.
./setup.sh fetches the right ones for your machine.Automatic (demo repo)
setup.sh resolves the repository from BONSAI_MODEL and BONSAI_FAMILY and downloads into models/. With neither set, the default is Ternary Bonsai 2 27B. To fetch additional variants later without re-running full setup:
BONSAI_FAMILY takes bonsai2 (Ternary Bonsai 2, the default), ternary (Ternary-Bonsai), or bonsai (1-bit Bonsai).
Manual (Hugging Face CLI)
Repositories and sizes
Ternary Bonsai 2 27B
The default. The GGUF repo also carries anmmproj vision tower, needed for image input; the MLX package keeps its vision tower in full precision.
Shipped components, sizes on disk:
Both bands need the PrismML llama.cpp fork for now. A plain
./setup.sh pulls 7.8 GB: the PQ2_0 packing this demo defaults to, plus the 4-bit vision projector. The MLX package is 8.49 GB — its block format stores an FP16 bias alongside the scale, which pushes the packed rate from 2.125 to 2.250 bits per weight, and it keeps the vision tower at full precision.
Bonsai 27B
All 27B repos also carry anmmproj file for vision (+0.9 GiB), needed for image input regardless of format.
Bonsai 8B
Bonsai 4B
Bonsai 1.7B
What’s inside a GGUF repo
You usually want exactly one quant file from each repo:-
Ternary Bonsai 2 repos: two packings of the same ternary g128 weights.
*-PQ2_0.gguf(2.16 bpw, 7.25 GB) is what this demo downloads by default — a simpler 2-bit representation that is cheaper to unpack.*-PTQ1_0.gguf(1.76 bpw, 5.93 GB) is the smaller footprint. Neither is uniformly faster; both require the fork binaries. See Formats & runtime support. -
1-bit repos:
Bonsai-{SIZE}-Q1_0.ggufis the file to use. Q1_0 is merged upstream, so it runs on a stock llama.cpp build. -
Ternary repos: three files ship side by side, and each needs the right binaries:
Repos also carry the FP16 reference weights (large; only for comparison work) and, for speculative decoding, the
*dspark-dflash*drafter sidecar (~0.6 GB; the old*dspark-Q4_1*.gguffiles are the pre-migration packing). -
27B repos only: also grab the
mmprojfile alongside the quant file — it’s required for image input and isn’t optional the way the FP16 reference is.
mlx-lm model directories (config, tokenizer, safetensors); download the whole repo. The 1-bit packs need the PrismML-Eng/mlx fork (branch prism) until mlx#3161 merges; the 2-bit ternary packs run on stock MLX.
Next steps
Check runtime support
Make sure your runtime has native kernels for the file you downloaded.
Run it
Start the local server and connect your tools.