Skip to main content
This page gets you from zero to a running Ternary Bonsai 2 27B on macOS, Linux, or Windows. One script installs build tools, downloads the model weights and pre-built inference binaries, and leaves you ready to chat. Prerequisites: git, curl, and enough disk for the model you pick (see sizes). Everything else, including Python tooling and the inference engine, is installed by the setup script.
If you skip step 2 entirely, setup uses its defaults: BONSAI_FAMILY=bonsai2, BONSAI_MODEL=27B — Ternary Bonsai 2 27B, a 7.8 GB download including the vision projector. The two earlier families also come in 8B, 4B, and 1.7B; Ternary Bonsai 2 is 27B only.
Setting this up with an AI coding agent? Point it at AGENTS.md in the repo — hardware-specific knobs, defaults, and what to ask you.

Install and run

1

Clone the repo

2

Choose your model (optional)

Not sure which to pick? Ternary Bonsai 2 is the highest quality; ternary is the earlier mid-point; 1-bit is smallest and fastest.
3

Run setup

Setup installs and downloads only — it doesn’t start anything. Re-running it is safe: it skips steps that are already done. To download a different model later, change the environment variables and run ./scripts/download_models.sh (or re-run setup).
4

Chat

On Apple Silicon you can also run through MLX, which is typically faster on M-series chips:

Start a server instead

Prefer a browser UI or an API endpoint? The repo ships an OpenAI-compatible server:
That gives you chat, vision, and tool calling in one place. Most tools such as Hermes and OpenClaw connect to the same endpoint. For a full ChatGPT-style interface, ./scripts/start_openwebui.sh starts the server and Open WebUI together. See Run the server for flags, sampling defaults, and API examples.

What setup actually does

Even on a fresh machine, setup.sh / setup.ps1:
  1. Installs system build tools (Xcode Command Line Tools on macOS, build-essential on Linux)
  2. Installs uv and creates a Python virtualenv in .venv/
  3. Downloads the selected model from Hugging Face into models/ (public repos, no token needed)
  4. Downloads pre-built llama.cpp binaries for your platform into bin/, or builds from source if none match
  5. On Apple Silicon, builds MLX from source for GPU acceleration
  6. Installs Open WebUI and the code-interpreter venv (.venv-jupyter) for the agentic demo — a few GB more and most of the wait; skip with BONSAI_OPENWEBUI=0 and BONSAI_CODE_INTERPRETER=0
Nothing is installed globally except the system build tools; deleting the Bonsai-demo directory removes everything else.

Next steps

Run the server

The OpenAI-compatible endpoint, its defaults, and API examples.

Connect your tools

OpenClaw, Hermes, and anything OpenAI-compatible.

Pick the right model

Size, family, and format trade-offs with real numbers.

Troubleshooting

Port conflicts, slow generation, models that won’t load.