git, curl, and enough disk for the model you pick (see sizes). Everything else, including Python tooling and the inference engine, is installed by the setup script.
If you skip step 2 entirely, setup uses its defaults:
BONSAI_FAMILY=bonsai2, BONSAI_MODEL=27B — Ternary Bonsai 2 27B, a 7.8 GB download including the vision projector. The two earlier families also come in 8B, 4B, and 1.7B; Ternary Bonsai 2 is 27B only.Install and run
1
Clone the repo
2
Choose your model (optional)
3
Run setup
./scripts/download_models.sh (or re-run setup).4
Chat
Start a server instead
Prefer a browser UI or an API endpoint? The repo ships an OpenAI-compatible server:./scripts/start_openwebui.sh starts the server and Open WebUI together. See Run the server for flags, sampling defaults, and API examples.
What setup actually does
Even on a fresh machine,setup.sh / setup.ps1:
- Installs system build tools (Xcode Command Line Tools on macOS,
build-essentialon Linux) - Installs uv and creates a Python virtualenv in
.venv/ - Downloads the selected model from Hugging Face into
models/(public repos, no token needed) - Downloads pre-built llama.cpp binaries for your platform into
bin/, or builds from source if none match - On Apple Silicon, builds MLX from source for GPU acceleration
- Installs Open WebUI and the code-interpreter venv (
.venv-jupyter) for the agentic demo — a few GB more and most of the wait; skip withBONSAI_OPENWEBUI=0andBONSAI_CODE_INTERPRETER=0
Bonsai-demo directory removes everything else.
Next steps
Run the server
The OpenAI-compatible endpoint, its defaults, and API examples.
Connect your tools
OpenClaw, Hermes, and anything OpenAI-compatible.
Pick the right model
Size, family, and format trade-offs with real numbers.
Troubleshooting
Port conflicts, slow generation, models that won’t load.