Skip to main content
The demo repo ships an OpenAI-compatible server. Start it once and every tool in these docs (OpenClaw, Open WebUI, Hermes, your own code) connects to the same endpoint.

Start the server

The server loads whichever model BONSAI_MODEL / BONSAI_FAMILY select — Ternary Bonsai 2 27B unless you set them (see Quickstart) — and also serves a built-in web chat at http://localhost:8080, with image upload, an MCP client, and a per-chat reasoning-effort picker. On Apple Silicon there is a second option, an MLX-backed server on its own port:
Both speak the same API. The llama.cpp server is the default; the MLX server is usually faster on M-series Macs.

The endpoint

The scripts bind to 127.0.0.1, so the server is reachable only from the same machine and there is no authentication. Setting BONSAI_HOST to any non-loopback address (for example 0.0.0.0) exposes an unauthenticated server to your network — only do that on a trusted network, or firewall the port.

What the script configures

start_llama_server.sh doesn’t just launch llama-server; it applies Bonsai’s recommended settings: Use the same sampling values when you call the API directly or configure a third-party tool; they are what the published benchmarks assume. ‹TODO: confirm benchmark sampling settings match the demo defaults›
Thinking mode defaults differ by size. The 27B is a reasoning model and serves with thinking on by default. In the built-in chat UI, click the lightbulb in the message box and pick a Reasoning effort — Off, Low (512 tokens), Medium (2,048), High (8,192), or Max (unlimited); the pick persists per browser and is sent with every request, no restart needed. Over the API, toggle it per request with thinking_budget_tokens (0 disables it, -1 is unlimited), cap the default for clients that don’t specify one with --reasoning-budget N, or disable it entirely with BONSAI_THINKING=0. The 8B/4B/1.7B demo scripts still disable thinking (enable_thinking: false) by default for faster responses.On slower hardware, thinking is usually the bulk of the wait — pick a lower effort.
Extra arguments pass straight through to llama-server, e.g. ./scripts/start_llama_server.sh --reasoning-budget 2048 to cap the default reasoning budget, or -c 32768 to pin the context size.

Optional extras

Three off-by-default flags for the llama.cpp server:

Call it

Any client that accepts a custom base URL works the same way.

Images and tool calls

The 27B accepts image_url parts in messages and does native OpenAI-style tool_calls with full round-trips, so vision and tool use work over this endpoint with no special client. The built-in chat UI also ships an MCP client (Hugging Face and DeepWiki preconfigured, per-chat opt-in). Costs, the image-token cap, and adding your own MCP servers: VISION.md and TOOLS.md in the demo repo.

Connect a tool

Open WebUI

A ChatGPT-style interface; one script starts everything.

OpenClaw

Bridge Bonsai into a personal AI agent gateway.

Hermes

Run the Hermes terminal agent on local Bonsai.

All integrations

Everything that speaks OpenAI-compatible.