Skip to main content
Bonsai 4B targets edge deployments where the bigger Bonsai models don’t fit or where you need headroom for other workloads. At 0.57 GB (1-bit), the weights are small enough to power applications on wearables, and mobile devices.

Specifications

Artifacts

The FP16 reference weights (8.05 GB) are in the ternary GGUF repo.

Run it

Through the demo repo:
Or directly with llama.cpp / MLX: