Specifications
Variants
The Ternary model is the recommended default. It preserves more of the original model’s visual quality while reducing the diffusion transformer to 1.21 GB. The 1-bit model brings the diffusion transformer below 1 GB and is optimized for devices where memory capacity and bandwidth are the primary constraints.
Artifacts
All model repositories are available in the Bonsai Image collection on Hugging Face.
The MLX repositories are intended for Apple Silicon. The Gemlite repositories use low-bit CUDA kernels for NVIDIA GPUs.
How to run it
Clone the Bonsai Image demo repository:- MLX on Apple Silicon - Gemlite on Linux with an NVIDIA GPU
Generate an image
Run a single generation from the command line:Launch the local studio
Start the generation API and browser interface:
Send a generation request from another terminal:
Performance
For a 512 × 512 image, Bonsai Image 4B generates in approximately:
On an M4 Pro, Bonsai Image 4B is up to 5.6× faster than the full-precision MFLUX pipeline. Mean active memory during 512 × 512 generation is approximately: