Specifications
Artifacts
The FP16 reference weights (16.38 GB) are also in the ternary GGUF repo for comparison work.
Run it
Through the demo repo:Ternary GGUF needs the PrismML llama.cpp fork; 1-bit runs on upstream llama.cpp. See Formats & runtime support.