bonsai 2 27b

bonsai 2 27b

An open ternary 27B-class GGUF reasoning model for local llama.cpp inference on CUDA, Metal, and CPU.

bonsai 2 27b cover

Overview

Bonsai 2 27B — GGUF is Prism ML's open ternary transformer model for local reasoning and text generation. The Hugging Face model card describes it as a 27B-class model for llama.cpp on CUDA, Metal, and CPU, with a Qwen3.8-27B hybrid-attention backbone and Apache-2.0 licensing.

Key Features

  • Ternary g128 weights using values {-1, 0, +1} with FP16 group scaling and an effective 1.72 bits per weight.

  • Two GGUF packings: PTQ1_0 at about 5.95 GB and PQ2_0 at about 7.21 GB, with custom ternary hybrid-attention kernels.

  • About 98.2% of the FP16 benchmark average retained in the model card's 14 thinking-mode benchmarks.

  • 262K-token context inherited from the Qwen3.8-27B hybrid-attention architecture.

  • Optional Q8_0 vision-tower packs for image input, while text-only inference uses the language model alone.

  • Local backends including llama.cpp with CUDA, Metal, and CPU; an MLX companion is also linked for Apple Silicon.

How to Use

Download a GGUF pack from Hugging Face and use the Prism ML llama.cpp fork or the tested Bonsai demo setup. The model card documents serving and CLI workflows, including model loading, GPU offload, context configuration, sampling parameters, and OpenAI-compatible local servers. Stock llama.cpp is not sufficient for these ternary files; the model card directs users to the Prism ML fork with the required low-bit kernels and Hadamard activation runtime.

Use Cases

Bonsai 2 27B is intended for local reasoning, coding, agentic tool calling, long-context text generation, and on-device experimentation where a much smaller footprint matters. The published measurements include desktop GPUs and Apple Silicon throughput, but actual performance depends on the selected packing, backend, hardware, quantization, and runtime.

Pricing and Availability

The model files are openly available on Hugging Face under Apache-2.0. It is not deployed through a Hugging Face Inference Provider, so the documented workflow requires local hardware and a compatible runtime.

Screenshots

Featured Products

Neogenio SEO & GEO AI Engine

SEO and GEO AI visibility for Businesses and Agencies. Explore its features, pricing, and fit for practical seo tools workflows.

Rankora

SEO simplified. Rankora tells you exactly what to do to rank. Explore its features, pricing, and fit for practical seo tools workflows.

BlogSEO

Rank #1 on Google & ChatGPT on Autopilot. Explore its features, pricing, and fit for practical seo tools workflows. Compare its capabilities and use cases...

Synapse

Find missing topics and outrank competitors in Google AI Overviews. Explore its features, pricing, and fit for practical seo tools workflows.

OranGEO

AI Search Optimization & GEO Analytics Platform. Explore its features, pricing, and fit for practical seo tools workflows.

Conductor AI

Purpose-built AI for AEO and SEO growth. Explore its features, pricing, and fit for practical seo tools workflows. Compare its capabilities and use cases...