PrismML-Eng/Bonsai-demoPublic

Bonsai Demo

AI summary: A comprehensive demo repository for running the highly efficient 1-bit Bonsai and Ternary-Bonsai language models locally.

Stars
3.3K
+17 today
Forks
365
Watchers
29
Open issues
12
Open PRs
5
Contributors
~50
Commits
178
Branches
6

ShellApache-2.0Created Mar 25, 2026Last push 4d ago+156 stars this week+994 this month

Quick answers

What is Bonsai-demo?
A comprehensive demo repository for running the highly efficient 1-bit Bonsai and Ternary-Bonsai language models locally.
What does Bonsai-demo do?
Bonsai Demo provides the tools, interfaces, and instructions needed to run the highly compressed Bonsai language models on consumer hardware. It supports the latest Bonsai 27B vision-language models, as well as 1-bit and ternary variants. The repository offers implementations for Mac via Metal, and Linux/Windows via CUDA, Vulkan, or CPU, ensuring broad hardware compatibility. It showcases advanced features like native OpenAI-style agentic tool calling and direct integration with Model Context Protocol (MCP) servers. The demo serves as the primary entry point for developers looking to deploy extremely resource-efficient LLMs locally.
Who is Bonsai-demo for?
AI researchers, local LLM enthusiasts, and developers looking to run highly efficient, quantized models on consumer hardware. It is ideal for those wanting to explore the cutting edge of 1-bit and ternary model performance.
How do I get started with Bonsai-demo?
https://github.com/PrismML-Eng/Bonsai-demo
How popular is Bonsai-demo on GitHub?
PrismML-Eng/Bonsai-demo has 3,266 stars and 365 forks on GitHub, and gained 156 stars in the last 7 days.
What license does Bonsai-demo use?
PrismML-Eng/Bonsai-demo is released under the Apache-2.0 license.

Star history

since Jul 28, 2026
01K2K3KJul 2026Aug 2026Sep 2026Oct 2026
3.3K stars as of Oct 4, 2026. Measured daily since Jul 28, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks
SepOctNovDecJanFebMarAprMayJunJulAugSepMonWedFri2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 1 commit2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 2 commits2026-03-31: 3 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 1 commit2026-04-05: 0 commits2026-04-06: 5 commits2026-04-07: 2 commits2026-04-08: 2 commits2026-04-09: 1 commit2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 9 commits2026-04-14: 0 commits2026-04-15: 23 commits2026-04-16: 7 commits2026-04-17: 1 commit2026-04-18: 5 commits2026-04-19: 0 commits2026-04-20: 4 commits2026-04-21: 0 commits2026-04-22: 1 commit2026-04-23: 3 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 1 commit2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 1 commit2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 0 commits2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 0 commits2026-07-03: 0 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 0 commits2026-07-09: 0 commits2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 3 commits2026-07-14: 2 commits2026-07-15: 1 commit2026-07-16: 0 commits2026-07-17: 9 commits2026-07-18: 1 commit2026-07-19: 3 commits2026-07-20: 4 commits2026-07-21: 0 commits2026-07-22: 4 commits2026-07-23: 1 commit2026-07-24: 0 commits2026-07-25: 3 commits2026-07-26: 0 commits2026-07-27: 0 commits2026-07-28: 0 commits2026-07-29: 1 commit2026-07-30: 1 commit2026-07-31: 0 commits2026-08-01: 0 commits2026-08-02: 0 commits2026-08-03: 0 commits2026-08-04: 0 commits2026-08-05: 2 commits2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 0 commits2026-08-09: 0 commits2026-08-10: 0 commits2026-08-11: 0 commits2026-08-12: 0 commits2026-08-13: 0 commits2026-08-14: 0 commits2026-08-15: 0 commits2026-08-16: 0 commits2026-08-17: 0 commits2026-08-18: 0 commits2026-08-19: 0 commits2026-08-20: 0 commits2026-08-21: 0 commits2026-08-22: 0 commits2026-08-23: 0 commits2026-08-24: 0 commits2026-08-25: 1 commit2026-08-26: 0 commits2026-08-27: 2 commits2026-08-28: 9 commits2026-08-29: 0 commits2026-08-30: 0 commits2026-08-31: 0 commits2026-09-01: 0 commits2026-09-02: 0 commits2026-09-03: 1 commit2026-09-04: 0 commits2026-09-05: 0 commits2026-09-06: 0 commits2026-09-07: 0 commits2026-09-08: 0 commits2026-09-09: 0 commits2026-09-10: 1 commit2026-09-11: 0 commits2026-09-12: 0 commits2026-09-13: 0 commits2026-09-14: 1 commit2026-09-15: 0 commits2026-09-16: 0 commits2026-09-17: 2 commits2026-09-18: 2 commits2026-09-19: 0 commits2026-09-20: 7 commits2026-09-21: 10 commits2026-09-22: 4 commits2026-09-23: 9 commits2026-09-24: 12 commits2026-09-25: 0 commits2026-09-26: 0 commits
168 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Permissive license

    Apache-2.0

  • Continuous integration

    Automated checks passing

  • Repeat trending

    5 trending appearances

What Bonsai-demo does

Bonsai Demo provides the tools, interfaces, and instructions needed to run the highly compressed Bonsai language models on consumer hardware. It supports the latest Bonsai 27B vision-language models, as well as 1-bit and ternary variants. The repository offers implementations for Mac via Metal, and Linux/Windows via CUDA, Vulkan, or CPU, ensuring broad hardware compatibility. It showcases advanced features like native OpenAI-style agentic tool calling and direct integration with Model Context Protocol (MCP) servers. The demo serves as the primary entry point for developers looking to deploy extremely resource-efficient LLMs locally.

AI researchers, local LLM enthusiasts, and developers looking to run highly efficient, quantized models on consumer hardware. It is ideal for those wanting to explore the cutting edge of 1-bit and ternary model performance.

  • Broad hardware support: Runs efficiently across Apple Silicon (Metal), NVIDIA (CUDA), AMD (ROCm), Vulkan, and standard CPUs.
  • Vision-language capabilities: Supports the new Bonsai 27B models for analyzing photos, screenshots, and PDFs natively.
  • Agentic tool calling: Implements native, OpenAI-compatible tool calling with full round-trip execution for autonomous tasks.
  • Model Context Protocol support: Integrates seamlessly with MCP servers, allowing the models to access external data sources.
  • Extreme quantization: Showcases 1-bit and ternary quantized models that drastically reduce memory usage without destroying performance.
  • Multiple demo interfaces: Includes both terminal-based and graphical UIs for interacting with the local models immediately.

Where teams use it

Low-resource local inference

Run capable, multi-billion parameter language models on older laptops or devices with limited unified memory.

Local vision-language analysis

Use the Bonsai 27B model to privately analyze sensitive PDF documents or screenshots entirely on local hardware.

Agentic workflow testing

Deploy the models to act as local, fast-executing agents that use standard tool calls to interact with MCP servers.

Quantization research

Experiment with the provided 1-bit and ternary models to understand the performance trade-offs of extreme LLM compression.

Getting started: https://github.com/PrismML-Eng/Bonsai-demo

README

main branch

Bonsai Demo

Backend and model format compatibility: BACKEND-SUPPORT.md.

Bonsai

Website  |  GitHub  |  Discord

Models: Bonsai 2 27B GGUF · Bonsai 2 27B MLX · Bonsai 27B · Ternary-Bonsai · Bonsai (1-bit)

Whitepapers: Bonsai 27B · 1-bit Bonsai 8B · Ternary-Bonsai 8B


Using this demo repository you can run Bonsai 2, Bonsai (1-bit) and Ternary-Bonsai language models locally on Mac (Metal), Linux/Windows (CUDA, Vulkan, ROCm), or CPU.

🌱 New: Bonsai 2 27B

Bonsai 2 27B is this demo's default. Full 27B-class reasoning in ternary weights, at 5.9 GB, running on a laptop or a single GPU.

  • 98.2% of FP16 intelligence retained at roughly a ninth of the size, with the reasoning core intact: math within half a point of full precision, coding level with the baseline.
  • Vision: send photos, screenshots and PDFs and ask about them, on both llama.cpp and MLX (see VISION.md).
  • Agentic tool calling: native OpenAI-style tool_calls with full round-trips, plus MCP servers in both demo UIs (see TOOLS.md).
  • Thinking: a reasoning model; pick the reasoning effort per chat in the UI or budget it per request.
  • 262K-token context, kept practical on-device by the hybrid-attention backbone.
  • 1.75 bits per weight in the PTQ1_0 packing. A second packing, PQ2_0, trades 1.3 GB for faster prompt processing and is what this demo downloads by default. See MODEL-FORMATS.md.

Bonsai 2 needs this demo's llama.cpp binaries, from the PrismML fork; stock llama.cpp cannot run these files. ./setup.sh fetches the right ones for your machine.

Quick Start below gets you there in two commands: ./setup.sh downloads Bonsai 2 27B, then ./scripts/start_llama_server.sh gives you chat, vision and tools at http://localhost:8080.

The earlier Bonsai families are still here, in smaller sizes too. See Models.

Best Practices (Bonsai 2 27B)

These match the model card. The start scripts already apply the thinking-mode sampling, so you only need these when calling the model from your own client.

Generation Parameters

Thinking mode (default) Instruct / non-thinking
temperature 1.0 0.7
top_p 0.95 0.80
top_k 20 20
min_p 0.05 0.0
presence_penalty 0.0 1.5
repetition_penalty 1.0 1.0

min_p=0.05 drops tokens far less likely than the top choice; in our tests it scored at least as well as min_p=0.0 and followed instructions more reliably. It is also llama.cpp's default. Give the model a generous output limit (-n 16384 or more, max_tokens for the API): it reasons before it answers, and a small cap ends generation mid-thought.

The model uses xhigh reasoning effort by default; use medium for shorter responses and a balance of speed and accuracy. low reasoning effort is not supported and when selected the model will behave close to xhigh.

System Prompt

A simple system prompt works well:

You are a helpful assistant

Quick Start

For repeated conversations and prompt-cache troubleshooting, see Prompt reuse and context checkpoints.

Setting things up with an AI coding agent? Point it at AGENTS.md, a guide written for agents (hardware-specific knobs, defaults, and what to ask the user).

macOS / Linux

git clone https://github.com/PrismML-Eng/Bonsai-demo.git
cd Bonsai-demo
./setup.sh

That installs and downloads only. To chat, start the server yourself, then open http://localhost:8080:

./scripts/start_llama_server.sh

Or ask one question from the terminal without a server:

./scripts/run_llama.sh -p "What is the capital of France?"

setup.sh fetches the llama.cpp binaries for your machine and the Bonsai 2 27B weights, 7.8 GB in the PQ2_0 packing this demo defaults to plus its vision projector. It also sets up Open WebUI and the code interpreter, which add a few GB more and most of the wait. Skip those with BONSAI_OPENWEBUI=0 and BONSAI_CODE_INTERPRETER=0.

Windows (PowerShell)

git clone https://github.com/PrismML-Eng/Bonsai-demo.git
cd Bonsai-demo
Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass
.\setup.ps1

Then start the server and open http://localhost:8080:

.\scripts\start_llama_server.ps1

Speed Benchmarks

See community-benchmarks/ for results on different hardware and templates to submit your own.

Models

Bonsai 2 27B is the default: plain ./setup.sh downloads and runs it. Two earlier families remain available in sizes 27B, 8B, 4B, and 1.7B. Every 27B model is a vision-language model: it accepts images as well as text.

Both earlier formats are landing in mainline llama.cpp: Q1_0 (1-bit) is fully merged upstream, and Q2_0 (ternary) now runs on mainline CPU, Metal, Vulkan, and CUDA. Details and mainline-compatible files: binary status and ternary status below.

Bonsai 2 (ternary, default)

Available in GGUF (llama.cpp) and MLX 2-bit formats. Both bands need our llama.cpp fork for now, which ./setup.sh installs.

Model Format HuggingFace Repo
Bonsai-2-27B GGUF prism-ml/Ternary-Bonsai-2-27B-gguf
Bonsai-2-27B MLX (2-bit) prism-ml/Ternary-Bonsai-2-27B-mlx-2bit

On Apple Silicon, Bonsai 2 text, images, and tool calls are supported by native mlx-vlm==0.7.2 in .venv-vlm. Re-run ./setup.sh to upgrade an older environment, then use ./scripts/start_mlx_server.sh or BONSAI_BACKEND=mlx ./scripts/start_openwebui.sh. The server keeps thinking enabled; use ./scripts/start_mlx_server.sh --thinking-budget 8192 --max-tokens 32768 for an example bounded server run (these are server flags, not run_mlx.sh flags). API clients use thinking_budget (MLX), not llama-server's thinking_budget_tokens. Use the model-card sampling settings in API requests: temperature: 1.0, top_p: 0.95, top_k: 20, min_p: 0.05. Open WebUI refuses to reuse an existing server on its MLX port for Bonsai 2 because its loader cannot be verified; stop it first so the launcher can start the tested runtime. This does not update LM Studio's separately bundled MLX runtime.

Set BONSAI_FAMILY=ternary or BONSAI_FAMILY=bonsai for the earlier families and their smaller sizes.

Bonsai (1-bit)

Available in GGUF (llama.cpp) and MLX 1-bit formats.

Model Format HuggingFace Repo
Bonsai-27B GGUF prism-ml/Bonsai-27B-gguf
Bonsai-27B MLX prism-ml/Bonsai-27B-mlx-1bit
Bonsai-8B GGUF prism-ml/Bonsai-8B-gguf
Bonsai-8B MLX prism-ml/Bonsai-8B-mlx-1bit
Bonsai-4B GGUF prism-ml/Bonsai-4B-gguf
Bonsai-4B MLX prism-ml/Bonsai-4B-mlx-1bit
Bonsai-1.7B GGUF prism-ml/Bonsai-1.7B-gguf
Bonsai-1.7B MLX prism-ml/Bonsai-1.7B-mlx-1bit

Set BONSAI_MODEL to choose which size to download and run (default: 27B).

Ternary-Bonsai

Available in GGUF (llama.cpp) and MLX 2-bit formats.

Model Format HuggingFace Repo
Ternary-Bonsai-27B GGUF prism-ml/Ternary-Bonsai-27B-gguf
Ternary-Bonsai-27B MLX (2-bit) prism-ml/Ternary-Bonsai-27B-mlx-2bit
Ternary-Bonsai-8B GGUF prism-ml/Ternary-Bonsai-8B-gguf
Ternary-Bonsai-8B MLX (2-bit) prism-ml/Ternary-Bonsai-8B-mlx-2bit
Ternary-Bonsai-4B GGUF prism-ml/Ternary-Bonsai-4B-gguf
Ternary-Bonsai-4B MLX (2-bit) prism-ml/Ternary-Bonsai-4B-mlx-2bit
Ternary-Bonsai-1.7B GGUF prism-ml/Ternary-Bonsai-1.7B-gguf
Ternary-Bonsai-1.7B MLX (2-bit) prism-ml/Ternary-Bonsai-1.7B-mlx-2bit

Set BONSAI_FAMILY=ternary to use this family.

Environment variables

Both variables are optional. If you set neither, the default is Bonsai-2-27B: that's what plain ./setup.sh downloads and runs.

Every launcher is configured through environment variables. The most common ones:

Variable Default Values Purpose
BONSAI_MODEL 27B 27B, 8B, 4B, 1.7B Model size. Bonsai 2 is 27B.
BONSAI_FAMILY bonsai2 bonsai2, ternary, bonsai Model family (bonsai2 = Bonsai 2, ternary = Ternary-Bonsai, bonsai = 1-bit Bonsai).
BONSAI_NGL auto-detect int; 0 = CPU-only GPU layer offload.
BONSAI_CTX auto (RAM-tiered) 0, or ≤ 262144 Context length (0/unset = automatic safe size).
BONSAI_HOST 127.0.0.1 any bind address Server bind address. A non-loopback value exposes the server — see the security note in the full reference.
BONSAI_SPECULATIVE 0 1 DSpark for previous-generation ternary/bonsai 27B only. No official Bonsai 2 drafter yet; warns and runs without speculation (SPECULATIVE.md).
BONSAI_KV4 0 1 4-bit KV cache for long contexts (KV-CACHE.md).

Full reference — all 24 variables (model/setup, server, MLX, Open WebUI, tools, and platform coverage): environment_variables.md.

Combine them freely:

./setup.sh                                                  # Bonsai-2-27B (default)
BONSAI_FAMILY=ternary ./setup.sh                            # Ternary-Bonsai-27B
BONSAI_FAMILY=ternary BONSAI_MODEL=1.7B ./setup.sh          # Ternary-Bonsai-1.7B
BONSAI_FAMILY=bonsai ./setup.sh                             # Bonsai-27B (1-bit)
BONSAI_FAMILY=bonsai BONSAI_MODEL=4B ./setup.sh             # Bonsai-4B
BONSAI_FAMILY=ternary BONSAI_MODEL=all ./setup.sh           # All 4 Ternary-Bonsai sizes
BONSAI_FAMILY=all BONSAI_MODEL=all ./setup.sh               # Full matrix
BONSAI_FAMILY=bonsai BONSAI_SKIP_GGUF=1 ./setup.sh          # Bonsai-27B, MLX only (macOS, saves disk space)

Upstream Status for Binary

Q1_0 is supported out of the box in upstream llama.cpp across many backends: CPU (generic, NEON, and optimized x86), Metal, CUDA, and Vulkan.

Runtime Status
llama.cpp (CPU, Metal, CUDA, Vulkan) ✅ Merged upstream, works out of the box
MLX (1-bit) ⏳ Pending upstream: mlx#3161; until it merges, use PrismML-Eng/mlx (branch prism, built automatically by setup.sh)

Upstream Status for Bonsai 2

Bonsai 2 needs the Hadamard activation transform, which is not upstream yet, so every band currently requires this demo's binaries from the PrismML fork.

Change Status Where
FWHT with F16 input (CPU) ⏳ Open ggml-org/llama.cpp#27779

More will be added here as they go up. Until this work lands, do not run Bonsai 2 on stock llama.cpp: PQ2_0 and PTQ1_0 are refused outright, but Q2_0 loads without a warning and outputs gibberish, which is why that band is kept in a separate repo.

Upstream Status for Ternary (Bonsai 1)

Ternary support has landed in mainline llama.cpp for CPU, Metal, Vulkan and CUDA, so the group-64 Q2_0 files run on a stock build with no fork needed. The x86 AVX-512-VNNI optimization is still pending, but x86 already works through the generic CPU path. MLX 2-bit runs on stock MLX.

Published files were deliberately not renamed, since too many things link to them. The result is three ternary formats on the current repos, and each needs the right binaries:

File Format Runs on
*-PQ2_0.gguf Group size 128 (2.13 bpw), our packing under its own ggml type. What this demo prefers: smallest file and usually fastest where the backend is optimized (CUDA, Metal, CPU, ROCm) This demo / fork binaries prism-b10658+
*-Q2_0_g64.gguf (27B file: *-Q2_g64.gguf) Group size 64 (2.25 bpw). The official llama.cpp Q2_0 format, widest backend coverage (adds Vulkan and SYCL) Mainline llama.cpp and fork binaries prism-b10658+
*-Q2_0.gguf (legacy, no g64) ⚠️ Deprecated. Pre-migration group-128 files stored under the ggml type id that now belongs to the official group-64 format Only the old prism-v5 releases; newer binaries refuse them with an error

Future releases drop the transitional suffix. The demo's setup scripts download PQ2_0 where the backend is optimized for it and the group-64 file otherwise (details: community-benchmarks/ternary-bonsai/README.md).

Speculative decoding: use this demo's binaries. Bonsai 2 27B has no official DSpark drafter released yet; the following applies to the previous-generation ternary and bonsai 27B models. Since the rebase, dspark rides on mainline llama.cpp's own DSpark implementation (ggml-org/llama.cpp#25173) with fork-side patches on top, and the drafter is the converted *dspark-dflash* sidecar (~0.6 GB; the old *dspark-Q4_1*.gguf files are the pre-migration packing). Use BONSAI_SPECULATIVE=1 with this demo's binaries — see SPECULATIVE.md.

To run the smaller ternary models directly on stock ggml-org/llama.cpp, use the group-64 files:

Model Repo File (mainline-compatible)
1.7B prism-ml/Ternary-Bonsai-1.7B-gguf Ternary-Bonsai-1.7B-Q2_0_g64.gguf
4B prism-ml/Ternary-Bonsai-4B-gguf Ternary-Bonsai-4B-Q2_0_g64.gguf
8B prism-ml/Ternary-Bonsai-8B-gguf Ternary-Bonsai-8B-Q2_0_g64.gguf
hf download prism-ml/Ternary-Bonsai-1.7B-gguf Ternary-Bonsai-1.7B-Q2_0_g64.gguf --local-dir models
hf download prism-ml/Ternary-Bonsai-4B-gguf  Ternary-Bonsai-4B-Q2_0_g64.gguf  --local-dir models
hf download prism-ml/Ternary-Bonsai-8B-gguf  Ternary-Bonsai-8B-Q2_0_g64.gguf  --local-dir models

What setup.sh Does

The setup script handles everything for you, even on a fresh machine:

  1. Checks/installs system deps: Xcode CLT on macOS, build-essential on Linux
  2. Installs uv: fast Python package manager (user-local, not global)
  3. Creates a Python venv and runs uv sync — installs cmake, ninja, huggingface-cli from pyproject.toml
  4. Downloads models from HuggingFace (all model repos are public; no token needed)
  5. Downloads pre-built binaries from the pinned GitHub Release (or builds from source if you prefer)
  6. Builds MLX from source (macOS only): clones our fork, builds it into the venv, installs the ML stack (mlx-lm, torch, transformers)
  7. Installs Open WebUI into the venv for the agentic demo (skip with BONSAI_OPENWEBUI=0)
  8. Builds the code-interpreter venv (.venv-jupyter): Jupyter + matplotlib / pandas / numpy / scipy / sympy / yfinance for the Open WebUI code interpreter (skip with BONSAI_CODE_INTERPRETER=0)

Re-running setup.sh is safe — it skips already-completed steps.


Running the Model

Every script runs Bonsai 2 27B unless you set BONSAI_FAMILY and BONSAI_MODEL to pick another one (Environment variables).

llama.cpp (Mac / Linux — auto-detects platform)

./scripts/run_llama.sh -p "What is the capital of France?"

These scripts run the llama.cpp backend and need GGUF weights. On an MLX-only setup (e.g. you used BONSAI_SKIP_GGUF=1), they stop with an error that points at both options — running the MLX script directly (run_mlx.sh / start_mlx_server.sh), or downloading the GGUF weights.

llama.cpp (Windows PowerShell)

.\scripts\run_llama.ps1 -p "What is the capital of France?"

MLX — Mac (Apple Silicon)

source .venv/bin/activate
./scripts/run_mlx.sh -p "What is the capital of France?"

Tested versions (reproducibility). The released MLX weights are plain safetensors and need no runtime patches. The 1-bit packs need an MLX build with 1-bit quantization support: the PrismML-Eng/mlx fork, branch prism, until mlx#3161 merges upstream. The 2-bit ternary packs run on stock MLX. The released 27B packs were validated with:

  • Python 3.11
  • mlx fork branch prism at commit 88c9c20
  • mlx-lm==0.31.2 (the version setup.sh pins)

setup.sh builds the fork from the branch tip. To pin the exact validated runtime instead, clone and check out the commit before running setup; setup reuses an existing ./mlx checkout:

git clone -b prism https://github.com/PrismML-Eng/mlx.git mlx
git -C mlx checkout 88c9c20
./setup.sh

Chat Server

Start llama-server with its built-in chat UI:

./scripts/start_llama_server.sh    # http://localhost:8080

For Windows PowerShell:

.\scripts\start_llama_server.ps1

The scripts auto-detect your GPU (Metal, CUDA, ROCm, Vulkan) and offload all layers. If the detection picks a GPU you do not want, for example Vulkan on a machine whose only GPU is a weak integrated one, set BONSAI_NGL=0 for CPU-only inference, or any layer count for partial offload (PowerShell: $env:BONSAI_NGL = "0").

Thinking

The 27B is a thinking model and serves with thinking enabled. To adjust it per conversation in the chat UI (no restart): click the lightbulb in the message box and pick a Reasoning effort: Off, Low (512 tokens), Medium (2,048), High (8,192), or Max (unlimited). The pick persists per browser and is sent with every request.

On slower hardware, thinking is usually the bulk of the wait; pick a lower effort in the UI. For API clients that don't specify a reasoning effort, you can cap the server-wide default by passing llama-server flags straight through the start script:

./scripts/start_llama_server.sh --reasoning-budget 2048

For API clients, the model's own reasoning_effort: "medium" is usually the better way to shorten thinking. At moderate output limits it thinks noticeably less than the default xhigh and is about as accurate, and it rarely runs out of budget mid-reasoning. Send it per request, or make it the server-wide default (a request can still ask for xhigh):

./scripts/start_llama_server.sh --chat-template-kwargs '{"reasoning_effort":"medium"}'

The chat UI's Reasoning effort levels are different: they are fixed thinking budgets (Medium is 2,048 tokens), which cut thinking off at that length rather than asking the model to think less.

Tool calling & MCP

The 27B does native OpenAI-style tool calling over the API, and the chat UI has an MCP client with Hugging Face + DeepWiki preconfigured (per-chat opt-in from the MCP selector in the message box, no prompt cost until you turn one on). Details, costs, and how to add your own servers: TOOLS.md.

Agentic demo

Bonsai 2 driving the Hermes agent end to end: from a two-line brief to a playable 3D skateboard game it verified in its own browser, then a round of plain-English feedback, all run with a fixed seed. Clip, pages, prompts, settings and the two scripts that run it: AGENT-DEMO.md.

Vision

Upload images in the chat UI (+ in the message box) or send image_url parts over the API; the scripts load the vision projector automatically and downscale very large images on slower backends. Costs, the image-token cap, and OCR tips: VISION.md.

Optional extras

Two experimental, off-by-default features for the llama.cpp chat server:

  • Speculative decoding: BONSAI_SPECULATIVE=1 pairs the previous-generation ternary or bonsai 27B with its DSpark drafter. Bonsai 2 27B has no official drafter yet; the launcher warns and runs without speculation for that family. Measured on an L40S (CUDA): 1.8-2.4x faster decode for the ternary 27B and 1.4-1.75x for the 1-bit 27B, workload-dependent (code/math best). On Apple Silicon (Metal) it only pays off for ternary code/math (~1.2x) and is a net slowdown otherwise, so leave it off on Macs. Needs this demo's binaries. Trade-offs and verification: SPECULATIVE.md.
  • 4-bit KV cache: BONSAI_KV4=1 cuts KV-cache memory roughly 3.5x for very long contexts, with an optional calibration bias for better quality (./scripts/make_kv_bias.sh). Details: KV-CACHE.md.
  • Vision projector in RAM: BONSAI_MMPROJ_CPU=1 keeps the 27B's vision projector in system RAM instead of VRAM (--no-mmproj-offload), freeing ~0.9 GiB of VRAM for KV/context on tight cards. The cost is a slower image prompt (the projector runs on CPU); text-only chat is unaffected.

Context Size

The 27B models support up to 262,144 tokens of context. The FP16 KV cache costs 64 KiB per token (~6.3 GiB at 100K), so 100K context fits on many consumer devices even without KV-cache quantization. The model's hybrid attention keeps the cache small for its size.

The launch scripts pick a default context sized to your machine's RAM, from 8K on small machines up to 131K for the 27B on machines with more than 71 GB (roughly 0.5 to 8 GiB of KV cache), so memory use stays predictable. Override with the BONSAI_CTX environment variable: pass any number up to 262144, or 0 (the same as leaving it unset) for the automatic RAM-tiered size. To force the model's full training context, pass the explicit number (e.g. BONSAI_CTX=262144) — only recommended on machines with plenty of headroom, since the scripts will not silently do this for you.

With the optional 4-bit KV cache (BONSAI_KV4=1) the cache drops to roughly 18 KiB per token, about 1.8 GiB at 100K, shaving ~4.5 GiB off the 100K figures below (for example, Ternary-Bonsai-27B on llama.cpp goes from ~13.7 to ~9.2 GiB).

Peak memory for the 27B (weights + activations + FP16 KV cache + ~1.2 GiB overhead; text-only, add ~0.9 GiB for the vision projector):

Model Format Weights 4K context 10K context 100K context
Bonsai-27B (1-bit) llama.cpp Q1_0 3.53 GiB 4.8 GiB 5.2 GiB 10.8 GiB
Bonsai-27B (1-bit) MLX 1-bit 3.92 GiB 5.5 GiB 5.9 GiB 11.4 GiB
Ternary-Bonsai-27B llama.cpp Q2_0 6.66 GiB 7.8 GiB 8.1 GiB 13.7 GiB
Ternary-Bonsai-27B MLX 2-bit 7.05 GiB 8.6 GiB 8.9 GiB 14.4 GiB
reference: 27B 16-bit GGUF BF16 47.73 GiB 49 GiB 49.6 GiB 55.2 GiB
reference: 27B "4-bit" llama.cpp UD Q4_K_M 15.73 GiB 17.2 GiB 17.6 GiB 23.2 GiB
reference: 27B "4-bit" MLX 4-bit 13.3 GiB 17.0 GiB 17.3 GiB 22 GiB

(The MLX packs are ~400 MiB larger than GGUF because MLX stores both scales and biases, GGUF only scales.)

Extra arguments pass straight through to llama.cpp, so ./scripts/run_llama.sh -c 8192 -p "Your prompt" also works for a one-off context override.

The older text-only sizes are smaller across the board; the 8B supports up to 65,536 tokens of context:

Estimates for Bonsai-8B (weights + KV cache + activations):

Context Size Est. Memory Usage
8,192 tokens ~2.5 GB
32,768 tokens ~5.9 GB
65,536 tokens ~10.5 GB

Open WebUI (Optional): the full agentic demo

Open WebUI gives you a ChatGPT-like interface on top of the local 27B: chat with images, tool calling against live tools, a server-side code interpreter (plots + market data), and a hidden-story sales database to investigate. Everything is configured automatically, no clicking through settings:

./scripts/start_openwebui.sh

setup.sh installs it for you; the script starts the backend, seeds the demo (tools, model settings, demo database), and opens http://localhost:9090. Backends, what to try, and customizing: OPENWEBUI.md.


Building from Source

If you prefer to build llama.cpp from source instead of using pre-built binaries:

Mac (Apple Silicon — Metal)

./scripts/build_mac.sh

Clones PrismML-Eng/llama.cpp, builds with Metal, outputs to bin/mac/.

Mac (Intel — CPU only)

./scripts/build_mac.sh

The script auto-detects Intel vs Apple Silicon. On Intel Macs, it builds with -DGGML_METAL=OFF (CPU only). MLX is also skipped automatically since it requires Apple Silicon.

Linux (CPU only)

./scripts/build_cpu_linux.sh

Builds a CPU-only binary with no GPU dependencies. Works on both x64 and arm64. Outputs to bin/cpu/.

Linux (CUDA)

./scripts/build_cuda_linux.sh

Auto-detects CUDA version. Pass --cuda-path /usr/local/cuda-12.8 to use a specific toolkit.

Linux (Vulkan)

# Install Vulkan SDK first (e.g. sudo apt install libvulkan-dev glslc)
git clone -b prism https://github.com/PrismML-Eng/llama.cpp.git
cd llama.cpp
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_VULKAN=ON
cmake --build build -j$(nproc)
# Binaries in build/bin/

Linux (ROCm / AMD GPU)

# Requires ROCm toolkit (hipcc)
git clone -b prism https://github.com/PrismML-Eng/llama.cpp.git
cd llama.cpp
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_HIP=ON
cmake --build build -j$(nproc)
# Binaries in build/bin/

Windows (CUDA)

.\scripts\build_cuda_windows.ps1

Auto-detects CUDA toolkit. Pass -CudaPath "C:\path\to\cuda" to use a specific version. Requires Visual Studio Build Tools (or full Visual Studio) and CUDA toolkit.

Windows (CPU only)

git clone -b prism https://github.com/PrismML-Eng/llama.cpp.git
cd llama.cpp
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release
# Binaries in build\bin\Release\

Requires Visual Studio Build Tools or full Visual Studio with C++ workload.


llama.cpp Pre-built Binary Downloads

All binaries are available from the pinned GitHub Release, also used by both setup scripts. The latest release may still be missing platform binaries while builds finish.

Platform
macOS Apple Silicon (arm64)
macOS Apple Silicon (KleidiAI)
macOS Intel (x64)
Linux x64 (CPU)
Linux arm64 (CPU)
Linux x64 (CUDA 12.4)
Linux x64 (CUDA 12.8)
Linux x64 (Vulkan)
Linux arm64 (Vulkan)
Linux x64 (ROCm 7.2)
Windows x64 (CPU)
Windows arm64 (CPU)
Windows x64 (CUDA 12.4)
Windows x64 (Vulkan)
Windows x64 (HIP/ROCm)
iOS (XCFramework)

Folder Structure

After setup, the directory looks like this:

Bonsai-demo/
├── README.md
├── TOOLS.md                        # Tool calling & MCP guide
├── OPENWEBUI.md                    # Open WebUI agentic demo guide
├── VISION.md                       # Image input: costs, caps, OCR tips
├── SPECULATIVE.md                  # Speculative decoding (experimental)
├── KV-CACHE.md                     # 4-bit KV cache (experimental)
├── AGENTS.md                       # Agent guide (hardware tuning knobs)
├── setup.sh                        # macOS/Linux setup
├── setup.ps1                       # Windows setup
├── pyproject.toml                  # Python dependencies
├── scripts/
│   ├── common.sh                   # Shared helpers + BONSAI_MODEL
│   ├── download_models.sh          # HuggingFace download
│   ├── download_binaries.sh        # GitHub release download
│   ├── run_llama.sh                # llama.cpp (auto-detects Mac/Linux)
│   ├── run_llama.ps1               # llama.cpp (Windows PowerShell)
│   ├── run_mlx.sh                  # MLX inference
│   ├── mlx_generate.py             # MLX Python script
│   ├── start_llama_server.sh       # llama.cpp server (port 8080)
│   ├── start_llama_server.ps1      # llama.cpp server (Windows PowerShell)
│   ├── start_mlx_server.sh         # MLX server (port 8081)
│   ├── start_openwebui.sh          # Open WebUI + auto-starts backends
│   ├── openwebui/                  # Open WebUI demo tools + seeding
│   ├── build_mac.sh                # Build llama.cpp for Mac
│   ├── build_cpu_linux.sh          # Build llama.cpp for Linux (CPU only)
│   ├── build_cuda_linux.sh         # Build llama.cpp for Linux CUDA
│   └── build_cuda_windows.ps1      # Build llama.cpp for Windows CUDA
├── models/                         # ← downloaded by setup
│   ├── gguf/
│   │   ├── 27B/                    # GGUF 27B model (+ mmproj for vision)
│   │   ├── 8B/                     # GGUF 8B model
│   │   ├── 4B/                     # GGUF 4B model
│   │   └── 1.7B/                   # GGUF 1.7B model
│   ├── Bonsai-27B-mlx/            # MLX 27B model (macOS)
│   ├── Bonsai-8B-mlx/             # MLX 8B model (macOS)
│   ├── Bonsai-4B-mlx/             # MLX 4B model (macOS)
│   └── Bonsai-1.7B-mlx/           # MLX 1.7B model (macOS)
├── bin/                            # ← downloaded or built by setup
│   ├── mac/                        # macOS binaries (Metal or CPU)
│   ├── cuda/                       # CUDA binaries (Linux/Windows)
│   ├── cpu/                        # CPU-only binaries (Linux/Windows)
│   ├── vulkan/                     # Vulkan binaries
│   ├── rocm/                       # ROCm binaries (AMD Linux)
│   └── hip/                        # HIP binaries (AMD Windows)
├── mlx/                            # ← cloned by setup (macOS)
└── .venv/                          # ← created by setup

Items marked with ← are created at setup time and excluded from git.


Appendix — FAQ

The model allocates huge memory or the machine freezes at startup

Older revisions defaulted to llama.cpp's -c 0, which uses the model's full training context (262K on the 27B) regardless of available memory and could exhaust it on constrained machines. The scripts now always use a RAM-tiered context instead; BONSAI_CTX=0 maps to that same safe default rather than -c 0. If you still hit memory pressure, pin a smaller context:

BONSAI_CTX=8192 ./scripts/start_llama_server.sh

M5 Mac on macOS 26.2/26.3: Metal compile errors, then out-of-memory

On M5 devices with certain macOS 26 point releases, the Metal tensor-API probe fails to compile at runtime (ggml_metal_library_init_from_source: error compiling source) and can leave the GPU in a bad state. This is an ecosystem-wide issue in the OS Metal headers, hitting every ggml-based project. Workaround, keeps full Metal speed and just skips the tensor-API path:

GGML_METAL_TENSOR_DISABLE=1 ./scripts/run_llama.sh -p "Hello"

Windows setup selects Vulkan or CPU instead of CUDA

Symptom: setup.ps1 reports [INFO] No GPU toolchain detected. Will use CPU build. or selects Vulkan, even though nvidia-smi runs successfully and detects your NVIDIA GPU.

Cause: Older versions of the setup script expected the header CUDA Version:. Some newer NVIDIA drivers instead report CUDA UMD Version:. The extra UMD prevented CUDA detection, causing setup to select another backend.

This can leave you with binaries under bin\vulkan or bin\cpu rather than bin\cuda. On affected Bonsai 2 setups, users reported prompts stalling without processing any tokens instead of producing a clear error.

CUDA source build runs out of memory or freezes

Symptom: cmake --build hangs, the system becomes unresponsive, or the build process is killed with an OOM error when building llama.cpp from source with CUDA enabled.

Cause: Compiling CUDA kernels is memory-intensive — each parallel compile job can consume several GB of GPU VRAM and/or system RAM. Running make -j$(nproc) on a machine with a low-VRAM GPU (< 16 GB) or limited system RAM can exhaust available memory.

How the build scripts handle this: build_cuda_linux.sh and build_cuda_windows.ps1 automatically detect the GPU's VRAM before building. If the maximum detected VRAM is less than 16 GB, the scripts cap parallelism at -j 2 instead of using all logical CPU cores. You will see a message like:

Detected GPU VRAM: 8.0 GB (< 16 GB) -- limiting CUDA build to -j 2

Manual override: If you still encounter OOM errors, reduce parallelism further by editing the build invocation in the relevant script, or close other GPU-heavy applications before building.

Metal fails to initialize on Apple M5 (macOS 26.2–26.4)

Symptom: On M5 Macs, run_llama.sh / start_llama_server.sh fail with Metal errors and produce no output, e.g.:

ggml_metal_library_init_from_source: error compiling source
ggml_metal_device_init: - the tensor API is not supported in this environment - disabling
ggml_metal_synchronize: error: command buffer 0 failed with status 5

Pre-M5 Apple Silicon (M1–M4) is not affected — those devices load the embedded, precompiled Metal library and never compile shaders at runtime.

Cause: On M5 (and A19) devices, ggml compiles its Metal library from source at runtime to enable the tensor API (Neural Accelerators). Some macOS 26 point releases ship stricter MetalPerformancePrimitives headers whose static_assert (bfloat/half type mismatch) breaks that runtime compile. This is an ecosystem-wide issue also seen in ollama and whisper.cpp; see #93.

Workaround: Disable the tensor API so the M5 uses the embedded library like pre-M5 devices — full Metal speed is kept, only the Neural Accelerator prefill boost is lost:

GGML_METAL_TENSOR_DISABLE=1 ./scripts/run_llama.sh -p "Hello"
GGML_METAL_TENSOR_DISABLE=1 ./scripts/start_llama_server.sh

This is much faster than falling back to CPU (BONSAI_NGL=0). If out-of-memory errors persist afterwards on lower-memory machines, additionally pin a smaller context, e.g. -c 16384 (extra args pass through to llama.cpp and override the default).

View on GitHub

Recent activity

commits and pull requests

Code frequency

additions and deletions
+9.4K-9.4KWeek of 2026-03-22: +208 linesWeek of 2026-03-22: -0 linesWeek of 2026-03-29: +9,361 linesWeek of 2026-03-29: -96 linesWeek of 2026-04-05: +225 linesWeek of 2026-04-05: -102 linesWeek of 2026-04-12: +4,030 linesWeek of 2026-04-12: -1,489 linesWeek of 2026-04-19: +782 linesWeek of 2026-04-19: -281 linesWeek of 2026-04-26: +0 linesWeek of 2026-04-26: -0 linesWeek of 2026-05-03: +1 linesWeek of 2026-05-03: -9 linesWeek of 2026-05-10: +0 linesWeek of 2026-05-10: -0 linesWeek of 2026-05-17: +0 linesWeek of 2026-05-17: -0 linesWeek of 2026-05-24: +0 linesWeek of 2026-05-24: -0 linesWeek of 2026-05-31: +0 linesWeek of 2026-05-31: -10 linesWeek of 2026-06-07: +0 linesWeek of 2026-06-07: -0 linesWeek of 2026-06-14: +0 linesWeek of 2026-06-14: -0 linesWeek of 2026-06-21: +0 linesWeek of 2026-06-21: -0 linesWeek of 2026-06-28: +0 linesWeek of 2026-06-28: -0 linesWeek of 2026-07-05: +0 linesWeek of 2026-07-05: -0 linesWeek of 2026-07-12: +5,308 linesWeek of 2026-07-12: -1,643 linesWeek of 2026-07-19: +623 linesWeek of 2026-07-19: -184 linesWeek of 2026-07-26: +402 linesWeek of 2026-07-26: -0 linesWeek of 2026-08-02: +68 linesWeek of 2026-08-02: -16 linesWeek of 2026-08-09: +0 linesWeek of 2026-08-09: -0 linesWeek of 2026-08-16: +0 linesWeek of 2026-08-16: -0 linesWeek of 2026-08-23: +1,030 linesWeek of 2026-08-23: -248 linesWeek of 2026-08-30: +160 linesWeek of 2026-08-30: -0 linesMar 22, 2026Aug 30, 2026
+22.2K lines added, -4.1K removed over the last year.

Commits per week

last 52 weeks
450Week of 2025-09-28: 0 commitsWeek of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 1 commitsWeek of 2026-03-29: 6 commitsWeek of 2026-04-05: 10 commitsWeek of 2026-04-12: 45 commitsWeek of 2026-04-19: 8 commitsWeek of 2026-04-26: 0 commitsWeek of 2026-05-03: 1 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 0 commitsWeek of 2026-05-24: 0 commitsWeek of 2026-05-31: 1 commitsWeek of 2026-06-07: 0 commitsWeek of 2026-06-14: 0 commitsWeek of 2026-06-21: 0 commitsWeek of 2026-06-28: 0 commitsWeek of 2026-07-05: 0 commitsWeek of 2026-07-12: 16 commitsWeek of 2026-07-19: 15 commitsWeek of 2026-07-26: 2 commitsWeek of 2026-08-02: 2 commitsWeek of 2026-08-09: 0 commitsWeek of 2026-08-16: 0 commitsWeek of 2026-08-23: 12 commitsWeek of 2026-08-30: 1 commitsWeek of 2026-09-06: 1 commitsWeek of 2026-09-13: 5 commitsWeek of 2026-09-20: 42 commitsSep 28, 2025Sep 20, 2026
168 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 1 commitsSun 9:00 — 0 commitsSun 10:00 — 0 commitsSun 11:00 — 1 commitsSun 12:00 — 0 commitsSun 13:00 — 0 commitsSun 14:00 — 0 commitsSun 15:00 — 0 commitsSun 16:00 — 0 commitsSun 17:00 — 0 commitsSun 18:00 — 3 commitsSun 19:00 — 1 commitsSun 20:00 — 0 commitsSun 21:00 — 2 commitsSun 22:00 — 0 commitsSun 23:00 — 3 commitsMon 0:00 — 0 commitsMon 1:00 — 0 commitsMon 2:00 — 1 commitsMon 3:00 — 0 commitsMon 4:00 — 1 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 2 commitsMon 8:00 — 1 commitsMon 9:00 — 1 commitsMon 10:00 — 2 commitsMon 11:00 — 0 commitsMon 12:00 — 1 commitsMon 13:00 — 1 commitsMon 14:00 — 1 commitsMon 15:00 — 1 commitsMon 16:00 — 2 commitsMon 17:00 — 2 commitsMon 18:00 — 2 commitsMon 19:00 — 4 commitsMon 20:00 — 8 commitsMon 21:00 — 2 commitsMon 22:00 — 1 commitsMon 23:00 — 6 commitsTue 0:00 — 0 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 2 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 0 commitsTue 9:00 — 1 commitsTue 10:00 — 0 commitsTue 11:00 — 1 commitsTue 12:00 — 1 commitsTue 13:00 — 0 commitsTue 14:00 — 1 commitsTue 15:00 — 2 commitsTue 16:00 — 1 commitsTue 17:00 — 0 commitsTue 18:00 — 0 commitsTue 19:00 — 0 commitsTue 20:00 — 1 commitsTue 21:00 — 0 commitsTue 22:00 — 2 commitsTue 23:00 — 0 commitsWed 0:00 — 1 commitsWed 1:00 — 3 commitsWed 2:00 — 1 commitsWed 3:00 — 1 commitsWed 4:00 — 0 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 0 commitsWed 8:00 — 10 commitsWed 9:00 — 4 commitsWed 10:00 — 2 commitsWed 11:00 — 2 commitsWed 12:00 — 3 commitsWed 13:00 — 1 commitsWed 14:00 — 1 commitsWed 15:00 — 2 commitsWed 16:00 — 0 commitsWed 17:00 — 3 commitsWed 18:00 — 1 commitsWed 19:00 — 0 commitsWed 20:00 — 4 commitsWed 21:00 — 3 commitsWed 22:00 — 2 commitsWed 23:00 — 1 commitsThu 0:00 — 3 commitsThu 1:00 — 2 commitsThu 2:00 — 1 commitsThu 3:00 — 0 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 1 commitsThu 7:00 — 0 commitsThu 8:00 — 0 commitsThu 9:00 — 1 commitsThu 10:00 — 1 commitsThu 11:00 — 3 commitsThu 12:00 — 1 commitsThu 13:00 — 1 commitsThu 14:00 — 2 commitsThu 15:00 — 2 commitsThu 16:00 — 2 commitsThu 17:00 — 4 commitsThu 18:00 — 3 commitsThu 19:00 — 2 commitsThu 20:00 — 0 commitsThu 21:00 — 0 commitsThu 22:00 — 1 commitsThu 23:00 — 1 commitsFri 0:00 — 2 commitsFri 1:00 — 0 commitsFri 2:00 — 0 commitsFri 3:00 — 1 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 1 commitsFri 8:00 — 0 commitsFri 9:00 — 1 commitsFri 10:00 — 1 commitsFri 11:00 — 2 commitsFri 12:00 — 1 commitsFri 13:00 — 4 commitsFri 14:00 — 0 commitsFri 15:00 — 0 commitsFri 16:00 — 0 commitsFri 17:00 — 2 commitsFri 18:00 — 0 commitsFri 19:00 — 1 commitsFri 20:00 — 1 commitsFri 21:00 — 2 commitsFri 22:00 — 0 commitsFri 23:00 — 2 commitsSat 0:00 — 0 commitsSat 1:00 — 0 commitsSat 2:00 — 0 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 1 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 0 commitsSat 11:00 — 1 commitsSat 12:00 — 0 commitsSat 13:00 — 2 commitsSat 14:00 — 0 commitsSat 15:00 — 1 commitsSat 16:00 — 0 commitsSat 17:00 — 3 commitsSat 18:00 — 1 commitsSat 19:00 — 1 commitsSat 20:00 — 0 commitsSat 21:00 — 0 commitsSat 22:00 — 0 commitsSat 23:00 — 0 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
Sep 30, 2026weekly#13+243
Sep 29, 2026weekly#13+243
Aug 15, 2026monthly#15+1,323
Aug 14, 2026monthly#15+1,323
Jul 15, 2026daily#15+7
  • obra/superpowers

    An agentic skills framework & software development methodology that works.

    295.2K stars · Shell

  • mattpocock/skills

    Skills for Real Engineers. Straight from my .agents directory.

    276K stars · Shell

  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    272.8K stars · JavaScript

  • NousResearch/hermes-agent

    The agent that grows with you

    251.2K stars · Python

  • firecrawl/firecrawl

    Supercharge your AI agents with data from the web and beyond. Building the library for superintelligence. 🔥

    188.6K stars · TypeScript

  • Significant-Gravitas/AutoGPT

    AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.

    187.7K stars · Python