JustVugg/colibriPublic

Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

AI summary: A pure C inference engine for running massive MoE models locally on consumer hardware using memory tiering.

Stars
39.3K
+644 today
Forks
4.3K
Watchers
321
Open issues
54
Open PRs
47
Contributors
~193
Commits
2.8K
Branches
11

CApache-2.0Created Jul 1, 2026Last push 1d agoLatest release v1.12.1+1.7K stars this week+12.6K this month

Quick answers

What is colibri?
A pure C inference engine for running massive MoE models locally on consumer hardware using memory tiering.
What does colibri do?
Colibrì is a highly optimized, zero-dependency C inference engine designed to execute massive language models, specifically Mixture-of-Experts (MoE) architectures, on consumer-grade hardware. It achieves this by employing an advanced AI memory multi-tiering system that treats disk storage, system RAM, and VRAM as a single, cohesive hierarchy. This approach allows users to run models ranging from 7B to 2.8T parameters by streaming experts directly from disk as needed. It currently supports six major model families, providing a unified command-line and web interface across all of them without the bloat of standard Python ML frameworks.
Who is colibri for?
Systems engineers, AI researchers, and power users who want to run extremely large language models locally.
How do I get started with colibri?
make -C c glm && COLI_MODEL=/path/to/model ./coli chat
How popular is colibri on GitHub?
JustVugg/colibri has 39,323 stars and 4,306 forks on GitHub, and gained 1,717 stars in the last 7 days.
What license does colibri use?
JustVugg/colibri is released under the Apache-2.0 license.

Star history

since Jul 28, 2026
010K20K30KJul 2026Aug 2026Sep 2026Oct 2026
39.3K stars as of Oct 3, 2026. Measured daily since Jul 28, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks
OctNovDecJanFebMarAprMayJunJulAugSepMonWedFri2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 0 commits2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 1 commit2026-07-02: 0 commits2026-07-03: 0 commits2026-07-04: 0 commits2026-07-05: 6 commits2026-07-06: 16 commits2026-07-07: 4 commits2026-07-08: 2 commits2026-07-09: 6 commits2026-07-10: 16 commits2026-07-11: 7 commits2026-07-12: 24 commits2026-07-13: 16 commits2026-07-14: 60 commits2026-07-15: 58 commits2026-07-16: 52 commits2026-07-17: 32 commits2026-07-18: 32 commits2026-07-19: 63 commits2026-07-20: 35 commits2026-07-21: 45 commits2026-07-22: 47 commits2026-07-23: 31 commits2026-07-24: 20 commits2026-07-25: 20 commits2026-07-26: 17 commits2026-07-27: 18 commits2026-07-28: 27 commits2026-07-29: 24 commits2026-07-30: 19 commits2026-07-31: 13 commits2026-08-01: 38 commits2026-08-02: 41 commits2026-08-03: 27 commits2026-08-04: 12 commits2026-08-05: 15 commits2026-08-06: 5 commits2026-08-07: 14 commits2026-08-08: 9 commits2026-08-09: 8 commits2026-08-10: 9 commits2026-08-11: 34 commits2026-08-12: 20 commits2026-08-13: 16 commits2026-08-14: 17 commits2026-08-15: 31 commits2026-08-16: 24 commits2026-08-17: 12 commits2026-08-18: 52 commits2026-08-19: 10 commits2026-08-20: 26 commits2026-08-21: 12 commits2026-08-22: 10 commits2026-08-23: 29 commits2026-08-24: 27 commits2026-08-25: 13 commits2026-08-26: 14 commits2026-08-27: 23 commits2026-08-28: 23 commits2026-08-29: 3 commits2026-08-30: 19 commits2026-08-31: 10 commits2026-09-01: 2 commits2026-09-02: 9 commits2026-09-03: 6 commits2026-09-04: 17 commits2026-09-05: 10 commits2026-09-06: 34 commits2026-09-07: 25 commits2026-09-08: 3 commits2026-09-09: 4 commits2026-09-10: 24 commits2026-09-11: 22 commits2026-09-12: 23 commits2026-09-13: 22 commits2026-09-14: 25 commits2026-09-15: 39 commits2026-09-16: 9 commits2026-09-17: 34 commits2026-09-18: 8 commits2026-09-19: 5 commits2026-09-20: 21 commits2026-09-21: 30 commits2026-09-22: 72 commits2026-09-23: 25 commits2026-09-24: 12 commits2026-09-25: 0 commits2026-09-26: 0 commits2026-09-27: 0 commits2026-09-28: 0 commits2026-09-29: 0 commits2026-09-30: 0 commits2026-10-01: 0 commits2026-10-02: 0 commits2026-10-03: 0 commits
1,795 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    39,323 stars

  • Breakout launch

    39,323 stars in 95 days

  • High momentum

    +644 stars today

  • Very active

    1,795 commits in 52 weeks

  • Community-driven

    ~193 contributors

  • Permissive license

    Apache-2.0

  • Repeat trending

    22 trending appearances

What colibri does

Colibrì is a highly optimized, zero-dependency C inference engine designed to execute massive language models, specifically Mixture-of-Experts (MoE) architectures, on consumer-grade hardware. It achieves this by employing an advanced AI memory multi-tiering system that treats disk storage, system RAM, and VRAM as a single, cohesive hierarchy. This approach allows users to run models ranging from 7B to 2.8T parameters by streaming experts directly from disk as needed. It currently supports six major model families, providing a unified command-line and web interface across all of them without the bloat of standard Python ML frameworks.

Systems engineers, AI researchers, and power users who want to run extremely large language models locally.

  • AI memory multi-tiering: Treats SSD storage, RAM, and VRAM as a unified hierarchy to run models larger than available memory.
  • Pure C architecture: Built entirely in C with zero external dependencies, ensuring maximum performance and minimal overhead.
  • MoE streaming: Dynamically loads only the necessary expert weights from disk during inference to conserve memory.
  • Frontier model support: Capable of running massive models like Kimi K3 (2.8T) and DeepSeek V4 Flash (284B) natively.
  • Unified interface: Provides standard chat, serve, and web commands across all supported model architectures via a single binary.

Where teams use it

Local massive model inference

Researchers run 700B+ parameter models on standard workstations without renting expensive cloud GPU clusters.

Inference optimization research

Systems engineers use the pure C codebase as a transparent platform to study and optimize memory tiering algorithms.

Edge deployment

Developers deploy highly capable MoE models on resource-constrained edge devices by leveraging the zero-dependency engine.

Private local chat

Privacy-conscious users run massive, state-of-the-art chat models entirely locally to ensure data security.

Getting started: make -C c glm && COLI_MODEL=/path/to/model ./coli chat

README

main branch

colibrì — tiny engine, immense model

Website Latest release

Website · Discord · English · 简体中文 · 繁體中文 · Italiano · 日本語

Tiny engine, immense model. Run frontier MoE models — 744B to 2.8T parameters — on consumer and heterogeneous hardware, in pure C with zero engine dependencies, by treating storage, RAM, and VRAM as a single inference hierarchy (AI memory multitiering).

Nine families run today: GLM-5.2/5.3 (744B), GLM-5.3-Flash (321B, with vision), Inkling (975B), Kimi K3 (2.8T), DeepSeek V4 Flash (284B), DeepSeek V4.1 Flash (552B, with vision), Qwen3.8-Flash-Next (125B + 51B n-gram), Qwen3.6 (35B-A3B) and OLMoE (7B) — one C file each, the same coli chat / coli serve / coli web front end. Full roster ↓

Colibrì is an inference engine you can run today, and an open research platform. Its primary goal is to pursue inference-side performance across the entire software/hardware boundary — model formats, memory hierarchy, storage I/O, placement, scheduling, kernels, speculation, and CPU/GPU overlap — so large models depend less on scarce hardware and cost less to run.

Colibrì treats VRAM, RAM, and storage as a single multitier hierarchy, and it is deliberately a place to test aggressive systems ideas — so there is no SLA on speed, and a hard guarantee on semantics: experiments must earn their place through reproducible end-to-end measurements, and the default policy never silently changes model precision or router semantics. Insufficient fast memory may reduce speed; it must not quietly redefine the model.

$ ./coli chat
  🐦 colibri v1.12.1 — GLM-5.2 · 744B MoE · int4 · streaming CPU
  ✓ ready in 32s · resident 9.9 GB
  › ciao!
  ◆ Ciao! 😊 Come posso aiutarti oggi?

See it running

colibrì web dashboard — live metrics, hardware panel, expert tiers

The web dashboard (./coli web), redesigned in 1.12.0: a workspace with a dock for the chat, Brio mode, the Brain page and Profiling, in a light or a dark theme. Here Qwen3.6 answering on a CPU box, experts streamed from disk.

the Brio page: a document read once, a probability for every allowed answer, and an entropy

Brio mode: the same model, told to stop writing. Give it a document and the only answers it may pick; it reads the probability of each one, generates nothing, and reports an entropy that says when it is not sure. Here: request changes at 99.9%, entropy 0.005, 4 tokens read, 0 generated.

the Brain page: the measured expert atlas of GLM-5.2 drawn as a cortex, ten regions to enter

The Brain page, Explore: the measured expert atlas of GLM-5.2 drawn as a cortex. 13,260 characterised experts in ten regions (Python, SQL, mathematics, poetry, law, Chinese…); position is measured routing affinity, not a learned embedding. Choose a region to enter it. Live routing switches to the model actually running: one cell per expert, colour is the storage tier, and every expert routed in a turn flashes white.

inside the Python region: 1,142 experts, one of them selected with its measured affinities

Inside the Python region: 1,142 experts as a constellation, each labelled by layer and index. The panel shows one of them, layer 17 expert 178: a generalist with entropy 3.13, whose measured affinity is 20.2% Python, 14.6% JSON, 14.2% conversation, 13.3% SQL.

the Profiling page: where the engine spends each turn

The Profiling page: where the engine spends each turn, by phase, with the last 30 turns as a trend. Here Qwen3.6 on a CPU box: 19.0 s of wall time for 36 prompt and 55 generated tokens, 2.9 tok/s, 11.4 s of disk service overlapped with compute.

The research mission

With Colibrì, private frontier model access is not limited by availability of hyperscaler-class hardware.

With its multitiering features Colibrì removes proprietary hardware dependencies aggressively optimizing functional inference engine pipelines.

Our operational mission includes changing how weights are represented and moved, deciding what lives in VRAM, RAM, or storage, overlapping heterogeneous compute, reducing launch and synchronization overhead, exploiting sparsity and reuse, and testing new decoding algorithms. Nothing is protected merely because it is conventional; nothing is adopted merely because a microbenchmark looks fast. The deciding result is end-to-end inference on real machines, with correctness and quality measured alongside throughput, latency, memory, and cost.

The practical consequence is accessibility: run a 744B-parameter model on hardware you already own, watch every expert fire in real time, and change the code that does it. Not renting intelligence behind an API — holding it: probing it, measuring it, improving it. The engine is deliberately small enough that the next useful optimization can come from anyone willing to measure it.

Core techniques and measured findings

  • One hierarchy, not limited by tier capacity. VRAM, RAM, and NVMe are placement tiers for the same weights; limited fast memory changes speed, not model semantics.
  • A JIT for weights. Measured routing heat drives a per-layer LRU, a learned pinned hot-store, and one-layer-ahead prefetch instead of loading every expert. It wins on repeatable workloads; history can overfit, and lookahead can lose on some hosts, so both remain measurable policies rather than promises.
  • I/O is part of the engine. Batched expert unions, overlapped reads and compute, O_DIRECT, and weighted dual-SSD striping attack the streaming path rather than pretending storage latency is free. O_DIRECT is drive-dependent, and dual-SSD still needs broader end-to-end community A/Bs.
  • Heterogeneous execution. CPU, CUDA, Metal, NUMA memory, and partial or full expert residency share one runtime and can be combined according to the machine; the profitable combination depends on compute, bandwidth, residency, and workload.
  • Compressed state without a different model. Token-exact forward validation, 57× smaller MLA KV state, persistent warm conversations, and faithful DSA keep optimization tied to correctness. These are memory, latency, and correctness properties — not a blanket throughput claim.
  • Speculation that must earn its keep. Native MTP and grammar-forced drafts are measured end to end and can be disabled when acceptance does not repay verification.

Open hypotheses, experiments, and how to help

Colibrì treats an optimization as a hypothesis until a controlled end-to-end A/B shows otherwise. These are the main questions now:

hypothesis evidence so far experiment still needed
Routing history can place experts better than plain LRU learned pins improve repeated workloads, but can overfit a prompt held-out, cross-session A/Bs across coding, chat, multilingual, and long-context workloads
Multiple SSDs can turn independent bandwidth into decode speed two independent NVMe drives measured +37.5% decode; a slower third drive was neutral after weighted striping (measurements) reproduce across drive speeds, controller layouts, and cache states
A hardware-aware planner can approach each machine's best configuration automatically RAM/VRAM budgets and several backends are detected today compare the generated plan with a controlled parameter sweep across laptops, workstations, NUMA hosts, and multi-GPU systems
Lossless or quality-bounded representations can reduce weight movement enough to matter format and quantization ablations exist, with correctness/quality gates reproduce quality, bytes moved, latency, and cost per useful token together — not compression ratio alone
Routing-aware speculation can pay before near-full residency MTP and grammar drafts work, but MTP has also measured a 32% loss around 85% expert hit map the break-even surface across acceptance, expert hit rate, batch union, and draft depth
CPU/GPU overlap can hide transfer and synchronization rather than merely move the bottleneck CUDA and Metal wins exist, but fast CPUs and low residency can erase them per-stage profiles and one-variable A/Bs across PCIe, unified-memory, and full-resident machines

Want to help? Pick one row and publish the negative results too. Record the hardware, commit, model/container, exact command, prompt, cache state, throughput, TTFT, expert hit rate, bytes read, and quality check; change one variable, repeat the run, and attach raw logs. Start with CONTRIBUTING.md, compare against the benchmark protocol, then open an experiment issue. A well-controlled failure is more valuable here than an unexplained fast number.

The idea

A 744B Mixture-of-Experts model activates only ~40B parameters per token — and only ~11 GB of those change from token to token (the routed experts):

only ~5.4% of parameters are active per token

So the model doesn't need to fit in fast memory — it needs to be placed:

  • the dense part (attention, shared experts, embeddings — ~17B params) stays resident in RAM at int4 (~9.9 GB);
  • the 19,456 routed experts (75 MoE layers × 256 + the MTP head, ~19 MB each at int4) live on disk (~370 GB) and are streamed on demand, with a per-layer LRU cache, a learned pinned hot-store, and an optional VRAM tier.

Think of the core algorithm as a JIT, but for weights. A compiler JIT never compiles the whole program — it watches what actually runs and compiles the hot paths, just in time. colibrì makes the same bet about a 744B parameter space: parameters are not resident state to be held, they are data to be staged across a heterogeneous storage hierarchy (VRAM / RAM / NVMe), exactly when the router proves they are needed. Measured routing heat decides which experts earn which tier, the router runs a layer ahead so prefetch hides the staging latency, and — like a JIT — the engine learns your workload: the more you run, the hotter the right experts get. It works because routing has measurable structure (see the expert atlas) — and structure is cacheable.

The engine is a single C file (c/colibri.c) plus small headers. No BLAS, no Python at runtime, no GPU required.

Local cluster mode

The coordinator keeps token generation, routing, and KV state local while disk-backed expert workers execute routed FFNs on other Macs. A layer's routed batch-union is sent as one persistent TCP request, so a token does not incur one round trip per expert.

Start the optional registration service:

./coli cluster coordinator --host 0.0.0.0 --port 8765

On each worker, with the same converted model available locally:

./coli cluster worker --model /nvme/glm52_i4 --port 9100 \
  --coordinator http://COORDINATOR:8765 --advertise-host WORKER_IP

Run the coordinator with discovery, or provide --cluster-workers HOST:PORT,... for a static setup:

./coli serve --model /nvme/glm52_i4 \
  --cluster-coordinator http://127.0.0.1:8765

The transport is disabled unless workers are configured, so the existing single-machine path remains unchanged. Dense-layer sharding and browser/WebGPU workers are separate follow-up seams.

How it works

The per-token path

route → union → place → overlap → learn

Every layer of every token walks the same five steps. The design goal is that placement only ever decides speed — the router's decisions and the weights' precision are the same whether an expert answered from VRAM or from disk.

One memory hierarchy instead of one memory requirement

VRAM / RAM / NVMe three-tier expert residency

Multiple SSDs: stream model copies from more than one drive

When decode is disk-bound, a second SSD can help: put a copy of the model on it and let the engine read from both drives. For GLM-5.2, from c/ in a source checkout (or from an unpacked release):

COLI_MODEL_MIRROR=/second/glm52_i4 python3 ./coli chat --model /fast/glm52_i4

The engine measures the drives at startup to weight the read split. Buffered reads use deterministic expert routing; eligible direct reads can stripe one expert across replicas. Independent drives provide bandwidth headroom, not a guaranteed token-rate multiplier: shared controllers, cache hits, and compute can limit the gain. See the multi-disk guide for Bash and PowerShell examples, measured gains and limits, and a single-drive comparison. Details worth knowing:

  • the mirror is validated at startup (per-file size + safetensors header must be byte-identical to the primary); divergent or missing files stay on the primary, so a partial mirror is fine — a smaller second SSD can serve the shards it holds;
  • the mirror is never written: .coli_usage, .coli_kv and all sidecars stay on the primary;
  • a read error on the mirror falls back to the primary (one warning, no crash), so unplugging the second drive mid-run degrades instead of killing the server;
  • routing never changes tokens — both copies are byte-identical; enable PROF=1 for the MIRROR: profile counters showing GB served per drive.

The same engine spans the whole range: on a 25 GB laptop everything streams from disk (slow but correct); on a large host the entire expert set becomes resident (CUDA_EXPERT_GB=auto PIN_GB=all) and disk drops out of the decode path entirely. Between the tiers sits a learning cache: the engine records which experts your workload routes to (.coli_usage, updated every turn) and pins the hottest ones automatically — colibrì literally gets faster the more you use it. On multi-socket hosts, COLI_NUMA=1 interleaves the resident weights across memory controllers (#82).

For a second drive that cannot hold the whole model, Colibri can rank a partial mirror from the expert history it already learns. Run a few representative prompts first so .coli_usage reflects the workload, then plan, stage, and verify the mirror:

./c/coli mirror plan  --model /fast/glm52_i4 --mirror /second/glm52_i4 \
  --budget-gib 200 --reserve-gib 20
./c/coli mirror stage --model /fast/glm52_i4 --mirror /second/glm52_i4 \
  --budget-gib 200 --reserve-gib 20
./c/coli mirror verify --model /fast/glm52_i4 --mirror /second/glm52_i4

The planner reads safetensors headers directly, follows split-model directories from COLI_MODEL_DIRS, and prioritizes shards that can serve the hottest routed experts. Staging never changes the primary model: it copies through temporary files, preserves the requested free-space reserve, verifies every shard with SHA-256, never deletes an existing mirror shard, and atomically publishes a receipt only after the selected mirror is ready.

Never wait for the disk twice

Misses are expensive, so the engine spends most of its cleverness avoiding and overlapping them: each expert's three matrices are stored adjacent and read in one pread; a bounded async I/O pool (PIPE=1, default) loads missing experts while resident ones compute; batched positions read each unique expert once (batch-union); and a router-lookahead thread (PILOT=1) prefetches the next layer's experts — routing is measurably 71.6% predictable one layer ahead. On GPUs, the resident pipeline (COLI_CUDA_PIPE=2) keeps the residual stream on-device across layers so the CPU expert loop runs uninterrupted; on Apple Silicon an experimental Metal backend does the batched expert math on the unified-memory GPU; and a Vulkan backend brings the expert tier, dense projections, and the MLA attention core to any GPU with a Vulkan 1.2 driver — including AMD cards via Mesa/RADV (the only backend for cards the vendor stacks no longer support, like the RX 580, and competitive with ROCm on RDNA4 — see the benchmarking notes).

On real NVMe, measure DIRECT=1. O_DIRECT bypasses the page cache and is often a large win on drives with DRAM cache and bandwidth headroom (+34% decode measured with PIPE=1 on a Blackwell/Windows box; 4.25→9.69 GB/s in iobench on a GB10) — but it is drive-dependent: QLC/DRAM-less or virtualised disks can be neutral to negative. Try it first; keep what your hardware rewards.

Faithful model, compressed state

The forward pass is validated against a transformers oracle (teacher-forcing typically 30-32/32; two tiny-oracle positions are floating-point near-ties and toolchain-dependent). MLA attention stores a compressed KV state — 576 floats/token instead of 32,768 (57× smaller) — and persists it across restarts (.coli_kv): conversations reopen warm with zero re-prefill, byte-identical to an uninterrupted session. DSA sparse attention (GLM-5.2's lightning indexer) is implemented faithfully and validated by forcing full-key selection to reproduce dense attention exactly.

Speculative decoding, honestly

GLM-5.2's native MTP head drafts tokens that the main model verifies in one batched forward — 2.2–2.8 tokens/forward when it pays. Two hard-won rules ship as defaults: the MTP head must be int8 (int4 heads collapse to 0–4% acceptance, #8), and draft and verify must compute the same function — SPEC_PIN=1 pins both to one kernel family (#163 is the full forensic story). Grammar-forced drafts (GRAMMAR=file.gbnf) add ~free acceptance on constrained JSON output. Whether speculation is a net win depends on your cache temperature — measure, and use DRAFT=0 when it doesn't pay.

Verify batches can also opt into an exact attention core with COLI_EXACT_VERIFY=1 (#689): the CPU MLA-absorb score and context dots accumulate integer products and round once, so a near-tie in a verify row resolves the same way on every host, at roughly 0.6x tok/s on a tiny oracle (the dot itself is ~5–7x the float loop). Two limits to know: with a quantised KV cache (tq1, TQ or int8 KV) the context dot keeps the float path, so exactness there is not provided; and a real near-tie flip has only been argued, not yet caught on GLM-5.2 at n=64.

What it achieves

measured decode speed by hardware class

Same engine, same int4 container — the hardware only changes where the experts live. Highlights from the full benchmark tables:

  • 6× RTX 5090, full residency: 5.8–6.8 tok/s decode, TTFT ~13 s (experiment log);
  • 128 GB CPU-only desktop: ~1.8 tok/s warm (#200);
  • single RTX 5070 Ti laptop-class box: 1.07 tok/s via the GPU-resident pipeline (#273);
  • 25 GB dev box: 0.05–0.1 tok/s cold — the proven floor where this project started, and still the honest baseline.

Quality is measured, not assumed: the int4 container's quantization cost and the scale-granularity/rotation ablations live in docs/benchmarks.md and #108/#81.

Get started

You need two things: the program (a few hundred KB) and the model (372 GB). Step-by-step for every platform in the Quick Start guide.

1. Get colibri

Download a prebuilt release — Linux, macOS and Windows, no compiler needed. Take the archive for your platform from Releases and unpack it:

mkdir colibri && tar xzf colibri-v1.8.0-linux-x86_64.tar.gz -C colibri && cd colibri
python3 coli info                         # engine ready ✓

Inside you get the engine (colibri, colibri.exe on Windows), the coli launcher and its Python helpers. Nothing to rename or configure — coli finds the engine next to itself. You only need Python 3 installed: the launcher and the API gateway are Python scripts, while the engine itself is pure C with zero dependencies.

Or build from source — needs gcc (or clang) with OpenMP:

git clone https://github.com/JustVugg/colibri && cd colibri/c
./setup.sh                                # checks gcc/OpenMP, builds, self-tests

Want coli on your PATH? From a checkout, pip install -e . registers it (the engine still lives in c/ — an editable install from the clone, not a wheel).

2. Get the model

A pre-converted GLM-5.2 int4 container is on Hugging Face — use the group-scaled (gs64) build with the int8 MTP head. It is about 372 GB, so put it on a disk with the room, ideally a fast one:

https://huggingface.co/mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp

GLM-5.3 is the same family and loads with the same engine. It has its own container, also group-scaled (gs64), about 419 GB. It ships without the MTP head, so speculative decoding stays off:

https://huggingface.co/Justvugg/GLM-5.3-colibri-int4-g64

⚠️ Use the gs64 container above, not the older per-row int4 mirrors (mateogrgic/…, jlnsrk/…): those measure ~9pp worse on quality and are the root cause of the original think-mode loops and never-terminating generations in #455. The gs64 container fixed those controlled per-row A/Bs, but it is not a general repetition or EOS-starvation guard. The MTP head must also be int8, not int4 (int4 → 0% draft acceptance, #8): ls -l <model>/out-mtp-* — int8 (correct) is 3527131672 / 5366238584 / 1065950496 as three files, or a single out-mtp-00000.safetensors of 9959321520 bytes (the current upload of the recommended container ships it as one file: same int8 tensors, 777 of them at one byte per element).

Or convert from the FP8 source yourself — one resumable command that never needs the full 756 GB on disk at once:

./coli convert --model /nvme/glm52_i4     # download+convert shard by shard (python, one-time)
Other supported models

GLM-5.2 is the reference model, but the same streaming approach runs six more families. Each is a sibling engine — one C file, its own architecture, the same coli chat / coli serve / coli web front end (the launcher picks the binary from the model's config.json):

What each one needs. These differ a lot, and reading two of them together has confused people into thinking the requirements contradict each other (#191). They do not — they are different models. None of them needs a GPU.

Model Disk for the weights RAM GPU
OLMoE ~7 GB (int8 container) 8 GB not needed
GLM-5.2/5.3 ~372 GB (5.2) / ~419 GB (5.3) 16 GB min, 24 GB comfortable not needed
GLM-5.3-Flash ~195 GB converted 25 GB (12 GB weights at int4 + expert cache) not needed
Inkling ~469 GB 25 GB with the int4 dense container, ~120 GB without not needed
Kimi K3 ~1.6 TB 32 GB+ not needed
DeepSeek V4 Flash ~167 GB (REAP 150B: ~85 GB) 16 GB min, 32 GB comfortable optional; any NVIDIA card from the GTX 10 series up (Pascal/Turing via CUDA_ARCH=portable-pre-ampere NO_TC=1, best on RTX 50) makes prefill 5-10x and decode ~2.5x faster
Qwen3.8-Flash-Next ~185.5 GB (official FP8 checkpoint) 16 GB min, 24 GB comfortable at the default context optional; the CUDA VRAM expert tier with dense trunk quantized to int8 in VRAM
Qwen3.6-35B-A3B ~20 GB (int4-gs64 container) 24 GB (needs full RAM residency) optional; the CUDA VRAM expert tier measured 1.44 -> 10.05 tok/s (7.0x) on two 8 GB cards, output bit-identical to CPU

A GPU only ever makes it faster. Speed is set by your disk, because the experts are streamed from it — expect a fraction of a token per second on a slow drive and a few per second on a fast one with the cache warm.

Family Total / active Weights Build Docs
GLM-5.2/5.3 744B / 40B mastouri/…-int4-g64-with-int8-mtp (372 GB) or Justvugg/GLM-5.3-colibri-int4-g64 (419 GB) make -C c glm this page
Inkling (Thinking Machines) 975B / 41B nbeerbower/Inkling-colibri-int4 (469 GB) make -C c inkling inkling.md
GLM-5.3-Flash (Z.ai) 321B / 40B zai-org/GLM-5.3-Flash — converted to int4-gs64 routed experts, dense stays BF16 and the precision is a load-time choice; vision included make -C c glm53 glm53-flash.md
Kimi K3 (Moonshot) 2.8T / 104B moonshotai/Kimi-K3 — original checkpoint, routed experts stay native MXFP4 make -C c kimi_k3 kimi_k3.md
DeepSeek V4 Flash 284B / 13B official sharded checkpoint — routed experts stay native fp4, dense stays fp8-e4m3; the REAP-pruned 150B (puwaer/DeepSeek-V4-Flash-0731-reap-150b, 85 GB, 132 of 256 experts) loads with the same engine and no conversion make -C c deepseek-v4 deepseek-v4.md
DeepSeek V4.1 Flash 552B / 16B official checkpoint, no conversion: experts are already fp4, dense is fp8-e4m3. 203 GB of it is an n-gram memory read from disk a few hundred bytes at a time, and the routed experts cost 4.5 GB per token against GLM-5.2's 12.7. Vision, tool calling and the DSpark draft head are all on make -C c deepseek_v41 deepseek-v41.md
Qwen3.8-Flash-Next (Alibaba) 125B + 51B n-gram / 6B Qwen/Qwen3.8-Flash-Next-FP8 — original checkpoint; PLE stays pageable and experts stay native block-FP8 make -C c qwen38 (CUDA=1 for the VRAM expert tier) qwen38.md
Qwen3.6 (Alibaba) 35B / 3B Kreuzzelg/qwen36-35b-a3b-colibri-i4-gs64 (~20 GB, recommended) — hybrid Gated Attention + Gated DeltaNet make -C c qwen36 (CUDA=1 for the VRAM expert tier) qwen36.md
OLMoE (AI2) 7B / 1B converted with c/tools/convert_olmoe_merged.py — int8 container, ~7 GB make -C c olmoe —

Qwen3.6 ships three pre-converted containers: int4-gs64 (recommended — measured cosine to the int8 anchor 0.98777 → 0.99313 and KL 0.109 → 0.080 against per-row, i.e. ~44% less quantization error), int4 per-row as the A/B baseline, and KAT-Coder v2.5, which the same engine runs unchanged — any architecture-identical checkpoint works without a code path of its own. With CUDA=1 the VRAM expert tier measured 1.44 → 10.05 tok/s (7.0×) on two 8 GB cards, output bit-identical to the CPU path.

Kimi K3 needs no conversion: its QAT-trained MXFP4 experts are streamed straight from the original Hugging Face shards, and the bf16 dense set is quantized at load time. Long agent sessions can opt into recurrent-state checkpoints (COLI_K3_CKPT=N slots in RAM, or parked on disk with COLI_K3_CKPT_DIR): an edited or follow-up prompt restores the deepest surviving checkpoint and re-prefills only the tail, instead of replaying the whole conversation through the SSM layers. On Vulkan hosts K3_VK_UP=auto sizes the expert tier upload from measured bandwidth. The engine's KDA and MLA paths are validated token-exact in CI against the vendor implementation.

Inkling ships int4 experts but bf16 dense weights (49.4 GB resident); on a host that cannot hold those, inkling.md has a one-pass tool that brings the dense set to 15.3 GB and lets the 975B run on a 25 GB box — with the honest trade-off written down.

3. Run it

COLI_MODEL=/nvme/glm52_i4 ./coli chat     # RAM budget, cache and MTP auto-detected
COLI_MODEL=/nvme/glm52_i4 ./coli plan     # inspect the planned VRAM/RAM/disk placement
COLI_MODEL=/nvme/glm52_i4 ./coli doctor   # read-only readiness check
COLI_MODEL=/nvme/glm52_i4 ./coli doctor --deep  # strict tensors/shards/index/mirror preflight
COLI_MODEL=/nvme/glm52_i4 ./coli tune     # measure and save this machine's fastest safe execution profile
./coli web  --model /nvme/glm52_i4        # API + dashboard, and opens a browser
./coli serve --model /nvme/glm52_i4       # API + dashboard, no browser (headless)
Brio mode: ask a closed question

Most of what people ask a model for is a choice, not a paragraph: which queue, which verdict, which of the four values a field may take. Brio mode hands the engine the options and reads the probability of each one instead of generating: completion_tokens is 0, no answer can fall outside your list, and every answer comes with an entropy, so "the model is not sure" is a number you can put a threshold on. It runs on all nine families, on the same server, and it is opt-in per request: chat is byte-identical for everyone who does not ask for it.

# in the TUI: the same model, told to stop writing
./coli chat --model /nvme/qwen36_i4_gs64
> /brio merge | request changes | close
> 340 lines, 8 files, no tests. CI is green but nothing covers that path.

# from anywhere: one JSON request on the running server
curl -s http://127.0.0.1:8000/v1/brio -H 'Content-Type: application/json' -d '{
  "model": "qwen36",
  "state": "340 lines, 8 files, no tests. CI is green but nothing covers that path.",
  "question": "What should the reviewer do?",
  "options": ["merge", "request changes", "close"]}'

questions asks many things about one document read once, and schema fills a JSON object one field at a time, valid by construction. Measured on Qwen3.6 against generating the same answer on the same CPU box: 2.4x on a four-field schema, 5.7x on four questions about one document. The whole mode, the request and reply shapes, and where it does not help: docs/brio.md. The dashboard has a Brio page as well.

On Windows a release archive ships coli.cmd: double-click it for the quick start, or run coli.cmd chat --model D:\glm52_i4 from cmd or PowerShell. From a source checkout the same commands work with python coli chat --model D:\glm52_i4. The .exe files are the engines, not the launcher: started on their own they have no model to load and exit immediately. The engine at runtime is pure C — python is only used by the one-time converter and the optional API gateway.

The same commands run any of the models

coli reads the model's config.json, picks the matching engine binary, and renders that family's chat template — so nothing about the command line changes between models. Build the engine you want once, then just point COLI_MODEL at the right directory:

make -C c glm                                     # GLM-5.2
make -C c inkling                                 # Inkling
make -C c kimi_k3                                 # Kimi K3

COLI_MODEL=/nvme/glm52_i4      ./coli chat        # TUI
COLI_MODEL=/nvme/inkling_i4    ./coli chat
COLI_MODEL=/nvme/kimi_k3       ./coli chat

./coli web --model /nvme/inkling_i4               # API + dashboard, opens a browser
./coli web --model /nvme/kimi_k3
./coli serve --model /nvme/inkling_i4             # API + dashboard, no browser

For the non-GLM engines coli chat starts the gateway locally and attaches the TUI to it, so the TUI, the API and the dashboard all go through the same arch-aware chat template — you never have to pass the template yourself.

Two things that differ per model, both documented in the per-model page:

  • Inkling on a RAM-tight host needs the int4 dense container and a small expert cache: ./coli chat --model /nvme/inkling_i4 --cap 2 (see inkling.md — the default --cap 8 wants ~14 GB of cache on top of the resident set).
  • Kimi K3 streams its MXFP4 experts from the original checkpoint, so there is nothing to convert — but the snapshot is ~1.6 TB (see kimi_k3.md).

4. Go deeper

topic doc
Benchmarks, community datapoints, quality measurements docs/benchmarks.md
Reproducible benchmark protocol and minimum report docs/benchmarking.md
Tuning knobs, policies, the learning cache, prefetch docs/tuning.md
Windows 11 native build (+ CUDA DLL) docs/windows.md
CUDA backend, VRAM expert tier, full residency docs/cuda.md
Vulkan backend (any GPU: AMD via RADV, incl. cards ROCm dropped) docs/vulkan.md
Apple Silicon Metal backend docs/metal.md
OpenAI-compatible API, KV slots, web dashboard docs/api.md
Brio mode: score a closed set of options instead of generating docs/brio.md
Experimental layer-segment embedding ABI docs/segment-runtime.md
Experimental tokenizer/embedding/head Edge ABI docs/edge-runtime.md
Grammar-forced drafts (structured output) docs/grammar-draft.md
Environment variable inventory docs/ENVIRONMENT.md

DeepSeek V4

DeepSeek V4 Flash streams the official checkpoint with no conversion: routed experts stay native fp4, the dense set stays fp8-e4m3 with UE8M0 block scales. MLA + DSA sparse attention, 43 layers, 256 routed experts plus one shared, top-6. Supported on x86-64/aarch64 Linux and Windows/MSYS2 (CPU), with an optional CUDA tier (Windows runtime DLL; Linux CUDA=1 direct link, verified under WSL2) that keeps every stage CPU-canonical and falls back per stage.

cd c
make deepseek-v4
python ./coli chat --model /path/to/DeepSeek-V4-Flash --ram 32
# also: coli run / coli serve / coli web
# Windows CUDA tier: make cuda-dsv4-dll CUDA_ARCH=portable  (+ make cuda-dsv4-dg-dll on RTX 50)

Two opt-in GPU levers are new and looking for community numbers, both default off and byte-identical when unset: DSV4_HYBRID=1 splits VRAM-tier misses between the GPU fill branch and the CPU branch using bandwidths measured at runtime, and COLI_CUDA_MOE_DOUBLE=1 (on top of COLI_CUDA_MOE_BATCH=1) prefetches the next layer's full expert set into a second VRAM bank while the current layer computes, falling back to the single bank when VRAM is short. The CUDA tier also runs on Pascal and Turing cards now (GTX 10 / RTX 20 series): build with CUDA_ARCH=portable-pre-ampere NO_TC=1.

Greedy decode and one KV slot. Tool calling is wired through the HTTP gateway with V4's native prompt and DSML call blocks; grammar is not supported. See the per-engine API matrix. Prefix checkpoints (in memory + on disk) make agent sessions and follow-up turns start in seconds after the first prefill of a system prompt. Measured on an RTX 5080 + 2 NVMe: 3324-token prefill 90 s, 8.3k-token first turn ~4 min once, later sessions/turns 6-9 s, decode ~1.6 tok/s at 3k context — see docs/deepseek-v4.md.

Give it RAM. 43 × 256 routed experts are ~137 GiB on disk and a token touches 301 of them, so the expert cache hit rate is what sets tok/s — --ram is the single most valuable knob, and it changes speed only, never output.

Speculative drafting exists and is off. DSpark's markov drafter and full MTP are both implemented and verified: a draft can save forward passes but never change a token, because every accepted token is still the target's own argmax. Measured on real multi-turn chat, they accepted 1 in 15 and 10 in 24, and the rejected-suffix replay of this engine's recurrent attention state cost more than the drafts saved — one 14-token answer took 495 seconds. So V4_DRAFT and V4_MTP default to 0 and the code stays, with the numbers beside it, for whoever retries this on faster storage.

See docs/deepseek-v4.md for the CUDA tier (build, DLL selection, GPU coverage), the environment reference, performance numbers, checkpoint validation, and the generated tiny independent oracle.

What's next

  • Inference-systems research is the product. The current hierarchy is LRU + a learned pin set; active work spans model formats, compression, placement, scheduling, I/O, CPU/GPU kernels, heterogeneous overlap, KV state, and routing-aware speculation. The objective is lower hardware requirements and lower cost per useful token. Everything lands the way this project works: measured end to end, reviewed, and developed in the open.
  • More open models. The tiering algorithm is model-agnostic: any MoE with routed experts can be staged the same way. Nine families run today (GLM-5.2, GLM-5.3-Flash, Inkling, Kimi K3, DeepSeek V4 Flash, DeepSeek V4.1 Flash, Qwen3.8-Flash-Next, Qwen3.6, OLMoE); further open-weight families — MiniMax among the candidates — earn an engine the way the first eight did: when someone measures one end to end.

Supporting the project

colibrì started as a one-person project on a 12-core laptop with 25 GB of RAM; today its numbers come from a community of real machines. If it's useful to you:

  • ⭐ star the repo and share it;
  • 🐛 open issues with benchmark numbers from your hardware — datapoints move this project more than anything else;
  • 💬 join the Discord community to discuss experiments, hardware results, and research directions;
  • 💬 reach out via GitHub issues to sponsor development or donate hardware.

Repo layout

Makefile                  root build/check entry point
c/
├── colibri.c             GLM-5.2 engine  (make glm)
├── inkling.c             Inkling engine  (make inkling)
├── kimi_k3.c             Kimi K3 engine  (make kimi_k3)
├── deepseek_v4.c         DeepSeek V4 Flash engine  (make deepseek-v4)
├── qwen38.c              Qwen3.8-Flash-Next text engine  (make qwen38)
├── qwen36.c              Qwen3.6 engine  (make qwen36)
├── olmoe.c               OLMoE engine  (make olmoe)
│
├── st.h                  safetensors index and range reads
├── quant.h               canonical container decoders
├── expert_ffn.h          routed-expert FFN kernel shared by the MoE engines (planar int4, layer runner)
├── tok.h, json.h         tokenizer and JSON parser
├── compat.h              Windows/macOS shims (POSIX names, one place)
├── expert_store.h        streaming expert cache
├── route_trace.h         routing telemetry and .coli_usage, engine-agnostic
├── kv_prefix.h           KV prefix reuse across turns
│
├── backend_cuda.*        optional CUDA tier   (CUDA=1)
├── backend_metal.*       optional Metal tier  (METAL=1)
├── backend_vulkan.*      optional Vulkan tier (VULKAN=1)
│
├── Makefile              build and local checks
├── coli                  user-facing CLI
├── openai_server.py      OpenAI-compatible HTTP gateway
├── resource_plan.py      RAM/VRAM planner behind `coli plan` and `coli doctor`
├── tools/                offline conversion, fixtures and benchmarks
├── scripts/              long-running conversion helpers
└── tests/                dependency-free C and Python tests
web/                      browser UI (pure OpenAI-API client)
desktop/                  Tauri v2 desktop shell wrapping the web UI
docker/                   container images
docs/                     reference docs, experiments, media

One .c per model family, over shared single headers. An engine owns its architecture and nothing else; anything two engines both need — the safetensors reader, the container decoders, the tokenizer, the expert cache — lives in a header they both include, so a fix reaches all of them at once. That rule is not decorative: the defects that keep recurring here are the ones where a mechanism landed in one engine and never reached its siblings.

From the repository root, make, make check and make clean delegate to the engine Makefile.

Why "colibrì"

The hummingbird weighs a few grams, hovers in place, and visits a thousand flowers a day. This engine keeps a 744-billion-parameter giant alive on hummingbird rations: 25 GB of RAM, twelve CPU cores, and a lot of disk patience.

Acknowledgements

colibrì is an engine; the minds it runs are a gift. Thank you to the teams releasing frontier-class weights in the open — Z.ai (GLM), Moonshot AI (Kimi), Alibaba Qwen, MiniMax, and Allen AI (OLMoE) — and to every contributor who benchmarked, bisected, replicated an atlas run, or sent a patch. This project is proof of what open weights make possible.

The project's expert placement, compression, and routing experiments also build on ideas and evidence from the following open research and systems work:

  • REAP and EASY-EP for output-aware and domain-specific expert importance.
  • SERE for similarity-based expert re-routing, and ReMoE for cache-locality-aware router fine-tuning.
  • MC-SMoE for routing-guided expert merging and compression.
  • MoBE and D²-MoE for shared expert bases and low-rank expert deltas.
  • HybriMoE for hybrid CPU/GPU expert scheduling, ScMoE for overlapping expert communication with computation, and OD-MoE for distributed on-demand expert loading.
  • vLLM, llama.cpp, and kTransformers for the open inference systems and expert-offload work that make comparisons reproducible.

The engine also stands on concrete engineering work, not only ideas. Each of these is used or reimplemented in the tree today:

  • safetensors — the container every engine reads (c/st.h), including its fp8 and I64 dtypes.
  • tiktoken — c/tok.h reimplements its byte_pair_encode exactly, merging the adjacent pair whose concatenation has the lowest vocab id, so a tiktoken-derived vocabulary needs no merges list.
  • llama.cpp — the GBNF grammar subset in c/grammar.h follows its syntax and its set-of-stacks PDA, and the Metal path borrows its newBufferWithBytesNoCopy residency trick.
  • vLLM — the reference for output semantics the engine matches position by position (e.g. where the final norm lands relative to the LM head).
  • transformers — the oracle: CI reproduces a random-init model token for token against it.
  • DietGPU — the GPU ANS codec behind the experimental compressed expert tier (COLI_ANS).
  • rocWMMA — the HIP backend maps CUDA's nvcuda::wmma fragment/mma_sync API onto it (c/backend_gpu_compat.h), which is what lets one .cu source compile for both vendors.

License

Apache 2.0, Copyright 2026 Vincenzo Fornaro. See LICENSE and NOTICE. GLM-5.2 weights are released by Z.ai under MIT.

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

19 total
  1. colibri 1.12.1v1.12.1Sep 24, 20261.4K downloads

    96 pull requests since v1.12.0, 80 of them from contributors. Two tokenizers brought back to the reference, brio on the ninth engine, `coli chat` working again at the default context on two families, and a placement decision that is now measured on the card in front of it instead of predicted. ### Tokenizers, measured against the reference - **#1654**: qwen36 tokenized differently from HF `tokenizers` in two ways. An added token right after punctuation was encoded as text (`X.<|im_end|>` was 7 tokens instead of 3, every chat turn ending in punctuation paid +4, #1653), and a whitespace run followed by a non-space was one piece where the regex's `\s+(?!\S)` leaves the last char to the next one, so every indented line of code tokenized differently. Measured on the real vocabulary: 2,803 lines and blocks of code, Markdown, Chinese and Japanese went from 757 identical to 2,803, with 5.4% fewer tokens. - **#1656**: OLMoE's `tokenizer.json` has no Split, a bare ByteLevel with `use_regex`, for which HF runs the original GPT-2 pattern; `tok.h` applied cl100k. A GPT-2 family in `tok.h`: 1,560/1,708 identical before, 1,708/1,708 after. The same measurement on GLM-5.2/5.3

  2. colibri 1.12.0v1.12.0Sep 20, 20262.6K downloads

    # colibri 1.12.0 81 pull requests since v1.11.0. A new way to ask a model a closed question, a redesigned dashboard and landing page, and a long run of small failures that used to answer 500 or die on a locale. ### Brio mode: score a closed set instead of generating - **#1632**: brio mode, on all nine engines. Hand the engine the options an answer is allowed to take and it reports the probability of each one instead of writing the answer: the option tokens are read, not sampled, so `completion_tokens` is 0 and no reply can fall outside the list. It is opt-in per request through two `SUBMIT` keys, `logprobs=k` and `pin=1`; without them every frame is byte-identical to before, which is asserted per engine rather than claimed. `max_tokens=0` became legal, but only together with `logprobs>0`. - The shared contract is three headers: `decode_batch.h` for the logprob tail, `serve_codec.h` for the wire, and `pin_pool.h` for nested state snapshots, so a request can keep one photograph of the shared instructions and a deeper one of the instructions plus the question. With one level the question is re-read once per option; with two, four items cost 176 tokens instea

  3. colibri v1.11.0v1.11.0Sep 13, 20269.3K downloads

    56 pull requests since v1.10.2. A ninth model family, five real bugs closed across four engines, and the two platforms the C tests never built on now building them in CI. ### A ninth engine: DeepSeek V4.1 Flash - **#1453**: DeepSeek V4.1 Flash (552B, 510 GB on disk) runs on a CPU box streaming experts from an SSD, with no conversion: the released checkpoint is read natively, fp8 dense with 32x32 ue8m0 tiles and fp4 experts whose layout is byte-identical to the mxfp4 the Kimi K3 engine already reads. Everything in the architecture that is not V4 is in: Engram (two n-gram memories of 384M rows, 203 GB, that never enter RAM), the DSA indexer with its two-level candidate source, hyper-connections, the compressor, the 32-layer vision tower with its aligner, DSpark speculative decoding (3 stages, blocks of 5, verified in one batched forward with a rollback for the rejected rows) and tool calling in the checkpoint's own DSML format. - Held token-exact in CI against a torch-only CPU reference (`c/tools/dsv41_ref.py`, written because the vendor's own forward needs tilelang GPU kernels), at three cache capacities, on a short and a 40-token prompt, under all three s

  4. colibri 1.10.2v1.10.2Sep 6, 20263.2K downloads

    Patch release. Three of these fixes answer reports made against 1.10.1 in the days after it shipped. No new engine; one behaviour change, opt-in. ## If you hit one of these on 1.10.1, this is the fix **`coli convert` produced a GLM-5.3-Flash container the engine refused** (#1368). It ran GLM-5.2's converter on everything; Flash nests its language model under the vision wrapper, so the embedding matched no rule, fell through to the generic fallback, and was quantized. The engine was right to refuse it, hours later, inside `coli web`. `coli convert` now reads the checkpoint's `config.json` and picks the converter from it. GLM-5.2's path is unchanged. An option the target converter does not take is refused, not silently dropped (#1369). And `tools/convert_glm53.py`, the right tool for Flash, is in the archive for the first time (#1364). **`coli doctor` said two core tensors were missing from a healthy GLM-5.3-Flash** (#1365). The check held three literal tensor names taken from GLM-5.2. It now matches roles, prefix-agnostically, so it holds for any family that nests, including ones not written yet (#1366). Nothing was ever missing from those downloads. **The archive was missing fi

  5. colibri v1.10.1v1.10.1Aug 31, 20261.6K downloads

    Packaging repair for the prebuilt archives. No engine changes: if you build from source, v1.10.1 is v1.10.0. ## What was wrong The v1.10.0 prebuilt archives were missing seven Python files, which broke four `coli` commands (#1296, thanks to simongonzalezdc for the report): - `coli convert`, `coli eval`, `coli mirror` and the benchmark fetch died with "No such file or directory" - images sent to Qwen3.8 through the server failed claiming Pillow and numpy were missing, when the real problem was a file we had not shipped ## What changed The release workflow no longer lists the Python files by hand. A packing script computes everything `coli` reaches, through imports and subprocess calls alike, and the same computation runs a second time as a check on the extracted archive before publishing. The old check only confirmed the one file the copy step had just written, which is why v1.10.0 went out green. Each archive now carries 14 Python files, verified by opening the archives rather than trusting the build log. ## Also in this tag - Engines now reject numeric arguments they can only read half of: `./qwen38 3x2` errors out instead of silently running with a cache of 3 instead of 3

Code frequency

additions and deletions
+79.8K-79.8KWeek of 2026-06-28: +202 linesWeek of 2026-06-28: -0 linesWeek of 2026-07-05: +13,871 linesWeek of 2026-07-05: -2,463 linesWeek of 2026-07-12: +36,477 linesWeek of 2026-07-12: -3,912 linesWeek of 2026-07-19: +54,972 linesWeek of 2026-07-19: -9,898 linesWeek of 2026-07-26: +26,378 linesWeek of 2026-07-26: -3,292 linesWeek of 2026-08-02: +79,784 linesWeek of 2026-08-02: -60,853 linesWeek of 2026-08-09: +11,040 linesWeek of 2026-08-09: -3,085 linesWeek of 2026-08-16: +24,806 linesWeek of 2026-08-16: -2,976 linesWeek of 2026-08-23: +55,608 linesWeek of 2026-08-23: -2,698 linesWeek of 2026-08-30: +46,732 linesWeek of 2026-08-30: -1,302 linesWeek of 2026-09-06: +19,707 linesWeek of 2026-09-06: -1,669 linesWeek of 2026-09-13: +16,267 linesWeek of 2026-09-13: -3,013 linesWeek of 2026-09-20: +28,295 linesWeek of 2026-09-20: -12,565 linesWeek of 2026-09-27: +0 linesWeek of 2026-09-27: -0 linesJun 28, 2026Sep 27, 2026
+414.1K lines added, -107.7K removed over the last year.

Commits per week

last 52 weeks
2740Week of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 0 commitsWeek of 2026-03-29: 0 commitsWeek of 2026-04-05: 0 commitsWeek of 2026-04-12: 0 commitsWeek of 2026-04-19: 0 commitsWeek of 2026-04-26: 0 commitsWeek of 2026-05-03: 0 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 0 commitsWeek of 2026-05-24: 0 commitsWeek of 2026-05-31: 0 commitsWeek of 2026-06-07: 0 commitsWeek of 2026-06-14: 0 commitsWeek of 2026-06-21: 0 commitsWeek of 2026-06-28: 1 commitsWeek of 2026-07-05: 57 commitsWeek of 2026-07-12: 274 commitsWeek of 2026-07-19: 261 commitsWeek of 2026-07-26: 156 commitsWeek of 2026-08-02: 123 commitsWeek of 2026-08-09: 135 commitsWeek of 2026-08-16: 146 commitsWeek of 2026-08-23: 132 commitsWeek of 2026-08-30: 73 commitsWeek of 2026-09-06: 135 commitsWeek of 2026-09-13: 142 commitsWeek of 2026-09-20: 160 commitsWeek of 2026-09-27: 0 commitsOct 5, 2025Sep 27, 2026
1.8K commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 14 commitsSun 1:00 — 11 commitsSun 2:00 — 7 commitsSun 3:00 — 9 commitsSun 4:00 — 9 commitsSun 5:00 — 4 commitsSun 6:00 — 2 commitsSun 7:00 — 3 commitsSun 8:00 — 2 commitsSun 9:00 — 14 commitsSun 10:00 — 20 commitsSun 11:00 — 8 commitsSun 12:00 — 18 commitsSun 13:00 — 23 commitsSun 14:00 — 20 commitsSun 15:00 — 18 commitsSun 16:00 — 18 commitsSun 17:00 — 8 commitsSun 18:00 — 10 commitsSun 19:00 — 8 commitsSun 20:00 — 13 commitsSun 21:00 — 22 commitsSun 22:00 — 18 commitsSun 23:00 — 31 commitsMon 0:00 — 16 commitsMon 1:00 — 7 commitsMon 2:00 — 6 commitsMon 3:00 — 4 commitsMon 4:00 — 2 commitsMon 5:00 — 2 commitsMon 6:00 — 6 commitsMon 7:00 — 2 commitsMon 8:00 — 8 commitsMon 9:00 — 10 commitsMon 10:00 — 4 commitsMon 11:00 — 4 commitsMon 12:00 — 5 commitsMon 13:00 — 9 commitsMon 14:00 — 5 commitsMon 15:00 — 10 commitsMon 16:00 — 11 commitsMon 17:00 — 15 commitsMon 18:00 — 26 commitsMon 19:00 — 22 commitsMon 20:00 — 28 commitsMon 21:00 — 11 commitsMon 22:00 — 12 commitsMon 23:00 — 25 commitsTue 0:00 — 13 commitsTue 1:00 — 19 commitsTue 2:00 — 7 commitsTue 3:00 — 7 commitsTue 4:00 — 2 commitsTue 5:00 — 4 commitsTue 6:00 — 6 commitsTue 7:00 — 7 commitsTue 8:00 — 12 commitsTue 9:00 — 12 commitsTue 10:00 — 5 commitsTue 11:00 — 11 commitsTue 12:00 — 14 commitsTue 13:00 — 17 commitsTue 14:00 — 21 commitsTue 15:00 — 21 commitsTue 16:00 — 19 commitsTue 17:00 — 20 commitsTue 18:00 — 15 commitsTue 19:00 — 19 commitsTue 20:00 — 42 commitsTue 21:00 — 19 commitsTue 22:00 — 25 commitsTue 23:00 — 26 commitsWed 0:00 — 13 commitsWed 1:00 — 16 commitsWed 2:00 — 11 commitsWed 3:00 — 3 commitsWed 4:00 — 3 commitsWed 5:00 — 1 commitsWed 6:00 — 2 commitsWed 7:00 — 10 commitsWed 8:00 — 9 commitsWed 9:00 — 7 commitsWed 10:00 — 5 commitsWed 11:00 — 6 commitsWed 12:00 — 13 commitsWed 13:00 — 13 commitsWed 14:00 — 11 commitsWed 15:00 — 18 commitsWed 16:00 — 11 commitsWed 17:00 — 7 commitsWed 18:00 — 7 commitsWed 19:00 — 17 commitsWed 20:00 — 11 commitsWed 21:00 — 10 commitsWed 22:00 — 18 commitsWed 23:00 — 17 commitsThu 0:00 — 12 commitsThu 1:00 — 15 commitsThu 2:00 — 9 commitsThu 3:00 — 6 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 4 commitsThu 7:00 — 4 commitsThu 8:00 — 6 commitsThu 9:00 — 12 commitsThu 10:00 — 11 commitsThu 11:00 — 15 commitsThu 12:00 — 13 commitsThu 13:00 — 8 commitsThu 14:00 — 12 commitsThu 15:00 — 21 commitsThu 16:00 — 17 commitsThu 17:00 — 8 commitsThu 18:00 — 7 commitsThu 19:00 — 12 commitsThu 20:00 — 30 commitsThu 21:00 — 17 commitsThu 22:00 — 8 commitsThu 23:00 — 7 commitsFri 0:00 — 9 commitsFri 1:00 — 7 commitsFri 2:00 — 5 commitsFri 3:00 — 5 commitsFri 4:00 — 2 commitsFri 5:00 — 1 commitsFri 6:00 — 0 commitsFri 7:00 — 6 commitsFri 8:00 — 7 commitsFri 9:00 — 12 commitsFri 10:00 — 9 commitsFri 11:00 — 6 commitsFri 12:00 — 14 commitsFri 13:00 — 14 commitsFri 14:00 — 15 commitsFri 15:00 — 9 commitsFri 16:00 — 10 commitsFri 17:00 — 13 commitsFri 18:00 — 12 commitsFri 19:00 — 4 commitsFri 20:00 — 6 commitsFri 21:00 — 6 commitsFri 22:00 — 15 commitsFri 23:00 — 7 commitsSat 0:00 — 8 commitsSat 1:00 — 14 commitsSat 2:00 — 7 commitsSat 3:00 — 9 commitsSat 4:00 — 7 commitsSat 5:00 — 1 commitsSat 6:00 — 0 commitsSat 7:00 — 3 commitsSat 8:00 — 2 commitsSat 9:00 — 4 commitsSat 10:00 — 3 commitsSat 11:00 — 3 commitsSat 12:00 — 7 commitsSat 13:00 — 12 commitsSat 14:00 — 14 commitsSat 15:00 — 21 commitsSat 16:00 — 10 commitsSat 17:00 — 11 commitsSat 18:00 — 11 commitsSat 19:00 — 10 commitsSat 20:00 — 6 commitsSat 21:00 — 10 commitsSat 22:00 — 7 commitsSat 23:00 — 9 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits900 (32%)
Community commits1,903 (68%)

2,803 commits in total over the last year.

DateListRankStars gained
Sep 25, 2026weekly#9+2,739
Sep 24, 2026weekly#9+2,739
Sep 23, 2026weekly#7+4,036
Sep 22, 2026weekly#6+5,565
Sep 21, 2026weekly#3+7,441
Sep 18, 2026daily#14+1,546
Sep 17, 2026daily#14+1,546
Sep 16, 2026daily#3+2,026
Sep 15, 2026daily#2+2,173
Sep 14, 2026daily#1+868
Sep 13, 2026daily#1+652
Sep 11, 2026daily#9+157
Sep 10, 2026daily#9+157
Jul 27, 2026daily#13+4
Jul 25, 2026daily#21+3
  • Genymobile/scrcpy

    Display and control your Android device

    151K stars · C

  • microsoft/PowerToys

    Microsoft PowerToys is a collection of utilities that supercharge productivity and customization on Windows

    139.2K stars · C

  • colbymchenry/codegraph

    Pre-indexed code knowledge graph, auto syncs on code changes, for Claude Code, Codex, Gemini, Cursor, OpenCode, AntiGravity, Kiro, CoPilot, and Hermes Agent — fewer tokens, fewer tool calls, 100% local

    73.2K stars · C

  • DeusData/codebase-memory-mcp

    High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.

    45.7K stars · C

  • facebook/zstd

    Zstandard - Fast real-time compression algorithm

    28K stars · C

  • antirez/ds4

    DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

    23.3K stars · C