TheTom/turboquant_plusPublic

AI summary: An experimental framework for compressing large language model KV caches using Walsh-Hadamard rotation and PolarQuant.

Stars
7K
+-6 today
Forks
922
Watchers
65
Open issues
36
Open PRs
11
Contributors
~3
Commits
342
Branches
14

PythonApache-2.0Created Mar 25, 2026Last push 2mo ago+-4 stars this week+11 this month

Quick answers

What is turboquant_plus?
An experimental framework for compressing large language model KV caches using Walsh-Hadamard rotation and PolarQuant.
What does turboquant_plus do?
TurboQuant+ is a research implementation and extension of the TurboQuant KV cache compression algorithm, designed to radically improve the memory efficiency of large language models during inference. By applying an asymmetric compression policy that utilizes Walsh-Hadamard rotations and PolarQuant codebooks, it achieves 3.8x to 6.4x compression of the transformer KV cache without significantly degrading generation quality. This allows developers to run massively longer context windows on consumer hardware or pack significantly more concurrent users onto a single cloud GPU. The core techniques developed in this repository have directly influenced upstream inference engines like vLLM and llama.cpp.
Who is turboquant_plus for?
TurboQuant+ is aimed at AI researchers, inference optimization engineers, and advanced LLM practitioners. It requires a deep understanding of transformer architectures, quantization mathematics, and low-level GPU programming.
How do I get started with turboquant_plus?
git clone https://github.com/TheTom/turboquant_plus.git
How popular is turboquant_plus on GitHub?
TheTom/turboquant_plus has 7,030 stars and 922 forks on GitHub, and gained -4 stars in the last 7 days.
What license does turboquant_plus use?
TheTom/turboquant_plus is released under the Apache-2.0 license.

Star history

since Jul 28, 2026
02K4K6KJul 2026Aug 2026Sep 2026Oct 2026
7K stars as of Oct 4, 2026. Measured daily since Jul 28, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks
OctNovDecJanFebMarAprMayJunJulAugSepOctMonWedFri2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 15 commits2026-03-25: 33 commits2026-03-26: 41 commits2026-03-27: 36 commits2026-03-28: 35 commits2026-03-29: 9 commits2026-03-30: 8 commits2026-03-31: 17 commits2026-04-01: 7 commits2026-04-02: 11 commits2026-04-03: 14 commits2026-04-04: 8 commits2026-04-05: 28 commits2026-04-06: 3 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 4 commits2026-04-10: 4 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 1 commit2026-04-21: 1 commit2026-04-22: 0 commits2026-04-23: 2 commits2026-04-24: 0 commits2026-04-25: 1 commit2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 6 commits2026-04-30: 12 commits2026-05-01: 5 commits2026-05-02: 10 commits2026-05-03: 5 commits2026-05-04: 1 commit2026-05-05: 4 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 1 commit2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 1 commit2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 2 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 1 commit2026-06-11: 1 commit2026-06-12: 10 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 0 commits2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 0 commits2026-07-03: 0 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 0 commits2026-07-09: 0 commits2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 0 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 1 commit2026-07-21: 0 commits2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 0 commits2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 0 commits2026-07-28: 0 commits2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits2026-08-02: 0 commits2026-08-03: 0 commits2026-08-04: 0 commits2026-08-05: 0 commits2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 0 commits2026-08-09: 0 commits2026-08-10: 0 commits2026-08-11: 0 commits2026-08-12: 0 commits2026-08-13: 0 commits2026-08-14: 0 commits2026-08-15: 0 commits2026-08-16: 0 commits2026-08-17: 0 commits2026-08-18: 0 commits2026-08-19: 0 commits2026-08-20: 0 commits2026-08-21: 0 commits2026-08-22: 0 commits2026-08-23: 0 commits2026-08-24: 0 commits2026-08-25: 0 commits2026-08-26: 0 commits2026-08-27: 0 commits2026-08-28: 0 commits2026-08-29: 0 commits2026-08-30: 0 commits2026-08-31: 0 commits2026-09-01: 0 commits2026-09-02: 0 commits2026-09-03: 0 commits2026-09-04: 0 commits2026-09-05: 0 commits2026-09-06: 0 commits2026-09-07: 0 commits2026-09-08: 0 commits2026-09-09: 0 commits2026-09-10: 0 commits2026-09-11: 0 commits2026-09-12: 0 commits2026-09-13: 0 commits2026-09-14: 0 commits2026-09-15: 0 commits2026-09-16: 0 commits2026-09-17: 0 commits2026-09-18: 0 commits2026-09-19: 0 commits2026-09-20: 0 commits2026-09-21: 0 commits2026-09-22: 0 commits2026-09-23: 0 commits2026-09-24: 0 commits2026-09-25: 0 commits2026-09-26: 0 commits2026-09-27: 0 commits2026-09-28: 0 commits2026-09-29: 0 commits2026-09-30: 0 commits2026-10-01: 0 commits2026-10-02: 0 commits2026-10-03: 0 commits2026-10-04: 0 commits2026-10-05: 0 commits2026-10-06: 0 commits2026-10-07: 0 commits2026-10-08: 0 commits2026-10-09: 0 commits2026-10-10: 0 commits
338 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Permissive license

    Apache-2.0

What turboquant_plus does

TurboQuant+ is a research implementation and extension of the TurboQuant KV cache compression algorithm, designed to radically improve the memory efficiency of large language models during inference. By applying an asymmetric compression policy that utilizes Walsh-Hadamard rotations and PolarQuant codebooks, it achieves 3.8x to 6.4x compression of the transformer KV cache without significantly degrading generation quality. This allows developers to run massively longer context windows on consumer hardware or pack significantly more concurrent users onto a single cloud GPU. The core techniques developed in this repository have directly influenced upstream inference engines like vLLM and llama.cpp.

TurboQuant+ is aimed at AI researchers, inference optimization engineers, and advanced LLM practitioners. It requires a deep understanding of transformer architectures, quantization mathematics, and low-level GPU programming.

  • Extreme KV cache compression: Achieves up to 6.4x reduction in cache memory footprint using advanced quantization techniques.
  • Walsh-Hadamard rotation: Implements fast, fused mathematical kernels on CPU, CUDA, and Vulkan to efficiently rotate cache representations.
  • Asymmetric compression policies: Applies different quantization strategies to the Key and Value caches to optimize the trade-off between speed and accuracy.
  • Upstream integration compatibility: Provides the underlying research and implementation details that have been merged into production systems like vLLM.
  • Near-native inference speed: Maintains prefill speeds comparable to standard 8-bit quantization while significantly extending context capacity.

Where teams use it

Long-Context LLM Inference

Researchers deploy models capable of analyzing entire books or massive codebases simultaneously without encountering out-of-memory errors on a single GPU.

High-Density Model Serving

Cloud providers utilize the compression techniques to drastically increase the number of concurrent user sessions a single inference node can support.

Edge Device AI Deployment

Engineers leverage the extreme memory savings to run complex, interactive chat models on hardware-constrained edge devices like high-end smartphones.

Quantization Algorithm Research

Academics use the repository as a baseline for developing new, even more aggressive asymmetric caching strategies.

Getting started: git clone https://github.com/TheTom/turboquant_plus.git

README

main branch

TurboQuant+

🚀 TurboQuant KV cache compression is now in vLLM (PR #38479, merged April 2026): --kv-cache-dtype turboquant_k8v4 and friends, with fused Triton store/decode kernels. The PR discussion drew on the asymmetric K/V findings from this repo. Upstream llama.cpp has merged the core idea too: Hadamard KV cache rotation (#21038, citing TurboQuant directly) with fast WHT kernels on CPU (#22631), CUDA (#23615), and Vulkan (#23687). Rotation + the stock q4_0 cache is essentially turbo4's rotation stage; the PolarQuant codebook, norm extraction, and asymmetric policies remain here and in the fork.

Implementation of TurboQuant (ICLR 2026) with implementation work, experiments, and follow-on findings beyond the base paper. Compresses transformer KV cache 3.8-6.4x using PolarQuant + Walsh-Hadamard rotation, at near q8_0 prefill speed and ~0.9x decode throughput at long context. Validated end-to-end from 1.5B to 104B at 128K context on a MacBook (turbo3, PPL 4.024, 74 GB peak memory).

This repository is the research home: the Python reference implementation, the validation papers, and the benchmark data. To run TurboQuant in an inference engine, pick from the table below. Pieces that prove useful and stable get upstreamed incrementally as small, reviewable patches.

Run It Today

Engine Platform Status Notes
vLLM CUDA / ROCm, datacenter Upstream, merged --kv-cache-dtype turboquant_k8v4 and friends (PR #38479)
llama.cpp All backends Upstream, rotation merged Hadamard KV cache rotation (#21038) + fast WHT kernels (CPU #22631, CUDA #23615, Vulkan #23687). Rotation + q4_0 cache approximates turbo4; full PolarQuant codec is in the fork below
llama-cpp-turboquant Metal, CUDA, HIP, CPU Production fork turbo2/3/4 KV cache + TQ3_1S/TQ4_1S weight formats; prebuilt binaries for Mac (Metal) and Windows (CUDA)
mlx-swift-lm Apple Silicon, Swift Upstream, merged Merged into Apple MLX (PR #232): full asymmetric family (turbo0v*/turbo8v*) + symmetric turbo4/3/2 with per-dimension key calibration, behind kvScheme. turbo8v3: 2.7x KV at affine8-class KLD across 6 families (1.7B to 32B, incl. 30B MoE); decode 0.7-0.8x fp16
vllm-swift Apple Silicon, Swift Active OpenAI-compatible serving built on mlx-swift-lm; no Python in the inference hot path
Atlas Rust, Metal Integrated turbo4 KV cache append/decode kernels
mlxcel Rust, MLX Community port Turbo KV cache docs. Thanks to @inureyes and the Lablup team for the careful port and upstream attribution
MLX Python fork Apple Silicon, Python Experimental TurboKVCache drop-in for mlx-lm and mlx-vlm; see docs/mlx-port.md
turboquant-pytorch PyTorch Independent implementation From-scratch PyTorch TurboQuant for KV cache; applies this repo's layer-adaptive compression and Sparse V findings
turboquant-vllm CUDA, vLLM package Community package TQ+ KV cache for vLLM with fused CUDA kernels; builds on this repo's block_size=128 and K-dominance research
Quansloth Local server product Community product Air-gapped local AI server built on the TurboQuant+ stack

Key Findings

Three follow-on findings, independently validated by multiple researchers across different hardware and backends:

  1. V compression is free. Compressing the value cache (even down to 2 bits) has zero measurable effect on attention quality when key precision is maintained. Confirmed on Metal (M5 Max), CUDA RTX 4090 (@sztlink), and CUDA RTX 3090 (@HyperionMS2040). See asymmetric K/V paper.
  2. All quality degradation comes from K compression. This is why asymmetric configs (q8_0-K + turbo-V) rescue models where symmetric fails. Validated across Qwen, Llama, Mistral, and Command-R+ families. See M5 Max stress test.
  3. Boundary layers are disproportionately sensitive. Protecting the first 2 + last 2 layers at higher precision recovers 37-91% of the quality gap. See Boundary V paper.

Additional experiments and writeups: Sparse V dequant (+22.8% decode at 32K, not TurboQuant-specific), block size optimization (5.12x compression), turbo4 resurrection (QJL hurts, PolarQuant works), EDEN optimal-S response (rotation is first-order, scale is second-order).

Quality at a Glance (M5 Max 128GB)

Cache Type Bits/val Compression PPL (wikitext-2, 512c) vs q8_0
f16 16.0 1.0x 6.121 -0.16%
q8_0 8.5 1.9x 6.111 baseline
turbo4 4.25 3.8x 6.125 +0.23%
q4_0 4.5 3.6x 6.142 +0.52%
turbo3 3.5† 4.6x† 6.176 +1.06%
turbo2 2.5 6.4x 6.507 +6.48%

turbo4 (4-bit PolarQuant) has the best quality after q8_0 — closer to q8_0 than q4_0, at better compression. turbo3 trades quality for maximum compression. turbo2 (2-bit) trades more quality for extreme compression — best used asymmetrically.

†turbo3 at default block_size=32. At block_size=128, turbo3 achieves 3.125 bits/val and 5.12x compression with identical PPL. See block size study.

Important: choosing the right config for your model. TurboQuant quality depends on your base weight quantization. Models with Q8_0+ weights work well with symmetric turbo (e.g., -ctk turbo3 -ctv turbo3). Some low-bit models with Q4_K_M weights may benefit from asymmetric K/V: use -ctk q8_0 -ctv turbo4 to keep K precision high while compressing V. K precision is the dominant quality factor because it controls attention routing via softmax. Bigger models absorb quantization stacking better (104B: +3.6% vs 70B: +11.4% for turbo3). Validate on your specific model. See Configuration Recommendations for the full tested matrix.

Everything else lives in docs/benchmarks.md: asymmetric K/V and Boundary V tables, prefill context scaling, MoE and dense decode speed, NIAH retrieval, KL divergence, 70B/104B stress tests, community results on RTX 3090, M1 Max, and AMD RX 9070 XT, and the speed optimization journey.

Getting Started

Python Reference Implementation

git clone https://github.com/TheTom/turboquant_plus.git
cd turboquant_plus
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"

# Verify — should print "141 passed"
python3 -m pytest tests/ -v

# Quick compression demo (no model needed)
python3 benchmarks/demo.py

# Validate on real model KV tensors (downloads Qwen3-1.7B, ~4GB)
pip install transformers torch accelerate
python3 benchmarks/validate_real_model.py

Requires Python >= 3.10, NumPy >= 1.24, SciPy >= 1.10. torch / transformers / accelerate are optional (real-model validation only).

Run Inference

The fastest path is a prebuilt binary, or build from the fork — see the llama-cpp-turboquant quick start for build instructions, supported backends, and usage details.

# Server mode with TurboQuant KV cache
./build/bin/llama-server -m models/your-model.gguf \
  --jinja -ngl 99 -c 262144 -fa on \
  --cache-type-k turbo3 --cache-type-v turbo3 \
  -np 1 --host 0.0.0.0 --port 8080
Flag Bits/val Compression vs fp16 Description
turbo3 3.5† 4.6x† 3-bit PolarQuant + WHT rotation. Best compression, q8_0 speed.
turbo4 4.25 3.8x 4-bit PolarQuant (16 centroids). Best quality.
q8_0 8 2.0x llama.cpp default quantized cache.
q4_0 4 4.0x llama.cpp 4-bit cache.

For per-model configuration guidance (symmetric vs asymmetric, Boundary V, large-model memory caps), see the Getting Started Guide and Configuration Recommendations.

Architecture

Input: KV cache vector x ∈ R^d (one attention head)
    │
    ├── Extract norm: γ = ||x||, x̂ = x/γ
    │
    ├── Random rotation: WHT + random sign flips
    │   coordinates ~ N(0, 1/d) after rotation
    │
    ├── Optimal scalar quantization (Lloyd-Max)
    │   turbo4: 16 centroids (4-bit), turbo3: 8 centroids (3-bit), turbo2: 4 centroids (2-bit)
    │
    └── Output: quantized indices + norm per block
        Compression: 3.8x (turbo4), 5.1x (turbo3), 7.5x (turbo2)

Note on QJL: reference only, not used in production.

The original TurboQuant paper (Zandieh et al. 2024, arXiv 2406.03482) includes a 1-bit QJL error-correction stage. The Python qjl.py here implements it for paper reproducibility.

Production drops QJL on both K and V. QJL eliminates reconstruction bias but amplifies variance, which softmax turns into attention noise. Five independent groups confirmed (buun, scos-lab, Arclabs001, +2). See turbo4-resurrection.md for the full ablation and mechanism.

If you're building on this repo: use TurboQuantMSE (V cache), or implement straight 4 to 8 bit PolarQuant on K. Only enable the QJL / TurboQuant (with QJL) classes if you are reproducing the original paper or doing K-side research below 8-bit.

Project Structure
turboquant/
├── rotation.py        # Walsh-Hadamard Transform + random sign flips
├── codebook.py        # Lloyd-Max optimal centroid computation
├── polar_quant.py     # PolarQuant — norm extraction + WHT rotation + scalar quantization
├── qjl.py            # QJL 1-bit quantizer (paper-faithful reference, see README §QJL). Not used in production.
├── turboquant.py      # Full TurboQuant pipeline
├── kv_cache.py        # KV cache integration layer
├── outlier.py         # Outlier channel strategy (2.5-bit, 3.5-bit)
├── lloyd_max.py       # Lloyd-Max quantizer implementation
├── utils.py           # Bit packing, memory measurement
├── isoquant.py        # IsoQuant (quaternion SO(4)) experimental comparison
└── rotorquant.py      # RotorQuant experimental comparison

tests/                 # 14 test files, 500+ tests
benchmarks/
├── demo.py                       # Quick compression demo
├── run_benchmark.py              # Server-based benchmark runner
├── benchmark_results.md          # Full benchmark report
├── benchmark_llama.sh            # llama.cpp benchmark script
├── benchmark_norm_correction.py  # Norm correction validation
├── benchmark_ppl_tq_vs_rq.py    # TurboQuant vs RotorQuant PPL comparison
├── temporal_decay_prototype.py   # Temporal decay experiment
├── test_with_llama.py            # Integration test at Qwen 3.5 dimensions
├── test_outlier_comparison.py    # Outlier strategy comparison
└── validate_real_model.py        # Real model KV tensor validation

docs/
├── benchmarks.md                 # Full benchmark and validation data
├── mlx-port.md                   # MLX framework port (Python)
├── changelog.md                  # v1 milestone history
├── turboquant-recommendations.md # Configuration guide (tested matrix)
├── windows-rdna4-setup.md        # Windows + AMD RDNA 4 build guide
├── papers/                       # Validation papers and experiment writeups
└── (25+ engineering docs, investigations, experiment logs)

Roadmap

Phase Status Details
Core algorithms (NumPy) ✅ 500+ tests across 14 test files
Distortion validation ✅ Matches paper bounds (Table 2)
Real model validation ✅ Rotation validated on Qwen3 KV tensors (kurtosis 900→2.9)
llama.cpp C port ✅ Metal GPU inference working on M1 through M5
Metal shader optimization ✅ q8_0 speed parity: prefill matches or beats q8_0
CUDA backend ✅ Community-tested on RTX 3080 Ti/3090/4090/5090, DGX Spark Blackwell
HIP/AMD backend ✅ RX 9070 XT (RDNA 4) validated, gfx1201 native
Asymmetric K/V ✅ q8_0-K + turbo-V rescues Q4_K_M models
Boundary V ✅ Layer-aware V compression, 37-91% quality recovery
Sparse V ✅ Attention-gated dequant skip, +22.8% decode on MoE. Upstream PR #21119
Block size optimization ✅ 32→128, 12% better compression, zero quality cost
vLLM upstream ✅ Merged as the TurboQuant attention backend (PR #38479)
llama.cpp upstream (rotation) ✅ Hadamard KV rotation + FWHT kernels merged upstream (#21038, #22631)
Upstream coordination 🔄 llama.cpp PR preparation for the full codec (#27)
TurboQuant+ extensions ⏳ Adaptive bits, temporal decay, MoE-aware compression
MLX Swift upstream ✅ Merged into Apple mlx-swift-lm (PR #232, 2026-07-20): JIT Metal kernels, asymmetric + calibrated symmetric schemes, 63 tests

Paper Reference

Docs

Contributing

Issues and PRs welcome. The main areas where help is needed:

  1. Upstream PR — prepare llama.cpp contribution (CONTRIBUTING.md requirements)
  2. CUDA kernel optimization — fused FA kernels, decode speed parity
  3. MLX memory recovery — implement FP16 KV drop + compressed-only attention for memory-constrained long context
  4. Quality metrics — multi-run statistics, additional task benchmarks (GSM8K, code gen, reasoning)
  5. Long context validation — 64K+ testing across architectures

Support

If you find this work useful, you can support it via GitHub Sponsors or BTC:

BTC: bc1qsfaaf6mkz2yxx2vavg2n0zgsf3qj25uh94t83rwuq7de67dey05sc3tgjx

Commercial support: For inference optimization and KV cache tuning engagements, DM @no_stp_on_snek on X.

License

Apache License 2.0 — see LICENSE.

Copyright 2026 Tom Turney.

Based on Google Research's TurboQuant paper (arXiv 2504.19874).

View on GitHub

Recent activity

commits and pull requests

Commits per week

last 52 weeks
1600Week of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 160 commitsWeek of 2026-03-29: 74 commitsWeek of 2026-04-05: 39 commitsWeek of 2026-04-12: 0 commitsWeek of 2026-04-19: 5 commitsWeek of 2026-04-26: 33 commitsWeek of 2026-05-03: 11 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 1 commitsWeek of 2026-05-24: 2 commitsWeek of 2026-05-31: 0 commitsWeek of 2026-06-07: 12 commitsWeek of 2026-06-14: 0 commitsWeek of 2026-06-21: 0 commitsWeek of 2026-06-28: 0 commitsWeek of 2026-07-05: 0 commitsWeek of 2026-07-12: 0 commitsWeek of 2026-07-19: 1 commitsWeek of 2026-07-26: 0 commitsWeek of 2026-08-02: 0 commitsWeek of 2026-08-09: 0 commitsWeek of 2026-08-16: 0 commitsWeek of 2026-08-23: 0 commitsWeek of 2026-08-30: 0 commitsWeek of 2026-09-06: 0 commitsWeek of 2026-09-13: 0 commitsWeek of 2026-09-20: 0 commitsWeek of 2026-09-27: 0 commitsWeek of 2026-10-04: 0 commitsOct 12, 2025Oct 4, 2026
338 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 1 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 3 commitsSun 8:00 — 1 commitsSun 9:00 — 3 commitsSun 10:00 — 0 commitsSun 11:00 — 1 commitsSun 12:00 — 3 commitsSun 13:00 — 2 commitsSun 14:00 — 5 commitsSun 15:00 — 2 commitsSun 16:00 — 2 commitsSun 17:00 — 1 commitsSun 18:00 — 6 commitsSun 19:00 — 4 commitsSun 20:00 — 2 commitsSun 21:00 — 5 commitsSun 22:00 — 0 commitsSun 23:00 — 1 commitsMon 0:00 — 0 commitsMon 1:00 — 0 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 0 commitsMon 8:00 — 0 commitsMon 9:00 — 0 commitsMon 10:00 — 2 commitsMon 11:00 — 1 commitsMon 12:00 — 0 commitsMon 13:00 — 1 commitsMon 14:00 — 1 commitsMon 15:00 — 2 commitsMon 16:00 — 1 commitsMon 17:00 — 0 commitsMon 18:00 — 0 commitsMon 19:00 — 0 commitsMon 20:00 — 0 commitsMon 21:00 — 2 commitsMon 22:00 — 2 commitsMon 23:00 — 3 commitsTue 0:00 — 0 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 0 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 1 commitsTue 9:00 — 2 commitsTue 10:00 — 1 commitsTue 11:00 — 8 commitsTue 12:00 — 4 commitsTue 13:00 — 4 commitsTue 14:00 — 0 commitsTue 15:00 — 0 commitsTue 16:00 — 1 commitsTue 17:00 — 0 commitsTue 18:00 — 0 commitsTue 19:00 — 1 commitsTue 20:00 — 3 commitsTue 21:00 — 12 commitsTue 22:00 — 0 commitsTue 23:00 — 0 commitsWed 0:00 — 2 commitsWed 1:00 — 1 commitsWed 2:00 — 0 commitsWed 3:00 — 0 commitsWed 4:00 — 0 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 2 commitsWed 8:00 — 1 commitsWed 9:00 — 2 commitsWed 10:00 — 2 commitsWed 11:00 — 4 commitsWed 12:00 — 1 commitsWed 13:00 — 0 commitsWed 14:00 — 1 commitsWed 15:00 — 1 commitsWed 16:00 — 3 commitsWed 17:00 — 3 commitsWed 18:00 — 5 commitsWed 19:00 — 4 commitsWed 20:00 — 7 commitsWed 21:00 — 4 commitsWed 22:00 — 3 commitsWed 23:00 — 1 commitsThu 0:00 — 1 commitsThu 1:00 — 0 commitsThu 2:00 — 0 commitsThu 3:00 — 0 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 0 commitsThu 8:00 — 0 commitsThu 9:00 — 3 commitsThu 10:00 — 12 commitsThu 11:00 — 7 commitsThu 12:00 — 7 commitsThu 13:00 — 3 commitsThu 14:00 — 4 commitsThu 15:00 — 11 commitsThu 16:00 — 4 commitsThu 17:00 — 8 commitsThu 18:00 — 3 commitsThu 19:00 — 0 commitsThu 20:00 — 0 commitsThu 21:00 — 1 commitsThu 22:00 — 4 commitsThu 23:00 — 5 commitsFri 0:00 — 2 commitsFri 1:00 — 3 commitsFri 2:00 — 0 commitsFri 3:00 — 0 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 7 commitsFri 8:00 — 8 commitsFri 9:00 — 8 commitsFri 10:00 — 7 commitsFri 11:00 — 2 commitsFri 12:00 — 4 commitsFri 13:00 — 6 commitsFri 14:00 — 6 commitsFri 15:00 — 2 commitsFri 16:00 — 2 commitsFri 17:00 — 0 commitsFri 18:00 — 0 commitsFri 19:00 — 2 commitsFri 20:00 — 2 commitsFri 21:00 — 3 commitsFri 22:00 — 1 commitsFri 23:00 — 5 commitsSat 0:00 — 1 commitsSat 1:00 — 0 commitsSat 2:00 — 0 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 4 commitsSat 8:00 — 2 commitsSat 9:00 — 4 commitsSat 10:00 — 2 commitsSat 11:00 — 0 commitsSat 12:00 — 2 commitsSat 13:00 — 0 commitsSat 14:00 — 10 commitsSat 15:00 — 5 commitsSat 16:00 — 3 commitsSat 17:00 — 2 commitsSat 18:00 — 9 commitsSat 19:00 — 1 commitsSat 20:00 — 3 commitsSat 21:00 — 1 commitsSat 22:00 — 3 commitsSat 23:00 — 2 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits338 (99%)
Community commits4 (1%)

342 commits in total over the last year.

DateListRankStars gained
Mar 31, 2026daily#12+233