Andyyyy64/whichllmPublic

Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.

AI summary: A CLI tool that aggressively profiles local hardware to recommend and rank the best-fitting local LLMs from HuggingFace.

Stars
6.7K
+11 today
Forks
371
Watchers
24
Open issues
6
Open PRs
8
Contributors
~26
Commits
267
Branches
25

PythonMITCreated Mar 4, 2026Last push 1d agoLatest release v0.5.19+35 stars this week+154 this month

Quick answers

What is whichllm?
A CLI tool that aggressively profiles local hardware to recommend and rank the best-fitting local LLMs from HuggingFace.
What does whichllm do?
whichllm is a Python command-line utility explicitly designed to eliminate the guesswork involved in selecting local Large Language Models. It actively profiles your specific machine's hardware, detecting exact GPU VRAM, system RAM bandwidth, and CPU capabilities. It then queries live HuggingFace databases and ranks the best available models based on constraints like partial RAM offloading and strict VRAM limits, entirely bypassing simplistic parameter count metrics. The tool also features powerful hardware simulation capabilities, allowing users to test hypothetical multi-GPU configurations before committing to expensive hardware purchases.
Who is whichllm for?
AI engineers, researchers, and hobbyists who run models locally and require an accurate, automated way to match HuggingFace models to strict hardware limits.
How do I get started with whichllm?
uvx whichllm@latest
How popular is whichllm on GitHub?
Andyyyy64/whichllm has 6,717 stars and 371 forks on GitHub, and gained 35 stars in the last 7 days.
What license does whichllm use?
Andyyyy64/whichllm is released under the MIT license.

Star history

since Jul 28, 2026
02K4K6KJul 2026Aug 2026Sep 2026Oct 2026
6.7K stars as of Oct 3, 2026. Measured daily since Jul 28, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks
SepOctNovDecJanFebMarAprMayJunJulAugSepMonWedFri2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 9 commits2026-03-05: 44 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 27 commits2026-03-10: 1 commit2026-03-11: 0 commits2026-03-12: 1 commit2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 9 commits2026-05-15: 16 commits2026-05-16: 12 commits2026-05-17: 23 commits2026-05-18: 9 commits2026-05-19: 2 commits2026-05-20: 5 commits2026-05-21: 2 commits2026-05-22: 3 commits2026-05-23: 0 commits2026-05-24: 1 commit2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 5 commits2026-06-04: 0 commits2026-06-05: 4 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 1 commit2026-06-09: 2 commits2026-06-10: 5 commits2026-06-11: 2 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 2 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 2 commits2026-06-18: 8 commits2026-06-19: 0 commits2026-06-20: 1 commit2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 2 commits2026-06-24: 3 commits2026-06-25: 1 commit2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 1 commit2026-06-29: 3 commits2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 0 commits2026-07-03: 5 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 1 commit2026-07-09: 1 commit2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 0 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 0 commits2026-07-21: 0 commits2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 0 commits2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 0 commits2026-07-28: 0 commits2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits2026-08-02: 0 commits2026-08-03: 0 commits2026-08-04: 0 commits2026-08-05: 2 commits2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 0 commits2026-08-09: 0 commits2026-08-10: 0 commits2026-08-11: 0 commits2026-08-12: 0 commits2026-08-13: 0 commits2026-08-14: 5 commits2026-08-15: 0 commits2026-08-16: 0 commits2026-08-17: 0 commits2026-08-18: 0 commits2026-08-19: 0 commits2026-08-20: 0 commits2026-08-21: 0 commits2026-08-22: 0 commits2026-08-23: 0 commits2026-08-24: 0 commits2026-08-25: 0 commits2026-08-26: 0 commits2026-08-27: 0 commits2026-08-28: 0 commits2026-08-29: 0 commits2026-08-30: 0 commits2026-08-31: 0 commits2026-09-01: 0 commits2026-09-02: 0 commits2026-09-03: 0 commits2026-09-04: 0 commits2026-09-05: 0 commits2026-09-06: 0 commits2026-09-07: 0 commits2026-09-08: 0 commits2026-09-09: 0 commits2026-09-10: 0 commits2026-09-11: 0 commits2026-09-12: 0 commits2026-09-13: 0 commits2026-09-14: 3 commits2026-09-15: 0 commits2026-09-16: 0 commits2026-09-17: 0 commits2026-09-18: 0 commits2026-09-19: 7 commits2026-09-20: 1 commit2026-09-21: 0 commits2026-09-22: 0 commits2026-09-23: 0 commits2026-09-24: 0 commits2026-09-25: 0 commits2026-09-26: 0 commits
231 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Actively maintained

    Pushed within 48 hours

  • Well documented

    High community health score

  • Permissive license

    MIT

  • Continuous integration

    Automated checks passing

What whichllm does

whichllm is a Python command-line utility explicitly designed to eliminate the guesswork involved in selecting local Large Language Models. It actively profiles your specific machine's hardware, detecting exact GPU VRAM, system RAM bandwidth, and CPU capabilities. It then queries live HuggingFace databases and ranks the best available models based on constraints like partial RAM offloading and strict VRAM limits, entirely bypassing simplistic parameter count metrics. The tool also features powerful hardware simulation capabilities, allowing users to test hypothetical multi-GPU configurations before committing to expensive hardware purchases.

AI engineers, researchers, and hobbyists who run models locally and require an accurate, automated way to match HuggingFace models to strict hardware limits.

  • Hardware profiling: Instantly and accurately identifies your machine's exact available GPU VRAM, system RAM bandwidth, and unified memory specs.
  • Recency-aware benchmarking: Ranks HuggingFace models using the most up-to-date performance metrics to prioritize truly usable inference speed.
  • Accurate partial offload calculation: Estimates the viability of running massive models by mathematically modeling partial layer offloads to system RAM.
  • Advanced hardware simulation: Allows users to override detected specifications with flags to accurately test hypothetical multi-GPU setups.
  • Configurable VRAM headroom: Lets researchers specify strict memory buffers to prevent out-of-memory errors during intensive context scaling.
  • Evidence-graded scoring: Rejects fabricated uploader claims and highly penalizes cross-family score inheritance to ensure true benchmark integrity.

Where teams use it

Local Inference Optimization

Developers setting up a local AI environment execute the tool to instantly identify the most capable model their current hardware can reliably support.

Hardware Purchase Planning

AI enthusiasts rigorously simulate running specific multi-GPU configurations to see exactly which models they will be able to run before buying expensive hardware.

Safe Model Selection

Researchers enable strict VRAM headroom flags to guarantee the specifically selected model will run stably without catastrophic out-of-memory crashes.

Automated Pipeline Integration

Engineers leverage the tool's JSON output directly in bash scripts to dynamically map the best HuggingFace IDs to their local Ollama inference configurations.

Getting started: uvx whichllm@latest

README

main branch

whichllm

PyPI version Python 3.11+ License: MIT Tests Sponsor

Andyyyy64%2Fwhichllm | Trendshift

Find the best local LLM that actually runs on your hardware.

Auto-detects your GPU/CPU/RAM and ranks the top models from HuggingFace that fit your system.

日本語版はこちら

Quick start

Run the recommendation command once, with no project setup.

uvx whichllm@latest

Simulate a GPU before you buy hardware.

uvx whichllm@latest --gpu "RTX 4090"

Install it when you use it often.

uv tool install whichllm
uv tool upgrade whichllm  # update an existing install

Other install paths.

brew install andyyyy64/whichllm/whichllm
pip install whichllm

Want a safer pick?

By default, whichllm is ambitious. It ranks the best model that looks runnable on your machine, including partial RAM offload and near-edge VRAM fits when they seem usable.

If you want a more comfortable LM Studio-style recommendation, start with:

uvx whichllm@latest --gpu-only --speed usable --vram-headroom 1GB

This keeps only models that fit fully in GPU VRAM, filters out slow estimates, and leaves extra VRAM for runtime overhead.

If LM Studio still says the model is slightly too large, increase the headroom:

uvx whichllm@latest --gpu-only --speed usable --vram-headroom 1.5GB

Common workflows

After install, run whichllm directly. For one-off runs, replace whichllm with uvx whichllm@latest.

# Best models for this machine
whichllm

# Pretend you have a specific GPU
whichllm --gpu "RTX 4090"

# Override detected iGPU/unified-memory limits
whichllm --vram 8 --ram-bandwidth 68

# Only show models that fit fully in GPU VRAM
whichllm --gpu-only
whichllm --fit gpu

# Simulate a multi-GPU workstation
whichllm --gpu "2x RTX 4090"

# Hide models that are technically runnable but too slow
whichllm --speed usable
whichllm --speed fast

# Pasteable GitHub / Slack / Discord output
whichllm --markdown

# Compare upgrade candidates
whichllm upgrade "RTX 4090" "RTX 5090" "H100"

# Find the GPU needed for a model
whichllm plan "llama 3 70b"

# Start a chat with a model
whichllm run "qwen 2.5 1.5b gguf"

# Print copy-paste Python
whichllm snippet "qwen 7b"

# Return JSON for scripts
whichllm --top 1 --json

demo

See it

$ whichllm --gpu "RTX 4090"

#1  Qwen/Qwen3.6-27B     27.8B  Q5_K_M   score 92.8    27 t/s
#2  Qwen/Qwen3-32B       32.0B  Q4_K_M   score 83.0    31 t/s
#3  Qwen/Qwen3-30B-A3B   30.0B  Q5_K_M   score 82.7   102 t/s

The 32B model fits your card fine — whichllm still ranks the 27B #1, because it scores higher on real benchmarks and is a newer generation. A size-only "what fits?" tool would hand you the bigger one. That gap is the whole point of whichllm. (Note #3: a MoE model at 102 t/s — speed is ranked on active params, quality on total.)

What can I run?

Real top picks (snapshot 2026-05 — your results track live HuggingFace data, this is not a static list):

Hardware VRAM Top pick Speed
RTX 5090 32 GB Qwen3.6-27B · Q6_K · score 94.7 ~40 t/s
RTX 4090 / 3090 24 GB Qwen3.6-27B · Q5_K_M · score 92.8 ~27 t/s
RTX 4060 8 GB Qwen3-14B · Q3_K_M · score 71.0 ~22 t/s
Apple M3 Max 36 GB Qwen3.6-27B · Q5_K_M · score 89.4 ~9 t/s
CPU only — gpt-oss-20b (MoE) · Q4_K_M · score 45.2 ~6 t/s

whichllm --gpu "<your card>" simulates any of these before you buy. By default, rankings include full-GPU, partial-offload, and CPU-only candidates when they are usable. Use --gpu-only or --fit full-gpu when you only want models that fit entirely in GPU VRAM. The default table shows memory, estimated generation speed, fit type, and published date. Speed is colored by practical usability: under 4 tok/s is red, 4-10 is yellow, 10-30 is green, and 30+ is bright green. ~ / ? still mark estimate confidence.

Why whichllm?

Fitting a model into your VRAM is the easy part. The hard part is knowing which of the models that fit is actually the best — and that is what whichllm is built to get right.

  • Evidence-based ranking, not a size heuristic — The top pick is chosen from merged real benchmarks (LiveBench, Artificial Analysis, Aider, multimodal/vision, Chatbot Arena ELO, Open LLM Leaderboard) — never "the biggest model that happens to fit."
  • Recency-aware — Stale leaderboards are demoted along each model's lineage, so a 2024 model can't outrank a current-generation one on an outdated score. The benchmark snapshot date is printed under every ranking, so a stale recommendation is self-evident instead of silently trusted.
  • Evidence-graded and guarded — Every score is tagged direct / variant / base / interpolated / self-reported and discounted by confidence. Fabricated uploader claims and cross-family inheritance (a small fork borrowing its much larger base's score) are actively rejected.
  • Architecture-aware estimates — VRAM = weights + GQA KV cache + activation + overhead; speed is bandwidth-bound with per-quant efficiency, per-backend factors, MoE active-vs-total split, and unified-memory vs discrete-PCIe partial-offload modeling.
  • One command, scriptable — whichllm prints the answer; add --json | jq for pipelines. No TUI, no keybindings to memorize.
  • Live data — Models fetched directly from the HuggingFace API, with curated frozen fallbacks for offline or rate-limited use.

Features

  • Auto-detect hardware — NVIDIA, AMD, Intel, Apple Silicon, CPU-only
  • Smart ranking — Scores models by VRAM fit, speed, and benchmark quality
  • One-command chat — whichllm run downloads and starts a chat session instantly
  • Code snippets — whichllm snippet prints ready-to-run Python for any model
  • Live data — Fetches models directly from HuggingFace (cached for performance)
  • Benchmark-aware — Integrates real eval scores with confidence-based dampening
  • Task profiles — Filter by general, coding, vision, or math use cases
  • GPU simulation — Test with any GPU: whichllm --gpu "RTX 4090"
  • Multi-GPU simulation — Repeat --gpu, use commas, or write 2x RTX 4090
  • Full-GPU filter — --gpu-only / --fit full-gpu hides offload candidates
  • Speed-aware filtering — --speed usable|fast hides slow rows by threshold
  • Markdown output — --markdown / -m prints pasteable GFM tables
  • Runtime memory budgets — --vram-headroom and --ram-budget avoid edge fits
  • Hardware planning — Reverse lookup: whichllm plan "llama 3 70b"
  • Upgrade planning — Compare your current machine with candidate GPUs
  • JSON output — Pipe-friendly: whichllm --json

Run & Snippet

Try any model with a single command. No manual installs needed — whichllm creates an isolated environment via uv, installs dependencies, downloads the model, and starts an interactive chat.

run demo

# Chat with a model (auto-picks the best GGUF variant)
whichllm run "qwen 2.5 1.5b gguf"

# Auto-pick the best model for your hardware and chat
whichllm run

# CPU-only mode
whichllm run "phi 3 mini gguf" --cpu-only

Works with all model formats:

  • GGUF — via llama-cpp-python (lightweight, fast)
  • AWQ / GPTQ — via transformers + autoawq / auto-gptq
  • FP16 / BF16 — via transformers

Get a copy-paste Python snippet instead:

whichllm snippet "qwen 7b"
from llama_cpp import Llama

llm = Llama.from_pretrained(
    repo_id="Qwen/Qwen2.5-7B-Instruct-GGUF",
    filename="qwen2.5-7b-instruct-q4_k_m.gguf",
    n_ctx=4096,
    n_gpu_layers=-1,
    verbose=False,
)

output = llm.create_chat_completion(
    messages=[{"role": "user", "content": "Hello!"}],
)
print(output["choices"][0]["message"]["content"])

Usage

# Auto-detect hardware and show best models
whichllm

# Simulate a GPU (e.g. planning a purchase)
whichllm --gpu "RTX 4090"
whichllm --gpu "RTX 5090"
# Specify variant
whichllm --gpu "RTX 5060 16"
# Override detected iGPU/unified-memory limits
whichllm --vram 8 --ram-bandwidth 68
# Simulate multiple GPUs
whichllm --gpu "2x RTX 4090"
whichllm --gpu "RTX 4090" --gpu "RTX 3090"
whichllm --gpu "RTX 4090, RTX 3090"

# Only show models that fit entirely in GPU VRAM
whichllm --gpu-only
whichllm --fit gpu
whichllm --fit full-gpu

# Avoid edge fits and background-RAM surprises
whichllm --vram-headroom 1.5GB
whichllm --ram-budget available
whichllm --ram-budget 8GB

# CPU-only mode
whichllm --cpu-only

# More results / filters
whichllm --top 20
whichllm --details          # show Downloads metadata instead of runtime columns
whichllm --speed usable     # minimum 10 tok/s
whichllm --speed fast       # minimum 30 tok/s
whichllm --min-speed 4      # exact tok/s floor
whichllm --markdown         # pasteable GitHub-Flavored Markdown table
whichllm --profile coding
whichllm --context-length 64k
whichllm --quant Q4_K_M
whichllm --min-speed 30     # exact tok/s floor
whichllm --evidence base   # allow id/base-model matches
whichllm --evidence strict # id-exact only (same as --direct)
whichllm --direct

# JSON output
whichllm --json

# Force refresh (ignore cache)
whichllm --refresh

# Show hardware info only
whichllm hardware

# Plan: what GPU do I need for a specific model?
whichllm plan "llama 3 70b"
whichllm plan "Qwen2.5-72B" --quant Q8_0
whichllm plan "mistral 7b" --context-length 32768

# Upgrade: compare your current machine against candidate GPUs
whichllm upgrade "RTX 4090" "RTX 5090" "H100"
whichllm upgrade "Apple M4 Max" --top 5

# Run: download and chat with a model instantly
whichllm run "qwen 2.5 1.5b gguf"
whichllm run                       # auto-pick best for your hardware

# Snippet: print ready-to-run Python code
whichllm snippet "qwen 7b"
whichllm snippet "llama 3 8b gguf" --quant Q5_K_M

Markdown output is intended for GitHub issues, READMEs, Slack, Discord, and blog posts:

whichllm --markdown
whichllm -m --top 5 --gpu "RTX 4090"

JSON model rows include fit_type, vram_required_bytes, vram_available_bytes, uses_multi_gpu, multi_gpu_effective_vram_bytes, estimated_tok_per_sec, speed_confidence, speed_range_tok_per_sec, speed_notes, benchmark_source, and benchmark_confidence. The speed range is a planning range, not a live benchmark.

Integrations

Ollama

Use JSON output to feed scripts that map HuggingFace IDs to your local Ollama model names:

# Pick the top HuggingFace model ID
whichllm --top 1 --json | jq -r '.models[0].model_id'

# Find the best coding model ID
whichllm --profile coding --top 1 --json | jq -r '.models[0].model_id'

Ollama model names do not always match HuggingFace repo IDs, so a small mapping step is usually needed before ollama run.

Shell alias

Add to your .bashrc / .zshrc:

alias bestllm='whichllm --top 1 --json | jq -r ".models[0].model_id"'
# Usage: ollama run $(bestllm)

Scoring

Each model gets a 0-100 score. Benchmark quality and size form the core; evidence confidence and runtime fit then scale it, with speed, source trust, and popularity as adjustments.

Factor Effect Description
Benchmark quality core Merged LiveBench / Artificial Analysis / Aider / Vision / Arena ELO / Open LLM Leaderboard, weighted by source confidence
Model size up to 35 log2-scaled world-knowledge proxy (MoE uses total params)
Quantization × penalty Lower-bit quants discounted multiplicatively
Evidence confidence ×0.55–1.0 none / self-reported ×0.55, inherited ×0.78, direct full
Runtime fit ×0.50–1.0 partial-offload ×0.72, CPU-only ×0.50
Speed -8 to +8 Usability gate vs a fit-dependent tok/s floor; reported with confidence and range metadata
Source trust -5 to +5 Official-org bonus, known-repackager penalty
Popularity tie-breaker Downloads/likes; weight shrinks as evidence strengthens

Score markers:

  • ~ (yellow) — No direct benchmark; score inherited/interpolated from the model family
  • !sr (bright yellow) — Uploader-reported benchmark only, not independently verified
  • ? (red) — No benchmark data available

Speed display:

  • red — Slow generation speed (<4 tok/s)
  • yellow — Marginal generation speed (4-10 tok/s)
  • green — Usable generation speed (10-30 tok/s)
  • bright green — Fast local generation speed (>=30 tok/s)
  • ~ (yellow) — Estimated tok/s range is available
  • ? (red) — Low-confidence speed estimate; backend/runtime sensitivity is high

Documentation

How it works

Data pipeline

  1. Model fetching — Fetches popular models from HuggingFace API:

    • Text-generation (downloads + recently updated)
    • GGUF-filtered (separate query for coverage)
    • Vision models (image-text-to-text) when --profile vision or any
  2. Benchmark sources — Current tier (LiveBench, Artificial Analysis Index, Aider) merged live when reachable, plus a curated multimodal / vision index; frozen tier (Open LLM Leaderboard v2, Chatbot Arena ELO). Tiers have separate caps and lineage-aware recency demotion so stale leaderboards stop over-rewarding older generations.

  3. Benchmark evidence — Five resolution levels, increasingly discounted:

    • direct — Exact model ID match
    • variant — Suffix-stripped or -Instruct variant
    • base_model — Base model from cardData
    • line_interp — Size-aware interpolation within model family
    • self_reported — Uploader-claimed eval (heavily discounted)

    Inheritance is rejected when a model's params diverge more than 2× from its family's dominant member, catching draft / MTP / abliterated forks that share a family_id with a much larger base.

  4. Cache — normally ~/.cache/whichllm/, or $XDG_CACHE_HOME/whichllm/ when XDG_CACHE_HOME is set to an absolute path:

    • models.json — 6h TTL
    • benchmark.json — 24h TTL

Ranking engine

  1. Hardware detection — NVIDIA (nvidia-ml-py), AMD (ROCm/dbgpu), Intel, Apple Silicon (Metal), CPU cores, RAM, disk
  2. VRAM estimation — Weights + KV cache + activation + framework overhead (~500MB)
  3. Compatibility — Full GPU / Partial Offload / CPU-only; compute capability and OS checks
  4. Speed — tok/s from GPU memory bandwidth, quantization, backend, fit type, and MoE active parameters
  5. Scoring — Benchmark (with confidence dampening), size, quantization penalty, fit type, speed, popularity, source trust (official vs repackager)
  6. Backend filter — Apple Silicon and CPU-only restrict to GGUF for stability; Linux+NVIDIA allows AWQ/GPTQ

Project structure

src/whichllm/
├── cli.py              # Typer CLI: main, plan, run, snippet, hardware
├── constants.py        # Backward-compatible exports for registry data
├── data/               # GPU, quantization, framework, and lineage registries
├── hardware/
│   ├── detector.py     # Orchestrates GPU/CPU/RAM detection
│   ├── nvidia.py       # NVIDIA GPU via nvidia-ml-py
│   ├── amd.py          # AMD GPU (Linux)
│   ├── apple.py        # Apple Silicon (Metal)
│   ├── cpu.py          # CPU name, cores, AVX support
│   ├── memory.py       # RAM and disk free
│   ├── gpu_simulator.py # --gpu flag: synthetic GPU from name
│   └── types.py        # GPUInfo, HardwareInfo
├── models/
│   ├── fetcher.py      # HuggingFace API, model parsing, evalResults
│   ├── benchmark.py    # Arena ELO, Leaderboard (parquet/rows API)
│   ├── grouper.py      # Family grouping by base_model and name
│   ├── cache.py        # JSON cache with TTL
│   └── types.py        # ModelInfo, GGUFVariant, ModelFamily
├── engine/
│   ├── vram.py         # VRAM = weights + KV cache + activation + overhead
│   ├── compatibility.py# Fit type, disk check, compute/OS warnings
│   ├── performance.py  # tok/s from bandwidth
│   ├── quantization.py # Bytes per weight, quality penalty, non-GGUF inference
│   ├── ranker.py       # Scoring, evidence filter, profile/match
│   └── types.py        # CompatibilityResult
└── output/
    ├── ranking.py      # Rich hardware and recommendation tables
    ├── json_output.py  # Ranking, plan, and upgrade JSON
    ├── plan.py         # plan command display
    ├── upgrade.py      # upgrade comparison display
    └── display.py      # Compatibility re-export shim

Development

git clone https://github.com/Andyyyy64/whichllm.git
cd whichllm
uv sync --dev
uv run whichllm
uv run pytest

Contributing

Contributions are welcome! See CONTRIBUTING.md for guidelines.

Support

If whichllm helped you find a model or avoid a bad hardware guess, sponsoring is appreciated. It helps keep the project maintained: hardware reports, packaging, test fixtures, benchmark updates, and support for more machines.

whichllm will stay open-source either way. Issues and PRs are always welcome.

Useful? A GitHub star helps other people find it, and I'd genuinely like to know what it picked for your rig. Drop it in Issues.

Star History

Star History Chart

Requirements

  • Python 3.11+
  • NVIDIA GPU detection via nvidia-ml-py (included by default)
  • AMD / Apple Silicon detected automatically

License

MIT

View on GitHub

Recent activity

commits and pull requests

Recent open issues

view all

Discussions

all 4

Releases and announcements

20 total
  1. v0.5.19v0.5.19Sep 19, 2026

    You can now use `whichllm -v` as a short form of `whichllm --version`. Thanks @Sophist-UK for the contribution (#159).

  2. v0.5.18v0.5.18Sep 19, 2026

    `plan owner/repo` can now fetch public Hugging Face repositories outside the cached catalog. Model grouping keeps organizations, model versions, checkpoint dates, and instruction/chat variants distinct. JSON `family_id` values change to preserve those distinctions. This release also fixes the ROCm detection case where a generic GPU name hides a recognized shared-memory SKU, and documents the limits of multi-GPU memory-placement estimates. Update with `uv tool upgrade whichllm` or `pip install --upgrade whichllm`. [Full changelog](https://github.com/Andyyyy64/whichllm/blob/v0.5.18/CHANGELOG.md)

  3. v0.5.17v0.5.17Sep 19, 2026

    This release adds newer Qwen models to the candidate list and includes general reasoning models in the math profile. It also fixes crashes when reading non-UTF-8 caches, respects the configured Apple GPU memory limit, and corrects bandwidth lookup for Ada laptop workstation GPUs. Ranking helpers have been split into smaller modules, and the README star-history chart works again. Update with `uv tool upgrade whichllm` or `pip install --upgrade whichllm`. [Full changelog](https://github.com/Andyyyy64/whichllm/blob/v0.5.17/CHANGELOG.md)

  4. v0.5.16v0.5.16Aug 14, 2026

    ## v0.5.16 sorry this release took longer than it should have - Escape Hugging Face model IDs, GGUF filenames, and quantization types in generated `run` and `snippet` scripts - Keep synthetic GGUF recommendations tied to direct quantizations of the selected checkpoint - Reject fine-tunes, merges, conflicting lineage, and artifacts without explicit provenance - Split benchmark fetching, caching, indexing, lineage, and lookup into focused modules - Lock the development tools used locally and in required CI checks ### Security v0.5.16 fixes [CVE-2026-58474](https://www.cve.org/CVERecord?id=CVE-2026-58474) ([GHSA-hfpc-7mr4-p297](https://github.com/advisories/GHSA-hfpc-7mr4-p297)), a code injection vulnerability affecting versions before 0.5.16. A malicious Hugging Face repository could use a crafted GGUF filename to alter Python generated by `whichllm run` or `whichllm snippet`. Exploitation requires the user to run the affected command or execute the generated snippet. Users on an earlier version should upgrade to v0.5.16 or later: ```bash python -m pip install --upgrade "whichllm>=0.5.16" ``` Thanks @RaghavRD for the benchmark refactor and @hannibal-lee for the generated-scri

  5. v0.5.15v0.5.15Jul 3, 2026

    ## v0.5.15 - Resolves ranked GGUF recommendations to the actual downloadable artifact repo and filename when the ranked base model and runnable GGUF live in different Hugging Face repos. - Retunes AA index normalization and refreshes the fallback snapshot for the reworked Artificial Analysis scale. - Validates ranking flags such as `--top`, `--min-speed`, and `--min-params` before ranking. - Splits the Hugging Face fetcher into focused modules while keeping the existing `whichllm.models.fetcher` import surface.

Code frequency

additions and deletions
+13.4K-13.4KWeek of 2026-03-01: +13,395 linesWeek of 2026-03-01: -7,726 linesWeek of 2026-03-08: +2,243 linesWeek of 2026-03-08: -443 linesWeek of 2026-03-15: +0 linesWeek of 2026-03-15: -0 linesWeek of 2026-03-22: +0 linesWeek of 2026-03-22: -0 linesWeek of 2026-03-29: +0 linesWeek of 2026-03-29: -0 linesWeek of 2026-04-05: +0 linesWeek of 2026-04-05: -0 linesWeek of 2026-04-12: +0 linesWeek of 2026-04-12: -0 linesWeek of 2026-04-19: +0 linesWeek of 2026-04-19: -0 linesWeek of 2026-04-26: +0 linesWeek of 2026-04-26: -0 linesWeek of 2026-05-03: +0 linesWeek of 2026-05-03: -0 linesWeek of 2026-05-10: +4,720 linesWeek of 2026-05-10: -674 linesWeek of 2026-05-17: +4,756 linesWeek of 2026-05-17: -1,007 linesWeek of 2026-05-24: +2 linesWeek of 2026-05-24: -0 linesWeek of 2026-05-31: +1,402 linesWeek of 2026-05-31: -478 linesWeek of 2026-06-07: +1,520 linesWeek of 2026-06-07: -126 linesWeek of 2026-06-14: +3,389 linesWeek of 2026-06-14: -1,069 linesWeek of 2026-06-21: +835 linesWeek of 2026-06-21: -38 linesWeek of 2026-06-28: +2,359 linesWeek of 2026-06-28: -1,303 linesWeek of 2026-07-05: +1,629 linesWeek of 2026-07-05: -1,473 linesWeek of 2026-07-12: +0 linesWeek of 2026-07-12: -0 linesWeek of 2026-07-19: +0 linesWeek of 2026-07-19: -0 linesWeek of 2026-07-26: +0 linesWeek of 2026-07-26: -0 linesWeek of 2026-08-02: +130 linesWeek of 2026-08-02: -40 linesWeek of 2026-08-09: +603 linesWeek of 2026-08-09: -27 linesWeek of 2026-08-16: +0 linesWeek of 2026-08-16: -0 linesWeek of 2026-08-23: +0 linesWeek of 2026-08-23: -0 linesWeek of 2026-08-30: +0 linesWeek of 2026-08-30: -0 linesWeek of 2026-09-06: +0 linesWeek of 2026-09-06: -0 linesWeek of 2026-09-13: +825 linesWeek of 2026-09-13: -108 linesWeek of 2026-09-20: +8 linesWeek of 2026-09-20: -2 linesMar 1, 2026Sep 20, 2026
+37.8K lines added, -14.5K removed over the last year.

Commits per week

last 52 weeks
530Week of 2025-09-27: 0 commitsWeek of 2025-10-04: 0 commitsWeek of 2025-10-11: 0 commitsWeek of 2025-10-18: 0 commitsWeek of 2025-10-25: 0 commitsWeek of 2025-11-01: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 53 commitsWeek of 2026-03-08: 29 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 0 commitsWeek of 2026-03-29: 0 commitsWeek of 2026-04-05: 0 commitsWeek of 2026-04-12: 0 commitsWeek of 2026-04-19: 0 commitsWeek of 2026-04-26: 0 commitsWeek of 2026-05-03: 0 commitsWeek of 2026-05-10: 37 commitsWeek of 2026-05-17: 44 commitsWeek of 2026-05-24: 1 commitsWeek of 2026-05-31: 9 commitsWeek of 2026-06-07: 10 commitsWeek of 2026-06-14: 13 commitsWeek of 2026-06-21: 6 commitsWeek of 2026-06-28: 9 commitsWeek of 2026-07-05: 2 commitsWeek of 2026-07-12: 0 commitsWeek of 2026-07-19: 0 commitsWeek of 2026-07-26: 0 commitsWeek of 2026-08-02: 2 commitsWeek of 2026-08-09: 5 commitsWeek of 2026-08-16: 0 commitsWeek of 2026-08-23: 0 commitsWeek of 2026-08-30: 0 commitsWeek of 2026-09-06: 0 commitsWeek of 2026-09-13: 10 commitsWeek of 2026-09-20: 1 commitsSep 27, 2025Sep 20, 2026
231 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 2 commitsSun 2:00 — 2 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 9 commitsSun 6:00 — 3 commitsSun 7:00 — 2 commitsSun 8:00 — 0 commitsSun 9:00 — 0 commitsSun 10:00 — 2 commitsSun 11:00 — 0 commitsSun 12:00 — 0 commitsSun 13:00 — 0 commitsSun 14:00 — 0 commitsSun 15:00 — 0 commitsSun 16:00 — 0 commitsSun 17:00 — 1 commitsSun 18:00 — 0 commitsSun 19:00 — 1 commitsSun 20:00 — 1 commitsSun 21:00 — 4 commitsSun 22:00 — 1 commitsSun 23:00 — 0 commitsMon 0:00 — 0 commitsMon 1:00 — 0 commitsMon 2:00 — 0 commitsMon 3:00 — 7 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 0 commitsMon 8:00 — 0 commitsMon 9:00 — 1 commitsMon 10:00 — 0 commitsMon 11:00 — 1 commitsMon 12:00 — 0 commitsMon 13:00 — 3 commitsMon 14:00 — 0 commitsMon 15:00 — 12 commitsMon 16:00 — 4 commitsMon 17:00 — 0 commitsMon 18:00 — 1 commitsMon 19:00 — 7 commitsMon 20:00 — 4 commitsMon 21:00 — 1 commitsMon 22:00 — 1 commitsMon 23:00 — 1 commitsTue 0:00 — 1 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 0 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 2 commitsTue 9:00 — 0 commitsTue 10:00 — 0 commitsTue 11:00 — 1 commitsTue 12:00 — 1 commitsTue 13:00 — 0 commitsTue 14:00 — 1 commitsTue 15:00 — 1 commitsTue 16:00 — 0 commitsTue 17:00 — 0 commitsTue 18:00 — 0 commitsTue 19:00 — 0 commitsTue 20:00 — 0 commitsTue 21:00 — 0 commitsTue 22:00 — 0 commitsTue 23:00 — 0 commitsWed 0:00 — 2 commitsWed 1:00 — 0 commitsWed 2:00 — 0 commitsWed 3:00 — 0 commitsWed 4:00 — 3 commitsWed 5:00 — 1 commitsWed 6:00 — 0 commitsWed 7:00 — 2 commitsWed 8:00 — 0 commitsWed 9:00 — 0 commitsWed 10:00 — 1 commitsWed 11:00 — 0 commitsWed 12:00 — 2 commitsWed 13:00 — 0 commitsWed 14:00 — 4 commitsWed 15:00 — 0 commitsWed 16:00 — 2 commitsWed 17:00 — 1 commitsWed 18:00 — 0 commitsWed 19:00 — 1 commitsWed 20:00 — 0 commitsWed 21:00 — 2 commitsWed 22:00 — 2 commitsWed 23:00 — 9 commitsThu 0:00 — 0 commitsThu 1:00 — 5 commitsThu 2:00 — 0 commitsThu 3:00 — 8 commitsThu 4:00 — 3 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 9 commitsThu 8:00 — 0 commitsThu 9:00 — 0 commitsThu 10:00 — 1 commitsThu 11:00 — 0 commitsThu 12:00 — 1 commitsThu 13:00 — 5 commitsThu 14:00 — 1 commitsThu 15:00 — 2 commitsThu 16:00 — 10 commitsThu 17:00 — 10 commitsThu 18:00 — 3 commitsThu 19:00 — 0 commitsThu 20:00 — 1 commitsThu 21:00 — 7 commitsThu 22:00 — 1 commitsThu 23:00 — 1 commitsFri 0:00 — 0 commitsFri 1:00 — 2 commitsFri 2:00 — 0 commitsFri 3:00 — 0 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 0 commitsFri 8:00 — 0 commitsFri 9:00 — 0 commitsFri 10:00 — 0 commitsFri 11:00 — 2 commitsFri 12:00 — 2 commitsFri 13:00 — 1 commitsFri 14:00 — 1 commitsFri 15:00 — 4 commitsFri 16:00 — 8 commitsFri 17:00 — 2 commitsFri 18:00 — 1 commitsFri 19:00 — 4 commitsFri 20:00 — 0 commitsFri 21:00 — 1 commitsFri 22:00 — 2 commitsFri 23:00 — 3 commitsSat 0:00 — 2 commitsSat 1:00 — 0 commitsSat 2:00 — 0 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 1 commitsSat 11:00 — 0 commitsSat 12:00 — 0 commitsSat 13:00 — 0 commitsSat 14:00 — 1 commitsSat 15:00 — 4 commitsSat 16:00 — 1 commitsSat 17:00 — 0 commitsSat 18:00 — 0 commitsSat 19:00 — 2 commitsSat 20:00 — 0 commitsSat 21:00 — 2 commitsSat 22:00 — 4 commitsSat 23:00 — 3 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits204 (76%)
Community commits63 (24%)

267 commits in total over the last year.

DateListRankStars gained
Jun 9, 2026daily#23+17