llm-fit

module
v1.6.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 11, 2026 License: Apache-2.0

README

llm-fit

CI Code Quality Security License OpenSSF Scorecard CI carbon

Which LLMs this machine can actually run, and how fast. Reads the hardware, does the memory and bandwidth arithmetic, and recommends a model, a quantization and a runtime — with an estimate of how many tokens per second you will get.

$ llm-fit suggest
Context 8k, KV f16. Ranked by capability among plans that run usable or better.

MODEL                              RUNTIME     QUANT         SIZE    DECODE  CTXMAX
Qwen3 30B-A3B (MoE)                llama.cpp   Q4_K_S    17.1 GiB     20/s     32k  good  29/48 layers on GPU
Qwen2.5 14B Instruct               ExLlamaV2   EXL2-4.0   8.6 GiB     33/s     14k  excellent
Phi-4 14B                          ExLlamaV2   EXL2-4.0   8.6 GiB     33/s     13k  excellent
Mistral Small 24B                  llama.cpp   IQ3_M     11.3 GiB     13/s     32k  usable  37/40 layers on GPU
Mistral Nemo 12B                   ExLlamaV2   EXL2-4.0   7.2 GiB     40/s     24k  excellent

Install

macOS, via Homebrew:

brew install fabiocicerchia/tap/llm-fit

Linux — a .deb, .rpm, .apk or Arch package from the latest release:

sudo dpkg -i llm-fit_*_linux_amd64.deb     # or rpm -i / apk add --allow-untrusted

Or with Go:

go install github.com/fabiocicerchia/llm-fit/cmd/llm-fit@latest

Or from a checkout:

make build      # -> ./bin/

Use

llm-fit detect                    # what is here, and what it implies
llm-fit suggest                   # models that run well, best first
llm-fit suggest -ctx 32768 -kv q8_0
llm-fit suggest -serving          # optimise for concurrency
llm-fit check qwen3-30b           # every quant × runtime for one model
llm-fit check -hf Qwen/Qwen3-14B  # any model, read from Hugging Face
llm-fit check ~/models/q4.gguf    # the file on disk: its own shape and quant
llm-fit engines                   # what this machine can run, and why not
llm-fit models                    # the built-in catalogue

-batch N for concurrent sequences, -engine NAME to restrict to one runtime, -top N for list length, and -min LEVEL / -min-quality N for the floors below which a result is not worth showing (default: usable, and no sub-3-bit).

Plan for hardware you do not have yet with -gpu, which takes VRAM, bandwidth and compute together from the spec table — overriding VRAM alone would model your card with someone else's memory capacity, which is nobody's hardware:

llm-fit suggest -gpu "RTX 4090"
llm-fit suggest -gpu "A100 80GB" -ram 256 -serving
llm-fit suggest -gpu "M4 Max"

-vram, -ram and -ram-bandwidth override individual figures. -json for everything.

System RAM bandwidth is measured, not assumed

It cannot be read without root, and it is what every CPU-offload estimate divides by. A table keyed on the CPU would be guessing about the memory fitted — the same Ryzen runs one DDR4 stick or four DDR5 ones, a 4× spread. So llm-fit measures it: a few hundred milliseconds of multi-threaded copy over buffers far larger than L3. On the machine this was written on that returns 43 GB/s against a 51.2 GB/s theoretical peak, where the assumption would have been 80.

Verify the download

Every release is signed with cosign, keyless: the identity is the workflow that published it, not a key anybody holds.

cosign verify-blob \
  --bundle checksums.txt.bundle \
  --certificate-identity-regexp 'https://github.com/fabiocicerchia/llm-fit' \
  --certificate-oidc-issuer https://token.actions.githubusercontent.com \
  checksums.txt
sha256sum --ignore-missing -c checksums.txt

Documentation

Full docs live in docs/. Runnable examples live in examples/.

License

Apache-2.0 — see LICENSE.

Directories

Path Synopsis
cmd
llm-fit command
llm-fit — which models this machine can actually run, and how well.
llm-fit — which models this machine can actually run, and how well.
internal
advisor
Package advisor turns the arithmetic into a recommendation.
Package advisor turns the arithmetic into a recommendation.
arch
Package arch describes a model's shape, which is all the memory and speed math actually depends on.
Package arch describes a model's shape, which is all the memory and speed math actually depends on.
catalog
Package catalog is the built-in list of models worth considering, with the architecture fields the math needs.
Package catalog is the built-in list of models worth considering, with the architecture fields the math needs.
engine
Package engine describes what each runtime can actually load and where it can run, which is half the advice: a plan that is arithmetically sound and names a format the runtime cannot read is not a plan.
Package engine describes what each runtime can actually load and where it can run, which is half the advice: a plan that is arithmetically sound and names a format the runtime cannot read is not a plan.
fit
Package fit is the arithmetic the tool exists for: does this model fit, and if it fits, is it fast enough to be worth running.
Package fit is the arithmetic the tool exists for: does this model fit, and if it fits, is it fast enough to be worth running.
gguf
Package gguf reads a model's shape out of a GGUF file's own header.
Package gguf reads a model's shape out of a GGUF file's own header.
hfapi
Package hfapi reads a model's shape from Hugging Face, for anything not in the built-in catalogue.
Package hfapi reads a model's shape from Hugging Face, for anything not in the built-in catalogue.
hw
Package hw works out what the machine actually has.
Package hw works out what the machine actually has.
quant
Package quant holds the bits-per-weight of every quantization format worth running, and which runtimes can load it.
Package quant holds the bits-per-weight of every quantization format worth running, and which runtimes can load it.
validate
Package validate compares predicted tokens/sec against measured ones.
Package validate compares predicted tokens/sec against measured ones.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL