
llmfit
Hundreds of models & providers. One command to find what runs on your hardware.
The Lens
By Erik Loyd, SaaS CEO and former COO/CFO of an AWS Premier Partner.
Updated Sep 2026
llmfit answers the question everyone running local AI asks first: will this model fit on my machine? It reads your CPU, RAM, and GPU, then scores every model in its catalog on memory fit, estimated speed, quality, and context length. No more downloading a huge model to find out it will not load. MIT licensed, free, one Rust binary.
Install with Homebrew, Scoop, MacPorts, uv, Docker, or a curl script, on macOS, Linux, or Windows. It detects NVIDIA, Apple Silicon, AMD ROCm, and Intel GPUs, understands GGUF, AWQ, GPTQ, and EXL2 quantization, and works with Ollama, llama.cpp, MLX, Docker Model Runner, and LM Studio. No account, and the README says nothing leaves your machine unless you ask. There is also a web dashboard and a REST API through llmfit serve.
Free at every size with no paid tier. Solo builders get the most out of it. Teams can run the API on shared GPU boxes to see what fits where. Once you have picked a model, ollama/ollama or ggml-org/llama.cpp does the actual running.
The catch: speeds are estimates from a formula unless you run the benchmark against a live runtime, and the GPU bandwidth table only covers about 80 cards. Windows and Intel Macs get weaker detection. The model catalog is baked in at compile time, so new models only show up when you update.
Free vs Self-Hosted vs Paid
fully freeFree: Everything. MIT licensed CLI and terminal UI, a web dashboard, a REST API, and a built-in catalog of hundreds of models sourced from Hugging Face.
Self-hosted: A single binary on your own machine. No account, no model download needed to get recommendations. The benchmark command needs a local runtime like Ollama already running.
Paid: None. No hosted version, no paid tier, no sponsor-gated features.
Completely free and open source.
What to do by team size
- Solo
- free
- Small team
- free
- Medium team
- free; run the API on shared GPU machines
- Large team
- free
Get tools like this every Wednesday
One featured tool, three on the radar. No fluff.
A low score is not a verdict on quality. Young and niche tools start low by design. How we calculate scores
Trust Signals
License: MIT License
Use freely, including commercial. Just keep the license.
Commercial use: ✓ Yes
About
- Owner
- Alex Jones (User)
- Stars
- 36,960
- Forks
- 2,346
Explore Further
More tools in the directory
everything-claude-code
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
264.9k ★hermes-agent
The agent that grows with you
247.9k ★ollama
Get up and running with Kimi-K2.5, GLM-5, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
181.2k ★