Tools/raullenchai/Rapid-MLX

Rapid-MLX

The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.

3.4kemergingPythonApache License 2.0trending

The Lens

By Erik Loyd, SaaS CEO and former COO/CFO of an AWS Premier Partner.

Updated Aug 2026

Rapid-MLX runs open models locally on Apple Silicon, and speed is the whole pitch. One command installs it (Homebrew, pip, or a curl script), one command serves an OpenAI-compatible API, and it also speaks the Anthropic messages format, so Claude Code, Codex CLI, and Aider can point at it as a drop-in backend. Apache 2.0, completely free.

The engineering focus is on the things that make local agent workflows fail in practice. A prompt cache brings claimed time-to-first-token under a tenth of a second on repeat prompts, and tool calling is handled across 17 parser formats, which matters because local models break agent loops on malformed tool calls far more often than on reasoning. It's in Homebrew core and publishes to PyPI, with releases landing near daily.

Mac-only by design. If you run agents against local models on an M-series machine, try it head-to-head with Ollama or LM Studio; the performance claims ship with benchmarks in the repo rather than vibes. On Linux or Windows, look at Ollama or vLLM instead.

The catch: it's young and moves extremely fast, the headline numbers are the project's own, and independent benchmarks are still thin. And no inference engine fixes the ceiling: a fast small model is still a small model.

Free vs Self-Hosted vs Paid

fully free

Free: Everything. Apache 2.0, installable via Homebrew, pip, or the project's install script. Chat CLI, OpenAI-compatible server, Anthropic-format endpoint, tool calling, prompt cache, and optional vision/audio extras all ship in the open package.

Self-hosted: The only mode, and it's one binary on your Mac. No accounts, no keys, no telemetry toll.

Paid: Nothing. There's no cloud tier or commercial edition. The spend is hardware: an Apple Silicon Mac with enough unified memory for the models you want to run.

Completely free and open source. Your only spend is the Mac it runs on.

Self-hosting ops:trivial

Get tools like this every Wednesday

One featured tool, three on the radar. No fluff.

Similar Tools

Score
74/100 · B+
Adoption17/30
Maintenance25/25
Community7/20
License15/15
Analysis10/10

A low score is not a verdict on quality. Young and niche tools start low by design. How we calculate scores

Trust Signals

Community discussions enabled

License: Apache License 2.0

Use freely. Patent grant included.

Commercial use: ✓ Yes

About

Owner
Raullen Chai (User)
Stars
3,441
Forks
394

Explore Further

More tools in the directory